Two hours of conversation produced five book recommendations the writer had never heard of and would actually read, without logging a single click. That is the entire empirical basis for a question worth taking seriously: if an LLM can elicit what you like by asking, does a recommender still need to watch what you do? The honest answer, based on the evidence available as of 2026-10-09, is that the architecture is worth prototyping for portability and data minimization, but the privacy benefits are the author’s own untested hypotheses, and the adversarial literature cuts in the opposite direction: preference data is itself an inference target, whether it sits in an analyst’s rating database, a server-held model artifact, or a federated update.
Everything about the motivating case comes from a single Hacker News thread posted on 2026-10-09 by a self-described writer, not an engineer, with no plans to develop the prototype further. There is no user study, no benchmark, and no replication. Treat every operational claim from it as one person’s self-report, not a finding.
The prototype: two hours of taste talk, zero behavioral logs
In the thread, the author describes sitting down with an LLM, “ChatGPT, Luna, for the curious”, and spending “about two hours explaining, in datail, my literary tastes- going into specifics like what sort of plot structure, characters, and humor I like, not just genre/specific titles.” The result, in their account: “the five recommendations I asked for were all ones that I hadn’t heard of, as well as books that I would actually read.” As a check, they then gave the model a list of books they had already read and asked it to guess whether they would have liked each one and why; the author reports the model succeeded.
This is a plausible demonstration of feasibility and nothing more. One person, one domain, self-scored results, no control condition. It does not show that conversational elicitation matches click-history personalization on quality, only that it can produce a satisfying result for at least one engaged, articulate user with two hours to spend. That is a real constraint on generalization: the author is a professional writer describing literary taste, which is close to the best case for verbal self-report.
What makes the thread interesting for builders is not the demo but the two ideas the author raises afterward, both flagged in the post as guesses. First: “it seems like the user data could potentially be minimized and pseudonymized before being sent to the LLM, though I lack the technical expertise to know what is realistic here.” Second: “if the LLM is conducting any searches, this would seem to make it possible to provide search services with less information about the user and their preferences (I’m not positive about this, so please do correct me if I’m wrong).” Both quotes are from the thread. Neither is a finding. They are the design hypotheses this article tests against the literature.
Who observes what: platform, LLM provider, search backend
The click-history architecture and the taste-profile architecture distribute visibility differently, and that distribution is the actual decision. The comparison below maps what each party observes under each design. Rows about provider behavior are architectural inferences, not documented facts: no cited source covers any LLM API provider’s retention or training-use policy; verify provider terms before shipping.
| Party | Click-history personalization | Conversational taste profile |
|---|---|---|
| Platform / analyst | Full behavioral log: clicks, ratings, dwell, sequences; research shows such data enables inference of political affiliation, sexual orientation, age, gender, and drug use | Sees only what crosses its servers; if the profile is user-held, potentially nothing beyond search traffic |
| LLM API provider | Nothing (unless platform uses an LLM layer) | The elicitation conversation, or a minimized/pseudonymized profile if the author pre-processes it; retention and training use [unverified] |
| Search backend | Raw user queries tied to an identity or session | Queries the LLM constructs from the profile; user intent is filtered through the model, per the author’s hypothesis |
| User | Little visibility into the inferred profile | Holds the raw profile; can read, edit, delete, or carry it to a competitor |
Two things follow from this mapping. The first is the author’s second hypothesis made concrete: profile-steered search genuinely changes what a search backend sees, because the query stream now reflects what the model decided to ask rather than what the user typed. That is a real architectural difference, though “less information” is doing quiet work in it; a well-constructed query derived from a taste profile still encodes the taste. The second is the one the thread underplays: the LLM API provider becomes the new concentration point. Every preference you would previously have leaked to a platform’s ranking system now crosses one API as an explicit, summarized, high-signal preference stream. You have not removed the observer. You have moved it, and arguably upgraded the fidelity of what it sees, because a stated profile is cleaner than inferred behavior.
Where the profile lives
Storage location is the axis where the literature is genuinely encouraging. A 2018 paper on privacy-conscious profiling in recommender systems shows that “a privacy-conscious user can learn her profile without revealing any information to the analyst. The protocol is practical and proven secure against semi-honest adversaries.” In other words, the technical core of the taste-profile idea, a user-held profile the platform never sees in the clear, is not speculative. A protocol exists, is proven secure under a stated adversary model, and is described as practical.
Two caveats keep this from being a green light. The protocol addresses a specific problem (letting a user learn her own profile without disclosure) inside a conventional recommender setting, not the full pipeline of an LLM-elicited profile that is then sent over an API to steer search. And the same paper supplies the strongest evidence for why the stakes are real: “Recent research has indeed demonstrated that this can be used by the analyst to infer private user attributes such as political affiliation, sexual orientation, age, gender, and even drug use.” That finding is about behavioral data like ratings, but the logic transfers uncomfortably well. A taste profile is preference data with the volume turned up. If ratings leak private attributes, a two-hour verbal profile of your humor, your politics-adjacent reading, and your emotional triggers is a richer target, not a safer one.
The portability angle, though, stands on its own merits even if the privacy angle wobbles. A profile the user holds is a profile the user can take elsewhere. That changes the economics of recommendation: behavioral exhaust is an asset the platform accumulates and the user cannot carry, while a taste file is an asset the user owns and can offer to a competitor on day one. That asymmetry is also the plausible reason incumbents are unlikely to build this for you. A platform whose discovery moat is its click history has no incentive to hand you a portable substitute. This is a structural argument, not a documented one: none of the sources here establish anything about any incumbent’s monetization or product strategy. Treat it as the article’s thesis, and weigh it accordingly.
Keeping it fresh: the elicitation problem
Click histories update themselves. Every session adds evidence, for better and worse. A taste profile goes stale the moment the conversation ends, and refreshing it costs the user another conversation. This is the taste-profile architecture’s least-discussed operational cost, and the conversational-recommender literature says the cost is not flat.
Research on conversational styles in preference elicitation finds that “adapting conversational strategies based on user expertise and allowing flexibility between styles can enhance both user satisfaction and the effectiveness of recommendations in CRSs.” Read against the HN prototype, this matters twice over. It suggests the author’s two-hour deep interview is one design point on a real spectrum, not the design. A novice user who cannot articulate plot-structure preferences needs a different conversational strategy than a professional writer, and the system has to know which it is talking to. It also suggests that elicitation quality is a tunable product surface: how the system asks determines what it learns, which means profile quality is partly an engineering problem and partly a skill-of-the-user problem. Click history makes no such demand on the user. For a product team, that is the honest tradeoff: you are replacing passive surveillance with active user labor, and some users will not or cannot do the labor well.
Quality control: the out-of-catalog failure mode
One controlled user study of a deployed LLM recommender supplies the clearest operational warning. In a study of an LLM-based conversational recommender on Videoland, a video-on-demand streaming platform, “For both recommenders, over 22% of tasks had user-selected recommendations that were not available on Videoland, suggesting that the system occasionally generated recommendations beyond the platform’s content availability, despite our attempts to control it.” This is a single-team study and the figure comes from one platform pairing, so treat it as a demonstrated failure mode rather than a universal rate. But it is the only measured number on out-of-catalog recommendations, and it points at a structural property of open-ended LLM recommendation: the model draws on what it knows, not on what your catalog holds.
Click-history recommenders cannot make this mistake. They rank items that exist. An LLM steering search from a taste profile will happily recommend a book that is out of print, a film on a rival service, or a plausible-sounding title that does not exist at all. For the HN author this was arguably a feature: unfamiliar books are the point. For a product team it is a catalog-matching pipeline you now have to build, with availability checks, fallback behavior, and a decision about whether “not on our platform” is a dead end or a referral. That cost belongs in the architecture comparison, and it is easy to miss because the demo feels magical before you count the misses.
A minimization checklist
If you build this, the evidence supports treating it as a data-minimization and portability exercise with privacy claims held at hypothesis status. Concretely:
- Minimize before the API call. The author’s first hypothesis is the right instinct: send the smallest profile that does the job, stripped of direct identifiers, not the raw two-hour transcript. What “realistic” pseudonymization means for preference data is unresolved; preference profiles are inherently identifying, so test re-identification rather than assuming the label does the work.
- Keep the raw profile under user control. The secure profile-retrieval protocol shows user-held profiles are technically tractable. Store the full profile on the user’s device or user-chosen storage; send derived, minimized views.
- Audit what each observer receives. Map your actual pipeline against the observer table above: what crosses the LLM API, what reaches the search backend, what you log yourself. Verify provider retention and training-use policies contractually; no cited source covers those terms.
- Test leakage empirically. Run membership-inference and attribute-inference evaluations against your model artifacts before claiming privacy superiority, using the methods in the white-box inference and attribute-inference literature as the bar.
- Handle the outlier users. If your best users have rare tastes, apply the federated privacy-disparity finding as a design constraint: distinctive profiles need stronger protection, not just the same protection. That finding comes from federated learning, where outlier clients face disproportionate inference risk; extending it to user-held taste profiles is an analogy, not a measured result.
- Budget for elicitation quality and refresh. Match conversational strategy to user expertise per the elicitation research, and decide how profiles expire or update. A stale profile is a wrong profile presented with confidence.
- Build the availability layer. Assume a sizable share of open-ended recommendations will point outside your catalog, per the VideolandGPT study, and decide what happens then.
Verdict
Build the taste profile if your goal is portability, user agency, or escape from a ranking moat that serves incumbents and buries non-obvious work. The feasibility case is real: one writer got five good, unfamiliar recommendations from a conversation, and user-held profiles have cryptographic backing. Do not build it on the claim that it is more private. That claim is the author’s own flagged guess, the counter-literature shows preference data leaks private attributes and model artifacts leak membership, and the distinctive users who benefit most face the most inference risk. No cited source covers provider retention policy, so even the trust-shift story rests on an unchecked link.
What would settle it is unglamorous: a user study comparing elicited profiles against click histories on recommendation quality, an inference-attack evaluation on a real taste-profile pipeline, and published provider data-handling terms. Until those exist, the honest position is the author’s own: a promising prototype, two good hypotheses, and an open question about who, in the end, owns what you like. The answer this architecture offers, you do, with caveats, is better than the default, and not yet proven.
Frequently Asked Questions
What is the main operational risk of using an LLM for recommendations?
For both recommenders, over 22% of tasks had user-selected recommendations that were not available on Videoland, suggesting that the system occasionally generated recommendations beyond the platform’s content availability, despite our attempts to control it.
How does a taste profile affect what search engines see?
profile-steered search genuinely changes what a search backend sees, because the query stream now reflects what the model decided to ask rather than what the user typed. That is a real architectural difference, though “less information” is doing quiet work in it; a well-constructed query derived from a taste profile still encodes the taste.
What is the primary benefit of a user-held taste profile?
A profile the user holds is a profile the user can take elsewhere. That changes the economics of recommendation: behavioral exhaust is an asset the platform accumulates and the user cannot carry, while a taste file is an asset the user owns and can offer to a competitor on day one.

Join the discussion
Share a useful perspective or ask a question about this article.