Articles — Whether that list is for you
Whether that list is for you
Say what you want the answer to do
You added something about yourself and still got a generic answer. Saying who you are is not the same as saying what the answer has to do.
You asked for a few days away and said you were going with your parents. The answer still suggested nightlife and a long day of walking, the same plan it would give someone going alone.
Without your context, the answer is for someone else is about the blank. This one is about what you fill it with. Not every kind of context does the same work.
What a 2025 study actually measured
In August 2025, researchers published Revealing Potential Biases in LLM-Based Recommender Systems in the Cold Start Setting. They tested models that had no purchase history, only the words in the prompt. They call that a cold start. The author names and the limits of the test are on Resources.
They did not test ChatGPT. They tested two open-source models, Gemma 3 and Llama 3.2, on music, movies, and colleges. Gemma was the more stable of the two in their tests. The job was re-ranking: the model had to pick and order items from a catalog it was given, not search the open web.
A neutral prompt, with no extra labels, was the baseline. Prompts that stated gender, age, or language produced different lists, including stereotype-shaped ones. Then, on the movie catalog with Gemma 3 12B, they added one task-relevant preference, “action movie fan,” next to those labels. The lists for different groups moved closer together. The shift was largest where the label-only gap had been biggest.
The model had a default. A preference the answer could satisfy outweighed a weaker stereotype signal.
A label is not a constraint
It is easy to read that as “always say more about yourself.” That is not what the test shows. What moved the recommendations was what the person wanted from the movies, not a demographic tag.
“I’m 42” is a label. “I want something that will not wear out on long training runs, from a shop in my city” is what you need the answer to do. “Short walks and an early dinner, because I am going with my parents” is the same kind of detail for a trip. “A stop on this route, not the tourist loop” is that detail for a drive. “Best things to do near me” stays vague until you say what the answer has to do, where, and for whom. Ask for a white t-shirt and you get Uniqlo is the shopping version of that blank ask. The glossary calls that match constraint fit.
The paper’s “I’m a girl” probe is not the same as stating where you buy, who you travel with, or what the route has to be. Those are still the preference kind of context. They are not the same as a stereotype label.
What this does not prove
The paper is an arXiv preprint. The lists came from Gemma and Llama re-ranking a closed catalog of music, movies, and colleges. It does not measure ChatGPT, Perplexity, or Gemini on a travel or product ask. The mechanism, a default that a stated preference can override, is reasonable to expect elsewhere. Reasonable is not the same as measured.
If you are the person asking
Look at what you typed. If it is only facts about you, the chatbot still has to invent the job. If it says what the thing has to do, who it is for, or where it has to work, the answer has something to match.
If you run a shop or a place
A page that only says who you are for (“for millennials,” “luxury,” “family-friendly”) and never names hours, route, size, or use competes on whoever the assistant imagines. One clear sentence the ask can match does more than a persona label.
Read next if: Ask the same thing in two tools · Why AI won’t just say it doesn’t know