Articles — How shortlists work

How shortlists work

Why AI won't just say it doesn't know

You ask where to buy a product in your country and get a shop and a price. When there is no current listing, the honest answer is that it does not know. These tools are built to answer anyway.

Illustrative example, not a logged test.

You ask where to buy Northloop creatine in your country. Two kinds of answer are both possible. One is honest. The other is what you usually get.

Honest

I don’t have a current shop or price for Northloop creatine in your country. Check Northloop’s own site for where it ships and what it costs today.

Filled in

Try SportHub — around €34 for a 500g tub, usually in stock. Same tone as a price check that actually happened.

SportHub may never have carried the line. The €34 may come from an old roundup. The filled-in version still reads like a finished recommendation. Wrong price, wrong shop, fake discount catalogs that failure mode. Here the focus is why the honest answer loses.

The four big assistants are built on different training stacks, so this is not one company’s design choice. The same tendency to fill a gap rather than hedge shows up across all of them. That still doesn’t tell you whether the particular answer you got today was wrong.

If you just asked the question, the practical takeaway is already here: a fluent, complete answer is something to verify, not something to trust on tone. The rest of this piece is why the tools are built that way.

The exam with no penalty for guessing

A September 2025 paper by researchers at OpenAI and Georgia Tech, Why Language Models Hallucinate (Kalai, Nachum, Vempala, Zhang), argues that much of what we call hallucination is predictable from how models get evaluated, rather than from one mysterious bug inside them.

Most public benchmarks score answers like a multiple-choice test with no penalty for a wrong answer and no credit for leaving a question blank. Under those rules, guessing is the winning strategy even when you expect to be right less than half the time, in the same way that a student who answers everything beats one who skips the hard questions when there is no negative marking.

On the creatine question, a wrong shop name and a right one both count as an answer. I don’t know, check the brand site counts as a blank. Training and public leaderboards mostly score the first two kinds the same way and give no credit for the third, so a model learns to name SportHub and a price even when both are stale.

The authors connect that to a statistical floor on rare facts. The singleton rate is the share of training prompts that appear exactly once with a real answer rather than an “I don’t know.” When a fact has no learnable pattern behind it, and the paper uses birthdays as its example, the model’s generative error rate cannot fall much below that singleton share, because a fact seen only once is hard to tell apart from noise. They also show that generative error runs at least about twice the rate at which the same model gets “is this statement valid?” wrong, so producing an answer is strictly harder than judging one.

Their proposed fix is behavioral calibration, which means scoring abstention explicitly: +1 for a correct answer, −2 for a wrong one, 0 for “I don’t know.” They are honest about how far that goes. A single hallucination-specific benchmark achieves little while hundreds of accuracy-only leaderboards keep rewarding guesses, which makes this an industry-wide scoring problem rather than something one fine-tuning run can fix.

The second cause: tuning rewards sounding sure

Benchmark design is one lever. RLHF (reinforcement learning from human feedback), the step that turns a base model into a polished assistant, is another.

Zhou et al. (Relying on the Unreliable) probed the reward model behind OpenAssistant, which is trained on human preference data. Plain confident phrasing scored 4.03 on average, phrases that hedge while still asserting (“strengtheners”) scored 0.82, and uncertainty language (“weakeners”) scored −1.86. Human annotators in the same study were far gentler toward those weakeners, scoring them 0.095. So the penalty for sounding uncertain sits mostly in the reward model that RLHF optimizes against, not in what people say they want when you ask them directly.

Those categories show up in the shopping answer. SportHub, €34, usually in stock is the confident phrasing the reward model liked. You might find it at SportHub for around €34 is a strengthener: still a guess, but softer. I’m not sure which shop stocks Northloop creatine in your country is a weakener, and the same study penalized that kind of line heavily in the model it measured. The scores come from Zhou et al.’s phrasing tests, not from a logged creatine prompt, but the pattern matches what you hear when a gap gets filled.

Sharma et al. at Anthropic (Towards Understanding Sycophancy in Language Models, ICLR 2024) found that preference data can favor responses which match what the user already believes over correct ones that push back, and that a convincingly written sycophantic answer beats a correct hedged one a non-negligible share of the time. If you add I think I saw it at SportHub last month, the same incentive can push the model to agree with you and lock in the wrong shop instead of correcting you.

Liang et al. (Machine Bullshit) borrow a distinction from Harry Frankfurt: someone lying knows the truth and denies it, while someone bullshitting simply doesn’t care about the truth as long as the answer sounds good. In their tests, the measure of truth-indifference climbs sharply after RLHF, and user satisfaction climbs with it.

Shapira, Benade, and Procaccia (How RLHF Amplifies Sycophancy) show how the effect compounds. Optimization does not merely preserve a slight human preference for agreeable answers, it magnifies it.

What this does and does not prove

This explains why a missing stock line or an unknown shop URL tends to come back as a filled-in sentence rather than silence. It does not mean every confident product fact is invented, because a stale page, the wrong source, and a confused identity all produce the same symptom on their own. For the difference between memory and a live lookup, see ChatGPT can answer without opening your page. For a model confidently restating a brand’s own copy, see the Morrowen case in ChatGPT recommended a deodorant that didn’t exist.

If you are the person asking, treat a sure-sounding shop, price, or packing list as a claim to check. If you run a shop, log the wrong shops, prices, and specs you see in real buyer questions. Clearer facts in the sources the assistant looks up help more than asking the model to sound less sure.

Read next if: AI can give stupid or dangerous advice · The assistant said the table was booked