Articles — Whether that list is for you
Whether that list is for you
The assistant said the table was booked
You asked for a table on Thursday. The reply said it was booked. No confirmation email arrived, and this site has not counted how often that happens.
Illustrative example, not a logged test.
You ask an assistant to book a table for four on Thursday at 19:00. A minute later it replies that the table is booked. No confirmation email arrives.
A reply can sound finished when the booking, the cart, or the payment never completed. Merj, a technical SEO firm that tests browser agents, lists this among the ways those agents fail: they declare success at the search results or the cart and never reach the confirmation page. A message that vanishes after a few seconds, or a validation error the page never keeps on screen, can leave an agent reporting success on something that failed. Merj gives no rate for this, and this site has not counted one. Treat it as a known way to fail, not as a measured frequency.
The same lab built a page where the Save button was labeled correctly and named Cancel for screen readers. ChatGPT’s agent, in a browser OpenAI has since retired, clicked Cancel first in all 25 runs. Claude and Perplexity clicked Save. One page, one task, from a firm that sells the monitoring. The useful part is that some assistants read the name the page gives the button, not only the picture. Write that name so it matches what the button does. That is for people who use screen readers first.
Ask to see the proof
Before you trust “booked,” look for a confirmation: a reference number, an email, or a page on the venue’s own system that shows the booking exists. If you cannot find one, treat the table as not booked and check with the venue. For anything you cannot undo, such as a payment, a deletion, or a message sent in your name, ask the assistant to stop at the last screen and show it to you before it presses the button.
Say how many, and which one
Merj describes a 2026 paper, OmnilingualGAIA2, that tested agents in ten languages inside a simulated contacts app. The task was to delete “my contact from the US” while the list held two US contacts. In English and Spanish the agent asked which one. In Chinese, Japanese, and Indonesian, which mark singular and plural differently from English, it deleted both. The paper counts that as an irreversible over-action.
Take this carefully. It is a simulated tool environment, not a web booking, and it is one case, reported through Merj rather than re-run here. The authors link the behavior to grammar cues. The practical line does not depend on that explanation. When an instruction could mean one thing or several, give the number and name the target. Ask for confirmation before anything is deleted.
A page can talk to the agent
Reviews, product descriptions, and forum posts are text an agent reads while it works. Anyone who can post a review can write an instruction into it — prompt injection in a page the agent loads. Merj’s example is a review that tells an agent to drop its earlier limits and add an extended warranty. That text arrives in the tool result for every visitor whose agent loads the page.
A smarter prompt does not fix this. Merj puts the controls outside the model: confirmation before irreversible steps, narrow permissions, and checks on the server. The part within your reach is the confirmation.
This is not the same case as the deodorant that didn’t exist. That brand was a thin page the assistant treated as a recommendation. A review that instructs an agent is text written to change an action, closer to companies using Reddit to change what AI recommends. The moment an agent also pays, you never get to apply that skepticism yourself. The habit underneath both is to verify before you trust a sure-sounding claim — see Why AI won’t just say it doesn’t know.
Read next if: Your catalog can be perfect and still never show up · How AI chatbots build a shortlist