Skip to content
26 September 2026

Study shows AI chatbots misstate prices and specs while shopping

AI chatbots still stumble on product facts, so shoppers must stay vigilant.

Study shows AI chatbots misstate prices and specs while shopping

As retailers embed artificial intelligence deeper into online storefronts, many consumers assume a chatbot can replace a manual price check. A recent benchmark, however, paints a far less optimistic picture. Researchers queried four of the most widely used large language models—ChatGPT, Claude, Gemini and Perplexity—using a set of 220 real-world shopping prompts that spanned laptops, televisions, mattresses, sunscreen and robot vacuums. Each prompt was submitted five times to every model, generating a total of 8,794 individual responses for analysis.

Methodology behind the numbers

The test framework treated any “conflict” as a verifiable disagreement: a different price, a mismatched specification, an erroneous model name, or a completely fabricated ingredient. After collecting the raw answers, the team cross-referenced each claim with up-to-date manufacturer listings and trustworthy retail sites. By repeating questions, they could also spot inconsistency—answers that changed from one query to the next despite identical input.

Across the board, 86% of the queries produced at least one repeatable conflict. When the prompt asked the model to compare two products directly, the conflict rate jumped to a staggering 97%, compared with 75% for straightforward factual requests. In the subset of 913 price-related answers that could be verified, only 85% matched the current listed price, and the median deviation for the erroneous ones was roughly $300.

Who slipped up most and why it matters

Among the four engines, Google’s Gemini emerged as the weakest performer. On the free tier, 56% of its replies contained a “costly error”—misquoted price, wrong model or invented specification. The paid Gemini-3.1-Pro preview reduced that figure only marginally to 54%. Claude showed a noticeable improvement when upgraded: costly errors fell from 44% on the free version to 21% on the paid tier. Perplexity fared best, with just 14% costly errors on its paid API and 15% on the free tier. ChatGPT trailed Perplexity but stayed ahead of Gemini, registering 17% costly errors on its paid plan and 19% on the free version.

Contradictions were another tell-tale metric. Gemini’s free service contradicted itself on 29% of repeated queries, while its paid tier still erred 27% of the time. By contrast, Perplexity’s paid tier displayed the lowest self-inconsistency at 13%. The study’s lead analyst noted that because these models generate answers probabilistically, the same question can yield different outputs within minutes—a clear warning sign for anyone treating an AI as a fully automated shopper.

Practical takeaways for holiday shoppers

The For price-sensitive purchases, a $300 median error can turn a “good deal” into a costly mistake. Users should therefore cross-check critical details on the manufacturer’s site or use multiple chatbots to triangulate information. Treat the AI as a brainstorming partner that can surface product ideas, compare feature sets at a high level, and point you toward relevant brands, but verify the final specs, prices and availability yourself.

As the technology evolves, accuracy is expected to improve, especially as developers feed models with continuously refreshed product graphs. Until then, the safest approach is a hybrid one: let the AI spark ideas, then let human judgment—or a trusted retailer—confirm the numbers. This strategy can keep holiday shopping both efficient and financially sound.

Author

Marcus Chen

Marcus Chen writes about consumer tech the way a friend who actually opened the device would describe it. Hardware-first, hype-skeptical, and fluent in benchmark numbers.