Key facts
- AI shopping assistants frequently provide conflicting information on prices, specifications, and product currency.
- A Product.ai study tested ChatGPT, Claude, Gemini, and Perplexity with 220 shopping questions.
- 86% of questions resulted in repeatable factual conflicts, according to Product.ai.
- 97% of head-to-head product comparison questions showed conflicts.
- 85% of verifiable price answers did not match current listings.
- Gemini's free tier showed the highest costly error rate at 56%.
A recent study by Product.ai indicates that current artificial intelligence models are not yet reliable enough to handle holiday shopping tasks, frequently providing inaccurate or conflicting information. The research tested free and paid versions of popular AI services, including ChatGPT, Claude, Gemini, and Perplexity, using 220 shopping-related questions. Across 8,794 responses, Product.ai found that 86% of questions led to repeatable factual conflicts, such as discrepancies in prices, product specifications, or whether a product was current. Head-to-head product comparisons were even more prone to errors, with 97% showing conflicts. When verifying prices, 85% of the answers did not match the current or listed price. Gemini's free tier exhibited the highest rate of costly errors at 56%, while Perplexity's paid version showed the lowest at 14%. The study also noted that AI models sometimes contradicted themselves when asked the same question multiple times. When errors occurred, the price discrepancy had a median difference of $300. Product.ai's head of search product, Dakota Nunley, advised consumers to use AI as a discovery tool but to verify information independently, rather than relying on AI agents for payment details or accepting their responses as definitive.
