Listen to this Post
As artificial intelligence chatbots continue to rise in popularity, one critical piece of advice remains constant: don’t rely on them for factual information. Despite their growing capabilities, AI systems like ChatGPT, Gemini, and Grok still struggle to deliver accurate answers. A recent study sheds light on just how often these tools provide incorrect or misleading information—sometimes with unwavering confidence. This article takes a closer look at this issue, highlighting why AI chatbots should only be used for inspiration, not for verified facts.
Study Reveals Serious Flaws in AI Chatbots
A new study conducted by the Tow Center for Digital Journalism tested eight popular AI chatbots that claim to conduct live web searches for accurate answers. These chatbots include:
– ChatGPT
– Perplexity
– Perplexity Pro
– DeepSeek
– Microsoft’s Copilot
– Grok-2
– Grok-3
– Gemini
The study’s objective was straightforward: present each chatbot with a simple task that could easily be solved with a web search. The task involved providing a link to an article, along with its headline, original publisher, and publication date, using a quote from that article.
What made the task achievable was that the excerpts chosen for the test were easily searchable on Google, with the original source usually appearing within the first three search results. However, despite the simplicity of the task, the chatbots’ performance left much to be desired.
The Results: Far from Reliable
The study found that most AI chatbots were incorrect most of the time. On average, these systems answered questions correctly less than 40% of the time.
- Perplexity, the most accurate chatbot, answered correctly 63% of the time, while Grok-3 was the worst, getting things right only 6% of the time.
- In terms of confidence, premium chatbots (the paid versions of AI models) often gave more confidently incorrect answers than their free counterparts.
- Chatbots were also more likely to fabricate citations—creating fake links or citing syndicated content without providing the correct source.
– Additionally, several chatbots ignored web
Despite these issues, the study did offer a silver lining—Apple made the right choice in partnering with OpenAI’s ChatGPT for Siri’s search functions. While it wasn’t perfect, ChatGPT provided the least-worst results compared to other AI chatbots, although it still missed key details.
What Undercode Says:
The findings of this study are a stark reminder that AI chatbots are not reliable sources for factual information—at least not yet. Their tendency to be wrong and overly confident about it points to a major flaw in their current design and functionality. What’s especially alarming is how these systems are sometimes more confident in their wrong answers than in those they get right.
For instance, many chatbots failed to admit when they didn’t have access to a specific article or couldn’t find a direct match. Instead, they filled in the gaps with speculative or fabricated information, which can be misleading for users who trust them blindly. This is a dangerous problem for anyone relying on chatbots for accurate data, especially for critical topics like news, science, and history.
An interesting aspect of the study was the inconsistent quality of the paid versus free chatbots. Premium versions of these models were often more overconfident in their inaccuracies, which could lead to a false sense of reliability for users who are paying for “premium” services.
Furthermore, the
What’s particularly concerning here is that users might mistakenly assume that chatbots are infallible simply because they can generate quick responses that sound authoritative. While AI can be useful for brainstorming, generating ideas, or offering general guidance, it’s not the right tool for precise, factual information.
Until these systems can better acknowledge their limitations and improve their accuracy, it’s crucial for users to remain skeptical of any definitive statements they make.
Fact Checker Results:
- Chatbots performed poorly in factual accuracy, with most answers being incorrect or incomplete.
- Confidence in incorrect answers was notably higher in premium versions of chatbots.
- Ethical issues, such as bypassing robot exclusion protocols, were observed in the study.
References:
Reported By: https://9to5mac.com/2025/03/11/ai-chatbots-cant-be-trusted-proves-study-but-apple-made-a-good-choice
Extra Source Hub:
https://www.facebook.com
Wikipedia
Undercode AI
Image Source:
Pexels
Undercode AI DI v2





