Skip to main content

Stop Expecting Certainty from Probability Machines

In March 2025, an independent audit of 11 leading chatbots found that when confronted with news-related prompts, these systems either repeated false claims or dodged the question in more than 41% of cases; in barely three-fifths of responses did they deliver a competent debunk. Performance has plateaued across months despite expanded web access and retrieval tricks, suggesting a ceiling on short‑term progress.