- AI advice cut participants’ accuracy from 27% to 9% while their confidence climbed from 30% to 76%.
- The AI advice experiment found that people nearly stopped admitting uncertainty, even when the model gave them false answers.
- Small cash incentives improved results, but participants with AI still performed far below people who answered without assistance.
- The findings raise particular concerns for students learning how to judge uncertainty before building independent critical-thinking habits.
Table of Contents
AI advice is making uncertainty disappear
Being able to say ‘I don’t know’ is not a failure of intelligence. It’s one of the few reliable signs that someone is actually thinking. A new experiment from researchers in France and Italy suggests AI advice can crush that instinct with startling speed: people became less accurate after consulting a chatbot, yet much more certain they were right.
The headline figures are ugly. Without a model in the loop, 44% of participants suspended judgment when they lacked enough information. Once they could ask AI, that rate fell to 3%. Accuracy dropped from 27% to 9%. Confidence went the opposite direction, rising from 30% to 76%.
That’s not the familiar story of a tool making occasional mistakes. My read is harsher: the experiment points to a social and psychological problem in how people receive machine-generated answers. The chatbot does not merely supply an error. It gives the error a suit, a tie, and a very convincing handshake.

The researchers picked questions where the model would fail
Valerio Capraro of the University of Milano-Bicocca worked with Chiara Marcoccia of École Normale Supérieure and Walter Quattrociocchi of Sapienza University of Rome on the study. Their design matters a great deal. Rather than testing participants on a domain where handing work to a capable system might be sensible, they chose visual-detail questions from films that AI models generally handled poorly.
One example concerned the colour of a team uniform in Bend It Like Beckham. The researchers used Step 3.5 Flash, described in the reporting as a model that was usually wrong on these questions. That choice strips away a convenient defense of AI advice: participants were not wisely outsourcing a task to a more dependable expert. They were following bad guidance because it was available and authoritative-looking.
Some people who had initially landed on the correct answer reportedly consulted the system and changed their minds to match its incorrect response. Anyone who has watched a colleague paste a chatbot answer into Slack with zero hesitation will recognize the dynamic. The problem is not that users have no judgment at all. It is that the presence of an answer can persuade them to stop using it.
‘People became much worse, the accuracy was only one third, but they were twice as confident,’ Capraro said.
The team also tested whether stakes would help. Offering money for correct answers did improve behavior, but only modestly: judgment suspension rose to 8%, and accuracy reached 16%. That is better than 3% and 9%, obviously. It is still nowhere near the 44% and 27% results from participants working without the tool.

Why AI advice lands differently than a bad Google result
A bad web search result usually leaves some work for the user. You see a list of links, scan competing sources, notice dates, and perhaps open a page. Chatbots package an answer as if the research has already happened. That convenience is the product. It’s also why AI advice can be unusually dangerous when the system is guessing.
This is closely related to what Wharton researchers have called ‘cognitive surrender’: people accepting incorrect AI output at high rates while reporting more confidence than people who worked independently. Capraro and his colleagues add a crucial detail. The availability of AI advice appears to interfere with the decision to withhold a judgment in the first place.
That distinction matters outside a lab. If an engineer asks a model about an obscure dependency, a manager requests a market estimate, or a teenager asks for help with homework, the right first move is often not a definitive answer. It may be a source check, a clarifying question, or a plain admission that the available information is thin.
Today’s mainstream products are not naturally built for that. Google has pushed generative summaries deeper into search, while OpenAI, Microsoft, Anthropic and others compete on assistants that feel fluid and decisive. You can see the appeal in Google’s Search Labs: the system aims to reduce friction between question and response. But friction was sometimes doing useful work. It made us pause.
The real risk is in schools and routine work
Capraro singled out children as a major concern, and I think he is right to do so. Adults at least remember a pre-chatbot internet, when searching meant comparing sources and getting lost in a few dubious forums along the way. Students growing up with AI advice may learn a different habit: ask, receive, repeat.
That does not mean schools should pretend the technology does not exist. That ship sailed the moment ChatGPT hit the classroom. Banning every tool tends to produce secret use, mediocre enforcement, and a lot of exhausted teachers. The better response is to teach verification as a visible skill. Ask students to identify what would change their minds. Require them to locate a primary source. Grade the reasoning trail, not merely the polished final paragraph.
Workplaces need the same muscle. If AI is used for low-risk drafting, brainstorming, or summarizing documents a human already understands, fine. If AI advice is influencing medical, legal, financial, security, or hiring decisions, organizations should demand provenance and human review rather than treating the model’s confidence as evidence. Confidence is an interface choice. It is not a measure of truth.
AI products need to learn the value of ‘maybe’
The lesson for companies building these systems is straightforward. A chatbot should not be rewarded only for producing an answer quickly and elegantly. It should be rewarded for calibrated uncertainty: stating when it lacks reliable grounding, surfacing competing possibilities, and making verification easy rather than burying it under a smooth paragraph.
Some systems now show citations, ask follow-up questions, or decline requests in narrow cases. Those features help, but they do not solve the broader issue exposed by this research. When AI advice arrives in an assured tone, people can mistake verbal polish for knowledge. Frankly, the industry has spent years training us to do exactly that.
The next contest in consumer AI should not be over who can sound most human. It should be over who can most reliably tell a user: ‘I may be wrong, and here is how to check me.’ Until that becomes normal, every confident chatbot answer will carry a small but real risk of making its user less informed than when they started.
Frequently Asked Questions
What did the AI advice study find?
Researchers from universities in Italy and France found that participants given AI input were less likely to admit uncertainty, less accurate overall, and far more confident in their answers. The experiment deliberately used questions on which the selected model was usually unreliable.
Why can AI make people overconfident?
In this study, access to AI advice reduced participants’ willingness to say they did not know while increasing their confidence, even though the model was usually wrong on the questions. Some participants who would have answered correctly on their own asked the AI and became wrong.
Does paying people for correct answers prevent overreliance on AI?
Not fully. Monetary incentives raised accuracy from 9% to 16% and increased admissions of uncertainty from 3% to 8%, but both figures remained well below the no-AI baseline measured in the experiment.

