Back

After 10 minutes of leaning on an AI assistant, people did worse and gave up more often once it was gone, a study finds

Chart from the study: the AI group solves more fraction problems while it has help, then drops below the group without AI once the assistant is removed
Chart: Grace Liu, Brian Christian, Tsvetomira Dumbalska, Michiel A. Bakker and Rachit Dubey

In experiments with 1,222 people, those who had an AI assistant for about 10 minutes did worse once it was taken away than people who never had it, and in two of the three tests they skipped more problems. The drop was concentrated among people who said they asked the AI for answers rather than hints. The researchers do not know yet whether the effect lasts beyond the session or builds up with daily use.

People who leaned on an AI assistant for about 10 minutes did worse, and often gave up more, once the assistant was taken away, according to a study by researchers at Carnegie Mellon University, MIT, the University of Oxford, UCLA and UC Berkeley. An early draft drew wide attention earlier this year. UC Berkeley said on Oct. 9 that the team presented the final version, checked by outside experts before it was accepted, this week at the Conference on Language Modeling, a meeting of AI researchers.

The researchers recruited 1,222 people in the U.S. through an online research site and split them into two groups at random, so any difference afterward could be traced to the AI rather than to who was in each group. In the first test, 354 people worked through 15 fraction problems. One group had an AI assistant built on GPT-5 in a sidebar for the first 12, and it would hand over the answer if asked. The other group worked alone throughout. Then the assistant was removed without warning for the last three. Everyone could skip a problem, and wrong answers cost nothing, so skipping showed who chose to stop trying.

With the assistant, the AI group solved more problems. Without it, they solved 57 percent of the last three problems, against 73 percent for people who never had help, and skipped 20 percent, against 11 percent. The researchers say a flaw in how they screened out weaker participants may have widened that first gap, so they ran a larger repeat with 667 people that fixed it. The AI group still did worse, solving 71 percent against 77 percent, but the difference in skipping was too small to rule out chance. A third test, with 201 people answering reading questions like those on the SAT, the U.S. college entrance exam, showed the same pattern as the first: 76 percent solved against 89 percent, and 8 percent skipped against 1 percent.

How people used the AI made a difference. In the larger repeat, about six in 10 people with the assistant said they mostly used it to get answers directly, and the drop was concentrated among them. People who said they used it for hints or explanations did about as well as people with no AI. But people chose for themselves how to use it, so this part of the study can't show that asking for answers caused the drop.

The researchers suggest two reasons. When AI answers in seconds, people start to expect tasks to take seconds, so working alone feels slow and hard by comparison. And handing off the work skips the struggle of working a problem through, which is how people learn what they are capable of. Brian Christian, a co-author and research fellow at UC Berkeley's Center for Human-Compatible AI, says AI could default to acting more like a tutor instead of handing over answers.

The study tested only short sessions, simple tasks and people paid to take part online, and it measured them right after the AI was removed. Its AI also gave full answers whenever asked, and whether a tutor-like AI would avoid the drop is still an open question. The researchers say they do not know whether the effect lasts for hours or days, or whether it grows or fades with daily use, and that answering those questions will take studies that follow people over many sessions.