Why Productive Struggle Makes Training Actually Stick
Productive struggle is why the best AI tutors withhold the answer. See the learning science, the new Stanford data, and how to build it into training.

When Khan Academy studied how students actually used its celebrated AI tutor, the verdict was blunt: "Too many students who had it available did not even try it," the company's chief learning officer admitted. Yet in a Stanford blind test this June, AI tutors beat law professors on roughly 75% of nearly 3,000 questions. Same underlying technology, opposite outcomes. The difference comes down to a decades-old idea from learning science called productive struggle, and it quietly decides whether your training builds real skill or just fills a screen.
This piece explains what productive struggle is, why even the most hyped AI tutors fell flat without it, and how to design training that makes people think instead of watch. The short version: learning that feels a little hard is learning that lasts.
What is productive struggle in learning?
Productive struggle is the effortful, slightly uncomfortable work of figuring something out before you are handed the answer, and it is one of the most reliable ways to make learning stick. The mental effort of retrieving a fact, attempting a problem, or explaining your reasoning is what encodes knowledge into long-term memory.
Psychologists have a broader name for the same idea: desirable difficulties. UCLA researcher Robert Bjork coined the term in 1994 to describe conditions that feel harder in the moment but produce stronger retention and transfer later. The catch is in the adjective. The difficulty has to be desirable, not just difficult.
The payoff is not vague. A Bellwether review of the research found that productive struggle strengthens four distinct things at once: memory and information processing, attention and engagement, motivation and mindset, and metacognition, per Bellwether's analysis of productive struggle.
There is even a physical story underneath it. When you practice a skill and get corrective feedback, your brain wraps the active neural pathways in myelin, a fatty sheath that speeds up signaling. A well-myelinated signal travels over 100 times faster than an unmyelinated one, according to Edutopia's summary of the neuroscience. Repetition at the edge of your ability lays that insulation down. Passively watching does not.

Why the first AI tutors made learning easier and worse
The first wave of AI tutors optimized for the wrong thing: removing effort instead of directing it. When a tool answers every question instantly, it feels helpful and quietly does the learner's thinking for them.
Khan Academy learned this the expensive way. Founder Sal Khan has acknowledged that the original version of its Khanmigo tutor "did not change student learning as much as many of us hoped it would." It sat next to the lesson as an optional helper, which meant students had to notice they were stuck and then formulate a good question, and most never did.
The research explains why that failure mode is so common. In one study cited by Bellwether, an AI assistant lowered the effort of a research task but the quality of students' final arguments dropped. A high school math study found learners using a chatbot as a "crutch" during practice, with worse performance later. Nearly 47% of student and AI conversations were just requests for answers with minimal engagement. Researchers gave the pattern a name: metacognitive laziness, offloading the thinking rather than doing it.
Bellwether frames the core question well: "When does ease enable greater learning, and when is ease a shortcut with a hidden cost?" For most first-generation AI tutors, ease was the shortcut.
What the Stanford AI tutor study actually proves
The Stanford result is not evidence that AI is smarter than professors. It is evidence that an AI tutor built to reason, rather than just retrieve, can meet an expert standard.
In the study led by Professor Julian Nyarko, 16 law professors wrote answers to 40 contracts questions, then commercial tutoring systems and Google's NotebookLM answered the same questions. Across nearly 3,000 blind, anonymized matchups, the AI answers won about 75% of the time, per Stanford Law School. The questions were not trivia. They demanded synthesis of competing arguments and a defensible conclusion.
Two numbers matter more than the headline. AI answers were flagged as pedagogically harmful or misleading just 3.5% of the time, versus 12% for the professors. And as first author Alejandro Salinas put it, "AI tutors can offer high-quality, on-demand support that complements classroom instruction, and may broaden access to expert guidance."
Read the two studies together and the lesson is clear. A capable model is necessary but not sufficient. The Stanford tutors won because they modeled good reasoning, and the value for a learner comes from being pushed to do that reasoning themselves.
How to build productive struggle into AI training
You build productive struggle into training by designing for retrieval, delay, and explanation, not passive playback. The mechanics are surprisingly concrete, and they map directly onto what Khanmigo changed in its rebuild.
- Ask before you tell. Open with a question the learner has to attempt, even imperfectly. Retrieval practice, the act of pulling an answer from memory, beats rereading or rewatching every time.
- Withhold the answer, coach the next step. When Bellwether researchers modified a chatbot to refuse direct answers and prompt problem-solving instead, learners saw nearly double the short-term gains and kept the long-term learning. Khanmigo's rebuild now guides students toward answers rather than handing them over.
- Make learners explain their reasoning. The rebuilt Khanmigo prompts students to justify their thinking. Khan calls the shift from "cognitive offloading" to "cognitive onloading." Explaining forces the encoding that watching skips.
- Interleave and space the practice. Mixing problem types and spreading practice over time feels harder and works better than massed, single-topic drilling.
- Weave the tutor into the work. Khan's blunt takeaway from the failure: "The AI could not just sit next to the content. It had to be woven into it." Help that appears inside the task gets used. Help that waits to be summoned does not.
This is exactly why interactive training videos that pause to ask a question, adapt to the answer, and respond in real time outperform a passive recording. The video is not the point. The moments where the learner has to do something are.

The difference between productive struggle and just being stuck
Productive struggle only works when the challenge stays within reach. Past that line it stops teaching and starts demoralizing. As the research puts it, the struggle must be productive, the difficulty must be desirable, and the zone of development must be proximal.
That last phrase is the guardrail. A task should sit just beyond what the learner can do alone but inside what they can reach with a hint or two. Set it too easy and there is no encoding. Set it too hard and you get frustration, guessing, and quitting. Struggle for struggle's sake is not a virtue.
Practically, that means the tutor needs to sense when a learner is stuck versus stretched, and offer a graded hint rather than the full solution or nothing at all. The goal is a learner who is working hard and making progress, not one who is staring at a wall. Good scaffolding does not remove the struggle. It keeps it productive.
FAQ
Does productive struggle mean I should make training harder?
Not harder for its own sake. The goal is the right kind of hard: tasks that demand active recall and reasoning but stay within reach. If learners are guessing blindly or giving up, the difficulty has crossed from desirable into counterproductive, and you need better scaffolding, not an easier task.
Will learners quit if the AI tutor refuses to just give answers?
They quit when the struggle is unproductive, not when it is hard. A tutor that offers a graded hint, confirms partial progress, and only then reveals the answer keeps people engaged. The research shows this design nearly doubled short-term gains while protecting long-term learning, so withholding the answer thoughtfully is a feature, not friction.
How is this different from adding a quiz to a video?
A quiz at the end tests whether people watched. Productive struggle builds the learning in the first place by interrupting the passive flow with retrieval, explanation, and feedback at the moment it matters. The difference is timing and interaction: thinking woven through the experience, not bolted on after it.
Passive video is where knowledge goes to die, and adding more of it will not fix a training program. What makes learning stick is the effort of doing, and the fastest path there is an AI tutor that adapts, asks, and makes each learner do the thinking. Build training around productive struggle, and completion stops being the metric that flatters you and starts being the one that means something.
Turn your training into an interactive experience
Nesoi transforms static content into interactive video experiences with AI tutors your team actually finishes.
Book a demo