Desirable Difficulties: Why Effortful Training Sticks
Desirable difficulties explain why smooth, easy training fails to stick. See the research and five ways to build productive struggle into learning.

Harvard students who sat through polished, easy-to-follow physics lectures were sure they had learned more than classmates who had to wrestle with the problems themselves. Their test scores said the opposite: the strugglers scored higher, and actual learning and the feeling of learning were strongly anticorrelated. Learning science has a name for this effect, desirable difficulties, and it explains why so much smooth, professional-looking training produces so little lasting skill.
This post covers what desirable difficulties are, why effortless training fools both learners and the people who buy it, what new AI research says about removing struggle entirely, and five practical ways to build productive struggle into your training programs.
What are desirable difficulties in learning?
Desirable difficulties are training conditions that feel harder and slow people down during learning, but produce far better long-term retention than easier alternatives. The term was coined in 1994 by cognitive psychologist Robert Bjork, whose UCLA Learning and Forgetting Lab has spent decades documenting the effect.
The lab's classic example is a math class. Students who drill one problem type for an hour feel a satisfying sense of mastery by the end. Students who practice the same problems shuffled together feel slower and clumsier. Two weeks later, on the test, the mixed-practice group wins.
The uncomfortable core of the research is that performance during training is a poor signal of learning. Fluency in the moment fades. Effort in the moment compounds.
Why easy training feels effective but fails
Easy training fails because the feeling of learning and actual learning are different things, and they often move in opposite directions. In the Harvard study published in PNAS, researchers randomly assigned students to either a polished traditional lecture or an active session where they worked through problems in groups. Same instructor quality, same content.
The results:
- Students in active sessions scored higher on tests of the material.
- The same students reported learning less than the lecture group.
- Course surveys and actual test results pointed in opposite directions.
Lead author Louis Deslauriers put it plainly: "Deep learning is hard work. The effort involved in active learning can be misinterpreted as a sign of poor learning."
This is exactly why bad corporate training survives. A slick video course that people watch passively earns great satisfaction scores, so it gets renewed. The scores measure comfort, not competence. If your only training metric is "did people like it," you are optimizing for the condition that produced the worst results in the study.
What happens when AI removes the struggle
The newest evidence comes from AI-assisted learning, and it points the same direction: when a tool removes the struggle, it often removes the learning too.
In a classroom experiment in Turkey covered by The 74, nearly 1,000 high school students practiced math with help from an AI tutor built on a leading large language model. During practice, the AI group looked great: they solved more problems. Then the tool was taken away for the real test, and students who had used the unrestricted version, the one that would hand over direct answers, scored 17% worse than students who never had AI help at all.
The most useful detail is the exception. A guarded version of the same tutor, designed to give hints instead of answers, avoided the drop entirely. The technology was not the problem. The removal of effort was.
A preliminary MIT Media Lab study found a similar pattern in writing. Researchers recorded brain activity while 54 participants wrote essays with an AI assistant, with a search engine, or with no tools. Neural connectivity scaled down as external support scaled up, and the AI-assisted group struggled to quote from essays they had written minutes earlier. The authors call the accumulating cost "cognitive debt." The study is a preprint with a small sample, so treat it as an early signal rather than settled science, but the signal agrees with 30 years of desirable-difficulties research.
For workplace learning the implication is direct. If your training lets people reach the answer without ever generating one themselves, practice performance will look wonderful and none of it will survive contact with a real customer, a real incident, or a real negotiation.
Which desirable difficulties work best in training
Four difficulties have the strongest evidence behind them, all documented extensively by the Bjork Lab:
- Spacing. Spread practice over multiple sessions instead of massing it into one long block. One of the most robust findings in cognitive psychology, and the reason a one-day onboarding dump is a retention disaster.
- Retrieval practice. Have learners pull knowledge out of memory with questions instead of re-reading or re-watching. Successful retrieval strengthens memory more than additional study time does.
- Interleaving. Mix topics and problem types during practice rather than finishing one before starting the next. It feels chaotic and works better, especially for learning to tell similar situations apart.
- Generation. Make learners produce an answer, even a wrong one, before showing them the solution. Generated material is remembered better than presented material.

Two cautions from the same research. First, difficulties do not stack neatly: combining spacing with interleaving adds little beyond either alone, so pick the ones that fit your content rather than piling on all four. Second, the difficulty must be desirable, meaning the learner can actually succeed with effort. Struggle that ends in failure most of the time is just frustration.
How to design productive struggle into workplace training
You can retrofit desirable difficulties into almost any program with five changes:
- Ask before you tell. Open each module with a question or scenario learners must attempt before the content plays. A wrong first guess is a feature: it primes memory for the correction.
- Replace summaries with quizzes. End each section with two or three retrieval questions instead of a recap slide. The recap does the remembering for the learner, which is precisely the problem.
- Space the follow-ups. Send a five-minute refresher three days after the session, then a week later. Modest, cheap, and it targets memory at the moment retrieval is getting hard, which is when practice pays most.
- Shuffle the scenarios. In compliance, sales, or safety training, mix case types within one exercise so learners must first diagnose which rule applies. Recognizing the situation is the skill that transfers.
- Make AI a hint engine, not an answer engine. The Turkey experiment shows the same AI tutor can either erase learning or protect it depending on one design choice. Configure AI support to question, hint, and withhold, the way a good coach does.

Format matters here. A static video cannot ask, wait, or withhold: it just plays. This is why interactive training videos change the economics of effortful learning. When the video itself pauses to make the learner answer, adapts to what they got wrong, and spaces follow-up questions automatically, the desirable difficulties are built into the medium instead of bolted on by a facilitator.
FAQ
Doesn't making training harder just frustrate employees?
Only if the difficulty is the wrong kind. Desirable difficulties are challenges learners can overcome with effort, like recalling yesterday's material or attempting a scenario before seeing the model answer. Confusing instructions, missing context, and puzzles far above skill level are undesirable difficulties, and cutting those actually makes room for the productive kind.
How do I know if a difficulty is working or just annoying people?
Measure learning after a delay, not performance during the session. If a technique slows people down on day one but improves a quiz score one or two weeks later, it is working exactly as the research predicts. If both in-session performance and delayed retention drop, the difficulty is not desirable and should go.
Can AI tools and desirable difficulties coexist in a training program?
Yes, and the evidence says design decides which you get. AI configured to give hints, ask follow-up questions, and adapt difficulty preserved learning in the Turkish classroom study, while the answer-giving version erased it. Use AI to personalize the struggle, not to remove it.
The lesson from three decades of research is simple to state and hard to accept: training that feels smooth is usually training that slides off. Learning sticks when people retrieve, generate, and wrestle, with support that keeps the struggle winnable. That is the entire case for interactive learning over passive content: not that it feels easier, but that it makes the right kind of effort unavoidable.
Turn your training into an interactive experience
Nesoi transforms static content into interactive video experiences with AI tutors your team actually finishes.
Book a demo