John Sweller's cognitive load theory explains why instruction must respect the limits of working memory. Intrinsic load is the inherent complexity of the material; extraneous load is wasted by poor design; germane load is the cognitive effort invested in building schemas. Effective instruction minimizes extraneous load, manages intrinsic load, and maximizes germane load. The same framework explains the expertise reversal effect: instructional aids that help novices (worked examples, explicit prompts, heavy guidance) become redundant processing for experts and actually hurt their performance. A single study plan cannot suit everyone - design must adapt as the learner advances.
For novices, fully worked solutions teach more than unsolved practice. This worked-example effect fades with expertise. Fading bridges the gap by gradually removing steps from worked examples; a middle stage is the completion problem, in which most steps are shown and a single gap is left for the learner to fill. Self-explanation - explaining out loud or in writing why each step in a solution is needed and how it follows from the prior step - consistently outperforms passive study of the same examples, even without extra feedback.
Once a learner is solving problems, the surface organization of practice matters. Blocked practice, doing many problems of the same type in a row, feels easier but produces weaker long-term retention. Interleaving, mixing problem types within a session, feels harder and lowers in-session performance but produces stronger transfer. Variable practice adds variety in the same direction by changing non-essential features of problems (numbers, names, contexts) while preserving the underlying structure - a technique sometimes called parameter variation. Rote repetition of identical problems is the weakest form. Transfer-appropriate processing captures the same idea: memory is best retrieved when the conditions at recall match the conditions at encoding, so practice that resembles the eventual test or real-world task transfers better than practice in unrelated conditions.
Anders Ericsson named the gold standard "deliberate practice": effortful focus beyond current ability, immediate informative feedback, and repetition with progressive refinement. Quantity alone is not deliberate practice; merely doing something many times is closer to massed practice. Gladwell's popular "10,000-hour rule" oversimplifies the research by ignoring quality of practice, individual differences, and the fact that many skills show diminishing returns far earlier. The mastery model (don't move past a topic until accuracy is at least 80–90%) and spaced repetition are complementary rather than competing: mastery ensures initial encoding, spacing prevents later forgetting.
Expertise lives in schemas, rich mental structures that group domain patterns together and let experts recognize situations and chunk information almost instantly. A schema is a larger, conceptual cousin of a chunk; phone numbers are chunks, a chess position is a schema, a disease category is a schema. Building such schemas is the goal of most instruction. Transfer is the proof that schemas work: near transfer applies knowledge to a closely related context (an algebra word problem solved in physics), while far transfer applies it to a very different domain (music theory informing programming logic). Studying with varied contexts and pairing examples with non-examples helps learners understand the boundaries of a concept rather than memorizing a single pattern.