technique

Spaced Repetition for Mnemonics: When to Review

Tomorrow, then about a week, then let every clean recall stretch the gap. The forgetting curve with its real numbers, SM-2 worked out in days, and how to review a picture at a locus rather than a fact on a card.

hero illustration for Spaced Repetition for Mnemonics: When to Review

Spaced repetition is the whole art of when to review, and it fits in three sentences you can check against a calendar. Within an hour of building a mnemonic, recall it once with the answer covered, to find and repair the weak pictures; that is part of building, not a review. The first spaced review is tomorrow, and the next comes after roughly a tenth to a fifth of the time you need to keep it. From there, every clean recall stretches the gap by about two and a half times and every miss pulls it back to tomorrow. The rest is the forgetting curve those sentences rest on, the dates they produce, and what changes when the thing reviewed is a picture at a locus rather than a fact on a card, which is where every guide in the technique library leaves you.

The schedules the internet repeats (1-3-7-14-30, 2-7-30, 2357, 7-3-2-1, "an hour, a day, a week") are rules of thumb from one family, and none we could find traces to a primary study. The research is more forgiving: the exact days matter far less than recalling instead of re-reading, spacing at all, and pressing miss when you missed.

The forgetting curve, with the real numbers

In 1879 and 1880 Hermann Ebbinghaus memorised lists of nonsense syllables, relearned them after gaps from twenty minutes to a month, and timed how much faster the second learning went; he published in 1885. What he measured was savings, relearning time saved, not percent recalled: if a list took ten minutes to learn and six to relearn, the savings were forty percent.

In 2015 Jaap Murre and Joeri Dros repeated the experiment with one subject, Dros, and concluded that the curve had been replicated: the same shape, with lower values. Both sets of savings, from their Table 3, rounded:

Gap after learningEbbinghausDros (2015 replication)
20 minutes58%47%
1 hour44%37%
9 hours36%28%
1 day34%32%
2 days28%23%
6 days25%17%
31 days21%4%

The "seventy percent forgotten in a day" and "half gone in an hour" figures that circulate are not what Ebbinghaus's printed table shows, and his syllables were chosen to have no meaning; pegs with vivid pictures are not nonsense syllables. What transfers is the shape: steep, then flat.

One detail in that table is the foundation of every good schedule. Read the replication's 9-hour and 1-day rows again: savings rose overnight, from 28 to 32 percent, a small rise at the one-day point. Murre and Dros write that the research on sleep and memory would predict such a boost and attribute it to sleep, as Jenkins and Dallenbach did, while noting that for this kind of experiment the cause remains to be established. Whatever it is, "tomorrow" is the right first spaced review for almost anything: you test after the bump, where the curve has flattened, instead of fighting the steep part at midnight.

A review is a recall, not a re-read

Re-reading notes or a peg list is not a review. A review is an attempt to produce the thing with the answer covered, then a check.

Roediger and Karpicke (2006) showed why. Students studied prose passages and then either restudied them or took free-recall tests. Five minutes later the restudiers were ahead; two days and a week later the tested students remembered substantially more, and the restudiers were the more confident group, because re-reading feels fluent and fluency feels like knowing.

Here, a review of a Major peg is: cover the word, see the number, produce the picture. A review of a locus is: stand at it, in your head, and see what is there. Every drill on the practice hub tests the same way: a prompt, your attempt, and only then the answer.

How long to wait depends on how long you need it

The experiment on gap length to know is Cepeda, Vul, Rohrer, Wixted and Pashler (2008). More than 1,350 people learned a set of facts, reviewed them once after a gap of up to three and a half months, and were tested up to a year later. At every test delay, lengthening the gap first improved final recall and then reduced it. The best gap grew in days as the horizon grew but shrank as a share of it: roughly 20 to 40 percent of a one-week horizon, 5 to 10 percent of a one-year horizon, so call it 10 to 20 percent for the months in between. Its abstract ends:

The interaction of gap and test delay implies that many educational practices are highly inefficient.

Turned into dates. Day 0 is the build with its within-the-hour check; day 1 is the first spaced recall, as in the table in the studying guide; and the Cepeda gap is the wait after day 1, never the wait after building:

You need itCepeda gap, counted from day 1Then, in practice
Exam in 7 days20 to 40% of 7 days: 1.5 to 3 daysday 1, day 4, the day before
Presentation in a month10 to 20% of 30 days: 3 to 6 daysday 1, about day 7, about day 20, two days before
Keep for a year or longer5 to 10% of 365 days: 3 to 5 weeksgaps of about 1, 6 and 16 days, then monthly: the SM-2 shape below

So 1-3-7-14-30 suits an exam in a month and wastes reviews on something you want for years. The same group's 2006 meta-analysis (839 assessments from 317 experiments) found the rule at every scale.

The fixed schedules people search for

Every named schedule is a rough draft of the same curve.

ScheduleReviews atGood forWeakness
1-3-7-14-30, 2-7-30days 1, 3, 7, 14, 30 (or 2, 7, 30)a month of exam material; a new peg liststops at 30 days; no grading
2357, 7-3-2-1days 2, 3, 5, 7, or counted back from the examthe last week before an examnothing outside that week
Rule of Five (O'Brien)at once, within a day, a week, a month, three to six monthsmnemonics you want for yearsfixed; no grading
SuperMemo on paper (Wozniak 1985)gaps of 1, 7, 16 and 35 days, then doublingpages of 20 to 40 itemswhole pages, not items
12 h then times 1.5 (Blocki et al. 2015)gaps of 12 h, 18 h, 27 h, 41 h and onmnemonic images in scenesmany early reviews; needs a timer
SM-2 (Anki's default)1 day, 6 days, then each gap times about 2.5anything; adapts per itemneeds software or bookkeeping

O'Brien, eight times World Memory Champion, teaches five reviews at widening gaps: immediately, then within a day, a week, a month, and three to six months, each one buying a longer stretch than the last. The name Rule of Five is the summaries' rather than his own. The 1-7-16-35 on study blogs is Wozniak's 1985 paper method, whole pages repeated at those gaps and then doubled; SM-2 replaced it once a computer could track each item. The 12-hours-then-times-1.5 schedule is the clearest published test of a rehearsal schedule on mnemonic images and gets its own section.

SM-2 in numbers

SM-2, devised by Piotr Wozniak in 1987 and set out in his 1990 thesis, is the algorithm behind Anki's default scheduler and, on this site, behind the palace, words and custom drills and the pi journey. It runs by hand.

  1. First successful recall: next review in 1 day.
  2. Second: 6 days.
  3. Each later one: the previous gap times the item's ease factor, which starts at 2.5.
  4. Each review moves the ease: a perfect recall adds 0.1, a hard one subtracts 0.14, a miss subtracts 0.8; it never falls below 1.3.
  5. A miss (any grade under 3 on the 0 to 5 scale) restarts the item at step 1.

The drills here collapse six grades to three buttons: miss (0), hard (3), easy (5). This is what they do to the calendar, computed from the code that runs the palace review. A miss restarts the count, so in the last column review 4 is a first recall again.

ReviewAll easyAll hardMiss on review three
11 day1 day1 day
26 days6 days6 days
316 days13 daysmiss: 1 day
445 days27 days1 day
5131 days52 days6 days
6393 days94 days13 days
Ease after six3.11.662.2

Three things to take from it. Six clean recalls carry an item past a year. A hard item is barely punished on any one review, but its ease drifts down, so after six reviews it is due in three months where the easy item is due next autumn. And a miss is expensive: back to tomorrow, ease down by 0.8, two more reviews to reach the six-day gap. That is correct, a missed item is fragile, which is why the miss button has to mean a miss. The 2.5 is the value Wozniak started with, not a law.

At the far end the numbers stop meaning anything: the seventh all-easy interval is more than three years, and Anki caps at a hundred. Items at multi-year intervals are simply things you know; stop tracking them.

What "review" means when the thing is a mnemonic

Mnemonics change one thing: what you recall is not the fact but the picture that leads to it, and the picture is what gets graded.

Take peg 16. The review is not "what is the word for 16" but: see the number, let the picture arrive, then read the word off it. If the dish is shattering on the tiles, the T and the Sh come with it: easy. If the word arrives by rote with no picture: hard, because the mnemonic is what you are maintaining and it is fading. Nothing: miss.

illustration for 16
16majorTop pickdishA ceramic plate slipping from soapy hands and shattering on tile. T is the edge hitting first, Sh is the spray of white shards.

Grading the image has a consequence people find odd: the image is allowed to dissolve. After enough reviews, 16 will simply mean dish, the way "cat" means a cat with no picture in between. That is the mnemonic finishing its job; keep grading easy, let the scaffolding come down, and rebuild the picture only if the direct link weakens.

The other consequence: a miss is a diagnosis. The reset is the algorithm's job; yours is to ask why the picture failed. Usually it was vague (a plate, rather than your mother's blue plate hitting the tiles), or lacked interaction, or merged with a similar picture. Fix it before the next review rather than re-seeing the old one and hoping: the broken join of the link method and the vague image of the keyword method are repaired at the build step, never in the calendar.

When to walk a memory palace again

A palace is loci in a fixed order, and the order changes the review.

Early, review it as a unit. For the first two or three reviews (the day after, then about a week after), walk the whole route from the fixed start and grade every locus. Walking end to end makes the route durable, and the route is what finds a dark locus: the loci either side tell you where to look.

Later, review loci individually. Once the route is solid, a full walk every time is wasteful. The palace review keeps SM-2 state per locus and puts every filled locus in front of you, most overdue first; you flip each to reveal, grade it, and get the loci you missed listed at the end so you know which images to rebuild. It does not walk the route for you, so keep an occasional whole-route walk in your head as well, always from the same start; a review that begins in the middle is a different test.

One of the few controlled studies of mnemonic images on a review schedule is the password experiment of Blocki, Komanduri, Cranor and Datta (2015). Participants memorised four Person-Action-Object stories in photographed scenes on different rehearsal schedules. The best rehearsed at 12 hours and then at gaps growing by 1.5 times: 77 percent recalled all four stories in all nine tests over 102 days, and on average 95 percent of those who remembered in one round remembered in the next. Everyone given only one or two stories recalled them every time (the preprint puts two stories at 90 percent). Failures rose with the number of stories and clustered at the first 12-hour check: an image that survives its first honest recall very probably survives.

Which palaces to keep alive is a separate decision. Foer describes competitive memorisers reusing a stock of palaces from event to event and deliberately walking through them to empty out last week's images before a contest, because those images linger. A palace holding a shuffled deck is meant to be cleared; the only walk it needs is the one that checks it is empty. A palace holding the presidents or a speech you give every quarter is a reference palace and gets the full schedule.

How many reviews are enough

Rawson and Dunlosky (2011), across three experiments with 533 students and more than 100,000 recall responses, found that retention one to four months later was best served by recalling each item correctly three times in the first session, then relearning it to one correct recall three more times at widely spaced sessions; the relearning had, in their words, pronounced effects on long-term retention with a relatively minimal cost. Against the SM-2 table that lines up: the learning session, then days 1, 7 and about 23, by which point the item is heading past a year on its own.

Whether those reviews should widen is less settled than every schedule implies. Expanding retrieval practice was proposed by Landauer and Bjork in 1978; Karpicke and Roediger (2007) compared it with equal spacing across three experiments and found that expanding won ten minutes later, equal spacing did better two days later, and what improved long-term retention in every case was delaying the first retrieval, however the rest were spaced. So make the first spaced review an honest recall after a night's sleep, do enough of them, and space them at all; a better curve gains you less than pressing miss when you missed.

Self-grading bias, and why the miss button is the whole game

Kornell (2009) had students learn flashcards either as one stack of twenty (each card returning after nineteen others, which is spacing) or as four stacks of five (massing), and compared both with cramming the day before. The big stack won, for 90 percent of participants. After the first session, 72 percent believed the small stacks had worked better.

Massed practice feels better because the answer is still warm when the card returns. The same warmth corrupts self-grading: a card whose answer you saw a moment ago feels easy, and a peg you produced with a hint feels known. It is not. The words drill here counts a hinted answer as a failed review whichever button you press, and the palace review reveals the answer only after you flip it; both exist to make the warm feeling harder to act on. Grade what you produced before the reveal, not how it felt after.

A 30-day plan for one mnemonic set

Here is the schedule applied to twenty Major pegs (00 to 19). The same plan works for the presidents, a PAO deck or a twenty-locus palace.

Day 0. Build the pictures for 00 to 19 from the Major System chart, saying each aloud. Within the hour, cover the words and recall all twenty from the numbers, cycling until every peg has come up correctly three times, and rebuild on the spot any picture that would not come. That is building, not reviewing, and it is the day's only work.

Day 1. First spaced review. Before looking at anything, recall all twenty. Easy if the picture arrived, hard if the word came without it, miss if nothing. Rebuild every miss with a more specific image.

Days 2 to 6. Nothing for this set; the gap is the treatment. If you want work, start 20 to 39 on day 2, a few days out of phase.

Day 7. Second spaced review. Expect the hard ones from day 1 to be hard again; that is the information you are after. Repair any misses.

Day 23 or thereabouts. Third spaced review. Most pegs now arrive as direct number-to-word links with the picture fading, which is what you want; rebuild the few still hard.

Day 30. Hand over. Adopt the Major set below, then open the custom drill on the Major deck, set the card count to 20, press "ready to recall" to skip the memorise timer, and grade each peg. The session puts due pegs first, then unseen ones, and each grade updates that peg's own SM-2 state; from here the calendar is per peg. The numbers drill is a different test, digit strings typed back against a clock with no schedule behind it; use it to check the pegs cold once a week.

adopt the starter system
Adopt Major System

The canonical Major System ships with Major Dictionary as a starter system you can edit, override, and drill against. One click to adopt; the system joins your dictionary and the practice surfaces light up.

To see the effect rather than take it on trust, run the free grocery challenge (ten items, no account) before day 0 and again after day 30.

Missed days and backlogs

You missed a day. Nothing is lost: an item due yesterday and recalled today has survived a longer gap than planned, which is slightly stronger evidence than the planned review. Review it, grade honestly, carry on.

You missed a week or more. Do not clear the backlog in one sitting; a hundred stale items reviewed at speed produce a hundred untrustworthy grades. Take the oldest twenty or thirty a day until the queue is normal.

Paper box, Anki, or the drills here

A paper Leitner box is five compartments: a card you recall moves one along, a card you miss goes back to the first, and compartment one is reviewed daily, the last monthly. It is SM-2 by hand with coarser gaps and no ease factor, and it works; its weakness is that nothing tells you what is due, and the last compartment tends to go unreviewed forever.

Anki runs SM-2 properly on any material. Its defaults are learning steps of 1 and 10 minutes, a graduating interval of 1 day, an easy interval of 4 days, starting ease 250 percent, a lapse back to the start, twenty new cards a day, and a maximum interval of 36,500 days. Since version 23.10 it also offers FSRS, which fits a model to your review history and targets 90 percent retention by default. The decks on the Major System page (100 two-digit pegs; 1,000 three-digit pegs) import straight into that cadence, number on the front, peg on the back. Nothing on an Anki card knows whether you saw a picture or recited a word; that discipline is yours.

The drills here know what each item is. The palace review schedules per locus, the words drill runs the same three grades over vocabulary from adopted packs with the keyword and story shown on reveal, and the custom drill runs them over any deck you own, due cards first. The pi journey treats each chunk of digits as one SM-2 card and will not unlock the next stage until the current one reaches 80, 75, 70 and 65 percent chunk coverage for stages one to four; coverage counts the chunks you have graded as recalled, so the honest check is the journey's separate typed test of the whole sequence. For a workflow that puts these together with notes and reading, see memory techniques for studying. Use whichever one you will actually open tomorrow.

How often should you review flashcards? Every day?

Whenever they are due, and only then. Open the queue daily, because items come due daily; that is different from reviewing every card every day, which is massing. Anki's default adds twenty new cards a day on top of the reviews; the 30-day plan above adds twenty pegs at a time for the same reason.


If you have a palace with loci in it, open the palace review now and grade every locus it hands you honestly, then walk the whole route once in your head from the start. If you do not, adopt the Major set above, build pictures for 00 to 19 today, check them within the hour, and come back tomorrow before you look at anything.