The Keyword Method for Foreign Vocabulary
A sound-alike keyword and one interacting image put a foreign word in your head in thirty seconds. Fifteen worked examples in Spanish, Mandarin and Japanese, the research, the review routine that keeps them, and a ten-minute test you can score tomorrow morning.
Ten minutes from now you will know ten words in a language you do not speak, learnt from nothing but their sound and a trick called the keyword method: find a word in your own language that sounds like the foreign one, then picture the two of them doing something together. Spanish tener, to have, sounds like tenor, and a tenor can hold a note in his fist. Tomorrow morning you can count how many of the ten are still there. Most will be, because each word will have something to hang on, which is the one thing a foreign word arrives without.
The keyword method is one of the most heavily tested mnemonics there is for vocabulary, fifty years of experiments behind it and a place among the ten techniques Dunlosky and colleagues assessed in their 2013 review. It is also only the front half of the job: it puts a word in your head in thirty seconds and does not, on its own, keep it there, and the review routine that does is here too. The parent page for the family is the guide to mnemonic devices, the technique index lists the rest, and this is the one for words.
Why foreign words refuse to stick
A word in your own language is a sound wired to a meaning by ten thousand uses. A word in a new language is a sound wired to nothing. The pairing of tener with "to have" is arbitrary, and arbitrary pairings are what memory is worst at.
Forgetting is steep too, and the classic numbers are usually misquoted: Ebbinghaus measured savings, how much less time a list of nonsense syllables took to relearn, not how many syllables were gone. The spaced repetition guide prints his series and Murre and Dros's 2015 replication in full, and the "seventy percent forgotten in a day" and "half gone in an hour" figures that circulate are not values in either series. What matters here is what was on the list: meaningless syllables learnt by rote, which is what a foreign word is until you give it a meaning. Giving it one is the whole of the keyword method.
The keyword method in two steps
The method was named and tested at Stanford by Richard Atkinson and Michael Raugh in the mid-1970s, and it has exactly two links.
The acoustic link. Find a word you already know that sounds like the foreign word, or like its first or stressed syllable. That is the keyword. Spanish tener (tay-nair) sounds like tenor. Mandarin nǐ sounds like knee. Japanese mizu sounds like mee-zoo.
The imagery link. Picture the keyword doing something to the meaning. The tenor has the high note, gripping it in his fist. You greet someone ("you") by tapping their knee. You are at the zoo, drinking water.
At recall the foreign word gives you the keyword, the keyword gives you the picture, and the picture contains the meaning: three hops you can see, instead of one arbitrary leap. Here are those three as they appear in the site's own packs, keyword and story on the card.
Tenor covers both syllables, and a tenor is a person, so he can be handed the meaning: he has the high note, visibly, in his fist.
Knee is the whole syllable and a body part, so the image can involve you. The tone is a separate job, handled below.
Mee-zoo covers both syllables in order, and a zoo gives the water somewhere to be.
If this looks like the substitute-word trick for names, it is. Harry Lorayne was teaching it for foreign words in The Memory Book in 1974, the year Raugh and Atkinson's first Stanford technical report named it, and the Linkword language courses incorporate it. The academic version stuck because it came with numbers: in one of Raugh and Atkinson's four Spanish experiments, students given keywords scored 88 percent on the final test against 28 percent without them, and Russian, whose sounds are much further from English, gave 72 against 46. Three decades later Sagarra and Alba compared methods across 778 beginning learners of Spanish; the keyword method again produced the best retention, with rote learning of translation pairs second and semantic mapping (grouping words by meaning) last.
Fifteen worked examples in Spanish, Mandarin and Japanese
Twelve more from the same packs, each with the reason its keyword was chosen; the last column is where the method lives.
| Word (sound) | Meaning | Keyword | Image | Why this keyword |
|---|---|---|---|---|
| estar (ess-tar) | to be (state) | STAR | a film star slumped in a chair, asked how she is doing | the stressed syllable; a person, not the abstract "be" |
| hablar (ah-blar) | to speak | blah-blah | a woman at a microphone saying nothing but "blah, blah" | the stressed syllable, and the keyword is already about talking |
| noche (noh-cheh) | night | nacho | a nacho bar that only opens at night, neon sign, moon above | first syllable plus the ch; you can taste the image |
| puerta (pwair-tah) | door | porter | a porter wedged in a door, holding it open with his back | a porter's whole job is doors, so the interaction is free |
| calle (kah-yeh) | street | kayak | a kayak paddled down a flooded street between parked cars | Spanish ll is a y sound, so "kayak" beats any "call" word |
| hǎo (third tone) | good | how | "how are you? all good", with a thumbs up | a rhyme that covers the whole syllable |
| bù (fourth tone) | not, no | boo | a ghost shouting boo with its arms crossed: no | exact sound; the falling fourth tone is a shout |
| mǎi (third tone) | to buy | my | "it's mine now": buying it and clutching it to your chest | exact sound; the dip and rise of the tone fits grab and lift |
| shū (first tone) | book | shoe | a shoebox packed with books, lid bulging | exact sound; two objects that fit inside each other |
| anata | you | Anita | "Anita, that's you", pointing straight at her | a name you already own covers all three syllables |
| neru | to sleep | Nero | Nero sleeps in his bed while Rome burns outside | a famous name whose story is already a picture |
| tsukue | desk | skewer | a kebab skewer driven straight through a desk | tsu is hard for English mouths; "skewer" catches the s-k-u shape |
You will disagree with some of these, and a keyword you make yourself is fine. But the evidence gives self-made keywords no advantage over handed-down ones: Campos and colleagues found keywords made by peers beat the learners' own for vivid words, and Thomas and Wang found that generating your own did nothing to slow forgetting. Take whichever arrives first.
How to choose a keyword that works
Beginners polish the image and rush the keyword, which is backwards: a bad keyword cannot be rescued by a good picture. Six rules.
Match the stressed syllable, or the first. The syllable your ear lands on is the one that cues the keyword. For hablar that is -blar, so blah beats hab. A keyword that covers only the first letter ("net" for noche) is a hint that offers you every Spanish word starting with n.
Prefer a concrete noun. "Tenor" is a person; "tenure" is a contract. Both sound like tener; only one can hold a note in its fist. Ellis and Beaton found imageable noun keywords helped learning while verb keywords impeded it.
Accept near sounds. Star for estar is not exact and does not need to be; you are building a cue, not a transcription.
Use a word you already own, from your own language or from target words already learnt. Once you know Spanish casa, it is a fine keyword for Japanese kasa (umbrella): a house held over your head.
One keyword per word. Craft advice rather than a finding: two candidates become, at recall, two half-formed images that cancel each other. Pick one and throw the other away.
When nothing sounds alike, split it into one keyword per chunk (mee-zoo), use a name (Anita, Nero), or take the first syllable and put the rest of the word into the picture as a label. Or borrow one: the peer-keyword result above is the case for a shared dictionary of them.
How to build an image that survives a week
The picture has to get you to the meaning a week later with no help. Four rules, each preventing a specific failure.
They must interact. The porter wedged in the door, not a porter beside a door. It is the same join the link method makes between list items.
Exaggerate the part that carries the meaning. The tenor holds the note until his face is purple; the shoebox of books splits.
Attach a real picture. Thomas and Wang found that supplying a picture of keyword and meaning together during study improved long-term retention, probably because a picture adds visual detail the mental image would otherwise lack. Draw it, badly, on the card.
Say the foreign word aloud while you look. The keyword is an approximation; saying the real word over the image binds the true sound to the scene, so the keyword can later drop away and leave the word.
Should you build a memory palace for vocabulary?
Usually not, and the advice that opens with "build a palace for Spanish" is answering a different question. A palace stores things in an order you have to reproduce; a vocabulary list has no order, because you need puerta when someone says door, not third on the left in the hallway. A word and its meaning are a pair, which is what the keyword is built for, and the parent page's when-to-use table draws the line the same way. The exception is a small set you must produce in sequence, a phrase list for a trip or the numbers one to twenty, which sits on a route perfectly well: the method of loci guide builds one.
What the research says, and what it does not
Fifty years of testing have documented the limits as well as the strengths.
It forgets faster than rote unless you review. Across four experiments with 218 students, Wang, Thomas and Ouellette found the expected keyword win on an immediate test and then greater long-term forgetting for the keyword group than for plain rote rehearsal. What answered it, in the same authors' later work, was not a cleverer keyword but a better image, which is the picture rule above. The other half of the answer, a recall before the loss happens, comes from the testing work below, not from them. That is a fixable failure, not a verdict on the method.
It works for recognition first, and for production only with good images. Ellis and Beaton found in 1993 that keywords helped learners go from foreign word to meaning, while rote repetition was better for going from meaning to foreign word. Beaton and colleagues showed in 2005 that with adequate keyword images the method improved both directions and the earlier result reversed.
It does little for abstract words. Campos, Amor and González had 363 students learn sixteen Latin words, half vivid and half not, by rote or by keyword. For the vivid words the keyword groups won; for the low-vividness half the method made no measurable difference. Give abstract words a concrete stand-in (a set of scales for "justice") or rote them.
It was rated "low utility" for general study. Dunlosky and colleagues' 2013 review of ten techniques declined to recommend the keyword mnemonic: narrow range of material, low efficiency once you count the training and the keyword-making, and learning that may not last. They rated practice testing and distributed practice high, and judged practice testing superior to the keyword mnemonic even on foreign vocabulary, where the head-to-head comparison they cite found cued recall no better after keywords than after testing, or lower a week out. We think the two combine rather than compete, and the routine below tests the keywords instead of trusting them: that is our reading, not theirs.
What the keyword does not cover in each language
"Does it work for Chinese or Japanese?" Yes; both are built from a small set of simple syllables, which suits it. The adjustment in every language is the same, English included: work out what the keyword does not cover (a gender, a tone, a character, a definition) and give that its own place in the picture.
| Language | What the keyword covers | What it misses | How to cue the missing part |
|---|---|---|---|
| Spanish | the sound of the words that are not cognates | gender, and the false friends whose sound hands you the wrong meaning | a fixed masculine prop and a fixed feminine prop in every image; a deliberate image for each false friend |
| Mandarin | the consonant and vowel of a pinyin syllable | the tone, and the written character | a fixed action per tone inside the image; a separate image for the character, built from its components |
| Japanese | the sound, and through the kana the reading | which reading a kanji takes, and the character itself | keyword the reading in the word you are learning; treat the character as its own task |
| English academic and technical terms | the sound of a long Latinate word | the definition's precision, and the word's register | keyword the stressed syllable, then write the real definition under the image |
Spanish: skip the cognates, then put the gender in the picture. A good share of the words you meet first need no keyword at all: -ción is -tion (nación), -dad is -ty (ciudad), -mente is -ly, and hundreds of nouns are near copies (hospital, animal, doctor). Read those; do not memorise them. Spend keywords on the words that are not cognates, and on the false friends, the cognates that lie: embarazada is pregnant, éxito is success, sensible is sensitive. Those deserve an image precisely because the sound hands you the wrong meaning for free.
For the rest, keyword the sound, not the spelling: the ll is a y, the j a breathy h, the h silent, b and v the same soft sound. Then put el or la into the image instead of learning it as a separate fact: choose one fixed extra element for every masculine picture and another for every feminine one, and make it appear every time. If your markers are a boxer and a bottle of perfume, the porter wedged in la puerta has perfume spilling down his uniform, and el libro is being read by a boxer between rounds. Ten words in, the marker arrives with the image and the gender with the marker.
Mandarin: a keyword for the syllable, a separate cue for the tone. A keyword gives you the consonants and vowel of a pinyin syllable, not the tone, and mǎi (buy) with the wrong tone is mài (sell). Give each tone a fixed action and build it into every image: first tone, the object floats level; second, it lifts; third, it dips and bounces back; fourth, it slams down, so the ghost shouting boo slams a door. The character is a third job: the keyword links sound to meaning, and the written form needs its own image, built from its components, once you have a hundred spoken words.
Japanese: the easiest case. Nearly every syllable is a consonant plus a vowel, and there are only about a hundred of them, so words break into keyword-sized chunks by themselves (mi-zu, ne-ru), and the kana are phonetic, so one keyword covers sound and reading. Kanji can have several readings, so keyword the reading in the word you are learning (mizu for 水) and treat the character as its own task.
English vocabulary, and the technical terms of a subject. Nothing about the method is foreign-language-only. Apoptosis, programmed cell death, sounds like "a pop, toes", and a cell shaped like a foot popping its toes off one at a time on a timer is programmed death: same two links, same thirty seconds. That is how a long Latinate term goes in, whether it is an SAT word, a drug name, an anatomy label or the jargon of a new job. What it misses here is not a gender or a tone but the definition: a keyword hands you a gloss, and a gloss is not what an exam marks, so write the real definition under the image and check the image against it. The guide to memory techniques for studying works that example out in full and puts it next to the other things worth memorising in a course.
Keywords plus spaced review: the routine that lasts ten years
Keyword experiments use short lists and short intervals: writing in 1995, Beaton, Gruneberg and Ellis put the typical study at fewer than forty words and a month or less, and set a single case against it. Their subject had worked through a Linkword Italian course, about 350 words in roughly ten hours, and had not touched Italian in the ten years since. Cold, he produced 35 percent of the test words with perfect spelling and more than half allowing minor slips. Ten minutes reading through the list took him to 65 and 76 percent, and a further hour and a half of revision brought recall back to virtually the whole list, which held for at least a month. That is what a keyword buys: not permanence, but a memory that relearns almost for free.
The second half of the routine is testing. Roediger and Karpicke had students learn short passages and then either restudy them or take a recall test: after a week the tested group led 56 percent to 42, and studying once with three recall tests gave 61 percent a week later against 40 for studying four times. Reading a list feels like learning and mostly is not; retrieving, and failing some, is.
How much testing, and when? Rawson and Dunlosky, across 533 students learning conceptual material, found the best trade between retention months later and time spent was to recall each item correctly three times in the first session, then relearn it to one correct recall in three widely spaced later sessions. Put that together with the keyword and you have the routine:
- Encode each new word with a keyword and an interacting image, and say the word aloud.
- Within the hour, before you close the session, test the batch until each word has come up correctly three times, and test it both ways: word to meaning, then meaning to word, said aloud and then written, because the keyword carries the sound and not the spelling. The words drill has a toggle for exactly those two, "word → meaning" and "meaning → word"; run the harder meaning-to-word direction at least a third of the time. This is where a weak image gets found and rebuilt, so it is part of building the word, not a review.
- The first spaced review is the next day. Then recall the batch at about six days, about sixteen days, and monthly after that. For an exam a month away: the next day, about day 7, about day 20, and a final pass two days before.
- Whatever you miss, rebuild the image rather than re-reading the pair.
The gaps need not be exact. The spaced repetition guide explains why the first spaced review is tomorrow rather than tonight, and why recalling at all matters more than the exact gaps; the words drill runs the SM-2 shape for you. It also covers the apps, if you would rather software kept the dates: Anki schedules with SM-2 by default and with FSRS since version 23.10, and its default of twenty new cards a day is double the rate the plan below assumes.
A 30-day plan for your first 300 words
How many words a day? No study sets a rate; ten is what we find survives the review load. Twenty is possible for a fortnight, and then the reviews pile up.
How long to 1,000 words? At ten a day, about a hundred days. That is not yet a newspaper: Nation estimated that unassisted reading at 98 percent coverage needs around 8,000 to 9,000 word families, and listening 6,000 to 7,000. The first thousand is still the most valuable, because it is the most frequent. The plan below is the first 300.
| Days | New words | Reviews |
|---|---|---|
| 1 to 3 | 10 a day, concrete nouns and top verbs | in session until each word has come up three times; next-day recall from day 2 |
| 4 to 7 | 10 a day | next-day recall; productive direction from day 4 |
| 8 to 14 | 10 a day | day 8 adds the six-day recall of day 1's words; next-day plus six-day from there, about 20 reviews a day |
| 15 to 21 | 10 a day; hold at 5 if reviews pass 15 minutes | next-day and six-day both running at full width |
| 22 to 30 | 10 a day, plus function words by rote | day 24 adds the sixteen-day recall of day 1's words, the heaviest stretch; on day 30 test the first fifty cold, both ways, and aim for 80 percent |
Three hundred words in a month at half an hour a day, most of it reviewing. Function words (the, and, of) get no keyword: they are abstract and frequent, and they arrive by themselves from reading. The plan works on paper or in the words drill; if the list is your own, from a textbook or an Anki deck, the import page takes CSV, TSV and .apkg files.
Retiring the keyword, and what replaces it
The keyword is scaffolding. After a few successful recalls the word arrives before the image does: you see puerta and think "door" without visiting the porter. When that happens, stop rehearsing the image. The keyword falls away on its own, and what is left is the word, which is what a native speaker has. If a retired word later goes dark, the porter can be rebuilt in a second.
What the keyword never had is the rest of the word. It cues a sound and a gloss and nothing else: not the spelling, not the gender or the conjugation, not the words this word normally travels with, not which of two near-synonyms a speaker would actually reach for. Knowing that puerta means door is a long way from producing a sentence with puerta in it that sounds like Spanish.
Only input supplies that, so the honest answer to "should I memorise vocabulary at all" is: for a while. Memorising earns its keep over the first thousand or two words, where the frequency is high enough that waiting to meet each one in the wild would take years. After that, reading and listening to material slightly easier than you think you need does more per hour, and it is the only thing that supplies the half a keyword cannot.
A ten-minute drill you can run now
Do this before deciding whether the method is for you.
- Write down ten words you do not know, with their meanings. The JLPT N5, HSK 1 and Spanish 5K pack pages each preview a handful of entries with keywords; take those, skip any article or particle, and top up from the table above. No account is needed.
- Set a timer for five minutes. For each word pick your own keyword (ignore the pack's if yours comes first), build an interacting image, and say the word aloud once while you hold it.
- Cover the meanings. Go down the foreign words and say each meaning. Score it.
- Cover the foreign words. Go down the meanings and say, then write, each word. Score that too; expect it to be lower.
- Close the page. Tomorrow morning, before looking at anything, repeat steps 3 and 4 and write the scores next to the first two.
If you build the images rather than read about them, expect most of the ten in the word-to-meaning direction on day one and most of those still there the next morning. Raugh and Atkinson's keyword group scored 88 percent in the lab, so eight out of ten is a fair target rather than a promise; the meaning-to-word score will run lower. For a timed version, import your own list and use the custom drill, which times the memorising and then tests you card by card. To watch the images fire in real sentences, paste a paragraph into the reading page, where every word from an adopted pack shows its keyword and story on hover.
Ten words, ten minutes, a score you can check tomorrow. The keyword method gets the words in; testing and spacing keep them there. When you want the app to keep the score, adopt a pack and open the words drill: it prompts each word, you recall and grade yourself, SM-2 sets the next date, and any answer you needed the hint for counts as a failed review whichever button you press.