Memory
Verbal Memory Test
Words appear one at a time. For each one you say whether it has already turned up in this run. Your score is d′ — how far apart “seen before” and “new” actually were for you — printed with the standard error your run can support. Not the streak everyone else reports, because the streak goes up when you simply answer “seen” more often.
Verbal memory
One word at a time. For each one, say whether it has already appeared in this run. Half of them have. Every word you judge is also a word you now have to remember, so it gets harder as it goes. Three mistakes end it — and answering “seen” when it was new costs a life exactly like missing a repeat, which is the point.
What the instrument contributed
- Trials this run
- 00 repeats, 0 new
- Band on d′
- —standard error, from the trial counts
- Word pool
- 249common concrete English nouns
Your hardware contributes almost nothing here, and for once that is not a consolation. Nothing on this page is timed, so there is no frame period to quantise and no input lag to cancel — lib/apparatus.js is unused, as it is on our hearing test. The instrument that limits this reading is the run itself, and we can put an exact number on it. d′ is estimated from two rates, and each rate is estimated from a finite number of trials, so the band above is arithmetic rather than a caveat — it is the standard error your own trials support. Two other limits have no number and we invent none. The pool is 249 everyday concrete nouns, and rarer or more vivid words are recognised differently, so a different pool would give you a different score. And the whole task assumes English is a language you read fluently: if it isn't, this measures your English, and no correction we could apply would change that.
How this is measured
What a trial is
One word, and one two-way answer: seen before, or new. Half the trials are repeats of words already shown this run and half are fresh, which keeps the two kinds of trial balanced — d′ is estimated from a rate on each side, and a run made mostly of repeats would have a false-alarm rate built out of almost nothing. Nothing on this page is timed, so take as long as you like on any word.
The four outcomes, and why two of them cost a life
Every answer lands in one of four boxes: a hit (a repeat you caught), a miss (a repeat you called new), a false alarm (a new word you called seen) and a correct rejection. A miss and a false alarm each cost one of your 3 lives — the same price, deliberately. Making them cost the same is what stops “answer seen more often” from being a winning strategy, and all four counts are printed with your result.
How the four counts become d′
d′ is z(hit rate) − z(false-alarm rate), where z is the inverse normal — the standard signal-detection measure of sensitivity, and the one the research this page cites is expressed in. It answers a different question from “how many did you get right”: it asks how far apart the two kinds of word felt, in standard deviations, and it is constructed so that shifting how willing you are to answer “seen” moves both rates together and leaves the difference alone. Alongside it we report c, the bias itself — which way you leaned when unsure. That number is not a fault; it is simply the thing a streak score silently pays you for.
The correction for a flawless run, and which way it errs
A run with no misses and no false alarms has a hit rate of 1 and a false-alarm rate of 0, whose z-scores are infinite. Every recognition study meets this, and we use the log-linear correction: add 0.5 to the hit and false-alarm counts and 1 to each trial total, always, not only when the rates come out extreme. It is applied to every run so that no two runs are scored by different rules. Hautus (1995) compared it against the alternative and found it both less biased and biased in one consistent direction — it under-states d′. We chose it for that direction: it puts the error where this site always puts it, on the side that makes you look slightly worse than you are rather than slightly better.
The band is arithmetic, not a hedge
d′ comes with a ± figure, and it is the standard error your own trial counts support, propagated from the binomial variance of each rate. This is the one page on the site where the limiting apparatus is not the hardware but the run — and unlike a display refresh rate, that limit has an exact value we can print. A run that ended after twelve words earns a band wide enough to say so. A long one narrows it. Below 5 trials of either kind we show no d′ at all and say why, because a number whose band is wider than itself is worse than no number.
What we cannot tell
We cannot tell a forgotten word from a mis-aimed tap, or from a word you knew you had seen and answered wrongly anyway. We cannot control how far back a repeat sits — it is drawn uniformly from everything you have seen, so it might be two words ago or eighty, and laboratory studies control that gap where we deliberately don't. And we cannot know whether English is a language you read fluently. If it isn't, this measures your English as much as your memory, and there is no correction we could honestly apply.
What your equipment contributes
Essentially nothing, and this is the second test on the site where that is true — there is no timing here at all. No performance.now(), no requestAnimationFrame, and lib/apparatus.js deliberately unused, exactly as on our hearing test. A display's frame period quantises when a stimulus can appear; it cannot touch a count of right and wrong answers. What limits this reading instead is the run length and the word pool: 249 everyday concrete nouns, chosen for even familiarity because rare and vivid words are recognised differently from plain ones. A different pool would give you a different score, and we state the pool rather than pretending it is neutral.
How your result compares
Recognition memory has been measured in d′ for decades, so unlike the streak score there is something real to compare against. The largest synthesis we found pools 501 experimental conditions from 232 experiments, with about 8,615 younger and 8,833 older adults. Words were 52.1% of the conditions (261 of them).
| Group | Participants | Mean d′ | SD |
|---|---|---|---|
| Younger adults | 8,615 | 1.85 | 0.8 |
| Older adults | 8,833 | 1.39 | 0.71 |
The age gap is 0.46 d′ units (95% CI 0.41 to 0.51), which the authors describe as roughly a 7% drop in the hit rate alongside a 7% rise in the false-alarm rate. That pairing is worth noticing: it is the same two quantities this page reports separately, moving in opposite directions — which is precisely what a single streak number cannot show you.
What we matched, and what we could not
Matched: the measure itself — d′ from hit and false-alarm rates, the standard in this literature; a balanced mix of repeated and new items; and unrelated common words as material.
Not matched, and this one is large: nearly all of that research uses a study phase then a test phase — you learn a list, then you are tested on it. Ours is continuous recognition: every word is a test item and a study item at the same moment, and the list you are being tested against grows under you as you go. Those are different tasks with different demands, and no constant separates them.
Also not matched: laboratory studies fix the number of trials in advance; ours ends when you make 3 mistakes, so the trial count depends on your own performance and a weaker run is a shorter one with a wider band. They control how far back a repeated item sits; we draw it uniformly from everything seen. And their materials are chosen per experiment, where ours is one fixed pool of 249 nouns.
Read 1.85 as a bearing on what typical looks like for a related task under controlled conditions — not as a line you have passed or failed.
No percentile is shown, and here the reason is unusually specific. That SD of 0.8 is not the spread of people — it is the spread of experimental conditions, 501 of them, differing in list length, delay, material and instructions. Placing you on a curve built from it would be comparing you against a distribution of study designs. Add the format difference above, add a trial count that your own performance decided, and add a band we print precisely because the estimate is not sharp, and a percentile would be the least defensible number on the page. One drawn from BlinkBench's own visitors would be worse still.
Reference range: Fraundorf, S. H., Hourihan, K. L., Peters, R. A., & Benjamin, A. S. (2019). Aging and recognition memory: A meta-analysis. Psychological Bulletin, 145(4), 339–371.
Method (d′ and c): Stanislaw, H., & Todorov, N. (1999). Calculation of signal detection theory measures. Behavior Research Methods, Instruments, & Computers, 31(1), 137–149.
Correction for extreme rates: Hautus, M. J. (1995). Corrections for extreme proportions and their biasing effects on estimated values of d′. Behavior Research Methods, Instruments, & Computers, 27(1), 46–51.
Last reviewed: July 2026
Frequently asked questions
Why isn't my score just how many words I got right?
Because that number can be raised without remembering anything. Answer “seen” more often and you stop missing repeats, so your run of correct answers gets longer — your memory is unchanged. Every verbal memory test on the web reports that streak. We report d′, which is built from both kinds of error at once: how often you said “seen” to a word that really had appeared, and how often you said it to one that hadn't. Shift your answering habits and d′ stays put. Our own self-check asserts exactly that, by scoring a cautious answerer and a liberal one with the same underlying memory and requiring the two d′ values to match.
What is d′, and what counts as a good one?
d′ is how far apart “I have seen this” and “this is new” felt to you, measured in standard deviations. Zero means the two were indistinguishable and every answer was a guess; higher means they were further apart. For context, a meta-analysis of 232 experiments found young adults averaging 1.85 and older adults 1.39 on old/new recognition. We deliberately do not turn that into a percentile for you — the reasons are set out in full on this page, and the shortest of them is that the 0.8 standard deviation describes variation between experiments, not between people.
Could I just answer “seen” every time?
You would last about six words. Saying “seen” to a word that is new is a false alarm, and it costs a life exactly like missing a repeat does — that symmetry is deliberate, and it is why 3 lives is enough to make the strategy fail quickly. If you somehow survived, your hit rate and your false-alarm rate would be equally high, which is the definition of d′ = 0. The streak strategy and the honest measurement disagree, and this page is built so the honest one wins.
Doesn't it get harder as it goes? Is that measuring me or the test?
It gets harder, and that is the test rather than you — which is why we do not report the run length as if it were a capacity. Every word you judge joins the list you then have to remember, and a repeat is drawn from anywhere in that list, so late in a run you may be asked about a word from eighty trials ago. Laboratory studies control that gap; we do not, and we could not without making the task something else. d′ absorbs some of this by pooling across the whole run, but it does not remove it, so a long run and a short one are not measuring quite the same thing. The band printed beside your score is honest about the trial count, not about this.