Close your eyes and picture a red apple. Some people see one. Roughly one in fifty see nothing at all, and until fairly recently most of them assumed the instruction was a figure of speech.
That condition has a name now. Aphantasia is the absence of voluntary visual imagery, and the interesting question for a site built around remembering colors is what happens to color memory when the picture is missing. The honest answer is that something else has to carry the color, and the most likely candidate is the word.
So I measured what a word is worth. Take the eleven basic color terms of English, hand each one the region of the sRGB gamut it owns, and ask how far apart two colors that share a name can be. The answer is 17.9 CIEDE2000. On the ten point scale this site scores guesses with, remembering only the word puts a ceiling of 6.30 out of 10 on your answer, before memory, attention or screen calibration have taken anything off the top.
That number turns out to be stubborn. Widening the vocabulary helps far less than you would expect, and the reason why is a piece of geometry rather than a fact about English.
What aphantasia is, and how it is normally measured
Adam Zeman and colleagues at Exeter proposed the term in 2015, after a patient who had lost his mental imagery following cardiac surgery prompted a wave of correspondence from people who reported never having had any. Prevalence estimates cluster around one to four percent depending on where the line is drawn. Zeman's own figures put complete absence of imagery near 0.7 percent and a broader weak imagery band near 2.6 percent.
The standard instrument is the Vividness of Visual Imagery Questionnaire, usually shortened to VVIQ. It asks you to picture a scene, a sunrise, a shop front, a relative's face, and rate the vividness of each image from one to five. It has been in use since 1973 and it works well enough that most aphantasia research rests on it.
It also has an obvious weakness, and the field knows it. VVIQ asks you to rate an experience only you can observe, on a scale whose anchors mean whatever you take them to mean. Somebody who has never had visual imagery has no reference for what a five is supposed to feel like. Somebody who has vivid imagery may still rate themselves harshly. Self-report is the only direct access anyone has to the mind's eye, and it is also the least trustworthy kind of data psychology collects.
This is why behavioural measures matter. If imagery is doing work, it should leave fingerprints on performance, and performance can be measured without asking anyone how vivid anything felt.
The drawing study that points at words
The most useful piece of evidence here comes from Wilma Bainbridge and colleagues, who in 2021 ran a large drawing study on aphantasia. Sixty one people with aphantasia and matched controls looked at photographs of real scenes, then drew them from memory, then later copied the same scenes with the images in front of them.
The results split cleanly in two. Spatial accuracy, meaning where things sat relative to each other, was essentially identical between the groups. Object memory was not. Participants with aphantasia recalled significantly fewer objects and drew them with far less visual detail.
The detail that matters most for color is what they did instead. They wrote. Rather than drawing a sofa, some wrote the word sofa in the space where a sofa had been. Bainbridge and colleagues describe this as increased reliance on verbal scaffolding, and they note that the same strategy left the aphantasia group making fewer false memory errors than controls. Words are a lossy way to store a scene, but they are a conservative one. You do not hallucinate a detail you never encoded.
If that is the strategy, then color memory in aphantasia should inherit whatever precision a color word carries. Which is a question you can put a number on.
How much color fits in a word
The method is simple enough to describe in a sentence. Sample the sRGB cube on a regular lattice, assign every color to the nearest of the eleven basic color terms in CIELAB, then ask how far apart two colors drawn independently from the same name are, measured in CIEDE2000.
That last step is the one that matters. It is not a measure of how big each named region is in the abstract. It models the actual failure. You were shown a color, all you kept was the word, and now you have to produce a color from that word. Both the color you saw and the color you produce are draws from the same region, so the expected distance between them is the error the word hands you for free.
Across 140,608 sampled colors, here is what each of the eleven basic terms costs, with the score this site would give a guess that far off.
- gray, 13.0 percent of the gamut, expected error 24.97, score 5.04
- black, 2.0 percent, expected error 20.93, score 5.72
- purple, 14.4 percent, expected error 18.83, score 6.12
- green, 13.4 percent, expected error 18.81, score 6.12
- white, 8.9 percent, expected error 18.68, score 6.15
- brown, 4.8 percent, expected error 17.98, score 6.29
- pink, 6.7 percent, expected error 17.46, score 6.39
- blue, 12.5 percent, expected error 15.03, score 6.91
- red, 6.2 percent, expected error 14.48, score 7.03
- orange, 4.7 percent, expected error 14.20, score 7.09
- yellow, 13.4 percent, expected error 13.91, score 7.16
Weighted by how much of the gamut each word covers, the floor comes out at 17.90 CIEDE2000, or 6.30 out of 10.
Gray is the worst word in English
The spread across those eleven entries is the part I did not expect. Yellow costs 13.91 and gray costs 24.97, so the least informative color word in English is nearly twice as lossy as the most informative one.
The reason is shape. Yellow occupies a small, tightly bounded pocket of color space. It has to be light, it has to be reasonably saturated, and its hue band is narrow, so saying yellow rules out a great deal. Gray does the opposite. It runs the entire length of the lightness axis, from nearly black to nearly white, and claims everything with low enough chroma along the way. Saying gray tells the listener that a color is not colorful. It says almost nothing about how light it is, which is the dimension the eye is most sensitive to.
This is consistent with something I measured while working through the shades of gray, where the named gray vocabulary turned out to lean cool by seventeen names to nine. English compensates for gray being a bad word by inventing a lot of more specific ones. Charcoal, slate, ash and silver all exist because gray on its own does not narrow anything down.
Does a bigger vocabulary rescue it
The obvious objection is that nobody remembers colors using only eleven words. People say teal, burgundy, mustard and sage. Designers reach for a few hundred names without thinking about it. So I ran the same measurement on wider vocabularies.
- 11 basic terms, floor 17.87, score 6.31
- 41 everyday words, the basics plus the shades people actually say out loud, floor 12.18, score 7.56
- 141 CSS named colors, a designer's working vocabulary, floor 7.74, score 8.62
Going from eleven words to a hundred and forty one, a nearly thirteen fold increase, cuts the error by less than half. That is a poor return, and it is worth understanding why, because the reason generalises past English.
Color words are cube root efficient
To take language out of it, I replaced the vocabularies with optimal ones. For a range of sizes, I ran k-means over the gamut in CIELAB to find the best possible set of N color categories, then measured the same within-category error. This is the best any language could do with N color words, rather than what English happens to do.
- 2 words, floor 29.98, score 4.34
- 4 words, floor 24.75, score 5.08
- 8 words, floor 20.45, score 5.81
- 11 words, floor 17.97, score 6.29
- 16 words, floor 15.13, score 6.89
- 32 words, floor 12.20, score 7.55
- 64 words, floor 9.38, score 8.23
- 128 words, floor 7.34, score 8.71
- 256 words, floor 5.80, score 9.07
- 512 words, floor 4.57, score 9.34
Fit a power law to that and the exponent is minus 0.347, with an R squared of 0.9982. That is a cube root law, and it holds almost perfectly across more than two orders of magnitude.
The geometry behind it is not complicated. Color space is three dimensional. Splitting a three dimensional volume into N cells shrinks each cell's diameter by the cube root of N, so error should fall as N to the power of minus one third. The measured minus 0.347 sits close enough to minus 0.333 that the small excess is explained by CIEDE2000 being a warped metric rather than a Euclidean one.
The practical consequence is brutal. Every doubling of your color vocabulary buys a 21 percent reduction in error. Extrapolating the fit, reaching a just noticeable difference of 2.3 would take roughly 3,700 color words, and getting to a CIEDE2000 of 1 would take around 41,000. No language has ever had anything like that. English has eleven basic terms and a few hundred usable modifiers, which by this measurement lands somewhere around a CIEDE2000 of 6 or 7 at absolute best, assuming perfect recall of the right word.
Language is simply not a high bandwidth channel for color. It never had to be, because for most of human history the person you were talking to could look at the thing.
What this says about testing for aphantasia
Here is where I think this becomes useful rather than merely tidy.
If verbal encoding is the fallback strategy, then it has a measurable ceiling. On this site's scoring, sustained performance from a purely verbal strategy should sit near 6.3 out of 10 with a basic vocabulary and somewhere below 8.6 even with a designer's vocabulary and flawless word recall. Scores comfortably above that band, held over many rounds, mean something other than a word is carrying the color.
That gives you something VVIQ cannot give you, which is an external reference point. You are no longer asking someone to rate the vividness of an experience against an imaginary scale. You are asking whether their performance exceeds what words alone can deliver, and the value of what words alone can deliver is computed rather than reported.
I want to be careful about how far to push this, because there are at least three ways it could mislead.
- Verbal is not the only non-imagery code. Someone might hold a color as a motor memory of where a slider sat, or as a relational judgement against the previous trial, without any picture and without any word. Bainbridge and colleagues found intact spatial memory in aphantasia, which is a reminder that the absence of pictures does not mean the absence of structure.
- The ceiling depends on the interface. A game with a continuous color picker allows fine adjustment that a multiple choice test does not, and any test that offers a shortlist hands back precision the verbal strategy never had.
- None of this is diagnostic. Aphantasia is a subjective condition, and a score is not a substitute for a clinical conversation. What a score can do is tell you that your color memory is or is not better than words, which is a genuinely different question from how vivid your imagery feels.
There is also a result in the aphantasia literature that should keep anyone honest here. People with aphantasia frequently perform normally on visual working memory tasks despite reporting no imagery at all. The subjective picture and the measured performance come apart. Whatever a color memory score measures, it is not vividness.
Why colors are the right test case
Most memory tasks are hard to score precisely. If someone describes a remembered face, there is no accepted unit for how wrong they were. Color does not have that problem. CIEDE2000 gives a defensible perceptual distance between what you saw and what you produced, calibrated so that a value near one sits at the threshold of noticing.
That makes color unusually well suited to this question. It is a continuous, three dimensional, precisely measurable stimulus with a well studied verbal vocabulary sitting on top of it. You can compute the exact cost of the word, which is what this article has done, and then measure how far a person beats it.
It also explains an experience a lot of players report. Holding a color for a few seconds feels easy and turns out to be hard, and the gap between the two is roughly the gap between having a picture and having a word. I looked at the decay side of this in how long you can remember a color, and at the encoding side in the science of color memory. This measurement fills in the floor underneath both of them.
Try it against your own numbers
The quickest version of this is the main color memory game. You see a color, it disappears, you reproduce it, and you get a score out of 10 built on the CIEDE2000 distance described above. Play a run of rounds and compare your average to 6.30, which is what pure verbal encoding with the basic color words would deliver.
If you want to isolate the verbal channel specifically, the Name That Color variant is the cleanest instrument on the site, because naming is the entire task. The Hex variant pushes the other way and asks for a numeric code, which is a different encoding strategy again and one that has no ceiling of this kind at all. The daily challenge is the fairest comparison over time, since everyone gets the same colors.
One prediction worth testing on yourself. If you have aphantasia and you are relying on words, your errors should be strongly clustered inside the named region, meaning you land on a plausible blue but the wrong blue. If you have imagery, your errors should be smaller and less tied to category boundaries. That difference is visible in your own results long before any questionnaire gets involved.