You have seen the chart. Ten swatches down the left, a color word beside each one, and a short list of feelings beside that. Red is passion. Blue is trust. Green is growth. It is one of the most reproduced diagrams on the internet, it appears in brand guidelines and pitch decks and interior design blogs, and almost every version of it prints a specific hex code next to the word.
That hex code is the problem. The research these charts point at is real research, and some of it is very good. But the biggest and best of those studies never showed anybody a color. It showed them a word. Somewhere between the paper and the chart, a word got converted into a swatch, and nothing in the data licensed that conversion.
So I measured the size of the gap. Below is what happens when you take the 701 named colors already published across this site, sort them into the ten words a color psychology chart uses, and ask how much color each word is actually standing for. The short version: blue covers 90 CIEDE2000 units from end to end, 101 hex codes are filed under two or more of these words at once, and one particular hex is officially red, purple and pink, which on a standard chart makes it passionate, luxurious and nurturing simultaneously.
The study everybody cites tested words
The largest piece of evidence in this area is Jonauskaite and colleagues, published in Psychological Science in 2020. It is a genuinely impressive piece of work: 4,598 participants across 30 nations speaking 22 native languages, each asked to associate 20 emotion concepts with 12 color terms.
The result is strong. Pattern similarity across nations came out at an average of r = .88, which is a high level of agreement for anything in cross-cultural psychology. Every one of the 30 nations agreed that black and red are the most emotional colors and that brown is the least emotional. Fourteen specific pairings held up everywhere, including black with sadness and fear, red with love and anger, and yellow with joy.
The local variation is just as interesting. Chinese participants linked sadness to white, which tracks funeral custom. Greek participants linked sadness to purple, which tracks mourning dress. Nigeria was the only nation to associate red with fear. Egypt was the only one that did not associate yellow with joy. Nations that were geographically or linguistically close gave more similar answers, with Germany and Switzerland almost indistinguishable.
Read the method again though. Participants associated emotion concepts with color terms. The stimulus was the word. Nobody in that study was shown #0000FF and asked how they felt about it. The finding is about the meaning of a vocabulary item, and the vocabulary item is a container. What the charts do is swap the container for one of the things inside it and keep the label.
How much room is inside a color word
Here is the ten words a chart uses, each drawn as the distance between its two furthest-apart named members. The bar is CIEDE2000, the same perceptual metric that the rest of this site runs on, where 2.3 units is roughly the point at which an ordinary eye starts to see a difference at all.
Blue is the widest chromatic word at 90.15. Its two extremes are Navy, #000080, and Aquamarine, #7FFFD4, which is to say that the word carrying the chart's most confident claim, trust and competence and calm, stretches from something close to black all the way to a pale mint. Brown runs 86.26 from Black bean to Beige. Purple runs 79.98. Gray runs the full 100, because gray as a word legitimately includes both black and white.
Divide any of those by the just noticeable difference and you get a number of perceptual steps, not a color. Ninety units is roughly forty visible steps end to end. The chart prints one square.
The swatch is one point in a cloud
You could argue that the chart's swatch is the representative case, the prototype blue that people picture when they read the word. Measure that and it does not hold either. For each word I took the swatch a chart typically prints, the basic CSS color, and measured its distance to every named member of its own family.
Blue is the worst offender. The pure #0000FF that the chart prints sits 33.7 CIEDE2000 units from the average named blue, and not one of the other 62 named blues is close enough to it to be mistaken for it. The same is true of green, purple, brown and teal: the square on the chart is perceptually alone in its own category, with no named color of that word sitting next to it. Brown is nearly as bad as blue on distance, at 30.7.
The swatch is not the middle of the word, in other words. In blue's case it is out near an edge, because pure #0000FF is an extremely saturated, extremely dark blue that almost nothing in the real vocabulary of blues resembles. Only red comes out looking reasonable, with Candy apple red and Scarlet sitting within a hair of #FF0000, which is probably why red is the row people find most convincing.
This is the same structural problem as the color wheel, where twelve spokes that are drawn 30 degrees apart turn out to sit between 2 and 85 degrees apart. A tidy diagram gets drawn once, gets copied a few million times, and nobody measures the thing it claims to depict.
The ten words are nowhere near the same size
The chart gives every word one row. That format quietly asserts that the words are comparable units. They are not. I sampled 16,000 sRGB codes at random, assigned every sampled code to whichever of the 701 named colors it was perceptually closest to, and counted how much of the cube each word ended up owning.
Green owns 26.2 percent of all sRGB codes. Orange owns 4.3 percent. Green and purple between them account for very nearly half the cube, while the entire warm end of the chart, red and orange and yellow and pink together, comes to just under a quarter. One row each.
Raw code counts overstate the dark end, because sRGB spends a lot of its 24 bits on differences too small to see. Correcting for that, by weighting every sample by how many other codes are perceptually identical to it, the ordering shifts but the inequality survives: purple and green still hold the largest shares of genuinely distinguishable color, at 18.5 and 17.4 percent, and orange holds the smallest at 4.2 percent. Against the roughly 380,000 distinguishable colors sRGB contains, blue's 12.3 percent share works out at somewhere around 46,000 distinguishable blues.
That is the number I would put on the chart if I were allowed one revision. Blue builds trust is a claim about 46,000 different colors.
The words overlap, which is the part that breaks it
Wide categories would be survivable if they were at least separate. They are not. Of the 591 unique hex codes in the corpus, 101, which is 17.1 percent, appear under two or more different color words. Nine appear under three.
The overlap is not confined to shared names. Take every named color and find its nearest neighbor anywhere in the corpus, excluding exact duplicates. For 192 of the 701, 27.4 percent, that nearest neighbor belongs to a different color word. The colors nearest to it in the world are not its own kind.
- Teal is the worst, at 59.3 percent. Three out of five named teals live closer to a green or a blue than to another teal.
- Brown is next at 55.4 percent, which matches what the shades of brown measurement found: brown is barely a hue at all, it is mostly a lightness and a low chroma.
- Gray is the only clean category, at 3.9 percent. It is clean because it is defined by the absence of chroma rather than by a region of hue, which is a different kind of category altogether.
And 48 cross-word pairs sit below the 2.3 threshold entirely, meaning the two colors are not reliably distinguishable by eye but are filed under different emotions. Amber the yellow, #FFBF00, and Mango the orange, #FDBE02, are 0.36 apart. Orange the orange, #FFA500, and Chrome yellow the yellow, #FFA700, are 0.71 apart. Maroon the brown and Barn red the red are 1.01 apart. Every one of those pairs straddles a row boundary on a chart that assigns them different meanings.
A category system where a quarter of the members sit closer to the neighboring category than to their own is not measuring a property of color. It is measuring a property of English.
What the 30 nations were actually agreeing about
Here is where the measurement turns constructive rather than destructive, because the Jonauskaite finding is robust and it deserves an explanation rather than a debunking.
Every nation agreed black and red are the most emotional colors and brown is the least. Notice what that ranking is not: it is not a hue ranking. Black has no hue. Red and brown have hues about 30 degrees apart, which is closer together than red is to orange. If emotional loading were a property of hue, red and brown would score similarly and black would not score at all.
Line the same three up against the measured properties instead and the pattern is immediate. Black is the floor of the lightness axis, L* 0. Red carries some of the highest chroma in the whole vocabulary. Brown, as the shades of brown piece measured, uses only 57.6 percent of the chroma sRGB can reach at its own hue and lightness, the lowest utilization of any chromatic family, next to orange at 90.1 and red at 85.3. Brown is what you get when you take a warm hue and drain it.
Our reading, and it is a reading rather than a result: what survives across 30 nations is not hue meaning, it is intensity meaning. Extreme colors carry emotional weight and washed-out colors do not, and the specific emotion attached to a specific hue is the part that turns out to be local, which is exactly the part the charts present as universal. Black means fear in every nation tested. Whether purple means mourning depends on whether you are in Greece.
That also explains why the chart format feels persuasive. It sorts by hue, which is the axis that carries the local, culturally-loaded content, and it holds intensity roughly constant by printing ten similarly saturated squares. It is organized along precisely the axis where the evidence is weakest.
The famous results that did not replicate
The other thing worth saying plainly is that the specific, actionable claims in this field have a poor track record. Three worth knowing:
- Baker-Miller pink. Alexander Schauss reported in the late 1970s that a particular pink reduced grip strength and aggression, and US correctional facilities painted holding cells with it. In 1988 Gilliam and Unruh ran the experiment again with 54 men and found no reduction in grip strength or heart rate. No subsequent work has supported the original result. The cells got repainted.
- Red and attractiveness. The 2018 meta-analysis by Lehmann, Elliot and Calin-Jageman pooled the literature and found d = 0.26 for men rating women across 2,961 participants, and d = 0.13 for women rating men across 2,739. Both small, both heterogeneous, and with evidence of upward publication bias in at least one direction. The authors' own summary was that the real effect could be very small and possibly nonexistent in real-world conditions.
- Red and test performance. Four replication attempts published in Collabra: Psychology in 2020, including two direct replications, failed to find the red effect on verbal reasoning scores that the original 2009 work reported.
None of that makes the whole field worthless. The cross-cultural association work is a different kind of claim, better powered, and it held up. But it does mean that the confident instruction on the chart, paint the call-to-action red and conversions rise, is not resting on anything as solid as the chart's tone implies. It is the same pattern we found looking at whether brain games work: a real and modest effect at the bottom, and an enormous amount of confident extrapolation stacked on top of it.
What an honest chart would look like
If you want to keep using color psychology, and there are reasonable design reasons to, here is what the measurement suggests you should change about how you read it.
- Treat the swatch as decoration. It is one sample from a region containing tens of thousands of distinguishable colors, and in blue's case it is not even a typical one. No study tested it.
- Trust the intensity claims over the hue claims. Dark and saturated reads as emotionally loaded in every nation tested. Which specific feeling attaches to which specific hue is the part that moved between Greece and Nigeria.
- Distrust any row near a boundary. Teal, brown and pink sit in regions where the majority of members are perceptually closer to the neighboring word. If your brand color is a teal, the chart row you are reading is not describing your color.
- Check the number, not the name. Two colors 0.36 units apart get different rows on the chart because one got called Amber and one got called Mango. Read the hex instead.
There is a nice demonstration of the underlying point available to anybody. Picture the blue from a color psychology chart, the trustworthy one. Now reproduce it. Most people, asked to match a color they saw thirty seconds ago, land tens of CIEDE2000 units away, which is a distance that spans several rows of the chart. That is the whole argument in one action: if the boundary between two emotional categories is smaller than your own reproduction error, the boundary is not doing the work the chart says it is doing.
You can test your own error on the color memory game, which shows you a color, takes it away, and scores how close you get in the same CIEDE2000 units used throughout this piece. Most first-time players score somewhere in the teens or twenties. Look back at the blue row above, the one spanning 90.15 units from navy to aquamarine, and ask how much of that range your own eye could reliably pin down.
Color psychology is not nothing. Across 30 nations and 22 languages, people really do agree about what color words feel like, and that is a finding worth having. It is just a finding about words. The chart is what happens when somebody needs a picture.