Color as Evidence: Encoding Certainty Without Miscalibrating Trust
In an analytic visualization, color has to encode an evidence state, not a mood. That one rule is a corner of a larger obligation: every design choice is an argument, and the honest target is a reader whose confidence matches how reliable the data actually is.
Color the same trust map two ways and watch what happens. The first version picks a palette that feels calm and premium, a gradient chosen because it matches the deck and photographs well. The second colors every node by one thing only: whether the claim it marks is verified, contested, or unbacked. The two look almost identical at a glance. But only the second one can be acted on, because only the second one tells you which of the worrying nodes is a confirmed problem and which is a hunch. The first is a decoration wearing the costume of analysis, and a skeptical technical reader smells that in about a second.
That gap is the whole subject. In an analytic visualization, color has to carry an evidence claim, or it shouldn’t be there. The narrow version of the rule is easy to state: if a color doesn’t encode an evidence state, it doesn’t go in. The wider version is the one that actually governs the desk’s work, and it is less comfortable. Every choice in a visualization is an argument, and most of them are invisible. The axis range argues. The baseline argues. What gets left out argues loudest of all, because nobody sees an omission. A chart borrows the authority of looking neutral while it makes a case, and the case is not neutral. So the design carries an obligation the palette hides: not to persuade as hard as possible, but to leave the reader with confidence that tracks how reliable the data really is.
Color is already a formal claim
The instinct to treat color as mood runs against a century of work saying it is a variable with rules. Jacques Bertin laid the groundwork in Sémiologie graphique in 1967, cataloguing the handful of retinal variables a mark can vary along (position, size, value, texture, color, orientation, shape) and showing that each has fixed perceptual properties (Bertin, 1967/2011). Some variables are selective: your eye can instantly pick out every instance of one value, which is why you can find all the red dots without scanning. Some are ordered: light-to-dark reads as low-to-high without a legend. Color hue is strongly selective and weakly ordered, so it is the right tool for marking categories and the wrong tool for showing magnitude. That is not taste. It is a property of the visual system, and it means the moment you assign a hue you have made a claim about what kind of thing you are encoding.
Cleveland and McGill sharpened this from theory into a measured ranking. Across a set of controlled experiments they ordered the elementary perceptual tasks by how accurately people read them: position along a common scale is read most precisely, then position on non-aligned scales, then length, angle, area, and near the bottom, color saturation and density (Cleveland & McGill, 1984). The practical reading is blunt. If you encode a number in color, you have chosen one of the least accurate channels available, so color should almost never carry the value a reader needs to read off precisely. It should carry the categorical judgment: this belongs to the verified pile, that one to the contested pile. Encoding choice is a statement about how well a value can be read, and pretending otherwise is the first place rigor leaks.
So the rule on this desk is strict, and it is the same rule Bertin’s selectivity implies. Across every artifact, color maps to one thing: the evidence state of what it marks. Bright means verified, with direct public evidence you could go check. Muted means unverified, plausible or asserted but not yet backed by the record. Violet-pink means contested, where the evidence points both ways or different audiences read it in opposite directions. Grey means noise, present but not load-bearing. And there is a legend every time. The legend is a contract. It is the promise that the color means exactly what it says and that you can hold the map to it.
If the color doesn’t carry a claim, it’s decoration, and decoration is where trust goes to die.
Every choice is an argument, and most are invisible
Color is the sharpest case of the wider problem, not the only one. Edward Tufte named the general failure four decades ago. His lie factor is a literal ratio: the size of the effect shown in the graphic divided by the size of the effect in the data, and a graphic with integrity holds it near one (Tufte, 1983/2001). A truncated bar axis inflates a two-percent difference into a doubling. A non-zero baseline turns a wobble into a cliff. Tufte’s point was that a graphic is something that can deceive, so it has to be built to tell the truth, and the deception usually hides in choices that look like formatting rather than argument.
The research tradition that followed made the argument explicit and named it rhetoric. Jessica Hullman and Nick Diakopoulos studied how narrative visualizations steer a reader and found that the steering is built into ordinary design decisions: what gets annotated, what gets omitted, how things are framed and encoded, all of it prioritizes particular readings over others (Hullman & Diakopoulos, 2011). They treat these as rhetorical techniques because that is what they are. A visualization is an argument made in a medium that looks like a neutral window onto data, and the look of neutrality is itself part of the persuasion. This is the academic version of the desk’s rule: the choices are arguments whether or not the designer meant them to be, so the only honest move is to own them.
The clearest case we have is our own. In the Bitcoin four-year-cycle piece, the same price history plotted on a log axis and a linear axis tells two different stories: one a steady climb, the other a sequence of violent booms (GoodGlyph Studio, Lab #15). Neither axis is a distortion. Both are legitimate. But choosing between them is an evidence decision that changes what the reader concludes, and the responsible version of that chart defends the choice in the open rather than picking the flattering one and hoping nobody asks. Once you accept that the axis is an argument, you cannot go back to treating the palette as innocent.
What color decides before the reader decides to look
There is a reason color deserves the harshest scrutiny of all the choices, and it sits below deliberate reading. Christopher Healey and James Enns synthesize the perception research on what the visual system processes preattentively, in the first fraction of a second, before attention is consciously directed (Healey & Enns, 2012). Certain color and luminance contrasts pop out on their own: a single saturated node in a field of muted ones is seen before the reader has decided to look at anything. That is exactly why a mood palette is not harmless. If the brightest thing on the map is bright because it looked good in the layout, the design has aimed the reader’s involuntary attention at the wrong node, and it did so before any conscious judgment could correct it. Color placement is an argument about what deserves notice, and preattention means the reader loses that argument before they know it was made.
Cartography solved the constructive half of this and packaged it. Mark Harrower and Cynthia Brewer built ColorBrewer around a simple, testable distinction: sequential schemes for ordered data, diverging schemes for data with a meaningful midpoint, qualitative schemes for categories, each matched to the structure of what it shows (Harrower & Brewer, 2003). Their tool turned “which colors” from a matter of preference into a decision you can get right or wrong against the data type, and it bakes in perceptual constraints like colorblind-safe ramps. The lesson for a trust map is direct. An evidence state is categorical, so it needs a qualitative palette where the hues are distinguishable and none of them implies more or less by accident. Use a sequential ramp for categories and you have quietly claimed an order that the evidence does not have. The palette is a perceptual instrument, and instruments have correct uses.
The honest target is calibrated trust, not maximum persuasion
Naming the choices as arguments raises the obvious question: arguments toward what? The answer that makes the desk’s work coherent comes from a strand of recent visualization research on trust, and it is worth stating precisely because the intuitive goal is wrong. The intuitive goal is to make the reader believe you. The right goal is calibrated trust: confidence that tracks the actual reliability of what is shown, neither blind acceptance nor blanket dismissal. Hamza Elhamdadi and colleagues define it in exactly these terms and argue it is the thing visual data communication should be measuring and designing for (Elhamdadi et al., 2022). A visualization that makes a reader more confident than the data warrants has failed even if it persuaded, because it moved trust in the wrong direction.
That reframes what a good design does. Lace Padilla and colleagues give the cognitive scaffolding: a cross-disciplinary model in which the visual encoding drives the reader’s decision process directly, so a design choice is not cosmetic but an input to what the reader ends up doing (Padilla et al., 2018). If encoding shapes cognition and cognition drives the decision, then coloring a shaky estimate in confident bright feeds the reader’s decision a signal the evidence cannot back. The color has stopped describing the evidence and started driving it. The evidence-state palette exists to keep the encoding honest at the exact point where it becomes a decision.
The uncomfortable finding in this literature is that readers cannot do this correction themselves. Oen McKinley, Saugat Pandey, and Alvitta Ottley, taking the viewer’s side of the problem, find that viewers’ own trust is under-studied and that most trust research stays theoretical and never reaches the people building charts (McKinley et al., 2025). Readers arrive without the tools to calibrate their own trust in a graphic, which means the calibration burden falls on the design. The chart has to do the work of telling the reader how much to believe it, because the reader has no reliable way to work that out alone. That is the whole argument for a legend that states an evidence state: it is the design taking on a job the viewer cannot.
The failure mode: manufactured confidence
If the target is calibrated trust, the characteristic way analytic charts fail is by manufacturing confidence the evidence does not support, and the mechanism is usually omission. Jessica Hullman documented why practitioners leave uncertainty out, and the reasons are structural rather than lazy: it takes effort, it is perceived as complex, and authors fear that showing doubt will undermine the message they are trying to land (Hullman, 2020). The result is a default toward false confidence. The uncertain estimate gets the same crisp treatment as the solid one, the error bars come off because they muddy the story, and the reader is handed a picture more certain than the underlying data. This is the exact failure the evidence-state rule is built against, seen from the inside of why it keeps happening.
The constructive answer is that color can be built to carry uncertainty rather than hide it. Michael Correll, Dominik Moritz, and Jeffrey Heer designed value-suppressing uncertainty palettes, where the more uncertain a value is, the more its color collapses toward an indistinguishable middle (Correll, Moritz & Heer, 2018). A reader literally cannot over-read a shaky estimate, because the palette refuses to give it a confident hue. This is the desk’s thesis made operational in the pixels: color carrying an evidence state directly, so that low evidence looks like low evidence and the design forecloses the false confidence rather than trusting the reader to supply the caveat. The muted end of the evidence palette does the same job by a cruder route.
Here is where the rule stops being tidy, and it should. Showing uncertainty is not automatically more honest. Varun Srivastava and colleagues, testing uncertainty visualization on thematic maps, found that adding it shifts trust in ways that depend on how it is shown, and not always downward toward appropriate caution (Srivastava et al., 2026). Depending on the encoding, surfacing uncertainty can leave a reader more trusting, or confused, or falsely reassured that the map is being rigorous simply because it looks complicated. So “make the uncertainty legible” is a lever, not a guarantee. The survey literature on uncertainty visualization catalogs how many representations exist and how differently they land, precisely because none of them is a free honesty win (Padilla, Kay & Hullman, 2021). The evidence-state palette is a bet that a small, legible set of categories calibrates better than a rich uncertainty encoding a reader can misread, and it is a bet the desk holds provisionally, not a solved problem. That is the real limit of the method, and pretending it away would be its own kind of false confidence.
The same confusion map colored two ways, side by side. Left: a pleasant, meaningless gradient chosen for mood. Right: the identical nodes and layout colored by evidence state (bright/verified, muted/unverified, violet-pink/contested, grey/noise), with the legend shown. Only the right one can be acted on, and the caption states exactly what each color claims.
The standard
All of it collapses into one test for anything that leaves this desk, and it is the line worth ending on.
If you can’t write the legend, you haven’t earned the color.
The moment you can state exactly what each color claims and stand behind it, you have a research artifact. Until then, however good it looks, you have a decoration. The same test extends to every other choice once you take the argument seriously: if you cannot defend the axis, the baseline, the scale, and the omission the way you defend the palette, those are decorations too, and they are lying more quietly. The discipline is the credibility. A skeptical reader with money on the line is going to ask why the bright node is bright, and on this desk there is an answer, written down, that survives the question.