| name | tufte-visual-confections |
| description | How to build "confections" — single compositions that juxtapose heterogeneous, real-and-imagined image-events to make an argument; use when designing explanatory posters, mixed-media dashboards, infographic narratives, onboarding scenes, kiosk/interface screens, or title/cover art that must reason rather than decorate. |
| tags | ["tufte","data-visualization","information-design","infographics","interface-design","composition","annotation"] |
Visual Confections: Juxtapositions from the Ocean of the Streams of Story
Overview
A confection is a single picture that gathers many separate image-events — some real, some imagined — and juxtaposes them on the flat page to make a point, narrate, explain, or list. It is neither a faithful scene (map, photograph) nor data poured into a standard format (chart, table); it is concocted on purpose, "showing all at once what never has been together." The discipline matters because confections are the native form for explanatory posters, mixed-media dashboards, infographic narratives, and content-first interfaces — and they live or die on the quality of the thinking behind them, not the cleverness of the assembly.
§1. What a confection is — and is not
Tufte's model (the "Ocean of the Streams of Story," borrowed from Rushdie):
- An event = the intersection of a noun and a verb (subject + action) — something happens.
- A plane of events = one time-slice holding every noun-verb combination at that instant.
- A story = a progression of noun-verb incidents — one strand running through time.
- An image sequence (Marey, Muybridge) = a short stretch of a single strand.
- A confection = an assembly of many image-events selected from different strands, lifted out and juxtaposed on the still flatland of paper.
What a confection does with that material: illustrate an argument, enforce visual comparisons, combine the real and the imagined, and tell another story.
Confection vs. the two things people confuse it with:
| Form | Source | Relation to reality | Job |
|---|
| Map / photograph | a pre-existing scene | direct representation | depict what is there |
| Statistical chart / table / map-as-data | data in a conventional format | encode measured quantities | quantify |
| Confection | many strands, gathered | concocted universe (real + imagined) | make a point, argue, list, narrate |
"What collage is for art, confections are for the design of information." — Tufte, Visual Explanations
Do/don't — the gateway test:
- Do (confection): assemble heterogeneous images to display information — usually expressible in words, often derived from words.
- Don't call it a confection if you are reproducing one real scene, or just formatting a dataset. That is illustration or charting; different rules apply.
§2. The two organizing structures
Every confection arranges its gathered images by one of two strategies — or both at once.
| Structure | What it is | Tufte's examples |
|---|
| Imagined scene | one concocted "universe" where unrelated objects coexist in a single space | Pugin's 25 churches massed as a skyline; Dürer's Melencolia I (1514) — symbolic objects gathered; Babar's Dream; Rousseau's Liberty Inviting Artists |
| Compartments | a grid or set of call-out cells, each holding one image-event | the call-out circles of the "Ultimate Weed"; the lower grid of Leviathan |
| Both at once | imagined scene fused with a compartment grid | Hobbes' Leviathan (sovereign-body above, grid below); Scheiner's Rosa ursina (compartmentalized imagined scenes); the Traps and Pitfalls poster |
Leviathan as the model of "both" (Hobbes, 1651): top half = an imagined scene (a giant artificial sovereign built from a multitude of tiny citizens, towering over mountains and spires to imply supreme power); bottom half = a grid of compartments cross-linked by both format and content. The compartments read downwards and across — each cell sized to differ from the one below but to pair with its opposite number (temporal rule left, ecclesiastical right; castle ↔ church, coronet ↔ mitre, cannon ↔ thunderbolt). It "adds up to something more than a statement of themes and something less than an argument," reasoning by analogy, metaphor, and visual-verbal parallelism. Concrete scale: a full verbal description of its elements ran ~31,200 words — an intensity of 91 words per square cm (579 per square inch).
Multifunctioning is the mark of a good confection: in Leviathan, citizens add up to the larger body, compartments link both vertically and horizontally, and the sovereign's head is a portrait of the author — single elements doing several jobs at once.
§3. Confections as visual lists
A confection can be an inventory — a list rendered spatially.
- Pugin's churches (frontispiece, 1843): 25 churches, chapels, and schools gathered into one grand skyline, "ranged like a Gothic New Jerusalem" (Kenneth Clark). A scenic inventory.
- Scheiner's seven sunspot methods (Rosa ursina, 1630): a visual list of seven viewing techniques (darkened glass, projections, reflections) staged on a terrace of astronomers under five suns.
Lists can carry verbs, not just nouns — the "Ultimate Weed": call-out circles around a central plant each state what the weed does (spreads quickly, resists removal, triggers allergies, poisons wildlife, resists herbicides). Drawings comment on a drawing; words and images blend into a coherent account of an imagined plant. The lesson: a confection portraying acts, verbs, and consequences is richer than a static parts-list.
Do/don't for visual lists:
- Do let the spatial arrangement itself carry meaning (Rousseau's rows of near-identical painters express order in space — a queue — and order in time — a flow — simultaneously).
- Don't make the reader work to decode the list. Pugin's failure: building names sit in a legend six pages away, printed vertically while the frontispiece runs horizontally, so identifying one building means detecting a tiny number buried in engraving lines, turning pages, and rotating the book. A list whose key is unusable is a decoration, not information.
§4. Confection vs. collage
The distinction is the whole point of the chapter: same technique (cut, paste, juxtapose), opposite purpose.
| Collage | Confection |
|---|
| Domain | art | design of information |
| Goal | pleasing or provoking visual experience | display visual information |
| Words | hardly expressible in words; rarely based on words | often expressible in words; often derived from words |
| Test of success | aesthetic experience, "the commonplace made miraculous" | how deeply it illuminates ideas and their relations |
| Example | Cornell's boxes (Medici Princess) — 3-D theaters of reverie | Burton's Anatomy of Melancholy title page; the Potomac graphic |
A confection is a "miniature theater of information" — a cognitive art that illustrates an argument, makes a point, explains a task, shows how something works, lists possibilities, or narrates a story. (El Lissitzky's The Constructor, 1924, sits on the seam: a photomontage that is also self-exemplifying — it depicts the process of graphic thinking, each overlapping image acting as a verb linking the nouns mind/eye/hand/compass/grid/paper.)
§5. Annotation, labels, captions, and the instructed viewer
Images alone are under-determined; text fixes the reading.
- Confections from texts (title pages of Rosa ursina, Anatomy of Melancholy, Leviathan) can sketch out complex writing and make visible what is "textually invisible, obscure, or beyond words." Some images are extreme reductions (a single emblem of a story); others enlarge the text, adding figures, details, and settings the source never specified.
- The under-determination problem: Genesis says Cain killed Abel but does not say how — so the picture must invent what the text omits.
- The instructed viewer (Meyer Schapiro): a few pictured elements — one or two figures, a single attribute — can evoke an entire known story for a viewer who already holds the text (Noah in the ark, Daniel between lions). Connotations not present in the bare text get fixed by surrounding commentary, ritual, captions.
Named failure mode — the out-of-towner: a display that relies on instructed viewers to supply the exegesis will only mystify "those viewers from out of town." Insider knowledge is not a substitute for legible annotation.
Do/don't for annotation:
- Do integrate labels, captions, and surrounding text directly into the image (Burton's stanzas keyed to compartments; the Anatomy couplet announces its own method). Rousseau weaves words in — the lion's scroll names individual artists, giving a verbal account inside the visual one.
- Don't assume the audience carries the source text. Annotate as if for the stranger; reward, don't require, prior knowledge.
§6. Confections argue — they don't decorate
The recurring theme: a confection's value equals the value of its underlying idea.
- Confectionary titles match confectionary images: Rosa ursina sive sol ("the bear-rose, or the sun") is a verbal melange — roses + bears + sun — because roses and bears were emblems of Scheiner's patron, the Orsini (Ursinus = bear) family; the title page concocts roses, bears, and suns to match. Title and image argue the same point.
- A confection can be its own lecture: Tansey's Myth of Depth (1984) stages Pollock walking on water, Greenberg lecturing on flatness, Motherwell studying the surface — the painting is Tansey's argument about flatness vs. depth, illusion vs. reality, complete with a numbered key (1. Noland … 7. Pollock). Descriptions of confections "seem to provoke the language of miracles."
- Babar's Dream (de Brunhoff, 1933): winged-elephant virtues (each carrying its emblem — flowers of hope, candle of knowledge, saw of perseverance, clock of patience) drive out demon-vices; an imagined moral universe that argues for personal, mind-and-individual virtue rather than any corporate or nationalist cause.
"Excellence in the display of information is a lot like clear thinking." — Tufte, Visual Explanations
Closing principle: as 15th-century perspective let the mind see diverse objects in a correct spatial context, a confection places diverse images into the narrative context of a coherent argument — making reading, seeing, and thinking one act.
§7. Failure modes (named)
Tufte's explicit list of how confections fail:
| Failure mode | Description | Antidote |
|---|
| Thin content | nothing worth assembling; assembly as substitute for substance | start from an intriguing concept, not a layout |
| Flimsy logic | the juxtapositions don't actually reason; connections are decorative | make every adjacency earn its meaning (Leviathan's paired cells) |
| Poor annotating text | captions/labels too sparse, absent, or buried to anchor the images | integrate legible annotation; design for the out-of-towner (§5) |
| Heavy-handed structure | the arrangement gimmick overwhelms the content | let structure be transparent, ordinary, conventional |
| The cutter's art (tendentious selection) | evidence cherry-picked to win a debate rather than illuminate | select to illuminate ideas and relations, not to score the point |
| Out-of-towner reliance | meaning depends on insider exegesis the viewer lacks | annotate the chain of actions explicitly |
Burton's own self-aware joke about bad assembly — "Marke well: If 't be not as 't should be, / Blame the bad Cutter and not me." (Robert Burton, The Anatomy of Melancholy) — names the cutter as the point of failure. Diagrams themselves are not the problem: Tufte insists reading errors come from specific local explanatory failures or untruths, not from any inherent defect of the diagram (the Chemical Atlas CO₂-cycle confection, 1854, wraps itself in apologetic text fearing "unduly literal readings" — the fear is misplaced; fix the local annotation instead).
§8. The interface as one-time confection
A computer can sort huge stockpiles and assemble a one-time confection for an immediate, local purpose — the museum kiosk being Tufte's worked example.
Principles of the content-first ("flat") interface:
- Information becomes the interface. The opening panel shows the scope of available information immediately; only a small corner is computer administration (touch-screen hint, language options). ~90% of the screen is substance — no decorative logotypes, no navigation chrome.
- Distribute in space, not time. Surface ~45 options at once rather than sequentially unveiling little bits down a decision tree.
- Each technology does its own job. The computer selects, organizes, and customizes; paper then prints a high-resolution, portable, permanent record (map + directions + a video snapshot of the visitor) — a memory the screen cannot be.
- Borrow real-world gesture. A red pointer on the live video image, linked to a red line and footprint on the map, mimics the human gesture of "go around and down that way."
Quantitative interface measures (use these to audit any screen):
| Measure | What to count | Target |
|---|
| Content share | % of screen for content vs. administration vs. nothing | maximize content |
| Typographic density | character count vs. printed material and other interfaces | approach print density |
| Commands available | number of commands immediately offered | more is better — if clearly and minimally displayed |
Worked numbers — the "Yearbook" book-metaphor screen (From Silver to Silica, 1991): only 18% of the screen showed substance (photographers and their work); 82% was administrative debris or nothing. The spread carried 53 typographic characters where a real book spread carries 1,000–50,000. A 1990s screen already had ~5–10% the resolution of a printed map; a wasteful metaphor squanders what little it has.
Named bad-interface failure modes ("television-disease": thin substance, contempt for audience and content, short attention span, over-produced styling):
| Anti-pattern | What goes wrong |
|---|
| News-broadcast | 30-sec director video, 20-sec curator videos; the information architecture mimics the org chart of the bureaucracy that built it (same sin as magazine frames marking each sub-editor's turf) |
| "Follow standard computer practice" | a tedious binary decision tree ("YOU LIKE ART? OR NOT?"); software logic exposed; context and overview lost; junk like fake drop-shadow lighting sanctified as a "standard" |
| Interface-as-billboard | the interface itself is the conspicuous visual statement — styling that masks a data dump; content becomes trivial and incidental |
The Pioneer plaque — the eternal confection (1972/1973): on a 15 × 23 cm (6 × 9 in) gold-anodized aluminum plate aboard Pioneer 10 and 11, an intensely quantified assembly — hyperfine hydrogen transition as the base unit of time/distance, a map of 14 pulsars locating the sun, planet scales spanning atoms to galaxies, outline human figures (heights given as the binary of decimal 8 × the 21.11 cm hydrogen wavelength = 169 cm). It demonstrates the ceiling of the form: a self-contained confection legible to an unknown viewer across "the eons and light-years," carrying its own complete annotation because there is no instructed viewer to assume.
§9. Application — building your own confection
Decision order when composing an explanatory poster, mixed dashboard, onboarding scene, or cover:
- Find the idea first. Confections stand or fall on how deeply they illuminate ideas and the relations among them — name the argument before sketching the layout (§6, §7).
- Pick a structure (§2): imagined scene, compartment grid, or both. Use both when you must show a unifying whole and itemized parts (Leviathan, Traps and Pitfalls).
- Gather heterogeneous image-events — maps, photos, diagrams, drawings, real and invented — from different "strands." Mixing media is the form working, not a flaw.
- Show verbs, not just nouns (§3): depict acts and consequences, not a static parts list.
- Make elements multifunction (§2): one image earning several jobs is the signature of a good confection.
- Annotate for the stranger (§5): integrate labels/captions; never require insider knowledge.
- Audit content share (§8): keep substance dominant; strip chrome, decoration, and org-chart-shaped structure.
- Check for the cutter's art (§7): did you select to illuminate, or to win? If to win, you have propaganda, not explanation.
Reference gallery (paraphrased examples):
| Confection | Year | Structure | What it teaches |
|---|
| Burton, Anatomy of Melancholy title page | 1638 | compartments | 10 cells ↔ 10 poem stanzas; the design reproduces the book's method (cut-and-paste of quotations and paraphrases), announced in its own couplet |
| Hobbes, Leviathan title page | 1651 | both | imagined sovereign-body + cross-linked grid; multifunctioning elements; ~31,200 words to describe |
| Scheiner, Rosa ursina | 1630 | both | confectionary title matches confectionary image; seven-method visual list |
| Pugin's churches | 1843 | imagined scene | scenic inventory; cautionary tale of an unusable legend |
| de Brunhoff, Babar's Dream | 1933 | imagined scene | an argued moral universe; virtues carry emblems |
| Rousseau, Liberty Inviting Artists | 1906 | imagined scene | space = queue, repetition = time; words woven in via the lion's scroll |
| Tansey, Myth of Depth | 1984 | imagined scene | the picture is its own lecture; numbered key |
| Lissitzky, The Constructor | 1924 | overlay/montage | self-exemplifying; images as verbs linking nouns |
| Potomac River danger graphic | 1985 | both | cut-away views + sequenced actions + 20–50 words per the 9 story-pictures; beats video, which is fixed-order and low-resolution |
| National Gallery kiosk | 1990s | flat interface | information is the interface; ~90% substance, 45 options at once; computer + paper each do their job |
| Pioneer plaque | 1972/73 | compartments | fully self-annotating confection for an unknown viewer |
The Potomac lesson on medium choice: the printed confectionary diagram explains why the river is deceiving — dangers beneath a tranquil surface — using cut-away and multiple views with focused annotation (20–50 words tied to each of 9 collaborating story-pictures). It beats a TV account, which would be jumpy, low-resolution, cover under a third of the material, and impose a fixed one-dimensional order; the printed page lets the reader control order, pace, and entry point. Tufte's rule: whenever possible give the audience words and images on paper, even just to supplement speech. And to activate viewers, add fine detail they can search and edit — e.g., 57 small portraits of the drowning victims and 57 dots on the map locating each death, so readers can particularize the danger by seeing where the victims walked.