Browse all of Kata

Design calibration — how the design language was decided

Source docs/design/2026-08-22-design-calibration.mdMarkdown

On this page

Date: 2026-08-22 · Status: decided; feeds konjo-design-language.md

Superseded in part (2026-09-29). The founder reopened every round below in the Kata interview. Rounds it changed (R1, R2, R4's implementation, R8's mascot rule, R12, R13) are marked there; where the two disagree, the interview wins. Everything else here still stands as the record of why.

Why this file exists: a design language that only records conclusions gets re-litigated the first time someone disagrees with one. This records what was studied, what was asked, what was chosen, and — most importantly — what was rejected and why. If you want to overturn something in the design language, start here.

Method

Two halves, run in parallel.

Market research. 46 apps, 315 App Store screenshots pulled at source resolution and read image by image (not from memory of the brands). Seven domains, each staffed by a researcher and an adversarial critic: athletic/fitness consumer apps, Apple-grade craft, operational software people demonstrably love, community and events, the martial-arts category baseline, craft fundamentals and the accessibility floor, and an audit of Konjo's own tree. The critic's job was to attack: verify that "users love this" claims had a real source, check that principles were derived from the screenshot rather than from brand reputation, reject anything an engineer could not build from, and flag whatever fights React Native, light/dark parity, or the 44pt floor.

Founder calibration. Twelve rounds. Each round was a rendered comparison sheet of real product screens (or, where the question was Konjo-specific, mockups built from Konjo's real token values), followed by a forced choice. Screens, not adjectives — "bold athletic" means nothing until you're looking at Nike next to Things 3.

A deliberate constraint: almost nothing was treated as fixed. The diagonal band, the no-depth law, #D62828, Bebas Neue and gold-for-achievement were all on the table. Only the name, the 19-rank belt ladder, light/dark parity and the 4+1 tab structure were held constant.

What was chosen

Round Question Chosen Rejected
1 Overall vibe Bold athletic (Nike Run Club) Things 3 quiet restraint · WHOOP dark instrument · Timepage warm editorial
2 Home theme Light-first (dark at parity) Dark-first · neither-is-hero
3 Separation Tonal fills — the no-depth law survives, the border dependency does not Hairlines only · soft elevation · gradient and glow
4 Type voice Heavy condensed caps for display System sans · editorial serif · geometric bold sans
5 Signature block Full-bleed hero — the diagonal band is retired Keep the band · one strip only · no block at all
6 Tab-home density One hero, then a little A working list · a dashboard of cards · dense operational
7 Progress form Streaks and milestones One enormous number · a ring · charts over time
8 Imagery No imagery in chrome Full-bleed photography · editorial photo + type · illustration and mascot
9 Belt progress Streak and badge grid The 19-rank ladder · progress-to-next-rank · one big number
10 Empty states Explain and offer one way in Centred void (current) · suggest from a source · never-empty
11 Voice Plain and direct Warm and encouraging · dojo-formal · coach-blunt
12 Red budget One red thing per screen Red for action / ink for structure · red as brand wash · red for urgency only
13 The promotion moment A certificate — the conferring instructor is the subject Ink takeover · badge-unlock burst · no moment at all
14 Instructor roster density None of the four — a tile grid with status glyphs in the corners 44pt rows · 64pt rows · 88pt cards · plain tap grid
15 Body text size Stay at 15pt 16pt · 17pt (iOS standard)
16 Studio Shared core, two dialects One language everywhere · two separate systems

Rounds 13–16 came out of the full research synthesis, which identified forks the first twelve never touched. Six lower-priority rounds remain unasked and are listed in the synthesis notes: card visibility, depth exceptions, social-post containers, class-start state, event cards with no cover image, and whether every number gets a sentence.

The three tensions, and how they resolve

These are the places where two answers pull against each other. The resolutions are the load-bearing part of the design language.

1. Bold athletic (R1) + heavy condensed caps (R4) vs "obviously easy to use"

Loud and legible are in tension only if loudness is unlimited. Nike Run Club is loud and trivially simple: one enormous number, three supporting stats, one chart, nothing else. The volume is spent once.

→ Loudness is a budget, and each screen gets one purchase. One loud block, one red element, one filled primary, one uppercase register. This became the organising idea of the whole document, and the four-budget check in the skill.

The specific failure it names: the current Train home spends the budget three times (three band strips, all uppercase 900, all competing), which is why nothing on it can be second-most-important.

2. Streaks and badges (R7, R9) vs premium and dignified (R8 rejected the mascot)

The founder chose game-shaped progress twice, over the belt ladder, a progress bar, and a big number. That is a signal, not a slip. But Duolingo's mechanics with Duolingo's tone would be wrong for a dojo, and R11's "plain and direct" is the antidote.

→ Duolingo's mechanics, Konjo's voice. Badges and streaks are the display; belts are the highest tier of badge, because a rank is a badge a human conferred. The 19-rank ladder survives as a deeper screen. Copy states facts rather than cheering. No mascots in progress UI, no confetti for attendance, and no loss-aversion mechanics — no "don't lose your streak!", no streak freezes. A dojo does not guilt people into attending.

I had initially written the opposite ("the belt ladder IS the milestone system") after round 7; the round-9 answer overruled it. Recording that rather than quietly rewriting history.

3. Full-bleed hero (R5) vs no imagery in chrome (R8)

Compatible, and deliberately so: the hero is a solid ink block carrying condensed type, not a photograph. It works on day one for a dojo that has never uploaded an image, which is every dojo at launch. Photography appears only where users put it.

4. All six statuses (R14) vs a count worth trusting

Round 14 is the only round where the founder rejected all four options and specified something better: "a grid with status icons in the corners of each tile. different icons mean different things (declining student, payment issue, etc.)" — which turns attendance-taking into triage. Asked which statuses earned a corner, they picked all six.

That creates a signal problem: six statuses across twelve tiles means most tiles carry a glyph, and "N need attention" stops meaning anything.

→ Tier, don't trim. Safety flag, waiver, declining attendance and payment past due are the attention tier and feed the count. Test-ready and birthday are good news — they show their glyph, in gold and neutral, and never inflate the number. All six stay visible; the count stays worth trusting.

Two defects found by building it rather than describing it: a tinted status chip is 1.14:1 on a white tile and vanishes, so glyphs are bare; and a white belt dot on a white tile is 1.2:1, so every belt dot now carries a hairline ring.

Where the research overruled or confirmed taste

Two of the founder's answers were made before the research landed and were then independently confirmed by it:

  • R3, tonal fills. Material 3 ships six elevation levels and then instructs designers to prefer tonal elevation, using shadow "only for elements that need more focus". WCAG 1.4.11 makes it enforceable: a boundary that identifies a component needs 3:1, which a soft drop shadow essentially never reaches and a surface delta always can. Separately, Konjo's own no-depth law was already broken 13 times in-tree — every violation a case where something had to float. Tonal fills give those cases somewhere legal to go.
  • R5, retiring the band. The skew costs (width/2)·tan 2° of vertical space computed at runtime — 6.8pt on a phone, 22.3pt at 1280pt web, which is the ~90pt of dead space under Events home. It only reads as intentional with ≥2 strips. It cannot hold data, cannot nest, and forces negative margins on every host. Rotated text also cannot use sub-pixel antialiasing, so the tilt softens every glyph.

One place the research corrected a rule rather than confirming it:

  • Uppercase. The literature says all-caps is read 10–20% slower (Tinker 1955, replicated 2019 at >13%), that dyslexic readers slow a further 13–18%, and that readers 55+ were 29% more likely to misunderstand terms set in capitals (Arbel & Toler, 2020). That reads like an argument against R4. But NN/g's write-up of the MIT AgeLab study found the opposite for a glanced single word: lowercase needed 26% more time than uppercase. Both are true, and they describe different tasks. → Uppercase is a label treatment: ≤3 words, ≤20 characters, never a sentence. And always via textTransform, because VoiceOver reads a literal "ADD" as "A. D. D.".

The spacing ramp, in full

The single most consequential number in the document, so here is the arithmetic. Measured across src/ and apps/studio/: 3,782 raw numeric spacing declarations vs 800 token-based — 83% magic numbers.

Candidate Survives Migration Verdict
Strict 8pt {0,8,16,24,32,40,48,64,80} 31% 2,739 declarations No. Kills 12, the most-used value in the app (858 uses). Every 12 becomes 8 or 16 — a ±4px change that is visible in a 15px list row. That is a rewrite of the app's rhythm, not a tokenization.
4pt {0,4,8,12,16,20,24,28,32,...} 64% 1,447 Close. But losing 2 removes the only step available for optical alignment, and Carbon and Atlassian both ship one.
Atlassian's table {0,2,4,6,8,12,16,20,24,32,40,48,64,80} 75% ~1,155 Chosen. 8 is the spine, 4/6/12/20 are legal half-steps (6 for gaps inside a component), 2 is optical correction only. The bulk of the migration is two values — 10 (420 uses) and 14 (338) — each moving 2px.
Full 2pt {0,2,4,6,8,10,12,14,16,...} 98% 96 No. 4,6,8,10,12,14,16 is a number line, not a scale. Seven near-identical choices in the 8–16 band is precisely what produced the 83% magic-number rate.

Migration risk is compounding: a 2px shift on one declaration is below the noticing threshold, but three nested containers each shifting 2px is 6px. Mitigation is a directional codemod (6→8 for gaps, 6→4 for insets; 10→8 for gaps, 10→12 for padding; 14→16 for section padding, 14→12 for inline gaps) followed by eyeballing the 40 densest screens.

And the ramp is only ~20% of the fix. packages/design-tokens/tokens.ts exposes seven semantic spacing names and zero numeric steps, so there is nothing to reach for when you need "the gap between a chip and its icon". There is also no Button, Card, Input, Text or Badge primitive, so all 128 screens re-declare the same paddings — 16+ bespoke *Card.tsx files, 132 files defining their own button style. Ramp plus primitives, or neither works. See the follow-up plan.

Evidence quality — read this before citing anything here

  • App Store screenshots are marketing, not product, and you cannot measure them. Many are tilted device mockups with headline overlays; all are downsampled. The critic pass independently re-measured a sample of the researchers' figures and found errors of 35% (an Apple Health margin), >2× (a Monarch progress bar), and 3–4× (a Flighty fill opacity) — plus several claims contradicted by the very screenshot cited for them. A keystoned mockup at ~1px per point cannot distinguish 0.5pt from 1pt, and the researchers reported precision they did not have. → No number in the design language is sourced from a reference screenshot. Every token value comes from Konjo's own tree, from a normative standard (WCAG, Apple HIG, Material), or from a computed contrast ratio. Reference measurements survive only as approximate ratios, explicitly labelled as indicative.
  • Ratings are context, not proof. Popularity is not love — the founder's own example is that Canvas is popular and widely hated. Where the research claims users love a product's experience, it cites a review or forum quote; where no such evidence was found, the researcher reported an empty result rather than inventing one.
  • No user has ever praised Konjo's design in writing. The audit looked and found nothing, in the repo or anywhere else. Recording that honestly. Everything in the design language is an argument from evidence and taste, not from user research on Konjo itself — and it should be revisited once real users have opinions.
  • Konjo's own screens were read from legal/site-img/screen-*.png, which are real shipped screenshots, not reconstructions — but at least screen-train.png is a stale build: it shows two band strips where the code renders three, plus a section and a component that no longer exist. Treat all five as indicative of the era, not of master, and re-capture them before citing them as current evidence. Rounds 2 and 5 of the calibration used reconstructions built from packages/design-tokens/tokens.ts because the real screenshots hadn't been found yet; the real ones confirmed the diagnosis rather than contradicting it.

What the adversarial pass changed

The critic teams' most useful output was not agreement. Three things changed because of them:

  1. Every measured reference value was demoted to indicative (above). This is the single biggest correction and it is why the design language contains no "Nike uses 64pt" style specs.

  2. brand.red on brand.ink was caught at 3.77:1. Red text on the ink hero — the first line of the new signature block — fails AA. Fixed by routing red-on-dark through accent.onHero #FF5252, and accent.onHero on surface.hero is 5.91:1 light. A red fill with white text was verified fine — brand.onRed on brand.red is 5.00:1.

    (This entry named accent.red until round 8 read it back. accent.red is #C42020 in light — red text on a page, not on ink — and accent.red on surface.hero is 3.21:1. An agent following the record as written would have reintroduced the exact failure the record exists to describe.)

  3. The proposed text.tertiary was caught failing on the card fill. #767676 clears 4.5:1 on white (4.54) but not on surface.muted (4.35) — and L3 makes #FAFAFA the standard card background, so secondary text sits on it constantly. Moved to #737373, the lightest grey that clears both (4.74 / 4.54).

  4. The chosen separation mechanism was backwards, and the synthesis caught it after the language had already shipped. R3 picked tonal fills; the first draft implemented that as a #FAFAFA card on a white page — 1.044:1, weaker than the 1.225:1 hairline it replaced. Deepening the fill fails differently: by #EDEDED (1.171:1) text.tertiary on it has dropped to 4.05:1. The resolution is to invert — tint the page, keep cards white — which reaches 1.119:1 at zero cost to text contrast, and is what Apple Health (1.116:1), Attio, Notion and Strong all actually ship. The founder's answer stands; the implementation of it was wrong.

  5. A fourth contrast failure nobody had checked. feedback.successText #1F8A4C on successFill #E6F4EC is 3.86:1. One critic checked that colour against white and called it "just under"; nobody checked it against the fill it is painted on. → #1A7A42.

  6. Several codebase figures were inflated 2–4× and are now corrected. <Pressable> is 954, not 2,178. There are 128 screens, not ~135. Raw spacing declarations are 3,782. DiagonalBand has three consumers, which makes retiring it far cheaper than the language implied. ReflectPill is dead code. EmptyState already exists at 24 call sites, so the plan extends it rather than adding a parallel component.

Also flagged and folded in: a single cross-platform elevation token cannot exist in React Native (iOS takes shadowColor/Offset/Radius/Opacity, Android takes elevation, and the two cannot render the same thing), which independently supports keeping page content flat. And Ladder's acid-yellow trick does not transfer — a near-white-luminance colour can carry dark text on a full-bleed fill; #D62828 cannot, so Konjo's red always takes white.

The blind test — and what it broke

Once the language, the skill and the reference set were written, a fresh agent was given only the skill, the reference images, and one screen's source, and told to redesign Events home without asking a single clarifying question — every question it wanted to ask counted as a hole. It produced a buildable spec and a verdict of "no, not without a human in the loop", with 22 questions, 9 contradictions and 11 invented values.

Everything below was fixed as a result. This is the most valuable thing that happened to the document, and it happened after the document was already "finished".

Two of the holes broke the primary formula outright:

  1. The hero had no legal text colour. The tab-home formula puts type on an ink block, and the palette had no on-ink values — so every agent would hardcode hex, breaking the law the document states three times. Fixed with text.onHero / text.onHeroMuted / accent.onHero.
  2. The hero was invisible in dark mode. brand.ink #111111 on the dark page #0E0E10 is 1.021:1 — the formula deleted its own centrepiece in one of the two themes it demands you verify. Fixed by making surface.hero invert: near-black on a light page, near-white on a dark one, ~17:1 in both. The retired band.lead token already worked this way; the new formula had quietly lost the knowledge.

And a self-inflicted one: the formula prescribed a 46/900 statement while the type scale had only hero 64 and display 34, and tokens.md says "do not invent a new number". The formula ordered agents to break the token law. Fixed by adding statement 46/900 and splitting the roles — hero is a numeral, statement is words.

Other fixes it forced: an error-state formula (there wasn't one, though ## What never ships required it); a fifth persona, because no listed persona owned the Events tab; a test for when a tab home should be an index rather than a hero screen (applied mechanically, the formula turned the Events tab into a screen with no events on it); explicit "instances vs roles" counting on the red and uppercase budgets; the small metrics an agent had to invent (icon sizes, avatar sizes, chip spec, scroll inset, skeleton fill); a note that the doc's short spacing aliases are the target and the code's long ones are current; and .png → .jpg in the skill, which would have made agents conclude the reference directory was empty.

The target image was itself a source of four contradictions — it predated the tinted-page correction, so it showed a white page with tinted tiles while the prose said the opposite. Now regenerated in both themes.

What the test praised is worth recording too, because it says what to protect: "if a sixth section wants on the screen, it goes one tap away" restructured the screen with no judgement call required; naming the specific defect in the specific screen (anti-pattern #3 counts the duplicates on Events home) made the fix mechanical; giving the reason for the ink hero ("it has to work for a dojo that has never uploaded an image") ended an argument before it started; and the pre-computed contrast ratios meant it cited rather than guessed.

The second blind test

The same protocol, a different screen (Learn home), run after the first round of fixes. It confirmed the hero work landed — "for a hero-led screen the answer would now be close to yes" — and then found that the fixes had a shape of their own: everything invested in the hero, nothing in the alternative.

The gap it named: the index-led rule existed as one clause, against a hero formula with ASCII, token assignments, an error state and two calibrated mockups. Half the tab bar was undesignable. Fixed by writing the index formula at the same fidelity and shooting tab-home-index-light/dark.jpg.

Three contrast traps with no legal escape inside the rules as written:

  1. The dark card/page pair was weaker than the light one it was defended with. #18181C on #0E0E10 is 1.089:1 against light's 1.119:1 — and with borders and shadows banned there was no recourse. Now #1C1C22, 1.137:1. Publish both themes' ratios whenever this law is restated; a separation law with one theme's arithmetic is half a law.
  2. surface.tag chips on the tinted page are 1.017:1 — invisible. The chip spec and the separation law were each correct and jointly produced a control nobody can see. Fixed by splitting border.control (a tappable edge, which WCAG 1.4.11 genuinely governs) from border.hairline on surface.card (a row divider, which it does not).
  3. accent.red as text on the page was 4.476:1 — a rounding-level AA failure in the token whose entire stated purpose is "red as text on a surface". Now #C42020 (5.25:1 page, 5.87:1 card); brand.red #D62828 stays the fill.

Two contradictions I had written in myself:

  • Ceilings vs quotas. §2 said "1, or none"; the checklist said "each spent exactly once". An index screen structurally spends zero, so the checklist declared every legal index screen defective. They are ceilings. The tab bar is chrome and exempt.
  • The target mockup violated the red budget it was meant to demonstrate — red eyebrow, red CTA, red streak dots, red tab. The dots were a genuine second role and are now ink; the tab bar is exempt by the rule above; the eyebrow and CTA are one role because they are one utterance.

Plus: label was uppercase in one file and not the other (it is sentence case; eyebrow is the only uppercase register, with a CTA inside the hero as the single exception); surface.white was documented as a synonym for surface.card and is not — in dark it is the page colour; and the proposed hero tokens had a fallback rule that routed straight into a violation the skill names by name, so they are now flagged as having no legal fallback at all.

What the two tests together say about the method: the first found what was missing, the second found what the fixes had unbalanced. Neither could have been found by re-reading. If this document is revised again, run the protocol again on a screen type it has not been tested against — Studio, a form, or a detail screen.

The third blind test — forms

Same protocol, a screen type the document had never been tested against: the incident report form. Chosen because it stresses validation, error presentation, a long scrolling layout, serious-consequence copy, and the one-filled-primary budget against Submit + Save draft + Cancel.

The headline finding: a 30:1 fidelity gap. Tab homes had 35 lines and two ASCII sketches; forms had 22 words. The tester re-read the section twice looking for the form guidance, because it did not occur to them that one line was all of it. They invented ~22 values on one screen, three of them sitting exactly where accessibility failures land.

Two values the document demanded be computed, and never published:

  1. border.control was 1.37:1 — it failed the 3:1 the language itself demands for a control boundary. Every input, chip and secondary button in the app would have been drawn with a border that does not identify a control. It is now #88888C (3.16:1 page / 3.53:1 card). This is the single worst defect any of the three tests found: the rule was right, the value contradicted it, and only building a form surfaced it.
  2. The feedback.* pairs were never published at all — the reference said "unchanged except success" and gave no numbers. Worse, the borders were 1.2–1.3:1 while the fills are ~1.02:1 against the page, so a banner sitting on the page was identified by nothing. All four borders darkened to clear 3:1; the full table is now published with ratios.

And one that simply did not exist: focus. There was no focus state anywhere in the design system — a WCAG 2.4.13 failure the moment anything renders on RN-web, which is the only surface CI actually walks. Now a 2px border.focus, the one sanctioned exception to the 1px border rule, matching the standard's own 2px-perimeter definition.

Contradictions it caught:

  • label 13/700 was assigned to "Buttons" while anti-pattern #13 — in the same skill — names "a 48pt button carrying 12pt type" as a defect. Follow the token, ship the anti-pattern. Resolved with a button table: label is for chips, secondary buttons and links; a 48pt primary takes rowTitle 16/600.
  • Does a display screen title spend the loud-block budget? The index formula claims zero and its own mockup ships a title. Resolved: chrome is exempt — tab bar, status bar, header icons, and the screen title. Control states too, or the tile grid would be illegal by its own document.
  • The separation law inverts on a form. surface.card is both "the card" and "the input fill", so an input inside a card is white-on-white. Resolved: inputs sit on the page, and if a field group must be a card, the inputs invert to a surface.page fill.
  • SKILL.md used proposed tokens without marking them, while claiming to be "everything you need for a normal screen". Now stated inline — and as of this session most of them are shipped in tokens.ts.
  • Linear was cited approvingly for "sentence-case section labels, no uppercase" under a heading naming it the reference for the Teach tab, while the skill mandates caps eyebrows. Caption rewritten: we take Linear's monochrome-with-one-icon-colour discipline, not its label casing, and the divergence is deliberate.

Bugs it found in the shipped incident screen, worth having regardless of the test: a retry after a partial failure creates a duplicate incident row; placeholderTextColor is text.quaternary at 1.92:1; occurred_at is hardcoded to now, so an incident cannot be filed an hour late; raw Supabase error.message is shown to instructors; severity pre-defaults to "Low" on a legal record.

Verdict was still "no" for forms — and the fixes above are the answer to why. Still missing and named as the next gaps: a Select / Date / Sheet spec (real forms are 60% pickers and the system has none), and a reference/konjo/form-target.jpg, since every committed target is a tab home or a tile grid.

The completeness rounds — the trend

The protocol says to record every round here, "so the trend is visible rather than remembered". It was not kept: this file held three rounds while chapters elsewhere opened with sentences like "a completeness test on the Students roster found…", and round 8 had to reconstruct the history by grepping for that phrase. That is not a filing problem — it is the reason the same failure shape recurred three rounds running without anyone noticing it was the same shape.

Backfilled, and kept from here.

# Screen Defects Contradictions What it actually found
1 Promotion form (mobile) 16 12 The prose was close and the component layer was the gap: eight primitives the system specified and nobody had built.
2 Promotion form 13 7 The repairs moved the drift into the new primitives.
3 Promotion form 3 7 Defects collapsed; contradictions did not. Every remaining one was mechanical — a number in prose disagreeing with the token.
4 Promotion form 1 4 Three of the four were in the reference target I had drawn myself.
5 Promotion form 2 7 All mechanical. The tester called the remainder pedantry.
6 Studio Students roster 11 16 The shape changed and the score doubled. Five rounds of tuning on a phone form had produced a system that passed on a phone form. Studio's whole metric set was four adjectives; the accessibility contract was React Native with exactly one ARIA attribute in it.
7 Studio Students roster 8 18 The metrics chapter and the web contract closed their holes completely — "roughly twenty values that were invention last round and are citation now". Contradictions went up because the numbers landed: "~46pt rows" cannot be contradicted; "46, and the cell is body 15/400" can be, and was, by a shipped font-size: 14px. The hole moved from numbers to semantics.
8 Studio Home (a dashboard) 24 10 The shape changed and the score doubled again. The table work fixed tables: tables.md 158 lines, snapshot cards twenty words. Found the invisible focus ring on every ink surface in the system, three AA failures in the rail, and one genuinely new hole — see below.
9a Studio Home, after the chapter 26 13 "Round 8 found missing shapes; this round found missing seams." No invented layout numbers — the grid, span, card taxonomy, alert slot and three-state model were all citation. The product-blocker half came back clean: eight facts read from the catalog and used, two genuinely absent and named. What remained was seams: the pointer model, the web-app plumbing, and four contradictions inside dashboard.md's own ASCII sketch.
9b Studio Leads (a board) 33 8 Chosen by the rule above — the shape furthest from the last repair. A different kind of hole. Not a fidelity gap: a zero. The system has never described a direct-manipulation interaction and had no vocabulary for one, and the universal convention for a lifted object is a drop shadow, which L3 outlaws. Twelve of the 33 were the drag model.
10 Studio Leads, after the chapter 19 6 The first time a shape has been re-run and improved rather than doubled. ~25 decisions that were invention are citation, and the two that mattered most landed hardest: is this allowed to be a drag (it routed straight into an open product fact and stopped) and what does it say out loud. The product-blocker half was the healthiest in the record — six of seven read out of the catalog and used to change the design. What remained was the container: 11 of 19 defects were the column the repaired gesture happens inside. And three of its six contradictions were inside the round-9 repairs themselves.

The pattern the trend exposes, which no single round could: a repair lands in the shape that was tested and does not generalise to the class of shapes around it. Round 6 said it about mobile-vs-Studio. Round 8 said it about tables-vs-cards, in the same words, one level down — and studio/metrics.md, the chapter written to fix round 6, reproduced it exactly: it specified the table row, the table header row, the checkbox and the belt swatch, and had no row for the card row, the stat value, or the avatar.

So the rule for choosing the next screen is: pick the shape furthest from the last repair. Not the sharpest instrument — the least-covered one.

What round 8 found that no earlier round could

  1. border.focus is 1.14:1 on surface.hero, and 1.00:1 in dark — the same hex. Every filled primary button in the system and all ~25 items in Studio's rail had a focus ring nobody could see. The generated table had published it, marked ✗, three lines under a sentence saying "a ✗ is not a warning, it is a defect". Nothing failed on it. The generator now carries a table of the pairs the language states as absolutes and refuses to emit a palette that breaks one.
  2. A dashboard's emphasis is data-driven, and the budget system only knows how to be static. One loud block, one red role, one filled primary — all decided at design time, on a screen with a fixed subject. A dashboard's most important fact changes daily and loudness here is a property of a component, not of a fact. Answered in studio/dashboard.md by allocating the emphasis to a slot rather than a card, and by ruling that a quiet morning correctly has no loud element at all.
  3. Empty is not "not set up". "0 leads this month" from a dojo with no lead form is a false statement about the business. A new dojo's first week is almost entirely this state, and it is the only week where every screen is seen for the first time.

What round 9 found, running two shapes at once

Both rounds independently ranked the same finding first, which is the strongest signal the gate has produced: Studio was specified as a layout, not as a pointer product. Konjo had an exhaustive touch model — pressed opacity, hitSlop, gesture back, reduced motion — and hover appeared in exactly four places, all of them added incidentally by the tables chapter and the rail. It is invisible on a table, because a table is the one place it was written down.

  1. Three rules stated as universals only computed on their home surface. Hover (surface.muted), the skeleton fill and the row divider (border.hairline) were all tuned against surface.card and published as general rules — and on surface.page they are 1.07:1, 1.09:1 and 1.09:1. Which means there was no divider in this system that was legal on the page, and every skeleton outside a card was invisible. The checker now has SURFACE and THRESHOLD rules: a chapter prescribing a colour for a state must name the surface it is judged on, and a published pair under 3:1 must say so in words.
  2. The generator's coverage had fallen behind the prose again — the same failure round 3 found in the feedback borders. Four surfaces computed, six painted on. Nine now.
  3. The budget system assumes a fixed subject, and round 8's fix for that was written as if it were about cards. It is now a clause on L4: a screen whose subject is chosen by data spends its emphasis through a designated slot or spends none, and spending none is correct.
  4. The system had no direct-manipulation vocabulary at all — no drag, drop, reorder, keyboard equivalent, announcement or cancel, and no mention of SC 2.5.7, which is the criterion a board exists to satisfy. patterns/direct-manipulation.md is the answer, and its load-bearing law is that a drag must be reversible or it is not a drag: a gesture has no before-the-tap moment in which to state a consequence, so an irreversible move is a control with a confirm.

The trend rule held for the third time running. Rounds 6, 8 and 9b each changed shape and each roughly doubled the score. What is different about 9b is that its defects were not missing values — they were a missing axis: every pattern in the system was a state machine over taps, and a gesture is not one.

Round 10, and the first time a re-run went down

The board came back 19/6 from 33/8. Every earlier re-run of a repaired shape had held roughly level (rounds 2–5) or gone up because the numbers had landed and were newly contradictable (round 7). This one fell by nearly half, and the diagnosis says why: the repair was aimed at a class of question rather than at a screen.

Three things worth keeping from it:

  1. A repair still lands one level in. direct-manipulation.md specified the gesture exhaustively — three border states, a cursor table, a keyboard path, five announcement strings — and the container not at all. Eleven of nineteen defects were the column: no width, no gap to its neighbour, no place in the accessibility tree. studio/metrics.md did this in 6→8; dashboard.md did it in 8→9a; this is the third recurrence of one failure shape. patterns/board.md is the answer for this instance; the general answer is below.
  2. Three of six contradictions were inside the round-9 repairs, written days apart and already disagreeing — the pointer model against itself sixteen lines apart, and the two chapters that were written in the same push putting a drop outcome in assertive and polite respectively. A repair is a new surface for drift, not an end to it.
  3. The cheapest systemic fix in the whole record came out of this round, and it is a writing convention rather than a value: a rule in a shape-specific chapter must say whether it is shape-local or system-wide. "Never infinite-scroll" lives in the tables chapter and is universal; "an empty region still prints a line" lives in the dashboard chapter and deliberately overrides the detail-screen rule. Both read as universal truths in the voice they are written in, and neither said which — which is precisely the mechanism by which every one of these repairs failed to generalise. It is now in governance.

Named, and now built (studio/overlays.md) — a Studio overlays chapter — the drawer, the menu/popover, the dropdown and the transient confirmation. All four are named as shipped in the screen inventory, and the system has exactly one recipe for a floating surface (scrim plus hairline) which was written for a phone's bottom sheet and gives the wrong answer for a desktop dropdown. That is the same arithmetic-forces-a-new-answer situation the banned drop shadow created for the drag, and it deserves the same treatment.

What it turned out to need was a rule, not a token. The prediction above was that the arithmetic would force a new value, and the open question below pre-authorised one elevation token for floating surfaces if it did. It did not. border.control already clears 3:1 on all three Studio surfaces (3.53 on card, 3.15 on page, 3.38 on muted) — what was missing was anyone saying that a floating surface's boundary uses it rather than border.hairline, which is 1.22:1 on a card and cannot mark a plane change. The shadow ban stands, no token was added, and governance step one — prove the existing tokens can't do it — earned its place.

Open questions

Carried into the design language rather than silently decided:

  1. Floating surfaces. R3 was asked about page content. Bottom sheets, toasts and menus currently separate with a scrim plus a hairline — my call under the founder's law. If that proves insufficient in practice, the narrow fix is one elevation token for floating surfaces only, not a general lifting of the ban.
  2. What replaces band.*. The hero block's token names are proposed, not built.
  3. Dojo photography. Chrome carries no imagery, but event and dojo pages will carry user uploads. Scrim and fallback rules are sketched, not specified.
  4. text.quaternary. Now illegal as a text colour. Every current use needs an audit.
  5. Studio's dialect. Confirmed as deliberately different (denser, light-only, no hero blocks), but Studio was not separately calibrated with the founder. Its brief remains authoritative for IA and inventory.