The standing gate. This is the one that can come back green.
Part of the Konjo design language.
Why there are two tests
For a while there was one test, and it asked two questions at once: is the design system complete? and can an agent build an entire screen with no human at all? The second can never be yes — some of a screen is a business decision the founder owns. So a perfect design layer still scored "no", three times running, and the signal was useless as a gate.
They are now separate, and only one of them is a gate.
| Design-completeness test | Blind test | |
|---|---|---|
| Asks | Given the product facts, was every design decision derivable? | What does this system not cover at all? |
| The agent gets | The skill, the router, and product-facts.md | The skill and the router only |
| A failure is | Inventing a design value, using an off-token number, hitting a contradiction | Every question it wanted to ask |
| Can reach green | Yes. This is the gate. | No, by construction |
| Runs | Before a chapter is called done | Periodically, to find holes |
The blind test's "no" is the expected result, not an alarm. It is a smoke detector: you run it to find what is missing, and it earns its keep by going off. It found four error borders invisible in dark mode and eight of nineteen belt ranks invisible on a dark card, neither of which a passing test would ever have surfaced. Keep running it. Do not gate on it.
The protocol
Give a fresh agent, with no other context:
.claude/skills/design/SKILL.mdand permission to follow its router wherever it leads.packages/design-tokens/tokens.tsand.claude/skills/design/references/.docs/design/product-facts.md— this is what makes the test passable.- Images under
docs/design/reference/if it wants them. - One screen to specify, described in a few plain sentences.
Forbid clarifying questions. Ask for the spec, then for three lists.
What it reports
DESIGN DEFECTS — every place it had to invent a design value: a size with no token, a spacing it chose, a state nothing specified, a colour it picked, a control anatomy it made up. Each one is a gap in the system and each one is fixable.
PRODUCT BLOCKERS — every place it needed a business fact. These are not defects. There are three cases, and the agent must say which:
- The fact is answered in
product-facts.md. Then it is not a blocker at all — it should have been read and used, and reporting it as a blocker means the routing failed. That is a design defect. - The fact is in the catalog's "Not yet decided" table. Then it is a blocker, and naming it is exactly right. The catalog has done its job: the gap surfaced at build time instead of being invented. Say which row, and say what you did instead of guessing.
- The fact is not in the file at all. Correct behaviour: name it and stop. It becomes a new catalog row.
This said "two cases" for a while, and the two it named collapsed the middle one into the first — which made every correctly-named open fact read as a routing failure. The catalog itself has always had three states.
CONTRADICTIONS — two documents disagreeing, or one disagreeing with itself or with the token source. Always a design defect.
The pass condition
Green: zero design defects and zero contradictions. Product blockers may be any number, as long as each is one of the three cases above — answered and used, catalogued as open and named, or genuinely new and named rather than invented.
That is the honest ceiling for a design system, and it is reachable. It says: an agent handed this system and the business facts builds the screen correctly, and the only thing that stops it is a decision that was always the founder's to make.
Running it
- Vary the screen. A form, a detail screen, a Studio table, a tab home. A system tuned on one shape passes on that shape and nothing else.
- Prefer the screens with history. The promotion form has been specified three times and is the sharpest instrument in the set.
- Fix, then re-run. A single pass tells you what is missing; a second tells you whether the fix unbalanced something else. Both earlier rounds found defects that only appeared after the previous round's repairs.
- Record the result in the calibration record with the date and the screen, so the trend is visible rather than remembered.
What a green result does not mean
It does not mean the screen is good. It means nothing was invented. A screen can be fully derivable from the system and still be the wrong screen for the person using it — that judgement is the checklist at the end of the skill and the bar it closes with: the person you designed it for gets what they came for on the first screen, without reading a paragraph.
Mechanical completeness is the floor. It is not the ceiling.