Measure Design System Impact: Part 1: Quality and Alignment

Measuring Design System impact is not about adding up checks. It is about building a scorecard that connects quality, alignment, adoption, and the operating health around the system. Every metric below includes what it measures, the evidence you need, how to evaluate it, a formula, and the mistake to avoid.
Quick map: Part 1
How the metrics are classified:
Component quality covers metrics (Governance, accessibility, tokens, contract vs code and interaction tests)
Design and code alignment covers metrics (Figma alignment and documentation freshness)
Metric | Files you need | How to measure it | Measurement type |
|---|---|---|---|
| Check every required section; optionally assess the quality of the rule. | Script + optional AI review | |
Stories and axe results per theme | Run axe in every theme; review keyboard and screen-reader use. | Script + mandatory manual review | |
| Find literals, validate token references, and compare declared tokens. | Script + AI for edge cases | |
Component contract and TypeScript types | Extract and compare props, types, required status, and values. | Script | |
Tests, play functions, and documented behaviours | Run tests, then compare behaviour coverage with the contract. | Script + AI/manual coverage review | |
Figma library, mapping, and expected variables | Compare each code component with its matching Figma component: name, variants, properties, and bound variables. | Script + AI visual review | |
Git history and component file map | Compare the latest prop change with its contract, Storybook examples, and any component guide. | Script + AI confirmation |
Component quality
Governance

What it measures: Whether every component has written rules for when not to use it, why it exists, and what a product team may change on its own versus what it must request from the Design System.
Files you need: One component contract, such as
contracts/Button.jsonordocs/Button.md, with fixed sections:neverUseFor(“not for page navigation; use Link”),rationale, andgovernance. Also keep a template or schema such ascomponent.schema.jsonthat declares those sections mandatory.How it is evaluated: A script walks every contract and checks that each required section exists and is not empty. Optionally, ask an AI: “Read this component contract. Does it explain concretely when not to use the component and why it exists? Answer yes/no and why.”
Formula: Component contracts with all three sections complete ÷ total components.
Watch out for: A complete template is not automatically useful. Rules such as “use when appropriate” pass a structural check but do not help a product team or an AI make a decision. Review a sample of contracts for concrete rationale and clear boundaries.
Accessibility

What it measures: Whether components meet WCAG expectations (contrast, labels, roles, focus, keyboard support, target size, and reflow) in every theme.
Files you need:
Examples for every component, such as
Button.stories.tsx, covering every meaningful state: default, hover, focus, disabled, error, and loading.An accessibility section in each component contract, such as
contracts/Button.json, listing role, accessible name, supported keys, and targeted WCAG criteria.Output from every automated check, stored per example and theme - for instance
test-results/a11y/button--primary.json- so every failure can be traced to a story and theme.
How it is evaluated: Most of it can be automated; the contract is what makes this possible.
Static rules. Run
@storybook/test-runnerwithaxe-playwrightin every theme to catch missing labels, invalid ARIA, and low contrast.Keyboard and focus. Generate one interaction test for every key listed in the contract; after closing a dialog or menu, assert focus returns to its trigger.
Accessibility-tree checks. Compare Playwright
ariaSnapshotoutput - role, name, and state - against the contract. Guidepup can drive VoiceOver or NVDA for slower end-to-end checks.Visible focus, target size, and reflow. Check focus-ring contrast after tabbing; flag interactive elements below
24×24 px; render at320 pxwide and check for horizontal scrolling or clipped text.AI visual review. Use it as a second pair of eyes for focus, overlap, cut-off text, and disabled states - never as the source of truth.
People, where judgement is needed. Review whether labels make sense and whether the complete experience works with a real screen reader, once per new component and before major releases.
Formula: Examples that pass every automated check ÷ examples tested, counting each theme separately. Report keyboard coverage separately: keys tested and passing ÷ keys listed in contracts.
Watch out for:
Theme-specific failures. A component can pass in light mode and fail in dark mode because of contrast. Test every theme.
Treating axe as the whole job. Everything it cannot see needs either a contract-derived test or a person.
Contracts that do not list keys. Keyboard tests can only be as complete as the contract.
Tokens inside components

What it measures: Whether components use tokens for colour, spacing, size, radius, and typography rather than literal values such as
#2563ebor16px.Files you need:
A token source such as DTCG
tokens.jsonortokens.css.Component style code such as
Button.tsx,Button.css, orButton.module.css.Optionally, a
tokenssection incontracts/Button.jsonthat declares expected mappings, for example"background": "color.fill.primary"and"padding": "spacing.m".
How it is evaluated:
Find literals. Script the baseline: locate hex/rgb colours and
px/remvalues in colour, spacing, size, and radius properties.Validate references. Find usages such as
var(--ds-…)and verify that every referenced token exists intokens.json.Compare the contract. When a contract declares expected tokens, compare actual use with that declaration.
Use AI only for edge cases. Ask whether ambiguous literals such as
border: 0or a legitimate1pxborder should be a token, and which one.
Formula: Values coming from a token ÷ (values coming from a token + literal values), calculated per component and in total.
Watch out for: Do not count
0or100%as failures. Decide first which literals are allowed without a token.
Contract vs code

What it measures: Whether the documented component contract matches what code actually accepts: the same props, values, and required props.
Files you need: The contract with documented props, for example
contracts/Button.json, including name, type, required status, and allowed values; and code types such asButton.types.tsor PropTypes.How it is evaluated: Use
react-docgen-typescriptor the TypeScript compiler to extract props, then compare them with the contract: props present in code but absent from the contract; props in the contract that no longer exist; and differences in required status or values, such as a contract listingsize: s | m | lwhile code acceptsxs.Formula: Components with no differences ÷ total components.
Watch out for: Inherited props. A component that accepts all HTML attributes such as
idandtitledoes not have to list them one by one, but its contract should say that it accepts them.
Interaction tests

What it measures: Whether automated tests cover behaviour: clicking, typing, keyboard navigation, and focus placement.
Files you need: interaction tests in Storybook
playfunctions or standalone files such asButton.test.tsx; and a behaviour list from the component contract, for example “opens with Enter and Space”, “closes with Escape”, and “returns focus to the trigger when it closes”.How it is evaluated: run
vitestor@storybook/test-runnerand confirm that tests pass; count components with at least one interaction test. For real coverage, use AI or manual review to compare the behaviour list with what tests actually exercise: “Which keyboard behaviours are missing from these Modal tests?”Formula: basic: components with passing tests ÷ total. Better: tested behaviours ÷ behaviours listed in component contracts.
Watch out for: tests that only prove the component renders. They do not test interaction.
Design and code alignment
Figma alignment

What it measures: Whether each component exists in Figma and matches code: variants, properties, tokens, and variables.
Files you need: The published Figma component library; a code ↔ Figma map such as
figma-map.jsonor Figma Code Connect, defining the matching component and property translation, for example, codevariant= Figma “Type”; and expected values from code, such asfigma-expected/Button.json, which specify the Figma variable expected for each variant background.How it is evaluated: Through the Figma API, read library components; check that every code component exists; compare variants and properties using the map; check that colours and spacing are bound to variables rather than hard-coded, and that the variables are the expected ones. For visual review, give an AI a Figma capture and a Storybook capture and ask it to list differences in size, spacing, colour, and states.
Formula: Components that exist in Figma and match ÷ total components. Report “missing in Figma” separately from “exists but is out of date”.
Watch out for: Existence is not freshness. Store when the component was last reviewed.
Documentation freshness

What it measures: Whether code changed while its component contract, examples, or guide were left behind.
Files you need: Git history, which gives the last-change date for every file; and a map of code files such as
Button.tsxandButton.types.tsto documentation files such ascontracts/Button.json,Button.stories.tsx, and the Button guide section.How it is evaluated: For each component, compare the last change in its types, where props are defined with the contract and examples. If code is newer, documentation may be stale. Use AI to confirm: an older file is not automatically wrong. “This is the change in
Button.types.tsand this is its contract. Is the contract still correct?”Formula: Components whose documentation is newer than their latest prop change ÷ total components.
Watch out for: An internal refactor can be newer without affecting documentation. Focus on type or prop files, not every implementation change.
What to do first
Automate accessibility, tokens, contract checks, and documentation freshness first. They are repeatable signals. Add human review where it matters: the meaning of rules, interaction coverage, and visual similarity. Add the full set of metrics as a new Storybook story or dashboard view, so the team can see whether each metric improves or declines over time.