v2.0.3 · local-first · offline runtime with SHA-256 evidence

Design engineering,
not style recommendation.

An Agent Skill, MCP server, and deterministic quality gate that checks whether an interface supports real tasks, exposes coherent states, uses maintainable component boundaries, remains accessible, and produces verifiable evidence.

Explore Quality Gates
scoped MCP tools
20scoped MCP tools
workflow stages
9workflow stages
evidence tiers
4evidence tiers
network calls at runtime
0network calls at runtime
npm install -g @ztothez/design-engineering@2.0.3zz-design --version
The nine-stage workflow

One consolidated decision, nine bounded stages

Each stage links product task, truthful state, information hierarchy, interaction states, visual direction, semantic tokens, implementation, automated evidence, and attributable human review. Expand a stage to see the contract it validates.

Example profileresponsive-overview
Pass6Failure1Limitation1Unverified1

Status values above are illustrative example output from a sample profile run. They are not live results from your repository — run the gate locally to produce real evidence.

Capabilities

What actually gets checked

Scoped MCP access to architecture, Figma, design-system, UX-pattern, and usability guidance — backed by deterministic validators rather than opinion.

Contracts & briefs

Generation is blocked until intent is defined.

  • Versioned product design briefs that block generation when users, tasks, data behavior, recovery, assumptions, or acceptance evidence are materially undefined.
  • Design-deliverable manifests validating visual direction, semantic tokens, typography, composition, density, state, motion, charts, contrast, provenance, icons, and presentations.
  • Interface-trust contracts for data mode, connection, result origin, freshness, fallback disclosure, and provenance-preserving records.
  • Operational information-design contracts for decision metrics, evidence-backed findings, chart purpose, hierarchy, exceptional states, and scalable collections.

Repository audits

Static analysis against deterministic anti-slop rules.

  • Coupling and component-size thresholds.
  • Raw design values used instead of semantic tokens.
  • Mock production paths and placeholder interactions.
  • Missing network states and missing accessible names.

Browser verification

Evidence captured from a real Chromium session.

  • Responsive layout, clipping, overlap, focus order, contrast, and target sizing.
  • Keyboard behavior, text resizing, reflow, reduced motion, and media handling.
  • Console errors and network failures recorded as artifacts.
  • V2 contracts for data-mode disclosure, chart alternatives, masked dynamic regions, and checksum screenshot regression.

Corpus & benchmarks

Scored against maintained positive and negative cases.

  • Versioned corpus with provenance, per-dimension scoring, recommendation MRR, and explicit abstention checks.
  • Executable AegisOPS, SceneStart, and Azure Optimizer benchmark contracts backed by qualified delivery pilots.
  • Local-only portfolio registry with a disposable snapshot boundary for non-destructive cross-product benchmarking.
  • Anonymous comparison methodology with claim ledger and release-readiness validation.

Offline runtime & CI

Reproducible without network access.

  • Self-contained offline runtime with approved knowledge and a serialized retrieval index.
  • SHA-256 integrity evidence across production dependencies.
  • Exact provenance and dependency inventories with active-reference isolation checks.
  • GitHub Actions verification with retained fixture evidence.
Evidence model

Four evidence types, never conflated

Automated checks, AI-assisted review, attributable human review, and representative-user study each carry different weight. The system keeps them distinct.

  1. 01

    Automated evidence

    Deterministic source, contract, browser, network, and state checks. Reproducible and machine-verifiable.

  2. 02

    AI-assisted expert evidence

    Identifies likely usability risks. It does not represent observed user behavior and never substitutes for it.

  3. 03

    Human-expert evidence

    Requires an attributable reviewer. Open severity 3 and 4 findings become acceptance-criterion candidates.

  4. 04

    Representative-user evidence

    Records observed task performance with appropriate study context. The only tier that speaks for real users.

Hard rule. AI agents must never create human attestations or present generated observations as representative-user evidence. Open severity 3 and 4 heuristic findings become acceptance-criterion candidates that require review before contract integration.

Release 2.0.3

What the published release actually delivers

Every figure below comes from the published release artifacts. It describes automated and maintainer-run evidence only — not representative-user validation or a guarantee of production quality.

v2.0.3
published release
20
scoped MCP tools
9
workflow stages
3
qualified product pilots
168
automated tests passed

Included in 2.0.3

  • Evidence-gated product design briefs.
  • Deterministic design-plan compilation.
  • Contained React and TypeScript fixture generation.
  • Bounded repair with before-and-after evidence.
  • Interaction, recovery, visual composition, theme, state, and screenshot verification.
  • Qualified delivery pilots for SceneStart, AegisOPS, and Azure Optimizer.
  • Locked holdout evaluation.
  • Packed installation, offline operation, and clean-room independence checks.
Requirements

Runs on your machine, not ours

A self-contained offline runtime with approved knowledge, a serialized retrieval index, production dependencies, and SHA-256 integrity evidence.

  • Node.js 22 or newer
  • npm
  • Chromium for browser verification
  • Linux, macOS, or Windows with a Chromium executable supported by Playwright
Next step

Install from source or from the verified release archive, register the stdio entrypoint, then run your first audit and browser verification profile.