AI workflows for digital product design
The case, the operating model, and the playbook for designing in code with AI agents: from brand file to shipped PR without a handoff layer. Summarized from the full playbook in 13 slides.
Three messages: the translation layer is the cost, agents removed the barrier, and judgment is now the differentiator
Design in Figma, spec for engineering, rebuild from scratch: every stage loses fidelity, and two sources of truth drift permanently. Designing in code eliminates the translation layer entirely. The screen the designer builds is the screen engineering reviews and ships.
With an agent harness (Claude Code, Codex, Hermes), a designer goes from idea to live URL in an hour and a complete flow in a day. No HTML/CSS prerequisite: the designer directs in design language; the agent implements against a tokenized system.
Generating five directions costs minutes. Knowing which is right is the designer's moat. The workflow instruments judgment: audits, critique loops, decision logs, and quality gates keep taste, not typing speed, as the limiting factor.
The traditional pipeline rebuilds every screen twice and keeps two sources of truth in permanent drift
Every stage of the traditional pipeline is a translation, and every translation loses something. In the playbook there is nothing to translate: the artifact the designer builds is the artifact that ships.
Five shifts define the model: the codebase becomes the design file, and handoff becomes continuous
| Dimension | Traditional | This model |
|---|---|---|
| Artifact | Figma mockup that must be rebuilt | Running screens in a browser; the prototype is the product |
| Source of truth | Figma file and codebase, drifting apart | One codebase: tokens, components, decisions, screens in git |
| Handoff | A moment; designer involvement drops off | Continuous; designer submits PRs, engineering reviews and merges |
| Design system | Library in a tool engineering cannot import | tokens.css + shadcn components + skill files; imported directly |
| States & edge cases | Discovered during QA | Designed at build time, verified by the agent in a browser |
The workflow runs in three temporal modes: set up once, build in a loop, maintain weekly
- Brand input & skill files
- Instruction file (CLAUDE.md / AGENTS.md)
- Design system scaffold
- Token export for engineering
- Accessibility audit on tokens
- Build flows end-to-end, all states
- Canvas & prototype review views
- Extract & enforce patterns
- Log decisions as they happen
- Handoff via PR
- Git sync & system branch merge
- A11y, consistency & performance audits
- Token export refresh
- Changelog generated from the diff
- Scheduled design critique
Four tools run the entire workflow, and the most important one is deliberately replaceable
- Holds project context, builds everything
- Terminal CLIs preferred; desktop apps improving fast
- Swappable: context lives in files, not the vendor
- Every branch push auto-deploys a preview URL
- The shareable artifact for testing & review
- Hosts the living design system site
- Single source of truth for design and engineering
- PRs as the handoff mechanism
- CODEOWNERS as the review gate
- Exploration outside the system; disposable
- Lanes: v0 for React, Lovable full-stack, Figma Make in-Figma
- Explore there, build here
Output quality is a context problem: what the agent has loaded matters more than how you phrase the request
One markdown file per domain (brand, a11y, copy, performance, components): the rules, the exceptions, and what good looks like. Plain markdown, readable by any agent, versioned in git.
CLAUDE.md / AGENTS.md: stack rules, folder structure, absolute constraints ("never hardcode values, never push to main"), and current status. Read automatically at every session start.
The agent responds to visual, spatial, felt vocabulary: contrast, density, hierarchy, tone. Interaction patterns are complete instructions: "apply progressive disclosure" needs no elaboration.
Seventeen steps take a team from brand file to production rhythm
- 1. Prepare brand input
- 2. Scaffold the system (shadcn + Tailwind v4)
- 3. Export DTCG tokens for engineering
- 4. Accessibility audit (WCAG 2.2, token level first)
- 5. Publish the living design system site
- 6. Build major flows, all states
- 7. Canvas & prototype views
- 8. Localization as content tokens
- 9. Extract & enforce patterns
- 10. Decision log & changelog
- 11. Content & copy audits
- 12. App-specific skills
- 13. Performance budget
- 14. Handoff via PR, not spec
- 15. Design on real production code
- 16. Deploy & environments
- 17. Git sync & weekly design rhythm
Quality is enforced by layered gates, and the agent verifies its own work in a browser before you see it
- WCAG 2.2 AA minimum; contrast fixed at the token level
- Performance: <200kb JS, LCP <2.5s, CLS <0.1; breaches block the weekly merge
- 44px touch targets; reduced-motion on every animation
- Blocking: shared components, new dependencies, auth/payments/data, API shape
- Awareness: new screens on existing components, copy, token-scale adjustments
- Designer merges: design-repo tooling, skill files, decision log, docs
Portability is the operating model: any designer works from any harness and resumes in any other, with no seams
The cross-tool instruction standard. Write project rules once; every harness reads them. Claude Code bridges via a one-line import.
Vendor-neutral tool connections (Figma, shadcn registries, browser, analytics) supported across Claude, OpenAI, and Google. Config travels in the repo.
The W3C community token format, stable since 2025. Consumed by Style Dictionary, Tokens Studio, and Figma variables; legible to any agent.
All accumulated expertise is plain text in git. Works pasted into any model or read by any CLI. No export, no lock-in.
The system scales from solo to a five-person team because the documentation is the onboarding
| Team size | Governance model |
|---|---|
| Solo / 1–2 | One owner. Second designer contributes to skill files via PR. Instruction file doubles as the onboarding doc. |
| 3 | Ownership split by domain: brand + tokens, components, governance. One reviewer on shared files. |
| 4–5 | CODEOWNERS on /skills and system files; monthly consistency audit; quarterly ownership rotation. |
- Instruction file stable across sprints
- Skill files produce consistent output without correction
- Decision log has real entries, not templates
- At least one full flow shipped through PR review
The same foundations unlock the frontier: parallel agents, measured iteration, and agentic products
Build one flow while a second agent audits, documents, or localizes. Defined specialists, git worktrees for isolation, and agent teams: a lead session that delegates to independent teammate instances, each with its own context, for parallel critique.
The autoresearch pattern applied to flows: agents propose variants inside design system constraints, previews serve them, analytics scores them. ~700 experiments found ~20 real improvements in the reference run.
A separate competency clients now ask for: generative UI assembled at runtime, plus agentic UX patterns (planning visibility, tool disclosure, streaming states, recovery, memory surfacing).
Start Monday: five moves take you from zero to the loop
Write brand.md. Convert your brand guidelines to one plain-text file: colors, type scale, iconography, motion, voice.
Scaffold the system. One prompt: shadcn + Tailwind v4 themed from brand.md, with a living style guide at /design-system.
Write the instruction file and skills. Absolute rules, folder structure, a11y and copy standards. This is the onboarding doc for every future agent and designer.
Build one flow end-to-end. Every state, agent-verified in the browser, deployed to a preview URL, shared as a link.
Open the first PR. Let engineering review a diff instead of a spec. Then set the weekly rhythm: sync, audit, merge, changelog.


