- 01Never hardcode a value that belongs to a token.
- 02The artifact the designer builds is the artifact that ships.
- 03Handoff is a pull request, not a spec.
- 04Every screen ships all of its states.
- 05Budgets block the merge: accessibility, performance, consistency.
- 06Log the decision at the moment it is made.
- 07Context lives in files, not in the vendor.
- 08Taste, not typing speed, is the limiting factor.
Why this approach
The short version of why this workflow exists and what it's better at than the traditional Figma-based process.
What you get
What this workflow is, who it's for, and what it costs.
Tools overview
Four tools. Three daily drivers. One fills a specific gap.
Agent harness
Vercel
GitHub
v0
The full progression
Three temporal modes. Setup happens once. Designing & building is a repeating loop. Maintenance runs in the background every week.
Setup
Design & Build
Maintain
Skills & instruction files
Give your agent persistent, reusable context so you stop re-explaining the same things every session. These files are the foundation everything else builds on, and they're plain markdown, so they work with any harness. One reading note: this playbook says "CLAUDE.md" throughout because Claude Code is the daily driver. Read it as shorthand for your instruction file; the same content lives in AGENTS.md for Codex, Hermes, and everything else (see Keep it portable).

Skills: modular context files your agent reads on demand
/skills folder in your project root. One .md file per domain: brand.md, a11y.md, copy.md, performance.md, decisions.md, components.md. These are your reference skills: the rules, the exceptions, and what good output looks like, with concrete examples. Plain markdown, readable by any agent, versioned in git..claude/skills/[name]/SKILL.md whose frontmatter description tells the agent when the skill applies, so it loads automatically whenever the task is relevant. Other harnesses have equivalents (Codex and Cursor read the same skill format or rules files), and the universal fallback works everywhere: "Read /skills/brand.md before we begin."/new-screen or a weekly audit, packaged the same way. Build them after your reference skills stabilize; see the custom slash commands & skills appendix.# .claude/skills/brand/SKILL.md
---
name: brand
description: Brand and visual standards for this product.
Use when designing, building, or reviewing any screen,
component, or visual asset.
---
Read /skills/brand.md and follow it exactly. It defines
the color system, typography, iconography, illustration
style, and motion principles. Never deviate without
logging a decision.
/skills/*.md and the auto-invocation wrapper in .claude/skills/. The knowledge files are vendor-neutral: they work pasted into any model, referenced from Cursor rules, or read by another agent CLI. The wrapper is the only Claude-specific part, and it's three lines.Complete instruction file example
# CLAUDE.md ## What this project is [Product name]: [one sentence description]. Primary platform: mobile web (iOS Safari + Android Chrome). Target users: [brief description]. ## Tech stack Framework: Next.js (App Router), latest stable at project start Components: shadcn/ui; use only these. No custom components if a shadcn equivalent exists. Styling: Tailwind CSS v4 (CSS-first config via @theme) + CSS custom properties in /styles/tokens.css Language: TypeScript throughout. No plain JS files. Package manager: pnpm ## Folder structure /app Next.js routes and pages /components /ui shadcn components (themed, do not modify structure) /[feature] product-specific components /styles tokens.css single source of truth for all design tokens globals.css base styles and dark mode assignments /skills context files (read before relevant tasks) /assets /icons SVG icons /illustrations illustration assets /tokens generated JSON exports for engineering /decisions decision log and changelog ## Skill files read before relevant tasks Brand + visual standards → /skills/brand.md Accessibility rules → /skills/a11y.md Voice and copy → /skills/copy.md Performance budget → /skills/performance.md Component library → /skills/components.md Decision logging → /skills/decisions.md ## Absolute rules: never break these 1. Never hardcode color or spacing values. Use CSS variables only. 2. Never install a new npm package without asking first. 3. Never modify files in /components/ui directly; theme through tokens. 4. Never push directly to main. All work goes on feature branches. 5. Never ship the /canvas or /prototype routes; design tools only. 6. Always add prefers-reduced-motion to any animation. 7. All touch targets minimum 44x44px. ## Deploy Platform: Vercel Preview: auto-deploy on every branch push Production: merge to main via PR only; engineering reviews first Password-protected previews: ask Claude Code to add basic auth or use Vercel's built-in password feature. ## Current status [Update this section weekly: what's in progress, what's next, what decisions are currently open.] Active branch: design-system/week-of-[date] In progress: [flow or component name] Open decisions: [list or "none"] ## Designer preferences Prefers high-contrast, low-density layouts. When in doubt, remove rather than add. Favour conversational, warm UI over structured, form-like UI. Never use placeholder copy; generate real copy for every screen.
Keep it portable
This playbook runs on Claude Code, but nothing in it should depend on Claude Code. Models leapfrog each other, tools get sunset, and pricing changes. The workflow survives all of that if you build it on the four layers that are now vendor-neutral standards. And portability is not just insurance; it's the operating model. Done right, a designer works from whichever harness fits the problem in front of them and picks the work back up in another without losing the thread, because the thread was never in the tool.
AGENTS.md: the shared instruction file
AGENTS.md; the stack rules, folder structure, absolute rules, and skill file map are tool-agnostic anyway.# CLAUDE.md: the thin shim @AGENTS.md # Claude-specific additions only, e.g.: Skill wrappers live in .claude/skills/; read them automatically when relevant. # or skip the file entirely and symlink: ln -s AGENTS.md CLAUDE.md
MCP: the vendor-neutral tool layer
.mcp.json) so the whole team, and every agent, gets the same connections.Vendor-neutral artifacts: everything important is a file in git
The payoff: work from any harness, resume in any harness
Prompt like a Designer
Claude Code responds well to designer language: visual, spatial, felt. You don't need to translate your instincts into technical instructions. Describe what you see and what you want, the way you'd talk to another designer. One framing before the techniques: this is context engineering. What the agent has loaded (your tokens, skill files, real components) shapes the output more than how you phrase any single request. The skills and CLAUDE.md setup from the previous sections is the heavy lifting; the language below is how you steer once the context is right.
Structure: how to frame a prompt
Designer language: contrast & visual weight
Designer language: space & density
Designer language: type & emphasis
Designer language: feel & tone
Designer language: layout & composition
Interaction design language
Tips & tricks
Prompting techniques, workflow shortcuts, and ways to get more out of Claude Code that aren't obvious from the main process.
Generate multiple options before committing
Scheduled design critiques from Claude
Read /skills/brand.md, /skills/a11y.md, and /skills/copy.md. Read all files in /app. Act as a senior product designer doing a design critique. Be honest and specific, not encouraging. Review the current state of the product against: 1. Brand alignment: does it feel like the brand or generic? 2. Hierarchy: is it clear what the most important action is on each screen? 3. Consistency: where does the experience feel uneven? 4. Friction: where might a user hesitate or get confused? 5. Copy: anything that sounds off, vague, or un-human? 6. Missed opportunities: what's a better version of this that we haven't tried yet? Be specific. Cite screens and components. Don't summarize what's there; evaluate it.
Second opinions from a different model
Reframe before you refine
Ask Claude to pressure-test your own decisions
Use Claude to explain your work to non-designers
Teach Claude your taste over time
/memory) is what Claude happened to note; CLAUDE.md is what you decided matters.Spot-check with real device previews
Upload screenshots to show Claude exactly what you mean
Build your HTML/CSS literacy to improve your prompting
var(--color-action-primary) references a value defined in tokens.css lets you catch when Claude hardcodes a hex instead of using a variable and correct it immediately.Responsive web presentations from your design system
Use the harness: plan mode, checkpoints, hooks
Design the loop, not just the prompt
When things go wrong
The most common failure modes in this workflow: what causes them, what the symptoms look like, and how to recover. Each failure has a pattern: read it once so you recognize it fast when it happens.
Wrong style from the start
Design system drift
Stuck in an iteration loop
Visual problems need screenshots, not descriptions
Prepare your brand input
Claude Code needs everything in text or pasteable form. Convert assets before you open the terminal.
What to gather before opening Claude Code
Complete brand.md example
# /skills/brand.md ## Identity Product: [Name], [one-line description of what it does and for whom]. Positioning: [e.g. "the approachable alternative to X"]. Tone: [3 words, e.g. warm, direct, human]. Never: formal, legalistic, passive voice, jargon, corporate-speak. ## Logo Primary file: /assets/logo/logo.svg Wordmark only: /assets/logo/logo-wordmark.svg Clear space: 16px minimum all sides. Minimum size: 24px height on screen. Approved on: white, [primary bg color], dark backgrounds. Never: stretch, rotate, recolor, place on busy backgrounds, use the icon without the wordmark in UI contexts. ## Typography Primary typeface: [Font name] Use for: all UI headings, labels, buttons, body. Weights in use: 400 (regular), 500 (medium), 600 (semibold). Never use 700+ weight; too heavy at mobile sizes. Secondary typeface: [Font name or "none"] Use for: editorial moments only (marketing, splash screens). Never in: forms, navigation, data displays. Type scale (px): xs: 12 captions, legal, metadata sm: 14 secondary body, form hints md: 16 primary body text (base) lg: 20 section leads, card titles xl: 24 screen titles 2xl: 32 hero headings 3xl: 40 marketing only Line height: 1.5 body / 1.2 headings / 1.0 buttons and labels. Letter spacing: -0.01em headings / normal body. ## Color system # Primitives: raw values, never used directly in UI brand-50: [lightest tint] brand-100: [light tint] brand-500: [main brand color] ← primary identity color brand-900: [darkest shade] neutral-0: #FFFFFF neutral-50: [off-white main surface] neutral-100: [light gray] neutral-500: [mid gray] neutral-900: [near-black primary text] # Semantic: use these in components, never primitives action-primary: brand-500 ← CTAs, links, active states action-on-primary: neutral-0 ← text on action-primary bg surface-default: neutral-50 ← page and card backgrounds surface-raised: neutral-0 ← elevated surfaces text-primary: neutral-900 ← body text text-secondary: neutral-500 ← supporting text text-disabled: neutral-300 border-default: neutral-100 danger: [red hex] ← errors, destructive actions warning: [amber hex] ← caution states success: [green hex] ← confirmations ## Color proportions (per screen) Neutral surfaces: ~80%, backgrounds, cards, inputs. Brand primary: ~10%, CTAs, key highlights, active states only. Accent/warm: ~10%, empty states, illustrations, celebrations. Rule: if a screen feels colorful, there's too much brand color. ## Accessibility Standard: WCAG 2.2 AA minimum on all text. AAA preferred for body. Approved pairs (tested): [neutral-900] on [neutral-50] ✓ AAA [neutral-0] on [brand-500] ✓ AA [neutral-900] on [neutral-0] ✓ AAA Failing pairs (never use): [brand-500] on [neutral-50] ✗ fails AA [neutral-500] on [neutral-0] ✗ fails AA for small text ## Iconography Library: Lucide React Stroke weight: 1.5px, never change this. Sizes: 16px (inline/compact) or 24px (standalone) only. Style: always outline, never filled. Custom icons: /assets/icons/, must match Lucide stroke style. When no Lucide equivalent: use a simple geometric SVG from Claude, not an image or emoji. ## Illustration Style: flat, minimal, warm; think editorial not technical. Palette: brand palette only, no additional colors in illustrations. Detail level: low, clarity over complexity at small sizes. Subjects: objects and abstract shapes, avoid depicting people. Files: /assets/illustrations/, WebP + SVG where possible. Usage: empty states, onboarding, error screens, celebrations. ## Photography Mood: natural light, real environments, candid not posed. Subjects: [relevant to your product context]. Avoid: stock photo aesthetics, artificial staging, excessive filters. Format: WebP, max 200kb, explicit width/height attributes always. ## Motion principles Default: no animation unless it communicates state or transition. Duration: 150–250ms for micro-interactions, 300ms for transitions. Easing: ease-out for entrances, ease-in for exits. Always add: prefers-reduced-motion override (see /skills/a11y.md).
Scaffold the design system
One prompt to go from brand file to a working, themed shadcn scaffold.
The initial Claude Code prompt
Read /skills/brand.md carefully before starting.
Scaffold a design system using shadcn/ui and
Tailwind CSS v4 (CSS-first config, no tailwind.config file)
based on the brand defined in brand.md. Generate:
1. /styles/tokens.css
CSS custom properties for all colors, type scale,
spacing scale (4px base unit), border-radius, shadows.
Define values in :root and .dark.
2. /styles/globals.css
@import "tailwindcss";
@theme inline block mapping Tailwind + shadcn tokens
(--color-primary, --color-secondary, etc.)
to our CSS variables from tokens.css.
Base typography rules.
3. /components/ui/ themed shadcn components:
Button, Input, Card, Badge, Alert, Select, Dialog.
Install via the shadcn CLI; do not hand-write them.
4. /app/design-system/page.tsx
Living style guide showing every token and component
in context. Renders at /design-system route.
Follow color proportions in brand.md.
Flag any accessibility conflicts you find.
tailwind.config.ts is legacy. The entire theme lives in the same tokens.css + globals.css pair that is already your source of truth. One less file for values to drift into.What to review after generation
--color-primary should map to your brand color, not shadcn's default. This mapping is where "generic shadcn" sneaks in./design-system in browser. Does it feel like your brand, or generic shadcn?.dark block; shadcn's default mappings need manual verification against your brand's dark colorway./design-system is your Figma replacement for system-level review. Share the Vercel preview URL; stakeholders see the real thing, not a mockup of it.The system you're building is agent context
Export design tokens for dev
Generate design-tokens.json before building any screens. This gives engineering the token contract early so both sides work from the same source of truth from day one.
Generate the token file
Read /styles/tokens.css and /styles/globals.css. Generate /tokens/design-tokens.json in W3C Design Tokens format (DTCG, stable spec 2025.10). Include: - color: all values grouped by role (primitive → semantic → component-level) - typography: families, size scale, weights, line heights, letter spacing - spacing: full scale from 4px base unit - border-radius: all radius values - shadow: elevation tokens - motion: duration and easing if defined For each token include: value, type, description, and the CSS variable name it maps to. Also generate: /tokens/design-tokens.ios.json (pt units for iOS) /tokens/design-tokens.android.json (sp/dp for Android)
Token structure: primitive → semantic → component
// design-tokens.json (DTCG format) { "color": { "primitive": { "brand-500": { "value": "#______", "type": "color", "description": "Base primary brand color" } }, "semantic": { "action-primary": { "value": "{color.primitive.brand-500}", "type": "color", "description": "CTAs, links, active states", "css-variable": "--color-action-primary" } } } }
button-bg-primary) are optional but valuable for native platforms. The three-layer structure is what makes the JSON useful to engineering, not just a color dump.Keep tokens in sync as the system evolves
design-tokens.json as a build artifact: regenerate it whenever tokens.css changes, not manually.Accessibility audit
Accessibility is built in from the token level up, not checked at the end. Run a proactive audit immediately after scaffolding, then again after every major build phase. The modern pattern is two layers: deterministic scanners (axe-core, Lighthouse) catch the mechanical violations in CI, and the agent sits on top doing what scanners can't: judging whether alt text is actually descriptive, whether focus order is logical, whether a screen is technically compliant but practically unusable. The prompts below are that judgment layer.

Token-level audit: before any screens are built
Read /styles/tokens.css and /skills/brand.md.
Run a full accessibility audit on the color token system:
1. Test every foreground/background color pair defined
in the semantic layer against WCAG 2.2 contrast ratios.
Flag: AA fail (<4.5:1 normal, <3:1 large text),
AAA fail (<7:1 normal); note both thresholds.
2. Identify any semantic token combinations that are
used together in components but have no contrast check.
3. Check that every interactive state (hover, focus,
active, disabled) has sufficient contrast in both
light and dark mode.
4. Output a report: passing pairs ✓, failing pairs ✗,
and a suggested fix for each failure (adjusted hex
value that passes AA while staying close to the brand).
Component-level audit: after scaffold, before flows
Read all files in /components/ui. For each component, audit: FOCUS Does every interactive element have a visible focus ring? Is focus order logical within the component? Can the component be fully operated by keyboard alone? ARIA Are roles, labels, and descriptions correctly applied? Do icon-only buttons have aria-label? Are error states announced to screen readers? TOUCH TARGETS Are all tap targets at least 44x44px on mobile? Is there sufficient spacing between adjacent targets? MOTION Does any animation respect prefers-reduced-motion? Output per component: pass ✓ | fail ✗ | needs review ⚠ For each failure, cite the line and suggest the fix.
The a11y skill file
# /skills/a11y.md ## Standards Target: WCAG 2.2 AA minimum. AAA where feasible. Platform: mobile-first. All touch targets 44x44px minimum. ## Color Never communicate meaning through color alone. Always pair color with a label, icon, or pattern. Approved contrast pairs: [generated from token audit] Failing pairs never to use: [generated from token audit] ## Focus Every interactive element must have a visible focus ring. Use: outline: 2px solid var(--color-action-primary); outline-offset: 2px Never remove outline without providing an alternative. ## ARIA Icon-only buttons always need aria-label. Form inputs always need associated label elements (not placeholder only). Error messages use role="alert" so they're announced immediately. Loading states use aria-live="polite". ## Motion All animations must respect: @media (prefers-reduced-motion: reduce) { animation: none; } ## Audit cadence Run token audit: after any token change. Run component audit: after any component update. Run full screen audit: weekly, on the system branch.
Ongoing screen-level audit: on the weekly branch
Read /skills/a11y.md. Read all files in /app. Run a full accessibility audit across all screens: 1. Re-check all color pairs in context (components on real backgrounds, not just isolated tokens). 2. Verify heading hierarchy per screen: no skipped levels. 3. Check that all images have meaningful alt text or are marked decorative with aria-hidden="true". 4. Verify every form has proper label associations, error handling, and success confirmation. 5. Check reading order matches visual order on mobile. Output: screen-by-screen report with pass/fail/warn. Group failures by severity: critical | moderate | minor. Critical issues block the weekly merge.
Publish your design system site
A living public site generated from your codebase, auto-updated on every merge. Engineering has a single URL to bookmark. The changelog writes itself. The system is never more than a week out of date.
Generate the minisite route
Read /skills/brand.md, /styles/tokens.css, and all files in /components/ui. Generate a design system documentation site at /app/design-system. Include: 1. Token reference: every color, spacing, radius, shadow, and typography token displayed in context with its CSS variable name and current value. 2. Component gallery: every shadcn component in use, rendered in our brand theme, with a code snippet showing how to use it. 3. Pattern library: any composite patterns in /components that combine multiple primitives (e.g. form layouts, card variants, empty states). 4. Changelog: pull from /decisions/changelog.md and render as a versioned list with dates and descriptions. The site should use our existing tokens and components. Deploy to /design-system route, accessible at the Vercel preview URL without authentication.
Auto-generate the changelog
Read the git diff between design-system/week-of-[date] and main. Write a changelog entry for /decisions/changelog.md. Format: ## [date] - Week of [date] ### Added - [new tokens, components, or patterns introduced] ### Changed - [tokens or components with updated values or behavior] ### Deprecated - [anything scheduled for removal, with replacement] ### Fixed - [a11y issues, visual regressions, or token drift resolved] Be specific: name the token or component, describe the change, and explain why if the decision log has context. Skip sections with no changes. Keep entries brief.
Weekly merge discipline keeps the site current
design-tokens.json and notify engineering that a new token version is available for the next sprint.Share it and replace Figma as the reference
Build major flows
Start with one flow end-to-end. Build it completely before moving to the next. Let inconsistencies surface naturally; you'll extract them into patterns later.
Build one flow completely before starting the next
Tweak in conversation, not in code
Build the next flow, then compare
Definition of done for a flow
Canvas & prototype views
Do this after your first flow is complete, not before. The canvas is how you spot structural problems: missing states, awkward transitions, flows that seem fine screen-by-screen but break down when seen as a sequence. It's a review step, not a build step: build first, then zoom out here before starting the next flow.
flows.config.ts: the register both tools read from
flows.config.ts. Adding a new screen = one line in the config. Both views update automatically.// flows.config.ts export const flows = [ { name: "[Flow 1 name]", screens: [Screen1, Screen2, Screen3, Screen4] }, { name: "[Flow 2 name]", screens: [EntryScreen, FormScreen, ReviewScreen, PendingScreen, SuccessScreen, ErrorScreen] } ]
Build the canvas view at /canvas
Build a canvas view at /app/canvas/page.tsx that reads from flows.config.ts. Requirements: Infinite canvas: pannable with click-drag, zoomable with scroll or pinch (CSS transform, no libraries). Each screen rendered at ~25% scale inside a labeled phone frame. Screen name shown below. Click any screen to expand full size in an overlay. Screens grouped by flow with flow label above each group. Arrange groups left to right. Mini-map in bottom-right showing canvas position. Toolbar: zoom in/out, fit-to-screen, flow-jump dropdown. Add /canvas to CLAUDE.md as a protected design-only route; never ships to production.
Build the prototype view at /prototype/[flow]
Build a prototype viewer at /app/prototype/[flow]/page.tsx that reads from flows.config.ts. Requirements: Centered 390px mobile frame on a neutral dark background. Tap/click advances to next screen. Back arrow goes back. Step counter at top of frame showing position in flow. Flow selector dropdown outside the frame. "Copy link" button copies URL with flow + step as query params; shared link opens at the exact screen. Left/right arrow key navigation. No UI chrome inside the phone frame. Deploy at /prototype/[flow-1-name], /prototype/[flow-2-name].
git tag v0.3-user-test. You can always share the exact version a stakeholder or tester saw, even after screens have changed.Localization
Treat copy the same way you treat color: as a token, not hardcoded text. Structure it from the start and localization becomes a content swap, not a rebuild.
Content tokens: the same pattern as design tokens
/locales/en.json, /locales/es.json, /locales/fr.json, structured the same way design tokens live in tokens.css. Keys are semantic, not descriptive: onboarding.cta.primary, not get-started-button-text.# /locales/en.json: content tokens
{
"onboarding": {
"hero": {
"headline": "Get started in minutes",
"subtext": "No setup required. Connect your account and go.",
"cta": "Create your account"
},
"steps": {
"connect": "Connect your tools",
"configure": "Set your preferences",
"launch": "You're ready"
}
},
"errors": {
"auth.expired": "Your session expired. Please sign in again.",
"form.required": "This field is required."
}
}
Language switcher in the canvas and prototype views
/canvas with a locale selector in the toolbar. Switching locale re-renders all screens with the selected language; you see layout impact immediately: German strings that break a button, Arabic RTL that reflows a card.# prompt to add locale switcher to canvas view
Add a locale switcher to /app/canvas/page.tsx.
- Render a dropdown in the top toolbar with all locales from /locales/
- Switching locale passes the selected language to all child screen components via context
- The switcher state persists in localStorage so it survives page reloads
- Default to 'en' if no preference is stored
Voice and tone per market
/skills/copy-[locale].md alongside your existing copy.md: "For es-MX: warm and conversational, avoid formal usted constructions in UI copy, contractions are fine, error messages should feel reassuring not clinical."Localize continuously, not as a phase
Extract & enforce patterns
Do this after your second major flow, not your first. Patterns only become visible through repetition; extracting from a single flow produces components that are too specific to generalize. Wait until you have enough surface area to see what actually repeats, then consolidate once rather than incrementally.
Ask Claude to find inconsistencies across all flows
Read all files in /app and /components. Identify inconsistencies across these dimensions: Spacing: are margins and padding consistent across similar components? Typography: are heading levels used consistently? Buttons: are primary/secondary/ghost variants used correctly and consistently? Cards: how many different card treatments exist? Should any be consolidated? Empty states: consistent handling? Error and loading states: same approach per flow? Output a report grouped by dimension. For each issue, cite the specific files and suggest the canonical pattern to adopt.
Consolidate into named canonical components
StatusCard, TransactionRow, StepIndicator, not BigCard or ListItemDark.Ask the system to enforce canonical components everywhere
These are now the canonical components: ContentCard → /components/ui/ContentCard.tsx StatusBadge → /components/ui/StatusBadge.tsx FormStep → /components/ui/FormStep.tsx EmptyState → /components/ui/EmptyState.tsx Scan all files in /app. Replace any inline implementations of these patterns with imports of the canonical component. Do not change visual output; only refactor structure. List every file you changed.
Decision log & changelog
Set this up before your second flow, not after you've shipped. Decisions made in the first flow are still fresh; once you're in the third or fourth flow, the context for why you built the system a particular way is already gone. The log is worthless if it's written retrospectively.

The decisions skill file
# /skills/decisions.md ## Purpose Log every significant design decision made during this project. A decision is significant if it: changes a token value, introduces or retires a component, establishes a new pattern, overrides a brand rule, or resolves a conflict between design and engineering. ## Format: one entry per decision Each entry must include: Date (YYYY-MM-DD) Area: token | component | pattern | flow | brand | a11y Decision: what was decided, in one clear sentence Rationale: why, the constraint, tradeoff, or insight Affected: which files, components, or screens this touches Status: active | superseded | under review Decided by: designer name or "Claude (auto)" ## When to log Log immediately when: a pattern is consolidated, a token is changed, a component is retired, an escalation is resolved, or a new canonical component is established. ## File location All entries live in /decisions/log.json (structured) and /decisions/CHANGELOG.md (human-readable summary).
Logging decisions as you work
# add to CLAUDE.md:
After completing any task that involves a design decision
(token change, component update, pattern consolidation,
escalation resolution), append an entry to
/decisions/log.json following the format in
/skills/decisions.md. Do this automatically; do not
ask for confirmation unless the decision area is "brand"
or "a11y", which require designer sign-off.
Generating the changelog
Read /decisions/log.json. Generate /decisions/CHANGELOG.md structured as follows: ## [Version / week tag] [date] ### Tokens [list token changes with rationale summary] ### Components [new, updated, or retired components with rationale] ### Patterns [pattern consolidations or new patterns established] ### Escalations resolved [decisions that required designer sign-off] ### Open / under review [anything logged as "under review" status] Keep each entry to one line. Link to the affected file where relevant. Mark designer-approved items with ✓.
Version history in the canvas view
Extend /app/canvas/page.tsx to add a version history panel. Requirements: Version panel (right sidebar, collapsible): Reads from /decisions/log.json Lists entries grouped by week/version tag, newest first Each entry shows: date, area badge, decision (one line), decided-by label Clicking an entry highlights the affected screens on the canvas with a colored overlay Filter by area: token | component | pattern | flow | all Version switcher (toolbar): Dropdown of all git tags with "design-system/" prefix Selecting a tag checks out that version and reloads the canvas so you can see exactly what the system looked like at any past milestone "Current" option always returns to HEAD Changelog tab (bottom drawer): Renders /decisions/CHANGELOG.md as formatted HTML Searchable by keyword Exportable as PDF for stakeholder sharing
Querying the decision history
Content & copy
Copy is a design material: it lives in the system alongside tokens and components, not in a separate doc. Without a copy skill, Claude generates placeholder language that never gets replaced and ships. Setting this up means copy gets reviewed on every screen, not at the end of the project.
The copy skill file
# /skills/copy.md ## Voice [Your product voice in 3 words, e.g. warm, direct, human.] Never: formal, legalistic, passive voice, jargon. ## Tone by context Onboarding: encouraging, low-stakes, welcoming. Errors: calm, specific, actionable. Never blame the user. Success states: brief, warm; don't over-celebrate. Empty states: helpful, not apologetic. Show the path forward. Destructive actions: clear and specific. State what will be lost. ## Button labels Always verb-first and specific: "Save changes", not "OK". Primary CTA: describes the outcome, not the action. Good: "Send money" / Bad: "Submit" Destructive: states what happens. "Delete account", not "Confirm". Cancel: always "Cancel", never "No" or "Go back". ## Error messages Structure: [what happened] + [why] + [what to do]. Never show raw error codes to users. Always offer a path forward: a retry, a support link, or next step. ## Character limits (mobile) Screen titles: 24 chars max. Subtitles / descriptors: 60 chars max. Button labels: 20 chars max. Toast notifications: 48 chars max. Empty state body: 80 chars max. ## Things to never say "Something went wrong" (too vague) "Please try again later" (no timeline) "Invalid input" (doesn't say what's invalid) "Are you sure?" (not specific enough for destructive actions)
Generating copy for screens
Copy audit: consistency sweep across flows
Read /skills/copy.md. Read all files in /app. Audit all copy across every screen for: 1. Voice violations: anything that sounds formal, cold, or uses passive voice. Quote the offending string and suggest a rewrite. 2. Button label inconsistencies: same action labeled differently across screens. List all button labels grouped by action type. 3. Error messages without a clear path forward: flag any error that doesn't tell the user what to do next. 4. Character limit violations: flag any string that exceeds the limits in copy.md. 5. Placeholder text still in production: any string containing "Lorem", "TBD", "TODO", or "[". Output: grouped by issue type, with file and line reference, and a suggested fix for each.
Copy as a design system export
Read all files in /app.
Generate /tokens/content-inventory.json, a structured
export of every user-facing string in the app:
{
"screen": "payment-error",
"element": "error-title",
"string": "Payment didn't go through",
"context": "shown after a failed card charge",
"character-count": 30,
"last-updated": "2026-01-15"
}
Group by screen. Flag any string over the character
limits defined in /skills/copy.md.
Build app-specific skills
You don't have to write skills from scratch; Claude can derive them from the codebase you've already built together. Write them after patterns stabilize, not before.
The skill files to build first
Ask Claude to generate skill files from what you've built
Read all files in /app and /components/ui. Generate /skills/components.md, a skill file documenting every component in /components/ui with: What it's for (one sentence) When to use it vs. alternatives Required and optional props A usage example Anything that should never be done with it Write it so a designer new to this codebase could read it and know exactly what to reach for and when.
Performance budget
AI-generated component code can be bloated: unused dependencies, heavy bundles, unoptimized assets. A performance audit runs before every weekly merge so issues are caught before they compound across the codebase.
The performance budget skill file
# /skills/performance.md ## Budget targets JS bundle (initial load): < 200kb gzipped CSS bundle: < 50kb gzipped Largest Contentful Paint (LCP): < 2.5s on 4G mobile Cumulative Layout Shift (CLS): < 0.1 Time to Interactive (TTI): < 3.5s on mid-range device ## Images All images served as WebP or AVIF. Max image weight: 200kb per asset. All images must have explicit width and height to prevent CLS. Use next/image (or equivalent) for automatic optimization. ## Dependencies No new npm package without justification logged to decision log. Prefer tree-shakeable libraries. No package that duplicates functionality already in the stack. ## Audit cadence Run before every weekly system branch merge. Any budget breach blocks the merge until resolved or a conscious exception is logged to the decision log.
Weekly performance audit prompt
Read /skills/performance.md. Run a performance audit on the current branch: 1. BUNDLE ANALYSIS Run: npx next build && npx @next/bundle-analyzer Flag any route whose JS exceeds 200kb gzipped. Identify the top 3 contributors to bundle size. 2. UNUSED DEPENDENCIES Scan package.json against actual imports across /app and /components. List any package imported nowhere. 3. DUPLICATE FUNCTIONALITY Flag any two packages that do the same job (e.g. two date libraries, two animation libraries). 4. IMAGE AUDIT Find any <img> tags not using the optimized image component. Find images without explicit width/height. Find any image over 200kb. 5. CSS AUDIT Find any inline styles that duplicate values already defined in tokens.css. Find any hardcoded color or spacing values that should reference a CSS variable. Output: pass ✓ | exceeds budget ✗ | warning ⚠ For each failure, cite the file and suggest the fix.
Catching bloat as it's generated
Handoff & dev collaboration
Your Claude Code output is the handoff. Engineers get a branch with working code, not a Figma link.
Handoff via code, not specs
DB & backend integration
Day-to-day dev collaboration
Designing on real code
Instead of handing off to engineering and waiting, you design directly on top of the production codebase: read-only access to the real repo, Claude Code doing the work, engineering reviewing and approving before anything merges. Design intent becomes production-ready code without a translation layer.
Read-only access to the engineering repo
# add to CLAUDE.md
Engineering repo: ../[repo-name] (read-only reference)
Before designing any component that engineering already owns,
read the relevant file in ../[repo-name]/src first.
Never create a file in ../[repo-name] directly.
All design work stays in this repo [your-design-repo].
Designing against real data shapes
# example session opener
Read ../[repo-name]/src/types/transaction.ts
and ../[repo-name]/src/api/transactions.ts.
Before I design the transaction history screen:
1. What fields are available on a transaction object?
2. What are the possible status values?
3. What edge cases should I design for (empty, error,
partial data, very long merchant names)?
4. Are there any fields I might want to show that
don't exist in the current data model?
The design PR: how you submit work for engineering review
# design PR description template: ask Claude to fill in
## What changed
[Visual description of the change: what users see differently]
## Implementation notes
[How it was built: component choices, token usage, state logic]
[Any deviations from the existing pattern and why]
## What to verify
- [ ] Visual: [specific thing to check in browser]
- [ ] A11y: [specific thing to check: contrast, focus, ARIA]
- [ ] Edge cases: [list from data shape analysis]
- [ ] Mobile: [anything to check specifically on small screens]
## Design decision log
[Link to relevant entries in /decisions/CHANGELOG.md]
Engineering review gates: what requires approval before merge
Using Claude to pre-check before submitting for review
Read ../[repo-name]/src/[path-to-affected-files]. Read my changes in /app/[screen].tsx. Before I open a PR, review my changes as an engineer would: 1. COMPATIBILITY: does my implementation use the same patterns, naming conventions, and abstractions as the surrounding engineering code? Flag any inconsistency. 2. COMPLETENESS: are there states, edge cases, or error conditions the engineering code handles that my design doesn't account for? 3. PERFORMANCE: any obvious inefficiencies? Unnecessary re-renders, missing memoization, heavy inline operations? 4. ACCESSIBILITY: does this meet the standards in /skills/a11y.md? Any focus, ARIA, or contrast issues? 5. BREAKING CHANGES: does anything I've changed affect other screens that use the same component or token? List every file that could be affected. 6. VISUAL VERIFICATION: run the app and walk the affected screens in the browser. Screenshot every state and compare against the current production version. Flag any unintended visual change, however small. Output: ready to submit ✓ | needs revision ✗ | review note ⚠ For each issue, cite the specific line and suggest the fix.
Keeping in sync as engineering changes the codebase
git log reads, one Claude summary, one picture of what changed on both sides.Deploy & environments
Designer-owned: local and preview. Engineering-owned: staging and production.
The environment ladder
# designer-owned local → vercel preview (feature branch URL) # engineering-owned staging → production # never push directly to main git checkout -b feature/name # build → push → Vercel auto-deploys preview URL
git tag v0.3-user-test, always recoverable to exactly what a user saw..env.example documenting which variables point to which environment: keeps staging and prod clearly separated.Git sync & design rhythm
A weekly and daily cadence for keeping your design system in sync with the team: pulling in changes, reviewing impact, refactoring patterns, and escalating decisions that need a human call.

Weekly sync: pull, review, merge
# weekly sync routine: run every Monday morning git fetch origin git checkout main git pull origin main # check what changed since your last session git log --oneline --since="7 days ago" # then ask Claude Code: Read the git log from the past week. Summarize what changed across /components, /styles, and /app. Flag anything that may affect the design system or existing screen layouts.
Daily branches: isolated work, clean history
# start of each day: branch from latest main git checkout main && git pull origin main git checkout -b design/2026-01-15-[task-name] # end of day: push and open a draft PR git add -A git commit -m "design: [what changed]" git push origin design/2026-01-15-[task-name]
design/YYYY-MM-DD-task-name. Makes it easy to find, sort, and understand the history at a glance.Weekly design branch: new styles, patterns & components
design-system/week-of-2026-01-15. Keep feature work and system work on separate branches.# weekly system branch git checkout main && git pull origin main git checkout -b design-system/week-of-2026-01-15 # then ask Claude Code: Read all changes made to /components and /styles across this week's feature branches. Identify any new patterns, components, or style values that should be promoted into the design system. List them with a recommendation for each: promote, consolidate, or discard.
Claude reviews incoming changes by scope
Read the diff between design-system/week-of-[date] and main. Classify every change by scope: SAFE: cosmetic or additive. No existing screens affected. → Apply automatically. No review needed. Examples: new component added, token value tweaked slightly, new utility class, copy fix. REVIEW: affects existing components or token values. Existing screens may look different. → Apply, then flag for visual review in the browser. Examples: spacing scale change, border-radius update, button variant modified, type scale adjusted. REFACTOR: a pattern has changed enough that screens using the old pattern need to be updated to stay consistent. → Apply refactor across all affected files, then flag a summary of what was changed for designer sign-off. Examples: card component restructured, form pattern updated, navigation component rebuilt. ESCALATE: breaking change or significant design decision that affects major flows or brand-level tokens. → Do not apply. Write a clear brief explaining what changed, what is affected, and what decision is needed. Flag directly to the designer for a human call. Examples: primary color value changed, type scale restructured, component removed, major layout pattern shift.
Refactoring screens for updated patterns
The [ComponentName] pattern has been updated. The new canonical version is in /components/ui/[ComponentName].tsx. 1. Scan all files in /app for usages of the old pattern. 2. Refactor each to use the updated component. 3. Do not change any logic, copy, or non-visual behavior. 4. After refactoring, output a list of every file changed with a one-line note on what was updated in each. 5. Flag any file where the refactor was ambiguous or where the old usage didn't map cleanly to the new pattern.
refactor: update all screens to [ComponentName] v2.Escalation: when Claude flags a decision for you
Weekly merge: closing out the system branch
# end of week: merge system branch to main git checkout design-system/week-of-[date] git rebase main # keep history clean git push origin design-system/week-of-[date] # open PR, ask Claude Code to write the description: Read the diff between design-system/week-of-[date] and main. Write a PR description that summarizes: New components or patterns added Components refactored and why Token changes made Any open decisions that were escalated What designers should visually verify after merge
design-tokens.json and notify engineering that a new token version is available.git tag design-system/v[week]. This gives engineering a stable reference point for each week's token set.
Appendix
Extended techniques, tooling, and team workflow →
▼
Meta improvements Untested
Every time you set something up from scratch, skill files, CLAUDE.md rules, token structure, you're building what could be a template. Don't wait until you're experienced to start systemizing. Capture these improvements while building your first project or two and every project after starts ahead.
A starter template pre-wired for this workflow
setup.md at the root: "Step 1: fill in /skills/brand.md. Step 2: run this Claude Code prompt." One page, no ambiguity.# repo structure of the starter template /skills/ brand.md # stub: fill in for each project a11y.md # pre-written universal rules copy.md # stub: fill in voice/tone per project performance.md # pre-written universal budget decisions.md # pre-written logging format components.md # auto-generated after scaffold /decisions/ log.json # empty, ready to receive entries CHANGELOG.md # empty /tokens/ # empty, generated by scaffold prompt AGENTS.md # pre-written with all framework rules CLAUDE.md # thin shim: @AGENTS.md setup.md # one-page quickstart
shadcn as the permanent base: never start from scratch
A standalone canvas & prototype app
design-viewer, that can be pointed at any project's flows.config.ts and renders the canvas and prototype views without any per-project setup.# design-viewer multi-project structure /workspaces/ [project-a]/ flows.config.ts # symlink or import from project decisions/ # symlink or pulled via API [project-b]/ flows.config.ts decisions/ # single deploy, all projects accessible design-viewer.vercel.app/[project-a]/canvas design-viewer.vercel.app/[project-a]/prototype/[flow] design-viewer.vercel.app/[project-b]/canvas
A shared skill library across projects
shared-skills repo with the universal skills: a11y.md, performance.md, decisions.md, shadcn.md. These rarely change and should be identical across projects.A cross-project pattern library
shared-components repo, stripped of brand tokens, themed purely through CSS variables so any project's token file skins them automatically.When you're ready to grow
The solo workflow is worth proving before you scale it. The right moment to bring in a second designer is when the system is doing the work, when CLAUDE.md and the skill files are stable enough to hand to someone else without explaining everything verbally.
Signs the workflow is ready to scale
What to look for in a second designer
Team & governance
At 2 designers, governance is a conversation. At 5, it needs structure. These rules define ownership by team size, branch authority, and how conflicting decisions get resolved. Add them gradually as the team grows.
Ownership at each team size
1–2 designers
3 designers
4–5 designers
Keeping the system consistent as the team grows
# monthly consistency audit
Read all screens in /src/screens.
Compare every color, spacing value, and component against:
- tokens.css
- /skills/components.md
- /skills/brand.md
List values used that aren't in the design system.
Flag any that appear across multiple files as candidates
for formalizing into the system rather than correcting.
Roles: what each person owns
Branch rules for teams
# branch naming convention for teams design/[designer-initials]/YYYY-MM-DD-[task] # e.g. design/kl/2026-01-15-checkout-form design-system/week-of-YYYY-MM-DD # one per week, owned by system owner only # CODEOWNERS: add to repo root /components/ui/ @design-system-owner /styles/ @design-system-owner /skills/ @design-system-owner /decisions/ @design-system-owner
Resolving conflicting design decisions
Read /decisions/log.json and the diffs from both design/[initials-1]/[branch] and design/[initials-2]/[branch]. Identify any components, tokens, or patterns where the two branches made different choices for the same problem. For each conflict: Describe what each approach does differently Note which is more consistent with existing decisions Note which better matches /skills/brand.md and /skills/a11y.md Recommend one approach with a rationale Do not merge. Output the conflict report only for the system owner to review and decide.
Onboarding a new designer
Read all files in /skills and /decisions/CHANGELOG.md. Generate a designer onboarding brief covering: 1. What this product is and who it's for (from brand.md) 2. The tech stack and tools (from CLAUDE.md) 3. The 10 most important design system rules 4. How branching and merging works on this project 5. The 5 most significant decisions made so far and why 6. What is currently in progress or under review 7. Where to find things: components, skills, tokens, flows Keep it under 2 pages. Write for a competent designer new to this specific project.
Beyond prototypes Untested
Once your first major flow is built, these three services let you design against something that feels like a real app (persistent data, real emails, live rate limiting) before engineering has built anything. Add them when you need to validate decisions that mock data can't answer. Claude Code scaffolds every integration from a single prompt.
Railway
Resend
Upstash
Unexplored design tools Untested
Tools I haven't used extensively enough to be prescriptive about. They're worth knowing exist and experimenting with, but I can't speak to them with the same confidence as the rest of this playbook.
AI imagery & video
AI custom SVGs
Custom illustrations
Motion & animation
prefers-reduced-motion from the a11y skill.Texture & custom graphics
Custom slash commands & capability skills Untested
The playbook is full of prompts that work. Two native mechanisms make them permanent: slash commands (you invoke them: one keystroke instead of a 10-line prompt) and capability skills (Claude invokes them: auto-loaded when the task matches). Same markdown, different trigger. Start with commands; promote a command to a skill when you notice you always want it to run without asking.
How to create a command
.claude/commands/. The filename becomes the slash command..claude/commands/ directory in your project. Each .md file in that directory becomes a slash command: new-screen.md becomes /new-screen.$ARGUMENTS, a placeholder for anything typed after the command. /new-screen onboarding passes "onboarding" as the argument..claude/commands/ directory to git; commands are project-level, shared with anyone who works on the repo.# .claude/commands/new-screen.md
Create a new screen for: $ARGUMENTS
Before building, read:
- /skills/brand.md
- /skills/components.md
- /src/stories/ if it exists
Build the screen using only components that exist in
the design system. If a new component is needed, flag
it before building; don't invent one mid-session.
After building:
- Add an entry to /decisions/log.json
- Update /src/stories/ with any new components
- Run the a11y audit on the new screen
Commands to build for this workflow
.claude/skills/[name]/SKILL.md so it fires without you remembering to type it.Multi-agent parallelism Untested
Claude Code can spawn sub-agents that run concurrently. Instead of doing things one at a time in a single session, you coordinate a team: each agent handles one job while the others run in parallel. Defined subagents, agent teams, and git worktrees for isolation are first-party infrastructure.
How sub-agents work
Agent tool. Each sub-agent gets its own context, runs its own task, and returns a result. The main session synthesises the results and decides what to do next..claude/agents/: a markdown file per agent (an auditor that reads a11y.md and brand.md, a copy reviewer that reads copy.md) with its own system prompt and tool permissions. Committed to git, shared with the team, invoked by name.CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1, so treat it as something to try deliberately rather than a default. You compose a critique from that: one teammate briefed on brand, one on a11y, one on conversion, each reviewing the same flow and pushing back on the others. That is the multi-agent version of the weekly design critique. Each teammate is a separate session, so a team costs several times the tokens of a single one; reserve it for decisions worth that spend.Design workflow patterns for parallel agents
components.md. Documentation is never a backlog item.# prompt pattern for parallel work
Run these two tasks in parallel using sub-agents:
Agent 1: Build the settings flow. Read /skills/components.md
first. Build /app/settings/page.tsx and its sub-pages.
Agent 2: Run the consistency audit on all screens in /app
except /app/settings. Read /skills/brand.md first.
Return a list of violations grouped by severity.
Wait for both to complete, then summarize what was built
and what the audit found.
Designing agentic products
Everything else in this playbook is about designing with agents. This section is about the other thing the phrase "AI design" gets used for: designing products that contain agents. The two get conflated constantly. They are different competencies, and clients now ask for the second one. A short field guide so you can tell them apart and know where to start.
Generative UI: interfaces assembled at runtime
experimental_evaluate for it and serves it through AI Gateway as typesafe-ai/jev. For generative UI that means "which component do we stream next" and "is this action safe to enable" become cheap, typed decisions the product owns, with the writing model kept for the sentences.Agentic UX patterns: the new interaction vocabulary
Metrics-driven flow refinement Untested
The AutoResearch pattern, propose a change, measure it, keep only improvements, applied to product flows. The design system defines what's proposable. The metrics define what's better. Claude runs the loop.
The design system as a constrained search space
What the loop looks like in practice
program.md that tells the agent what to optimize and what it cannot touch. "Optimize the onboarding flow for step-3 completion rate. You may vary: CTA copy, button variant (primary/secondary), field ordering in the form. You may not change: the number of steps, brand colors, typography, or component structure."# program.md: flow refinement instructions
Goal: increase completion rate on /onboarding/step-3
Metric: step3_completion_rate (from Amplitude, segment: new_users_7d)
Baseline: current production branch
What you may vary:
- CTA button copy (must be <5 words, present tense, action verb)
- CTA button variant: use ButtonPrimary or ButtonLarge from Storybook only
- Form field order (all fields must remain present)
- Helper text below each field (tone: reassuring, max 12 words)
What you may not change:
- Number of steps or step titles
- Any value in tokens.css
- Component structure or layout
- Anything in /skills/brand.md
Cycle: deploy preview → wait 48h for data → read metric → commit if >2% improvement
Maintaining design fidelity at scale
program.md is the design review; it defines the entire search space the agent will explore. A well-written program file means every variant that emerges is something you'd have approved anyway.What you need before this is viable
Write the eval: taste as a testable rubric
/skills/evals.md and make it a standing gate: before any screen is presented to you, an agent scores it against the rubric and cites every failure. You review work that already passed your own written bar.experimental_evaluate, so the gate can live inside the loop rather than in a separate service. Set the confidence floor in /skills/evals.md; anything below it comes to you.Data & analytics Untested
Connect your behavioral and revenue analytics to the canvas and prototype views. See real user data in the context of what you designed, then ask Claude what to change and why.
Analytics in design context
Wiring behavioral analytics to the canvas
analytics.config.ts. Each screen has a corresponding funnel step name in the analytics platform. This mapping is what lets the canvas fetch the right metric for each screen./api/metrics?screen=onboarding-step-3 and gets back the completion rate, time-on-screen, and exit rate for that screen.# analytics.config.ts: map screens to analytics events
export const analyticsMap = {
"onboarding-welcome": { funnel: "onboarding", step: "view_welcome" },
"onboarding-connect": { funnel: "onboarding", step: "view_connect_account" },
"onboarding-configure": { funnel: "onboarding", step: "view_configure" },
"onboarding-complete": { funnel: "onboarding", step: "view_complete" },
"checkout-cart": { funnel: "checkout", step: "view_cart" },
"checkout-payment": { funnel: "checkout", step: "view_payment" },
} satisfies Record<string, { funnel: string; step: string }>
Revenue context from Stripe
Querying Claude with live data
# canvas data export: paste into Claude prompt
Screen: checkout-payment
Period: last 30 days
Completion rate: 61.2% (↓ 8.1% vs prior period)
Exit rate: 38.8%
Median time: 2m 14s (↑ 47s vs prior period)
Top exit event: click_back_button (44% of exits)
Device split: mobile 71% / desktop 29%
Mobile completion: 54.1% | Desktop completion: 76.8%
Stripe conversion (completed → paid):
This period: 28.4%
Prior period: 31.1%
Gstack: live QA in the terminal Untested
A suite of Claude Code skills built by Garry Tan, CEO of Y Combinator: twenty-plus specialist roles (QA lead, design reviewer, security officer) of which the browser-automation skill, /browse, is the one this section covers. Point it at your running app and Claude can navigate, interact, screenshot, and diff without leaving the terminal. github.com/garrytan/gstack
What Gstack does
git clone --depth 1 https://github.com/garrytan/gstack.git ~/.claude/skills/gstack && cd ~/.claude/skills/gstack && ./setup, then invoke skills like /browse directly in Claude Code. Note: gstack's /browse and the Playwright MCP server solve the same problem (agent eyes on the running app); pick one, not both.How to use it in a design workflow
# end-of-session QA prompt
Use gstack to walk the onboarding flow.
Screenshot each screen. Flag anything that:
- looks visually broken
- has text that overflows or wraps badly
- uses a color or font that doesn't match the design system
- has an interactive element that doesn't respond correctly
The skills this section uses
| Skill | Your specialist | What they do |
|---|---|---|
| /browse Essential | QA Engineer | Gives the agent a real browser: navigate, click, fill forms, screenshot, diff. /open-gstack-browser shows the browser headed so you can watch. |
| /setup-browser-cookies | Session Manager | Import cookies from your real browser (Chrome, Arc, Brave, Edge) so the agent can test pages behind a login. |
| /qa | QA Lead | Walk the app, find bugs, fix them with atomic commits, re-verify. Generates a regression test for every fix. |
| /design-review | Designer Who Codes | Audits spacing, hierarchy, consistency, and AI-slop patterns, then fixes what it finds with before-and-after screenshots. |
/office-hours to /ship, is in the README. The list changes most weeks, which is why it is not reproduced here.

