The agent-control layer for the PawHaven monorepo. It turns "generate code when asked" into a repeatable engineering process: classify the task, follow a fixed workflow, verify against the real artifact, ship.
Orientation for any agent or session entering this repository. Read this file first; load the rest on demand.
.codebuddy is PawHaven's agent-control layer: the markdown files (plus settings.json) that define how AI agents work on this repo β who they are, what rules they follow, and which workflow each task runs. Its core idea: don't maximize code output β maximize verified engineering quality ("write less, but higher quality code").
It holds: a router (dispatcher.md), workflows, principles, agents, rules, docs, and settings.json. The files do nothing on their own β CodeBuddy's agent runtime reads them and follows them.
For PawHaven specifically β a React 19 + TypeScript monorepo (apps/, packages/, libs/) with a design system, i18n in zh-CN / en-US / de-DE, and multiple engineering roles β the .codebuddy keeps every agent and subagent on the same engineering discipline while each role stays scoped to its job.
Every task flows through the three layers end to end β nothing else in the repo runs:
flowchart TD
U["User prompt<br/>(or /pawhaven-mode, or a subagent dispatch)"] --> R
subgraph L1["LAYER 1 Β· Router β dispatcher.md"]
R{{"Classify the task"}} --> IDX["Read the principles index<br/>principles/ Β· 13 rules"]
IDX --> PICK{{"Match ONE workflow<br/>from the registry"}}
end
subgraph L2["LAYER 2 Β· Workflows β workflows/*.md (fixed step sequences)"]
PICK --> INV["Investigation<br/>read-only question"]
PICK --> BUG["Bug Fix<br/>reported defect"]
PICK --> FEA["Feature<br/>new / changed behavior"]
PICK --> REF["Refactoring<br/>behavior-preserving change"]
PICK --> PERF["Perf Issue<br/>measured slowness"]
PICK --> DES["Design Decision<br/>architecture / API choice"]
PICK --> ARC["Architecture Change<br/>crosses a boundary"]
INV & BUG & FEA & REF & PERF & DES & ARC --> TODO["Copy the workflow's steps<br/>VERBATIM into the todo list<br/>skipped step stays as: skip: reason"]
end
subgraph L3["LAYER 3 Β· Principles + Skills + Agents"]
TODO --> EXEC["Execute each step<br/>fires a principle Β· skill Β· subagent"]
EXEC --> EVID["Each step produces evidence<br/>repro Β· trace Β· green check"]
end
EVID --> VER{{"Verify on the REAL artifact<br/>typecheck Β· tests Β· build Β· rendered surface"}}
VER -- "fails" --> EXEC
VER -- "passes" --> HF["Review Handoff<br/>workflows/handoff.md"]
HF --> DONE["Ready for human review<br/>you review the diff Β· you open the PR"]
Each layer, one line (executed in order):
- Step 1 β LAYER 1: Router (
dispatcher.md) β decides which job this is. It never does the job; it picks the workflow and copies that workflow's steps into the todo list. - Step 2 β LAYER 2: Workflows (
workflows/) β the fixed step sequences. They don't know how to do anything either; each step just fires a skill or a subagent. - Step 3 β LAYER 3: Principles + Skills + Agents β the only layer that actually does work: principles decide, skills know the domain, agents execute and own the result.
All 8 workflows share this shape β numbered steps, each one firing layer-3, ending in a verification loop and the review handoff (workflows/handoff.md). Example: Bug Fix:
flowchart LR
subgraph WF["Bug Fix β steps copied verbatim into the todo list"]
S1["1. Reproduce it on the real surface"] --> S2["2. Binary-search the root cause<br/>seed with how + why"]
S2 --> S3["3. Plan the smallest root-cause fix"]
S3 --> S4["4. Delegate implementation to a subagent<br/>named data shape + success criteria"]
S4 --> S5["5. Verify on the SAME surface<br/>the failing repro now passes"]
S5 --> S6["6. Review handoff<br/>commits ready Β· nothing pushed Β· you open the PR"]
S5 -- "still fails" --> S2
end
Code review is not a one-shot gate β it is a loop that only exits when both passes are clean:
flowchart LR
IMPL["Implement<br/>frontend / backend agent"] --> TEST["Test<br/>testing agent"]
TEST -- "β failures" --> IMPL
TEST -- "β
pass" --> REV["Code Review<br/>code-review agent"]
REV --> TECH["TECH REVIEW<br/>is the code written well?<br/>best practices Β· anti-patterns<br/>Layer 1 scans + feature & logic"]
REV --> PAT["PATTERN REVIEW<br/>does it fit the project?<br/>architecture & design + contracts<br/>boundary + architecture doctors"]
TECH --> AGG{{"Blocking issues?"}}
PAT --> AGG
AGG -- "β blocking β routed back to<br/>the responsible implementer" --> IMPL
AGG -- "β
pass" --> HF["Review Handoff<br/>typecheck Β· tests Β· build green"]
Two verdicts, either can block: Tech Review asks "is the code written well" (best practices, anti-patterns, code quality); Pattern Review asks "does the change fit the project" (follows the project's overall development rules/patterns, fits the current architecture). A blocking finding routes back to the responsible implementer β Step 1 Fix β Step 2 Retest β Step 3 Re-review β until both verdicts pass.
| Folder | Holds | Role in the system |
|---|---|---|
dispatcher.md |
The router skill | The router. Classifies tasks, holds the non-negotiables, the 13-principle index, autonomy policy, subagent defaults, reply rules, workflow registry. |
workflows/ |
8 fixed step sequences | The plans. Copied verbatim into the todo list, executed step by step. |
principles/ |
13 leaf skills | Decision rules. Read in full when applied; a citation must trace to a choice it changed. |
agents/ |
Role definitions | Who does what (pawhaven, architect, frontend, backend, code-review, ...). Each has scope and escalation. |
skills/ |
Skill folders | Domain knowledge (react, styling, i18n, code-review doctors). Triggered by description in frontmatter. |
rules/ |
Always-on constraints | Repo facts: architecture, security, testing, documentation, git. |
docs/ |
Shared reference | Product strategy, system architecture, design spec, agent-communication protocol, standards. |
memory/ (runtime, git-ignored) |
Daily memory log | Today's memory file .codebuddy/memory/YYYY-MM-DD.md records the task section, phase digests, loops, and stuck points. The next stage references it; a stalled task is traced through it. |
Each folder ships a README.md β read the folder README before diving into the files inside.
Every task, regardless of entry point (user prompt, /pawhaven-mode, or a dispatched subagent), runs the same 6-step process:
- Open the todo list. First item in every multi-step task: read the Principles section in full. Non-negotiable.
- Match a workflow. Classify the task against the workflow set (Section 4) and pick one.
- Copy workflow steps verbatim into todos. This happens before any task-specific reasoning. A step you skip stays in the list as
skip: <reason>. Silently dropping named steps is the primary failure mode this guards against. - Execute each step, firing skills and subagents. Each step produces evidence (a repro, a trace, a green check).
- Name the principles. Every reply names each principle that shaped a decision and the specific choice it changed.
- Write the reply per the mode skill's reply rules: short declarative sentences, impact framed for consumer + maintainer, failing-then-passing output pasted verbatim.
"the search page takes 2 seconds to render, find out why and fix it"
- Task matched to the perf-issue workflow.
- Workflow steps land in the todo list: Step 1 baseline β Step 2 trace β Step 3 optimize β Step 4 re-measure.
- Baseline trace captured on the real surface (devtools, profiler).
- Profiling points at a re-render cascade; the domain-model principle moves the state to its right home.
- Re-measure on the same surface: the before/after numbers prove the win.
- Reply reports the numbers verbatim, the root cause, and the regression guard.
- Verbatim-todo discipline β workflow steps land in the todo list word-for-word. Each workflow ends in an evidence-gated feedback loop: Step 1 failing repro β Step 2 pass (bug-fix), Step 1 pinned behavior β Step 2 still green (refactoring), Step 1 baseline trace β Step 2 post-fix trace (perf).
- Verification is the religion β prove-it-works: verify against the real artifact (typecheck, tests, build, rendered surface), never against "it compiles".
- Memory log (daily runtime file) β all progress goes to today's memory file
.codebuddy/memory/YYYY-MM-DD.md(git-ignored). The orchestrator appends a## Task:section per task and a digest after every stage; the next stage references it, and a stalled task is traced through it (Step 1 Stuck Log β Step 2 last Phase β Step 3 Handoff). Memory files are append-only daily records β no clearing step. - Review is two passes β Tech Review (best practices, anti-patterns) and Pattern Review (fits the project's patterns and architecture). Two verdicts, either can block; blocking findings loop back to the implementer.
- Subagent discipline β every subagent routes through the shared agent type (
agents/pawhaven.md) so it inherits the same methodology; the parent owns every subagent's work, reviews the diff itself, and writes its own summary. - Principles are decision-forcing β laziness protocol (bias to delete), subtract-before-you-add, model-the-domain, prove-it-works, never-block-on-the-human, migrate-callers-then-delete-legacy.
- Never block on the human β if the answer is observable by running something, prototype it instead of asking; reserve questions for genuine product or preference calls.
| Workflow | Entry (trigger) | Core flow (in order) | Exit (result) |
|---|---|---|---|
| Investigation | how/why/are-we-sure question | 1. read the real code β 2. run what's runnable β 3. cite evidence | cited answer, no code change |
| Bug Fix | reported defect | 1. reproduce β 2. root-cause β 3. fix β 4. verify on the same surface | PR + failing-then-passing evidence |
| Feature | new or changed behavior | 1. data shape β 2. design explore β 3. checkpoint β 4. implement β 5. verify all states | PR + cross-state verification |
| Refactoring | behavior-preserving structure change | 1. pin behavior β 2. subtract β 3. move in units β 4. pin stays green | PR + green pin |
| Perf Issue | measured slowness | 1. baseline β 2. trace β 3. optimize β 4. re-measure | PR + before/after numbers |
| Design Decision | architecture / data-model / API choice | 1. model domain β 2. enumerate options β 3. observe forks β 4. decide + tradeoff | decision + docs record |
| Architecture Change | crosses a component/package/API boundary | 1. parallel design β 2. migration wave β 3. verify boundary | PR + boundary verification + docs record |
| Review Handoff | end of most other workflows | 1. ordered commits β 2. typecheck Β· tests Β· build green β 3. handoff summary β 4. stop | verified diff, PR opened by you |
The catalog above is the one-line summary β for the full steps and evidence gates, open the workflow's own file in workflows/ (one file per workflow). This README only points; the workflow files are the detail.
- This file β orientation.
docs/agent-communication-protocol.mdβ the structured output formats every agent uses to interoperate..codebuddy/memory/YYYY-MM-DD.mdβ the daily memory log: what each stage produced, and where to look when a task stalls.docs/README.mdβ the documentation index (product strategy, architecture, design spec) for domain context.- The folder READMEs β
workflows/Β·principles/Β·agents/Β·skills/Β·rules/(each directory ships one; read it before the files inside). The router (dispatcher.md) lives at the root, introduced in Section 2 and indexed in Section 7. rules/β the always-on constraints for the current task.- The relevant agent + skill β only what the task needs (progressive disclosure).
Every skill is a folder containing SKILL.MD:
- Frontmatter (
name,description,version) βdescriptionis also the trigger: what it does AND when to use it. - Body β purpose, core rules, examples.
references/β bulky lookup material (best-practices, decision trees, forbidden patterns). Loaded on demand, not every time.
Skills cross-reference each other via a ## Related section. Do not duplicate a sibling skill's content β link to it.
The router: dispatcher.md holds the non-negotiables, the 13-principle index, autonomy policy, subagent defaults, reply rules, and the workflow registry. The shared methodology every agent and subagent inherits. The mode's non-negotiables override workflow details when they conflict β they are the floor no workflow may go below.
laziness-protocol Β· subtract-before-you-add Β· experience-first Β· outcome-oriented-execution Β· model-the-domain Β· boundary-discipline Β· make-operations-idempotent Β· migrate-callers-then-delete-legacy-apis Β· prove-it-works Β· fix-root-causes Β· sequence-verifiable-units Β· guard-the-context-window Β· never-block-on-the-human.
Entry point: SKILL.MD (orchestrator). Two passes, two verdicts:
- TECH REVIEW β is the code written well?
figma-doctorgate +typecheck-doctorΒ·react-doctorΒ·style-doctorΒ·i18n-doctorΒ·backend-doctorΒ·test-doctor(every review) + Layer 3 feature & logic deep review (best practices, anti-patterns, code quality). - PATTERN REVIEW β does the change fit the project?
boundary-doctorΒ·architecture-doctor+ Layer 2 architecture & design + Layer 4 type contract deep review (follows the project's overall development rules/patterns, fits the current architecture).
The review is adversarial: attempt to break the change before confirming it, every MUST FIX backed by evidence, no silent polishing. Either pass can block the handoff independently; blocking findings loop back to the responsible implementer.
react Β· styling Β· i18n Β· component Β· redux Β· react-query Β· react-hook-form.
Each links to its references/best-practices.md.
- Keep state local; server state β TanStack Query, client state β Redux, forms β React Hook Form.
- All styles via
@pawhaven/design-systemtokens β no hardcoded colors, no magic numbers. - All user-visible text via
t()inzh-CN/en-US/de-DE. - Components graduate to a package only when 2+ features use them.
- Communicate in the structured formats from the communication protocol.
- Follow the glossary β one term, one meaning, across the whole system.