This constitution governs the development of skillet — the SKILL.md Evaluation Toolkit — a public, open-source, multi-harness Swift CLI that turns eval-driven development (EDD) for agent skills into software.
Scope: This repository only. It covers the Swift package (EDDCore and the effectful layers
above it), the skillet CLI, the boundary file formats, the harness adapters, and all supporting
documentation and CI.
Orientation: This document defines how we build skillet. It is intentionally orthogonal
to skillet-design.md, which defines what skillet does (product principles P1–P10, settled
decisions D1–D7). When this constitution and the design doc overlap, they reinforce each other;
when they appear to conflict, the design doc governs product behavior and this constitution
governs development practice. Appendix A maps the two so neither drifts from the other.
Philosophy: skillet is a deterministic-core CLI with frozen boundary contracts, multi-harness adapters, and a never-auto-commit safety rule. Correctness of the loop, trustworthiness of measurement, and simplicity take precedence over breadth of features.
Statement: Every feature MUST begin with a specification and proceed by test-driven development: tests written first, verified to fail, then implementation makes them pass.
Rationale: Specs align work with user-facing value and provide measurable success criteria.
TDD prevents regressions, enables confident refactoring of the net-new surface (gates engine,
evidence lifecycle, next), and documents expected behavior as executable assertions.
Practices:
-
MUST create a spec for every feature before implementation; the spec describes user-facing behavior and acceptance criteria, not implementation details.
-
MUST scope each spec to a single feature or small, independently testable increment.
-
MUST write tests before implementation (red → green → refactor) and verify they fail first.
-
MUST validate boundary codecs against golden fixtures and pipelines against recorded
HarnessReplayfixtures (see Principle III). -
MUST NOT combine multiple unrelated features in one spec.
-
MUST NOT mark a task complete while its tests fail, are missing, or pass without ever having failed.
-
MUST NOT have a test wait a length of time picked so that one thing finishes before another — a margin that can be tuned. Such a test does not report whether the code works, it reports how busy the machine is, and it passes everywhere until the one machine that matters is loaded. Concretely:
- Code whose behaviour depends on time MUST accept a clock (the standard library's
Clock, defaulting to the real one), and tests MUST supply one they control, wherever that code is reachable in the same process. - Where it is not reachable — a test that runs the built binary, whose clock is beyond reach — the
stand-in it is given MUST be one that never finishes on its own (
tail -f /dev/null, not a sleep long enough to win), so there is no race to lose, and the test MUST carry a time limit so that a limit which stops firing fails the test instead of hanging it. Real time passing is fine here: it is the behaviour under test, not a guess about speed. (sleep infinityis not portable — it fails on macOS.)
Waiting until a stated condition holds, re-checking at intervals and giving up after a bound, is allowed and preferred — it returns as soon as the condition holds, so a slower machine takes longer and still gets the right answer, and reaching the bound reports that the thing never happened rather than losing a race. That is the standard replacement for a fixed wait, not an exception to this rule.
This is the standard distinction between testing against something that never answers and something that is merely slow; only the second creates a race. Measured on 2026-08-29 on one machine: the identical behaviour checked by starting a real program that idles failed 5 times in 10 under load; checked with a controlled clock and no program it passed 15 out of 15 under twice the load (Specs/020 §10, rounds fifty-three to fifty-eight, which record four consecutive attempts to make a racing check reliable, and why none of them could be).
- Code whose behaviour depends on time MUST accept a clock (the standard library's
-
SHOULD develop outside-in, from the user's (or operating agent's) perspective first.
-
MAY add property-based tests for pure functions (gates engine, scorers, aggregation).
Compliance: PRs MUST include tests written first. Specs that bundle unrelated features MUST be rejected in review.
Statement: Behavior changes MUST be driven by discovered evidence, never imagined failure modes; and any delta the tool reports MUST rest on measurement that is demonstrably sound.
Rationale: Error analysis — looking at real traces and categorizing real failures — is the highest-ROI activity in improving any LLM system (Husain & Shankar). The corresponding hazards are equally well documented: LLM judges are systematically overconfident, and evaluation criteria drift as more outputs are reviewed. skillet exists to make the trustworthy path the default; its own development MUST hold to the same standard it sells.
Practices:
- MUST ground new evals, lint rules, and fix logic in captured evidence (the corpus, friction events, triage findings), not in speculation — "codify what you discover, not what you imagine."
- MUST treat a previously-failing eval that now passes as the proof obligation for any change that claims to fix a behavior; corroboration precedes codification.
- MUST keep the scorer↔judge contradiction join correct and surfaced first: a deterministic scorer and the judge disagreeing on the same expectation outranks the raw pass rate.
- MUST version judge prompts (
judge_prompt_version) and require an explicit judge model so results are reproducible and re-gradable offline. - MUST distinguish measurement from noise: infrastructure (no-terminal-event) failures are retried and stamped; judged FAILs are never retried; trial exit class is first-class in records.
- MUST NOT let TTY presentation or a favorable-looking number substitute for a sound estimate.
- SHOULD prefer free, deterministic gates (lint, scorers) before any paid model judging.
- SHOULD guard against criteria drift by recording the critique alongside binary verdicts and preserving corrective-turn signal where the loop depends on it.
- MAY add a judge↔human-label calibration harness as the rigorous extension of this principle.
Compliance: PRs that change graded behavior MUST cite the evidence and the proving eval. Changes to judging or contradiction logic require explicit review.
Statement: The core (EDDCore) MUST be pure, synchronous, and unit-tested, with everything
probabilistic or effectful isolated behind protocols above it; and skillet's external contracts
(boundary file formats, --json payloads, exit codes) MUST be stable and enforced by tests.
Rationale: Determinism is what makes the gates engine, scorers, aggregation, and trace parsing testable rather than aspirational. Frozen contracts are what let a corpus, a downstream script, or an operating agent depend on skillet across versions. Both are load-bearing for the whole tool.
Practices:
- MUST keep
EDDCorepure and synchronous; it spawns no processes and performs no I/O beyond reading the inputs handed to it. - MUST place all model calls, process execution, network, and filesystem effects behind protocols in the effectful layers (HarnessKit, JudgeKit, etc.), each record/replayable.
- MUST treat frozen boundary formats (
evals.json,trigger-eval.json,benchmark.json, SARIF 2.1.0, session bundles) as never-break contracts, enforced by golden fixtures in CI. - MUST enumerate all decoded fields in golden fixtures and round-trip unknown keys rather than dropping them on re-encode (an omitted field is an unenforced freeze).
- MUST carry a
schemafield on every--jsonpayload; JSON changes are additive within a major version. - MUST treat exit codes as a stable API.
- MUST NOT break a frozen format, rename/remove/retype a bundle field, or change an exit code without a major version bump and migration notes.
- SHOULD add new bundle files (additive) rather than mutating existing ones.
- MAY accelerate with a gitignored cache, but the cache MUST never originate state — everything MUST remain recomputable from committed repo files.
Compliance: CI fails if codecs produce or reject anything the goldens do not. PRs touching
EDDCore MUST include or update unit/property tests.
Statement: skillet MUST be usable by humans and by AI agents without memorizing every command: human-readable TTY output paired with machine-stable contracts, discoverable affordances, errors that teach, composable plumbing, and onboarding documentation agents can act on.
Rationale: skillet is a tool for maintaining agent skills, so it should itself be operable by an agent (and by a newcomer who has not read the manual). The CLI-design canon (clig.dev) and the agent ecosystem agree: ease of discovery, consistency, robustness, and a machine-stable surface are what let both a person and a program drive a tool confidently.
Practices:
- MUST offer human-first TTY output for people AND
--json(withschema) for programs and agents on every command (P7). - MUST support
-h/--helpon every command, lead help with examples, and surface the most common flags first. - MUST support
--versionand use standard flag names where a convention exists — including--no-input(disable all prompts; fail with the flag that supplies the input) as the agent/script complement to--yes, and-n/--dry-runfor preview. - MUST end command output by suggesting the next sensible command, so the workflow is self-revealing (progressive disclosure).
- MUST make every failure message state what went wrong, why, and the exact command that fixes
it; never fail silently (the
--skill-pathfalse-negative class is extinct by construction). - MUST keep porcelain verbs thin compositions of stable, independently usable plumbing — any
porcelain action MUST be reproducible with plumbing +
--json. - MUST maintain a root
AGENTS.md(plus scopedAGENTS.mdfiles where a directory needs deltas) so an operating agent can discover commands, contracts, and boundaries without reading source — kept in sync when commands, flags, or contracts change. - MUST NOT treat human TTY output as an API or require interactive prompts on a path an agent
or script must take (
--yesMUST exist for confirmations;--dry-runMUST exist to preview). - SHOULD follow established CLI conventions for naming, exit-on-error, and signal handling.
- SHOULD stay responsive on long operations — print something before any paid or network call
and show progress on a TTY — while never animating into piped output or
--json, and page large human output through$PAGERonly on a TTY (P7). - MAY ship shell/command completion to further lower the discovery cost.
Compliance: Reviews verify --json, -h/--help, teaching errors, and AGENTS.md updates for
every new or changed command.
Statement: skillet MUST build and pass its free test suite on every supported platform in CI, behave deterministically across environments, and gate merges on automated quality checks — with free, deterministic checks running before any paid one.
Rationale: "CI in this repo is the Linux user." Cross-platform reliability and determinism are core to a tool whose entire value proposition is trustworthy, reproducible measurement; spend discipline (free-before-paid) is both an economic and a correctness safeguard.
Practices:
- MUST build and test on macOS 14+ and current Ubuntu LTS in CI.
- MUST keep the default CI path free: pure unit/property tests, golden codecs, and
HarnessReplayfixtures — no paid model calls on every PR. - MUST gate merges on build, the free test suite, and lint/format checks all passing.
- MUST ensure deterministic behavior: identical inputs produce identical outputs across platforms for all pure components.
- MUST NOT merge code that breaks any supported platform or that makes the default CI path spend money.
- SHOULD run at most one opt-in, env-gated live smoke job per adapter; everything else runs free.
- SHOULD keep boundary goldens authoritative — a CI failure there is a real contract breach, not flakiness to be silenced.
- MAY add scheduled (non-PR) jobs for heavier or paid validation.
Compliance: The CI pipeline enforces all MUST-level gates; platform or golden failures block merge.
Statement: skillet MUST protect secrets, preserve user privacy, never write to a user's skill or git history without explicit human action, and keep its dependency and process-execution surface minimal and sanctioned.
Rationale: skillet captures real production sessions (which leak credentials) and commits a corpus; a single unsanitized bundle is a real breach. A tiny, audited dependency surface and a single sanctioned way to launch processes keep the supply-chain and execution risk legible.
Practices:
- MUST sanitize secrets before writing any captured bundle — redact in place with typed markers; the raw secret never enters the repo.
- MUST fail closed: if the secret scanner cannot run, capture refuses to write rather than emitting an unsanitized bundle, and offers a remedy.
- MUST run the secret scanner detection-only and fully offline; its network validation step MUST stay disabled.
- MUST NOT emit, log, or commit secrets, private keys, or tokens — not in output, error messages, or the corpus.
- MUST NOT auto-commit, ever, and MUST NOT edit a live
SKILL.mdexcept via the explicit, opt-insuggest --applycontent-anchored path (which refuses a dirty tree and stops short of the commit).iterateoperates only in throwaway worktrees. - MUST NOT add telemetry or make network calls except to providers the user configured.
- MUST launch every subprocess through
swift-subprocess— the single sanctioned launcher; noFoundation.Process, no rawposix_spawn. - MUST confine all process execution and effects to the effectful layers;
EDDCorespawns nothing. - MUST NOT auto-probe other applications' private caches or binaries (the harness ban policy).
- MUST NOT add a new runtime or development dependency without a constitutional amendment and
explicit justification. The sanctioned build/runtime library deps are
swift-argument-parser,swift-yaml,swift-subprocess, and the standard library; andbetterleaks(MIT) is the sanctioned runtime binary — a resolved external executable, not a SwiftPM build dep — for the secret scanner, resolved from an explicit[sanitize].scanner_path/SKILLET_BETTERLEAKS_BIN/PATH(detection-only, offline; skillet never invokes its network validation). Per-platform vendoring via a SwiftPM.artifactbundlebinary target is the intended delivery and is deferred pending that packaging infrastructure (tracked as a roadmap item — secret-scanner vendoring); until it lands, a resolvablebetterleaksis a hard prerequisite forcapture, which fails closed without it, and the resolved version is recorded in each bundle'ssanitizationprovenance. The cache MAY use the system SQLite library. The sanctioned test-only deps areswift-clocks(MIT, Point-Free) and the two packages it brings with it —swift-concurrency-extrasandxctest-dynamic-overlay(both MIT, same authors) — linked into test targets only, never into a shipped target, so nothing they contain reaches a released binary. It supplies a clock whose time the test moves by hand, which is how a wait can be exercised without waiting: shipped code takes the standard library'sClockand defaults to the real one, tests pass a controlled one. Justification: checks that assert a real wait finished inside some bound are really asserting the machine was not busy, and one such check turned the shared branch red on 2026-08-27 when a 60 ms wait took six seconds (Specs/020 §10, rounds fifty-one and fifty-three). There is no standard-library orswift-testingsubstitute — that proposal was still unprioritised as of December 2025. (Amended 2026-08-27 — addition only; every MUST above is unchanged.) (Amended 2026-07-10 — delivery only; the redact-before-write / fail-closed / detection-only-offline MUSTs above are unchanged. Rationale:.artifactbundlepackaging not yet available.) - SHOULD estimate paid trials up front and confirm before spending (TTY) or require
--yes(scripts) — spend is visible and consented. - SHOULD record redaction provenance (scanner, version, count) in bundle metadata.
Compliance: Code review MUST verify secret handling, the no-auto-commit rule, and the launcher discipline. CI scans for obvious violations (logging patterns, forbidden process APIs).
Statement: skillet MUST follow open-source best practices: clear documentation, welcoming contribution and disclosure paths, explicit licensing, and simplicity over cleverness.
Rationale: skillet is a public tool whose adoption depends on legibility. Good docs and simple, readable code lower friction for both human contributors and the agents that will operate it.
Practices:
- MUST maintain a clear README with install (GitHub Releases, Homebrew tap, Mint,
swift build), setup, and usage examples. - MUST maintain
AGENTS.md(root + scoped) as first-class onboarding for AI agents and humans (cross-referenced with Principle IV). - MUST rely on the org-level contribution, security-disclosure, and code-of-conduct documents
at
21-DOT-DEV/.github; add repo-local copies only if a skillet-specific deviation is needed. - MUST include a
LICENSEfile (currently MIT — see Governance › Deferred Decisions). - MUST document every public type and command and keep documentation in sync with behavior.
- MUST apply KISS and DRY; readability over cleverness.
- MUST NOT let documentation drift from shipped behavior (stale docs are a defect).
- SHOULD make documentation verifiable rather than aspirational — the doc analogue of TDD:
code and CLI examples in
README.mdandAGENTS.mdSHOULD be exercised in CI (compiled, or run with exit-code /--jsonassertions), and internal links SHOULD be link-checked, so a broken example or claim fails the build like any other test. - SHOULD keep
AGENTS.md's "commands true now" claims machine-checked — a CI job SHOULD confirm every documented command exists and exits sanely, since an agent acting on a false claim is the highest-cost documentation defect. - MUST NOT assert exact human/TTY output as a documentation contract (Principle IV, P7):
example verification checks exit codes and
--jsonpayloads only, never the prose a command prints — otherwise doc-tests calcify output that is explicitly not an API. - SHOULD carry a review threshold on narrative docs and treat past-threshold prose as a drift finding to re-review (the prose analogue of golden-test freezing for contracts).
- SHOULD provide issue/PR templates and respond to contributions promptly and respectfully.
- SHOULD document the stability tiers (frozen formats, exit codes,
--json, command/flag semver, no-promise TTY) prominently. - MAY count experiments rather than features when planning the roadmap.
Compliance: PRs MUST include documentation and AGENTS.md updates for new or changed
behavior. Reviews enforce readability and KISS/DRY.
- MUST point reporters to the org-level
SECURITY.mdat21-DOT-DEV/.githubfor reporting instructions and contact method; add a repo-localSECURITY.mdonly for a skillet-specific deviation. - MUST honor the org's acknowledgment timeline and coordinated-disclosure commitment.
- SHOULD acknowledge reporters in release notes (with permission).
Contract- and security-relevant changes (secret handling/sanitization, boundary-format
edits, exit-code or --json schema changes, the no-auto-commit/--apply path, the harness ban
policy, dependency or launcher changes):
- MUST document the implications in the PR description (audit trail).
- SHOULD allow a brief community-review window before merge for non-trivial cases.
- All workflow state MUST be derived from committed files under
evaluations/; the.skillet/cache MAY accelerate but MUST NOT originate state. Deleting the cache MUST lose nothing. - User-level preferences (XDG config, design §5.2) are configuration inputs — like flags and environment variables — not state; they tune defaults and MUST NOT originate evidence, gate, or run state. The gitignored cache stays repo-local.
- YAML is for human-editable policy, never behavior. A surface may be expressed as YAML
only when the design §7.6 litmus test holds; otherwise it stays Swift.
skillet config setrewrites targeted lines in place (swift-yaml does not preserve comments on re-emit).
Note: This constitution defines technology-agnostic development principles. This section records current choices, which may change without a constitutional amendment unless a principle binds them (e.g., the sanctioned-dependency list in Principle VI).
- macOS 14+
- Linux (current Ubuntu LTS)
- Language: Swift 6 (strict concurrency)
- Build: Swift Package Manager
- CLI framework:
swift-argument-parser - Config / evidence frontmatter:
swift-yaml(YAML 1.2; replaces Yams; no TOML dependency) — no tagged release yet (pin by revision); itsYAMLproduct needs C++ interop, so it is confined to an isolated codec seam to keep the pure core interop-free, and wired only when that codec lands - Process execution:
swift-subprocess(sole sanctioned launcher) - Transitive production dep:
swift-system(FilePath) — rides in withswift-subprocess, whose API surfaces it; declared asHarnessKit'sSystemPackagedependency (theProcessLauncherseam) and used by theIntegrationTestsharness. Sanctioned as part of theswift-subprocessadoption — no separate supply-chain entry, so the Principle VI list is unchanged - JSON / SARIF / frozen formats: Foundation
Codable(no added dependency) - Secret scanning:
betterleaks(MIT), offline detection-only — resolved from config/env/PATH(per-platform.artifactbundlevendoring deferred; see Principle VI) - Cache: system SQLite (cache only)
- Testing: Swift Testing / unit + property tests; golden fixtures;
HarnessReplayfixtures. One unit-test target per kit (each kit testable in isolation); anIntegrationTeststarget drives the built binary viaswift-subprocess(the command surface lives in the executable). Test files carry a 300-line soft cap (split suites; extract harness/fixture helpers)
EDDCore (pure) → TraceKit → { HarnessKit, JudgeKit, ScoreKit, LintKit, CorpusKit, ProjectKit } →
AnalysisKit / RunKit / IterateKit / RenderKit → the skillet executable (ALL ArgumentParser
commands + wiring, ~≤50 lines per command). There is no separate skilletCLI library target — the
executable itself is the top wiring layer. ProjectKit owns project discovery, config I/O, and
init scaffolding (filesystem effects kept out of the executable so they remain unit-testable).
claude-code, opencode, direct-api, replay.
This constitution supersedes ad-hoc development practices. Deviations MUST be explicitly justified
and approved. The design doc (skillet-design.md) remains authoritative for product behavior;
this constitution is authoritative for development practice.
Model: 21-DOT-DEV maintainer (BDFL). The maintainer may amend this constitution directly; the community proposes changes via GitHub issues.
- Propose the amendment with rationale and impact analysis.
- Update the version per semantic versioning:
- MAJOR: Backward-incompatible governance changes or principle removals/redefinitions.
- MINOR: A new principle or materially expanded guidance.
- PATCH: Clarifications, wording, or non-semantic refinements.
- Update dependent artifacts (
.specify/templates/*once they exist,AGENTS.md, README). - Record the change in the Sync Impact Report and Version History.
- Commit with a descriptive message.
| Trigger | Action |
|---|---|
| Adding/removing a runtime or dev dependency | Full Principle VI review + amendment |
| Adding or changing a harness adapter | Capability-flag + ban-policy + visibility-contract review |
Editing a frozen boundary format / exit code / --json schema |
Principle III contract review; semver impact |
Changing secret handling, sanitization, or the --apply/no-commit path |
Security review (Principle VI) |
| Changing judging, contradiction join, or gate thresholds | Principle II review |
| New or changed command/flag | Principle IV review + AGENTS.md/README sync |
| Tier | Contract |
|---|---|
| Frozen boundary formats | Never break (golden-tested) |
| Exit codes | Stable API |
--json payloads |
Versioned via schema; additive within a major |
| Command names & documented flags | Semver: breaking ⇒ major (pre-1.0: minor + CHANGELOG notes) |
| Human TTY output | No promise — explicitly not an API |
- Pre-1.0: breaking command/flag changes land as minor bumps with CHANGELOG migration notes.
- Post-1.0: strict semver; breaking changes require a major bump.
The repo-root AGENTS.md is the operational onboarding layer for AI agents and humans; this
constitution is the development-principle layer. They are kept in sync, not duplicated:
AGENTS.mdMUST reference this constitution as authoritative for principles and MUST NOT duplicate principle text (duplication causes drift); it carries only a distilled, action-time Boundaries list and current/Planned operational guidance.- When a principle, the sanctioned-dependency list (Principle VI), a boundary contract
(Principle III), or a command/flag (Principle IV) changes, the same PR MUST update
AGENTS.md. AGENTS.mddescribes current reality; as each roadmap phase lands, its "Planned" content is promoted to verified guidance and its status banner updated.
- License: the repo
LICENSEcurrently ships MIT; design §14 Q2 recommends Apache-2.0 (patent grant; Swift-ecosystem norm). Resolve before public launch and update Principle VII. - Trademark: a sanity check against Palo Alto Networks' "Skillet" family is required before a public launch (design D7 / Appendix C).
- PR reviewers verify constitutional alignment.
- CI enforces MUST-level gates (blocking) and surfaces SHOULD-level items (warnings).
- Three-tier enforcement: MUST blocks merge · SHOULD warns and requires override justification · MAY is informational.
Version: 1.5.0 Ratified: 2026-06-18 Last Amended: 2026-08-29
Changelog:
- 1.5.0 (2026-08-29): MINOR amendment (Principle II, spec- and test-driven development) — a new MUST NOT: no test may wait a length of time picked so one thing finishes before another. Time-dependent code takes a clock and defaults to the real one; tests supply a controlled one where the code is reachable in-process, and where it is not (tests driving the built binary) the stand-in must be one that never finishes, under a time limit. This follows the standard distinction between testing against something that never answers and something merely slow — only the second is a race. Rationale is measured, not asserted: the same behaviour failed 5 of 10 runs under load when checked by starting a real idling program, and passed 15 of 15 under twice the load when checked with a controlled clock and no program. Two checks were removed and one stand-in replaced to comply; no tunable timing margin remains anywhere. No principle added, removed, or weakened. Ripples the design doc §12 and AGENTS.md; the enabling dependency was sanctioned in v1.4.0.
- 1.4.0 (2026-08-27): MINOR amendment (Principle VI, dependency hygiene) —
swift-clocks(MIT) and the two packages it pulls in transitively,swift-concurrency-extrasandxctest-dynamic-overlay(both MIT), are sanctioned as test-only dependencies: three packages, not one. All three are linked into test targets only and absent from every shipped target, so the released binary's dependency surface is unchanged — verified by building the shipped product alone into an empty build folder: their sources are fetched while the package graph is resolved, and not one of them is compiled or linked, nor does any of them leave a module behind. Shipped code now accepts the standard library'sClockand defaults to the real one; tests supply a controlled clock so a wait can be checked without waiting. Rationale: a check bounding how long a real wait took is a claim about how busy the machine is, and one turned the shared branch red on 2026-08-27; no standard-library equivalent exists. No principle added, removed, or weakened; no MUST changed. Ripples Package.swift (test targets only), RunKit's trial timer and HarnessKit's launcher timeout (both gain a defaulted clock parameter). - 1.3.0 (2026-07-10): MINOR amendment (Principle VI, secret-scanner delivery only) — the sanctioned
betterleakscompanion changes from per-platform vendored to resolved from config/env/PATH, with per-platform.artifactbundlevendoring deferred pending SwiftPM binary-artifact packaging (tracked as a roadmap item). The redact-before-write / fail-closed / detection-only-offline MUSTs are unchanged;capturefails closed whenbetterleaksis unresolvable, and records the resolved scanner version in bundle provenance. Rationale:.artifactbundleinfrastructure not yet available; ships the security control now (secure-by-default) rather than deferringcapture. Ripples design §6.1/§13 (Specs/013, Specs/014). - 1.2.1 (2026-07-01): PATCH amendment — factual syncs from the Phase-1 completed-items audit
(M4, Roadmap/phase-1-review.md §5): the Technology Stack note now records
swift-systemas a transitive production dependency (HarnessKit'sProcessLauncherseam, via the sanctionedswift-subprocess) rather than test-only;RenderKitrestored to the Package Architecture layer diagram (matches design §11); the out-of-domain "onion/control credentials" phrase removed from Principle VI's secrets rule. No principle added, removed, or changed in force. - 1.2.0 (2026-06-20): MINOR amendment — Principle IV (Human- and Agent-First, Composable CLI)
gains concrete clig.dev-derived practices: support
--versionand standard flag names including--no-input(the agent/script complement to--yes) and-n/--dry-run; stay responsive with progress on long operations; and page large output through$PAGERon a TTY. Also clarifies Repository State Discipline — user-level XDG preferences are config inputs, not workflow state (design Q7 adopted: a preferences tier in §5.2, D3 scoped to workflow state). Syncs the development charter withskillet-design.mdv0.5's clig.dev conformance reconciliation (design Appendix B). No principle added or removed. - 1.1.0 (2026-06-19): MINOR amendment — Principle VII (Open Source Excellence) gains
verifiable-documentation practices:
README.md/AGENTS.mdexamples SHOULD be CI-exercised, AGENTS.md "true now" command claims SHOULD be machine-checked, documentation verification MUST NOT assert exact human/TTY output (P7-aligned), and narrative docs SHOULD carry a re-review threshold. Establishes the documentation analogue of the project's TDD discipline without forcing literal red-green TDD onto prose. No principle added or removed. - 1.0.0 (2026-06-18): Initial constitution. Seven development-governance principles with
three-tier enforcement, 21-DOT-DEV (BDFL) governance, contract/security-relevant change
protocols, and a stability-tier table. Principles refactored after best-practice research to
give Evidence-Driven Measurement and Human-/Agent-First CLI first-class billing. Contributing /
security / code-of-conduct delegated to the org
.github; companion repo-rootAGENTS.mdauthored with a reciprocal sync contract (reference-not-duplicate).
This constitution (development practice) is orthogonal to but reinforces skillet-design.md
(product behavior). The table shows where each development principle draws on the design doc's
product principles (P) / decisions (D) and external best practice.
| Constitution principle | Design doc (product) | External best practice |
|---|---|---|
| I. Spec-First & TDD | (process; complements all) | Speckit / spec-driven dev |
| II. Evidence-Driven & Trustworthy Measurement | P4, P10; §8 contradiction join; §9.3 | Husain & Shankar (error-analysis-first, criteria drift); LLM-judge calibration literature |
| III. Deterministic Core & Stable Contracts | P2, P8; D3, D5; §7.2, §11 | Frozen-contract / golden-file testing |
| IV. Human- & Agent-First, Composable CLI | P1, P3, P6, P7; D2 | clig.dev (human-first, discovery, robustness); AGENTS.md convention |
| V. Cross-Platform CI & Quality Gates | P8, P9; §10, §11, §12 | Flaky-test hygiene; free-before-paid spend discipline |
| VI. Security, Privacy & Dependency Hygiene | P5, P9; D6; §6.1, §11, §12 | Supply-chain minimalism; secret-scanning fail-closed |
| VII. Open Source Excellence | D1; §12 | Open-source norms; roadmap-as-experiments; verifiable-docs / doctest & link-checking; docs-as-code |
Settled decisions (D1–D7) remain product invariants owned by the design doc; this constitution assumes them and does not relitigate them.