This file is the canonical cold-start brief for any AI agent
working inside a baditaflorin fleet service repo. It is maintained in
services-registry/CLAUDE.md (the registry is the catalog) and
propagated to every fleet repo via fleet-runner inject.
If you find a stale copy that differs from this one, the registry copy wins — refresh and re-propagate, don't fork it.
Before you read, triage, build, bump, or deploy ANYTHING:
git fetch origin --tagsfirst, then operate againstorigin/main(viagit show origin/main:<file>or a freshgit worktree add … origin/main). The shared/root/workspace/<repo>/working tree is whatever the last (often parallel) agent left it — routinely months stale. Acting on it ships old code, opens PRs against already-fixed bugs, and lets concurrent agents stomp each other.fleet-runner build-testruns against the working tree and is NOT freshness-correct — confirm any "failure" against a freshorigin/mainworktree before triaging. Details below under "pull often, push often." DON'T SKIP THIS.
Per-service specifics (port, mesh, slug, version, category) live in
the repo's own service.yaml + deploy.yaml + README.md. This file
is intentionally generic — it explains the fleet, not any one
service.
Building a new service? See
services-registry/SERVICE-TEMPLATE.md — the
canonical per-service scaffold (file-by-file templates for main.go,
service.yaml, Dockerfile, etc., plus a paste-ready cold-start
prompt you can feed Claude / ChatGPT / Gemini). Propagated to every
fleet repo next to this file.
Most changes are low-risk: the fleet's own tooling (probe-first
deploy, rollback-on-/selftest-fail) is the safety net, so a plain
fleet-runner deploy <repo> is enough. A few categories aren't
covered by that net — check the relevant section before you act,
not after something breaks fleet-wide:
| Touching... | Do this first |
|---|---|
One go_<thing> service, no shared code |
Just fleet-runner deploy <repo> — you're covered. |
safehttp, middleware, auth, or the gateway |
fleet-runner canary <repo> first (bake + structured health/latency verdict) — see "Fleet-wide changes — modify 130 repos at once" in FLEET.md. Never a fleet-wide rollout on first touch here. |
go-common or a go-fleet-* primitive (has >1 consumer) |
See "Fleet-wide changes — change go-common, not consumers" below — one bad edit breaks every consumer at once. |
| DNS records or secrets | Read RUNBOOK-UNATTENDED.md first. DNS only via Hetzner Cloud API (HCLOUD_TOKEN) — never dns.hetzner.com. Secrets only live in go-fleet-secrets — never in env, repos, or services.json. |
proxy_egress in overrides.json |
Read the proxy_egress section in FLEET.md first — it's not simply "on = safer": some upstreams 403 the proxy IPs, some services need it on to get an internet route at all. Direction depends on the upstream. |
| A SQLite-backed service | The three mandatory rules under "SQLite safety" below aren't optional — read them first. |
Not sure which tier something is? Default one tier higher and reach
for canary before a broad rollout.
~220 service repos under github.com/baditaflorin/*. The canonical
catalog is services-registry/services.json; the canonical
conventions doc is services-registry/FLEET.md — read it first
for any fleet-wide task.
services.json is ~280 KB / ~250 entries / ~26 fields each. If you
only need IDs, names, ports, TRL, or URLs, fetch a slice instead
— it's the same raw.githubusercontent.com path with a different
filename. Sized for AI agents on a token budget.
| URL suffix | shape | size | use when |
|---|---|---|---|
services.ids.json |
["a11y-quick", …] |
~5 KB | "what services exist?" |
services.names.json |
[{id, name}] |
~13 KB | pickers / menus |
services.minimal.json |
[{id, name, mesh, kind, category, language, trl, url}] |
~44 KB | catalog overview |
services.urls.json |
[{id, url, health_url, example_path, auth_help}] |
~63 KB | building Open / smoke links |
services.trl.json |
[{id, trl, trl_ceiling, trl_assessed_at, …}] |
~31 KB | TRL audits |
services.ports.json |
[{id, host_port, container_port}] |
~12 KB | port allocation / conflict checks |
services.deploy.json |
[{id, mesh, kind, runtime, language, repo_url}] |
~40 KB | fleet-runner deploy targeting |
Base URL: https://raw.githubusercontent.com/baditaflorin/services-registry/main/<file>.
Only fall back to the full services.json when you need fields the
slices don't carry (auth surface details, vhost knobs, descriptions).
Slices are derived — never edit them; edit services.json /
overrides.json and run python3 bin/generate.py.
Every entry in the registry has three orthogonal classifying axes.
Don't conflate them — agent tooling gates behavior on kind, not on
mesh.
kind |
What it is | Has port? | /health? |
Workspace on LXC? | Bumpable version? | Counted in fleet-runner health / smoke / deploy? |
|---|---|---|---|---|---|---|
container |
Docker service on the dockerhost | yes | yes | yes | yes | yes |
static |
Static GitHub Pages site | no | no | no | no | no — has its own fleet-runner pages-audit |
If this repo's service.yaml (or registry entry) says kind: static,
stop looking for a Dockerfile, a port, or Go code. Pages services
are HTML/CSS/JS published by GitHub Pages CI — there is no container
to deploy and no /health to probe.
mesh |
Domain pattern | Auth | Typical contents |
|---|---|---|---|
mesh-0exec |
<slug>.0exec.com |
?api_key=… or X-API-Key header — keystore-gated |
proxy, search, ocr, security |
mesh-0crawl |
<slug>.0crawl.com |
Authorization: Bearer / X-API-Key / ?api_key=… — keystore-gated (same auth surface as 0exec) |
domains, recon, web-analysis |
mesh-pages |
*.github.io / custom |
none (static) | dashboards, catalogs, browser-only WASM apps |
Both container meshes are gated by the same keystore (see auth
section below). One revoke = killed everywhere. The 0crawl path-token
shape is preserved as a backwards-compat alias and feeds into the
same auth_request flow on the nginx side.
runtime |
What it means |
|---|---|
compose |
Default for kind: container. Docker-compose on the dockerhost; deploy = docker compose pull && up -d |
systemd |
Reserved — a service unit on a host; deploy = systemctl restart |
binary |
Reserved — a static binary run by hand or by a launcher |
k8s |
Reserved — managed by a kube manifest |
github-pages |
Default for kind: static. Built and served by GitHub Pages CI |
external |
Reserved — runs outside the fleet, included for reference only |
runtime is orthogonal to language. A Go service might be runtime: compose today and runtime: systemd tomorrow without re-classifying it as a different language or kind. fleet-runner deploy dispatches on runtime.
language |
When to use it |
|---|---|
go |
Default for kind: container in this fleet |
node |
Node.js services (a handful of proxies + Bing/Duck SERP scrapers) |
python |
Python services (currently 1: python-proxy) |
c |
C services (currently 1: c-proxy) |
rust |
Reserved for future use |
html |
Default for kind: static — plain HTML/CSS/JS Pages sites |
wasm |
Static Pages site whose primary payload is a WASM binary |
other |
Anything that doesn't fit |
fleet-runner --filter language=go converge (or --filter kind=container,language=go update-dep …) narrows bulk operations so
a Go-only dep bump never touches a Node, Python, or static service.
Look at service.yaml in this repo to see which axes apply.
Every services.json entry may carry a trl field 1–9:
| TRL | Band | Meaning |
|---|---|---|
| 1–3 | toy | single regex / no tests. Don't depend on it. |
| 4–5 | developing | curated lists, multi-step logic, partial tests. |
| 6–7 | real | RFC-compliant parsing, evidence trails, real test coverage. |
| 8–9 | production | battle-tested, cross-checks, SLA-grade. |
trl_ceiling marks services that structurally cannot advance
further (e.g. needs a browser engine, needs paid threat intel).
trl_assessed_at older than ~90 days is stale — re-audit.
| Repo | Role | Visibility |
|---|---|---|
services-registry |
canonical catalog (services.json + FLEET.md + this file) | PUBLIC |
go-common |
shared Go lib — SSRF-safe HTTP, jsbundle recovery, apikey client, ua, middleware | PUBLIC |
mesh-common |
shared TS/React runtime for the mesh-* P2P fleet (see "mesh-* P2P fleet" below) |
PUBLIC |
go-fleet-persona |
cross-app + cross-origin display-identity service (persona.0exec.com) |
PUBLIC |
go-apikey-service |
the keystore — issues/verifies/revokes API keys for mesh-0exec |
varies |
go-catalog-service |
renders services.json into catalog.0exec.com |
PRIVATE |
go_fleet_runner |
CLI to operate the fleet (health, smoke, inject, push, …) |
PRIVATE |
0crawl-platform |
nginx vhost templates (also embedded in fleet-runner) | PRIVATE |
fleet-state |
live operational state, runbooks, SSH topology | PRIVATE |
The mesh-* repos under baditaflorin/* are a distinct fleet from the
0exec/0crawl container services described above. They are browser-only,
rootless WebRTC apps published as static GitHub Pages sites (kind: static,
mesh: mesh-pages). They do not have container images, ports, /health
endpoints, or keystore-gated auth.
Each app depends on @baditaflorin/mesh-common via file:../mesh-common
so the bundle is fully self-contained on every GH Pages deploy — no npm
publish step, no runtime dependency on a registry.
Single source of truth for the look, feel, and capabilities of every
mesh-* app. Versioned via package.json and CHANGELOG.md; consumers
re-bundle on the next npm run build.
Headline primitives (current as of 0.10.x):
| Module | What it does |
|---|---|
MeshShell |
App chrome: ⚙ settings FAB + drawer, 📡 invite QR FAB, self-ref bar, beacon |
SettingsDrawer |
Room id + signaling/TURN overrides; injection slot for per-app extras |
createMeshConfig |
One-call config factory (app name, accent, version, signaling/TURN defaults) |
useYRoom |
{doc, provider, peerId, peerCount} for a Yjs room over WebRTC |
clockSync |
NTP-over-Yjs offset → mesh-median time (~10–30 ms stable) |
commitReveal |
SHA-256 commit/reveal for anonymous votes, fair RNG, role assignment |
identity + tofuRegistry |
Ed25519 keypair + TOFU pinned-pubkey registry (per-room crypto identity) |
moderator + ModeratorBadge |
Signed first-claim-wins role with 30-min auto-expire |
PersonalQR / QRExchange |
Inline-SVG QR (real-URL payload) + camera scanner |
useAwareness |
Typed wrapper around y-protocols/awareness (presence / cursors / typing) |
PeerAvatar |
Deterministic SVG avatar from peerId / pubkey — zero network, zero PII |
useTypedMap / useTypedArray |
Zod-validated Y.Map / Y.Array — hostile peers' writes filtered at the edge |
useRoomSeal / deriveRoomKey |
Room-wide AES-GCM seal via PBKDF2(passphrase, roomId) — opt-in E2E |
MeshErrorBoundary |
Drop-in crash containment for the <Feature> subtree |
useMeshLink |
Typed encoder/parser for the #r=…&p=…&x=… deep-link fragment |
useMultiRoom |
Run several Yjs rooms in one tab (facilitator dashboards, embeds, side-by-side) |
usePresenceCursors |
Figma-style live cursors built on useAwareness |
useThreadedMessages |
Y.Map<msgId, {parent, body, by, at, sig}> with post() / reply() |
useReadReceipts |
Per-peer monotone "last seen at message N" |
useOfflineQueue |
Buffer writes when isolated; replay through flush() on reconnect |
useFileShare |
Chunked file share through the Yjs transport |
SafeMarkdown |
Allow-list-sanitised Markdown via marked (no raw HTML pass-through) |
useFakeTime |
Test-only clock fixture; production collapses to Date.now() |
useFleetPersona |
Cross-app + cross-origin display identity (nickname + name + avatar) |
FleetAvatar |
Drop-in avatar for the current fleet persona; reuses PeerAvatar |
FleetIdentityPanel |
Drop-in settings UI; auto-mounted inside MeshShell by default in 0.10.1+ |
useFleetPersona({ appName, serviceUrl? }) resolves a FleetPersona
(nickname + name + avatarSeed + avatarVariant + paletteIndex) through
three tiers, falling back gracefully if any is unavailable:
| Tier | Where | Notes |
|---|---|---|
| L0 | per-app localStorage |
Always wins once the user types something |
| L1 | same-origin localStorage |
Free on GH Pages — every mesh-* app under baditaflorin.github.io shares one origin |
| L2 | https://persona.0exec.com |
Optional cross-origin sync; 2 s fetch timeout; fire-and-forget; service-down → silent |
The L2 fetch never blocks the UI; if the service is down or slow, L0/L1
keep the user's identity intact. Writes from app code propagate down the
stack respecting the per-app mode setting (off / local-fleet /
remote-fleet).
Field validation: strict ASCII allowlist ^[A-Za-z0-9_\- .]{1,32}$ on
every text field — identical on client and server, neutralises a class
of stored-XSS / homoglyph concerns. anonId and writeToken are
128-bit hex; writeToken is held only inside the module's closure so
app code cannot accidentally exfiltrate it.
Cross-origin handoff: fp.buildHandoffUrl(targetOrigin) returns a URL
with #fp=<base64> that carries anonId + writeToken + persona. Open
on the new origin → useFleetPersona auto-consumes the fragment on
mount, wipes the hash from the URL, and the new origin has the same
identity. Works with PersonalQR for cross-device transfer.
MeshShell mounts FleetIdentityPanel inside the settings drawer by
default in mesh-common ≥ 0.10.1. Apps that want to opt out (kiosks,
apps that own their own identity flow) pass fleetIdentityServiceUrl={null};
apps that want a staging endpoint pass their own URL.
Public URL: https://persona.0exec.com (canonical) and
https://fleet-persona.0exec.com (alias).
| Aspect | Value |
|---|---|
| Registry id | fleet-persona |
| Mesh / kind / runtime | mesh-0exec / container / compose |
| Port | 18209 (host) → 18209 (container) |
| Image | ghcr.io/baditaflorin/fleet-persona:<sha> (cosign-signed) |
| Auth | none — public read by design; writes argon2id-gated by client-held writeToken |
| Storage | pure-Go SQLite (modernc.org/sqlite) on a single docker volume |
| Tests | 21 Go unit + handler tests; testing/smoke.sh runs the full lifecycle end-to-end |
Wire surface:
GET /v1/health → 200 { ok, ver, ts, ro }
GET /health → same (fleet canonical alias)
GET /selftest → 200 { ok, ver, checks[] } (200 = all green, 503 = any fail)
GET /v1/persona/{anonId} → 200 { nickname, name, avatarSeed, avatarVariant, paletteIndex, updatedAt } | 404
PUT /v1/persona/{anonId} → 204 body: { …persona, writeToken }
DELETE /v1/persona/{anonId} → 204 body: { writeToken }
CORS is * by design — the read API is public; anonId itself is the
only read-secret (and display names are non-PII). Writes require a
separate 128-bit writeToken argon2id-hashed at rest.
Privacy stance: no IP addresses stored (only short-lived hashes in rate-limit buckets, daily-rotated salt); no User-Agent / Referer / app name kept; no mapping that lets the operator reconstruct which apps a browser opens.
Rate limits: 60 reads/min/IP and 30 writes/hour/(IP, anonId) by default;
4 KB body cap; global read-only kill switch via -read-only flag.
kind: static— no Dockerfile, no/health, no keystore. The ⚙ FAB rendersSettingsDrawerwhich auto-mountsFleetIdentityPanelonmesh-common ≥ 0.10.1. No per-app code needed to surface the panel.- HTML/CSS/JS published from
docs/. Build locally withnpm run build; commitdocs/and push. GH Pages serves it within ~1 min. - Pre-commit hook runs
bash scripts/smoke.sh(vitest unit tests +vite build+ sanity-checkdocs/index.html). Never--no-verify; if it fails, runnpm run fmtand re-commit. useFleetPersonais the canonical way to read/write the user's display name — never roll your ownmyNamelocalStorage key for new apps. Existing apps can migrate at their own pace; the panel is opt-in for the user via the sync-mode radio.
The keystore is the fleet's single point of compromise. Treat it
like a CA root: every 0exec and 0crawl service trusts whatever it
says. If this repo is on mesh-pages (i.e. kind: static), the
keystore does not apply — skip this section.
Three canonical request shapes (every mesh, every service):
Authorization: Bearer <key>— production canonical, what every SDK uses.X-API-Key: <key>— legacy header alias, same handler.?api_key=<key>— demo / browser-playground only (key leaks in logs).
A fourth legacy shape, https://<slug>.0crawl.com/t/<token>/..., was
deprecated on 2026-05-14. The gateway returns 410 Gone with
Location: /<rest>?api_key=<token> and a Deprecation header for any
caller still using it. After one deprecation cycle (~2026-06-14) the
410 block will be removed; /t/<anything> will return plain 404.
Request flow at the gateway:
- nginx vhost captures the key into
$api_key_in(Bearer regex → X-API-Key header → ?api_key query, in that order). Static fallback— sunset 2026-08-22 (security risk). This step previously accepted the universal demo key ($default_token, from/etc/nginx/conf.d/_default_token.conf) immediately and setX-Auth-User: demo, surviving keystore outages for the public demo path. A static, undifferentiated, rate-limit-only gate in front of every service was judged too broad a bypass and has been removed from the gateway.$api_key_in == default_tokennow falls through to step 3 like any other value and gets a normal 401 from the keystore. There is currently no public, unauthenticated demo path — every caller needs a real keystore-issued key. Don't referencedefault_tokenas a working example in service docs; if you find one, fix it the same way this passage was fixed (mark it sunset, point at real auth) rather than leaving it looking live.- Otherwise nginx POSTs
X-Verify-Key: $api_key_into the keystore's/verifyviaauth_request. - Keystore checks SQLite → returns 200 +
X-Auth-User/X-Auth-Scope, or 401. - On 200, nginx forwards the original request to the service container
with
X-Auth-*headers ANDX-API-Key: $api_key_inpopulated, so the upstreammiddleware.TokenAuthKeystoresees a positive auth signal regardless of which gateway auth path was taken.
Services do not call the keystore themselves — nginx already gated
the request. Trust the gateway-injected X-Auth-* headers. If you
genuinely need verification inside a service (admin tooling, internal
RPC), use the canonical clients — never handroll HTTP calls:
// Middleware (preferred — gateway header fast-path + keystore fallback + Cache + fail-closed 503):
import "github.com/baditaflorin/go-common/middleware" // ≥ v0.7.0
// Direct client (only for non-HTTP-handler code):
import "github.com/baditaflorin/go-common/apikey"
c := apikey.New() // reads APIKEY_SERVICE_URL + APIKEY_SERVICE_ADMIN_TOKEN
verifier := apikey.NewCache(c) // 15-min positive cache, no negative cache
result, err := verifier.Verify(ctx, userKey)Keystore outage behaviour (designed-in graceful degradation):
- Static fallback in nginx keeps the public demo key working.
apikey.Cachein each service keeps recently-verified callers working ~15 min.- Snapshot data in
fleet-state/state/snapshot.jsonflags the keystore as BROKEN once/healthfails — that's the alert. - Recovery procedures:
- WAL stuck readonly (HTTP 409 "attempt to write a readonly database"):
public —
go-apikey-service/docs/recovery-keystore-readonly-wal.md. - Full keystore outage / data wipe: private
fleet-state/RUNBOOK.mdunder "keystore outage".
- WAL stuck readonly (HTTP 409 "attempt to write a readonly database"):
public —
The admin token (X-Admin-Token on /issue, /revoke, /list,
/purge) is stored as ADMIN_TOKEN on the keystore container and
read by clients from APIKEY_SERVICE_ADMIN_TOKEN. Rotation playbook:
private fleet-state/OPS.md.
The section above covers INBOUND keystore auth (how a service authenticates its callers). For OUTBOUND auth — how a service identifies itself when calling another fleet service — see ADR-0027 — Fleet authentication canonical flow.
In one line: the canonical bootstrap is
fleet-runner key provision <slug> (atomic: issue keystore key +
write /opt/services/<slug>/.env on dockerhost + docker compose up -d).
Audit at rest with fleet-runner audit fleet-auth-scope —
flags services on default_token (will silently 401 against vault).
Code-side guard: every service that does outbound calls to a fleet
sibling MUST use apikey.MustResolveCritical(slug, "FLEET_API_KEY")
in main.go. The binary fail-fast-exits if FLEET_API_KEY is
empty, default_token, or has an unknown prefix — surfaces what
would otherwise be a silent run-time 401.
fleet-runner deploy pins the dockerhost compose to
ghcr.io/baditaflorin/<repo>:<short-sha> (the git short sha of the
commit the binary was built from). :<version> and :latest still
get pushed for human-readable discoverability but production traffic
only touches :<sha>. Drift detection reduces to "does the pin
match origin/main HEAD" — see
ADR-0028 — Image tagging + version-bump policy.
Legacy :latest / :<semver> pins on the dockerhost are auto-migrated
to :<sha> on the next fleet-runner deploy <slug> invocation, so
existing services flip over organically as they get touched.
Sunset on 2026-05-14. The gateway returns 410 Gone with
Location: /<rest>?api_key=<token> and Deprecation: version="v1".
Any SDK or client still using /t/<token>/... should follow the
Location header to the canonical shape. The 410 block itself will
be removed in the following deprecation cycle; after that
/t/<anything> returns 404.
Defense in depth: go-common/middleware v0.11.0 dropped path-token
extraction from extractToken, so even a caller bypassing the gateway
and hitting an upstream container directly with /t/<token>/... will
not be authenticated. The only paths that work are the three canonical
auth shapes documented above.
| Package | Import path | Purpose |
|---|---|---|
| safehttp | github.com/baditaflorin/go-common/safehttp |
SSRF-safe HTTP client, DNS-rebind guard |
| ua | github.com/baditaflorin/go-common/ua |
Standard User-Agent builder |
| jsbundle | github.com/baditaflorin/go-common/jsbundle |
source-map recovery for scanning JS bundles |
| apikey | github.com/baditaflorin/go-common/apikey |
keystore client (Verify, Cache, admin endpoints) |
| middleware | github.com/baditaflorin/go-common/middleware |
TokenAuthKeystore HTTP middleware (≥ v0.7.0) |
| loadshed | github.com/baditaflorin/go-common/loadshed |
non-blocking concurrency gate: cap calls to a slow upstream, fast-503 the excess (loadshed_shed_total) (≥ v0.65.0) |
import (
"github.com/baditaflorin/go-common/safehttp"
"github.com/baditaflorin/go-common/ua"
)
client := safehttp.NewClient(
safehttp.WithTimeout(10*time.Second),
safehttp.WithUserAgent(ua.Build(ServiceID, Version)),
// v0.16.0+ : auto-emit traces, auto-consult backoff, auto-tag degraded[]
safehttp.WithTraceCollector(os.Getenv("CALL_TRACER_URL")),
safehttp.WithBackoffCoordinator(os.Getenv("BACKOFF_COORDINATOR_URL")),
safehttp.WithDegradedSink(°raded),
)
// Errors: safehttp.ErrBlocked, safehttp.ErrInvalidScheme, safehttp.ErrMissingHost
u, err := safehttp.NormalizeURL(rawInput)safehttp is for OUTBOUND (public-internet) HTTP only. Calls to sibling
fleet services on the private docker mesh (http://go-fleet-*:<port> /
http://<slug>:<port>) MUST use a plain net/http.Client, NOT safehttp
— the SSRF guard correctly rejects private IPs and will return
safehttp.ErrBlocked immediately. Use a short timeout (1-3s),
fail-open semantics, and append "<primitive>-down" to a per-request
degraded []string slice on env-unset / timeout / 5xx. Surface
degraded[] in the response JSON envelope. Confirmed pattern across
every Phase 3 consumer migration (2026-05-17, ADR-0024) — every agent
that tried safehttp for intra-mesh calls hit ErrBlocked and pivoted to
plain http.Client.
Selftest and policy rule engines also live in go-common (v0.17.0+ /
v0.18.0+):
| Package | Import path | Purpose |
|---|---|---|
| selftest | github.com/baditaflorin/go-common/selftest |
Canonical /selftest suite consumed by go-fleet-selftest-aggregator |
| policyeval | github.com/baditaflorin/go-common/policyeval |
Small in-Go rule DSL: (fact, []Rule) -> decision + explanation — replaces ~5 custom rule engines |
- Port: from
PORTenv; fallback to a build-time constant; must matchservice.yaml, compose, anddeploy.yaml. - Health:
GET /health→{"status":"ok","service":"<id>","version":"<ver>"}. - Version:
GET /version→{"version":"<ver>"}. - Metrics:
GET /metrics(Prometheus). - Gateway health:
GET /_gw_healthis added by the nginx template, not by the service — don't re-implement. - User-Agent:
ua.Build(ServiceID, Version). - Docker image:
ghcr.io/baditaflorin/<id>:<version>(novprefix on the tag). - Tagging:
git tag <version>(novprefix), e.g.1.2.3. - service.yaml must keep:
id,name,version,port,category,healthblock,testblock.
The catalog is services.json (auto-derived). Per-service hand-curated
patches live in services-registry/overrides.json. Two shapes coexist:
Per-slug patches (current shape, unchanged):
{
"python-proxy": { "proxy_read_timeout": "300s", "trl": 6 },
"node-search-bing": { "vhost": { "proxy_buffering": "off" } }
}Per-slug container resource caps (cpus, mem_limit, pids_limit).
The fleet-wide backstop lives in host-conventions.yaml
(container_defaults, default cpus: 2.0 / mem_limit: "1g" /
pids_limit: 512). Heavy/browser services raise them per-slug in
overrides.json, which wins over mesh_defaults and
container_defaults (precedence rule #4):
{
"infrastructure-fetch-cache": { "cpus": 4.0, "mem_limit": "3g" },
"html-proxy": { "cpus": 8, "mem_limit": "6g" }
}cpus is a number (e.g. 4 or 4.0); mem_limit is a docker size
string (e.g. "3g"). These are honored by fleet-runner render-compose
(and deploy --render-compose) only with fleet-runner ≥ the version
from go_fleet_runner PR #82 — earlier binaries silently rendered the
host-conventions.yaml default. A service that hand-maintains its own
caps in its base docker-compose.yml opts OUT of overlay injection with
"compose_self_managed": true.
Operational gotcha: a live
docker update --cpus Non a running container reverts to the rendered value on the nextrender-compose/deployunless the per-slugcpusoverride is set inoverrides.json. Declare the cap there to make it durable.
Bulk rules (new, via reserved $rules key):
{
"$rules": [
{
"name": "phone-extractor-san-cert",
"match": { "mesh": "0crawl", "ids": ["a11y-quick", "broken-links", "…"] },
"patch": { "cert_domain": "phone-extractor.0crawl.com" },
"why": "46 vhosts share phone-extractor's SAN cert"
}
]
}Match clauses: any of ids (explicit list), mesh, kind, language,
runtime, category — combined with all-of semantics. Rules apply in
declaration order; per-slug entries win. Use rules to encode "47 services
share this cert_domain" as one line instead of 47.
Multi-service repos (via reserved $expand key) — one GitHub repo
emits N catalog entries when a compose project ships multiple
independently-addressable services (different host_ports, different
*.<mesh>.com hostnames) inside one repo. Each child gets its own
slug + url + host_port; $rules and per-slug overrides re-apply on
top. The parent's topic-derived entry is dropped iff
replace_parent: true. When the children later split into their own
repos with their own mesh-* topics, drop the $expand entry and
the per-repo topic-derived entries take over with the same slugs.
{
"$expand": [
{
"name": "go-fleet-metrics-hub-children",
"parent_repo": "go-fleet-metrics-hub",
"replace_parent": true,
"children": [
{ "id": "fleet-discovery", "host_port": 18201, "container_port": 8080, "category": "observability" },
{ "id": "fleet-grafana", "host_port": 18202, "container_port": 3000, "category": "observability" },
{ "id": "fleet-prometheus", "host_port": 18203, "container_port": 18203, "category": "observability" }
],
"why": "one compose project, three host_ports — register all so allocate-port sees them"
}
]
}External / third-party containers (via reserved $external key) —
third-party upstream containers (Plausible, …) that run on the
dockerhost but are NOT in the fleet repo set: no mesh-* GitHub
topic for the generator to latch onto, no fleet-runner ownership of
the compose file. Without a registry row their host_port is
invisible to allocate-port — and the next deploy will silently
re-claim it (the 2026-05-19 plausible incident). One row here keeps
the allocator honest. Defaults to runtime: external, which gates
fleet-runner deploy/build off so the upstream compose dir stays
untouched. $rules are NOT applied to $external entries (fleet
knobs like cert_domain / proxy_egress have no meaning for an
upstream-managed container). See ADR-0031.
{
"$external": [
{
"id": "plausible",
"name": "Plausible Analytics",
"host_port": 18204,
"container_port": 8000,
"repo_url": "https://github.com/plausible/community-edition",
"auth": { "type": "none" },
"external_compose_dir": "/opt/services/plausible/",
"external_image": "ghcr.io/plausible/community-edition:v3.2.1",
"why": "fleet-pixel's optional upstream; runs at /opt/services/plausible/ — register so allocate-port sees 18204 as taken"
}
]
}Audit surface — never grep overrides by hand:
fleet-runner overrides list [--filter mesh=0crawl] [--key cert_domain]
fleet-runner overrides explain <slug> # full breakdown per key + source
fleet-runner overrides audit # stale slugs, unused rules, key adoption counts
fleet-runner converge also surfaces overrides drift (stale per-slug
entries that reference removed services; rules with no matching
service).
The fleet has one TEMPORARY workaround active. It is tracked, has a separate fix in flight, and must be removed from this doc when the underlying issue ships. Treat it as a known degradation, NOT "this is fine".
Three known bugs: it double-prefixes the service name, references an
undefined safehttp.CheckURL, and --push doesn't actually push.
Separate fix in flight against go_fleet_runner. Until that fix ships
AND the binary on Builder LXC 108 is updated, the canonical scaffold
path is copy-from-peer:
# 1. Copy a recent successful peer in the same mesh + language.
cp -r go_domain_amp_detector go_domain_<new>
cd go_domain_<new>
rm -rf .git
# Search/replace identifiers (id, name, slug, port, description).
# Allocate the port first (via the fleet-runner shim, or the explicit
# bastion form documented in "How to invoke fleet-runner" below):
fleet-runner allocate-port --count 1
# 2. Init + create the GitHub repo + push.
git init && git add -A && git commit -m "initial scaffold from go_domain_amp_detector"
gh repo create baditaflorin/<repo> --private --source=. --remote=origin \
--description "<one-line description>" --push
# 3. Tag the first version.
git tag 0.1.0 && git push origin 0.1.0
# 4. Add the canonical GitHub topics so bin/generate.py picks it up.
gh repo edit --add-topic mesh-0crawl \
--add-topic kind-container \
--add-topic language-go \
--add-topic runtime-compose \
--add-topic category-<cat>Remove this section when fleet-runner new-service is fixed and the
LXC 108 binary is updated.
Binary at /usr/local/bin/fleet-runner on Builder LXC 108. From
any workspace dir on that LXC:
fleet-runner health [--insecure] # /health on all live container services (skips kind=static)
fleet-runner smoke [--insecure] # GET example_url on all container services
fleet-runner pages-audit # verify pages_url 200s for every kind=static entry
fleet-runner build-test # go test ./... in every kind=container,language=go workspace
fleet-runner update-dep <mod@ver> # bump dep across all language=go repos (or a subset: --repos a,b / --filter mesh=…,category=…,ids=a;b)
fleet-runner rollout --dep <mod@ver> [--grep REGEX] [--graph-callers-of S,…] [--depends-on S,…] [--clone] [--apply [--pr]] # blast-radius: discover EVERY affected service (grep ∪ graph ∪ depends_on), bump + build/test, land only the green (plan-only by default)
fleet-runner deploy-all # redeploy a filtered set (--repos a,b / --mesh / --framework); honors the exclude list (won't touch infra)
fleet-runner inject <src> <dest> # copy a file into every repo (still all kinds, on purpose)
fleet-runner exec "<cmd>" # shell command in every repo (filterable)
fleet-runner push "<msg>" # commit+push all dirty repos
fleet-runner nginx-render # regenerate vhosts from templates
fleet-runner rotate-default-token <value> # gateway-only rotation, zero repo edits
fleet-runner default-token # print the current gateway default token
fleet-runner overrides list # per service, which override keys apply (and via which rule)
fleet-runner overrides explain <slug> # one service: every override key and its source (slug vs rule)
fleet-runner overrides audit # stale per-slug entries, unused rules, per-key adoption counts
fleet-runner new-service <name> <port> [cat] # scaffold new service
fleet-runner scaffold-compose <repo>... # write canonical docker-compose.yml from registry (--apply [--commit --push] | --missing for every container repo lacking one)
fleet-runner scaffold-service-yaml <repo>... # write canonical service.yaml from registry
fleet-runner render-compose <repo> # print canonical docker-compose.yml to stdout (no write — read-only preview)
fleet-runner audit compose-drift # surface compose files diverging from the canonical fleet shape
fleet-runner audit registry-host-port-set # services without a registered host_port (silent squatter discovery)
fleet-runner stats # audit log + token usage summary
Bootstrap a service that's missing its compose at HEAD — the canonical
two-liner (replaces hand-rolled cat > docker-compose.yml heredocs that
got the shape wrong in the past):
fleet-runner scaffold-compose <repo> --apply --commit --push
fleet-runner deploy <repo> --bootstrap --force-build --skip-smokescaffold-compose --missing will find every container repo without a
compose at HEAD and render one for each in a single pass.
All commands accept --filter kind=container,language=go (and so on)
to narrow the set. The filter axes are mesh, kind, language,
category, and ids (the surgical "just these repos" axis — list
members separated by ; since commas separate k=v pairs). For
update-dep and deploy-all there's also a dedicated --repos a,b,c
convenience flag (accepts registry ids OR workspace dir names). So a
two-repo dep bump no longer needs hand-edits:
fleet-runner update-dep --push --repos go_foo,go_bar <mod@ver>.
update-dep keeps Kind=container + Language=go defaults, so a targeted
bump never reaches a static/non-Go repo, and the exclude list always
wins over an explicitly-named repo (you cannot --repos your way into
redeploying the keystore). All commands accept --tokens-used N --model NAME for LLM accounting. kind: static entries are skipped
by default on every container-shaped operation — don't try to deploy
or health-check a static Pages site.
| Target | SSH |
|---|---|
| Bastion | ssh root@0docker.com |
| Builder LXC 108 | ssh root@0docker.com 'pct exec 108 -- bash -lc "<cmd>"' |
| Dockerhost VM | ssh -J root@0docker.com ubuntu_vm@10.10.10.20 |
| Webgateway | ssh -J root@0docker.com florin@10.10.10.10 |
-
Builder LXC 108 is a Proxmox container on
0docker.com. Hosts per-service build workspaces at/root/workspace/<repo>/, thefleet-runnerbinary, and (pilot, ADR-0035) Woodpecker CI server+agent at/opt/woodpecker/—docker compose psthere to check status;.envholds the generated secrets, not committed anywhere.AI-agent rule — always use a git worktree, never the shared workspace directly. Multiple AI sessions (or a session + a human) routinely target the same repo concurrently; sharing
/root/workspace/<repo>/produces silent races (one session'sgit checkout/reset --hardclobbers the other's working tree mid-build, image tags get pushed in the wrong order, deploys flip to the loser's commit). Each session must isolate its checkout:cd /root/workspace/<repo> git fetch origin git worktree add /root/wt/<repo>-<short-purpose> origin/<branch> cd /root/wt/<repo>-<short-purpose> # do work, build, push, then: git worktree remove /root/wt/<repo>-<short-purpose>Worktrees share the same
.git(cheap; no extra clone), but each has its ownHEAD, working tree, andgit status. The shared/root/workspace/<repo>/stays as the canonical "long-lived upstream tracker" — operate on it only for read-only inspection (git log,git diff); nevercheckout/resetthere.Build images from inside the worktree with the same
docker buildx build --platform linux/amd64 --provenance=false …command; tag with a short purpose suffix (e.g.1.6.172-postmerge,1.6.171-traits-pr9c) so concurrent builds don't trample one canonical tag. Remove the worktree on exit so Builder LXC disk doesn't accumulate stale checkouts.AI-agent rule — pull often, push often. Never trust the working-tree state as truth. The shared workspace's HEAD is whatever the last session left it as, often months stale. Three failure modes this causes, all observed live 2026-05-16:
- Wrong-direction drift detection. Reading
/root/workspace/<repo>/service.yamlas "the intended version" surfaces a months-old version string and triggers a rebuild that would ship the wrong code if the build worktree weren't separately onorigin/main. - False triage rabbit holes. A "broken" repo per
fleet-runner build-testis often just stale workspace state —origin/mainis already green because a parallel agent pushed the fix you can't see. PRs get opened against problems that don't exist; obsolete-on-arrival. - Concurrent agents overwriting each other. Holding commits locally invites another session landing a competing fix; both branches diverge invisibly until one stomps the other.
The hard rules:
- Before reading any repo state (
service.yaml, source, tests,go.mod):cd /root/workspace/<repo> && git fetch origin --tagsfirst. Then read viagit show origin/main:<file>or a freshgit worktree add ... origin/main. Never the working-tree file directly. - Before triaging a "broken" repo: re-run the failure against
a fresh
origin/mainworktree. If green, the workspace is stale, not the code. Stop and verify before opening a PR. - Push every commit immediately.
git commit && git pushis one breath. Holding commits locally is what enables the concurrent-collision class above. - Fleet-runner subcommands that read workspace state (deploy,
audit, etc.) must
git fetch originfirst and read viagit show origin/main:<path>.readRepoServiceYAMLingo_fleet_runner/deploy_helpers.gois the canonical example. fleet-runner build-testis currently NOT freshness-correct — it runs against the working tree. Treat its output as a lower bound, not authoritative; confirm any "failure" against a fresh worktree before triaging.
- Wrong-direction drift detection. Reading
-
Dockerhost VM runs the service containers. Compose dirs:
/opt/services/<repo>/,/opt/security/<repo>/,/home/ubuntu_vm/pentest/<repo>/. -
OpenObserve LXC 106 (same SSH access pattern as Builder LXC 108 above, just a different
pct exectarget id; imageopenobserve/openobserve:v0.14.7+ bitnami/postgresql metastore) is the fleet's log aggregator. Root creds live in the LXC's owndocker-compose.yml— see privatefleet-state/OPS.mdunder "OpenObserve root credentials", never repeat them in a service repo. Retention isZO_COMPACT_DATA_RETENTION_DAYS = 30(dropped from 90 on 2026-08-23 — no SLA requires longer right now; it applies fleet- wide across every stream and takes effect promptly on restart, not gradually).Query it with
bin/oo, not hand-rolled curl.services-registry/bin/oois the deterministic CLI (logs,grep,errors,since-redeploy,context,compare,restarts,rate,versions,new,check,save/saved/run,summary,link,containers,hosts,tail,query,streams,stream-info) — runbin/oowith no args for full usage rather than duplicating it here; the script's own header comment is the source of truth for what each command does and its defaults, and this list will drift if a command gets added/renamed without this line being touched too. NeedsOPENOBSERVE_USER/OPENOBSERVE_PASSWORD/OPENOBSERVE_HOSTset (seebin/fleet-runner.env.example; real values infleet-state/OPS.md). Proxies every request through the bastion via SSH — LXC 106 is only reachable from inside the0docker.comprivate LAN.query/rundefault to compact JSON with OpenObserve's own response metadata stripped (pass--prettyfor indented + full metadata) — the other commands already print hand-formatted plain text, no JSON envelope. Every user-supplied value going into a WHERE clause is SQL-escaped (sql_escapein the script) and every request normalizes failures into a consistent, non-silent error shape (_oo_error— never a raw traceback, never indistinguishable from "zero results") — copy both patterns if you add a new command that talks to OpenObserve.Two streams, both queryable the same way:
default— OS-level syslog/journald, forwarded via plainrsyslog(omfwdin/etc/rsyslog.d/*.conf) from the dockerhost VM, the nginx proxy manager VM, and the LXC itself. sshd, kernel, cron, dockerd's own daemon events, and any systemd service that logs via journald.docker_logs— every container's stdout/stderr, fleet-wide, taggedcontainer_name/host/image/ compose labels. Added 2026-08-23 to close exactly the gap that bit an agent that night:docker logsonly shows the CURRENT container instance, so a redeploy silently discards history.docker logsand thejson-filedriver are UNCHANGED on every service — this is a second, independent path, not a replacement.
How
docker_logsgets populated: a small Vector (timberio/vector, pinned by digest) container namedvector-log-shipperruns at/opt/observability/vector-log-shipper/on every docker host (as of 2026-08-23: the dockerhost VM and the prod docker host). It reads every container's logs via the Docker socket — the same read pathdocker logsuses — and ships a copy to OpenObserve; it does not touch each container's own logging driver or config, so nothing about existing services changes, and no per-service compose edits were needed. New containers are picked up automatically. Gotcha found wiring this up: Vector 0.57.0's${VAR}env-var interpolation does not reliably substitute inside thehttpsink'suri/auth.*fields in this image — config loads fine but every request fails with "invalid uri character" / 401. Don't fight it: rendervector.tomlwith real values baked in before deploying (a plain template +sed/similar is enough) rather than relying on Vector's own interpolation for those fields.Adding a new docker host to the fleet should get a
vector-log-shippertoo, following the same compose shape (see the existing deployments as the reference) — this isn't automated byfleet-runneryet. -
Webgateway runs nginx (the public TLS terminator) and the keystore-aware
auth_requestflow. vhosts live as regular files in/etc/nginx/sites-enabled/<host>.{http,https}.conf(NOT symlinks to sites-available —sites-available/is unused). Editsites-enabled/directly for one-offs, or re-render viafleet-runner nginx-render --push. Surprised an agent for ~10min in the 2026-05-17 batch (ADR-0023 gap 5). -
Build + push:
docker buildx build --platform linux/amd64 --provenance=false -t ghcr.io/baditaflorin/<id>:<ver> --push .
Operational topology and credentials are in private
fleet-state/OPS.md — never commit SSH targets, IPs, or tokens to
service repos.
When a session gets long enough that context feels tight, hand off to a
new Claude with RESUME-PROMPT.md. It's a copy-paste
first-message that carries forward what's been built, what's deployed,
what's blocked, and where to pick up — so the user doesn't have to repeat
themselves. The prompt stays public (no topology / IPs / tokens); it
references this CLAUDE.md, RUNBOOK-UNATTENDED.md, and private OPS.md
for the rest.
Three primitives let any agent ship a brand-new service from "local code" to "live with DNS + scope + secrets" without operator intervention:
| Service | Role | Port |
|---|---|---|
go-fleet-secrets |
Encrypted vault for tokens (Hetzner, GitHub PAT, SMTP, platform API keys) | 18140 |
go-fleet-dns-sync |
Registry → Hetzner Cloud DNS reconciler (30-min ticker) | 18141 |
go-fleet-preflight |
Pre-deploy checklist (registry + DNS + port + secrets) | 18142 |
The full operational playbook — bootstrap, secret rotation, "how to add
a new service unattended", agent anti-patterns — lives in
RUNBOOK-UNATTENDED.md. Read that before
asking the user for tokens or where things live.
Key facts an agent needs to know:
- DNS API:
https://api.hetzner.cloud/v1(Bearer auth). The olderdns.hetzner.comConsole API is deprecated — don't use it. - Zone:
0exec.comis Hetzner Cloud zone id1285812. - Gateway IP:
176.9.123.221— every fleet A record points here; nginx terminates TLS and routes to the right upstream port. - Token env name:
HCLOUD_TOKEN(canonical, matches hcloud-cli + official Go SDK).HETZNER_TOKENkept as back-compat alias. - Secrets: live in
go-fleet-secrets, NEVER in env on dockerhost, NEVER in service repos, NEVER in services.json / overrides.json. Each secret has aconsumersallowlist (X-Auth-Userscope). - Before any deploy: call preflight; expect 200 (green) or 424 with a detailed checklist of what's red.
For any AI agent (Claude, Gemini, Haiku, GPT-anything) that lands in this repo and is asked to bump versions, allocate ports, or deploy. The fleet has canonical tooling — your job is to learn to invoke it. This section gives you the exact commands plus the manual fallback when the canonical tool isn't reachable.
fleet-runner lives on Builder LXC 108 at /usr/local/bin/fleet-runner.
The LXC is a Proxmox container on 0docker.com. From any host with SSH
access to the bastion:
# One-off invocation (works from your laptop, a CI runner, anywhere):
ssh root@0docker.com "pct exec 108 -- /usr/local/bin/fleet-runner <subcommand> [args...]"
# Examples:
ssh root@0docker.com 'pct exec 108 -- /usr/local/bin/fleet-runner converge'
ssh root@0docker.com 'pct exec 108 -- /usr/local/bin/fleet-runner allocate-port --count 1'
ssh root@0docker.com 'pct exec 108 -- /usr/local/bin/fleet-runner audit --all'
ssh root@0docker.com 'pct exec 108 -- /usr/local/bin/fleet-runner bump-version go_<repo> patch --push'If you don't have SSH access to 0docker.com, stop and ask the user
to run the command, copy-pasting the exact line above. Do not
substitute a different command. If you can't involve the user
(autonomous run), drop down to the "manual fallback" recipe in each
section below — but mark in your output that you used the fallback so
the user can verify nothing drifted.
On a fresh workstation, once SSH keys to the bastion are set up
(target identities are in private fleet-state/OPS.md), install
fleet-runner-shim as /usr/local/bin/fleet-runner and every recipe
below works with the bare command (drop the
ssh "$FLEET_BASTION" 'pct exec "$FLEET_LXC" -- "$FLEET_REMOTE_BIN"' prefix).
One-liner:
curl -fsSL https://raw.githubusercontent.com/baditaflorin/services-registry/main/bin/fleet-runner-shim \
| sudo tee /usr/local/bin/fleet-runner >/dev/null \
&& sudo chmod +x /usr/local/bin/fleet-runner
fleet-runner --help # smoke test — should print the remote binary's helpAfter install, the canonical examples shorten to e.g.
fleet-runner converge, fleet-runner allocate-port --count 1,
fleet-runner deploy go_<repo>. The shim is dumb — it just forwards
argv over SSH to LXC 108 — so output, exit codes, and prompts behave
exactly as on the LXC. Source: services-registry/bin/fleet-runner-shim.
Some verbs need env vars sourced (notably the key … and
audit fleet-auth-scope family need APIKEY_SERVICE_URL +
APIKEY_SERVICE_ADMIN_TOKEN; deploy needs HETZNER_TOKEN for the
DNS step — all from fleet-state/OPS.md). Drop
bin/fleet-runner.env.example into
~/.fleet-runner.env, fill in the admin tokens, then
source ~/.fleet-runner.env from your shell rc. Without it, every
admin op errors with keystore unavailable: dial localhost:18021
and every deploy skips the DNS step with $HETZNER_TOKEN not set.
Env-file MUST use export KEY=VAL lines, not bare KEY=VAL —
otherwise source populates the current shell but the vars don't
propagate to child processes (the fleet-runner binary). The
template in bin/fleet-runner.env.example is already in the right
shape; copy it verbatim. The LXC 108 copy at /root/.fleet-runner.env
was migrated to export form on 2026-05-19. If you maintain another
copy elsewhere, mirror the same shape.
A host_port collision does NOT fail loudly — it clobbers a live service.
fleet-runner deployresolves the dockerhost compose dir by host_port. If you hand-pick a port another service already owns,deploy <your-service>finds the other service's/opt/services/<that-repo>/directory, overwrites itsdocker-compose.ymlwith your image, and rolls your container in place — silently evicting the live service that owned the port. Observed 2026-05-31:fleet-piperegistered on a hand-picked18256took downsvg-icon-inventory(which already owned 18256), because the deploy deployed pipe into svg's compose dir. Always take the port fromallocate-port; never hand-pick.Recovery if you collide: (1)
docker compose downyour squatter in the victim's dir to free the port; (2)fleet-runner deploy <victim_repo> --bootstrap --force-build— a plain deploy is fooled into a no-op when the squatter happens to report the same version, so force it; (3) move your service to a free port in BOTH the registry (overrides.json→ regen → push) and the repo (Dockerfile, compose, service.yaml, deploy.yaml), then redeploy.
Canonical (preferred):
ssh root@0docker.com 'pct exec 108 -- /usr/local/bin/fleet-runner allocate-port --count 1'
# Output: a single integer like 18099 — that's your host_port
# Multiple at once:
ssh root@0docker.com 'pct exec 108 -- /usr/local/bin/fleet-runner allocate-port --count 3'Manual fallback (when canonical isn't reachable):
- Open
services-registry/services.jsonand find the highesthost_portcurrently in use in the reserved range (default18100–18999). - Pick the next integer above the max.
- Add an entry to
services.jsonwith bothhost_port(e.g.18099) andcontainer_port(what the service binds inside its docker container — usually8xxx). - Verify no clash:
grep -E '"(host\|container)_port":\s*<your-pick>' services-registry/services.jsonshould return only your line.
When you hit "port X is already taken" — the case Gemini got wrong:
The registry is the truth, not the running container. Find the squatter:
# Anyone claiming this port in the registry?
python3 -c "import json; d=json.load(open('services-registry/services.json')); print([e['id'] for e in d if e.get('container_port')==8313 or e.get('host_port')==8313])"
# Services WITHOUT a registered host_port (likely silent squatters):
ssh root@0docker.com 'pct exec 108 -- /usr/local/bin/fleet-runner audit registry-host-port-set'If the squatter has no registry entry, add one for it with
allocate-port. Your service keeps its original port. Only reallocate
your service's port if the squatter has a legitimate registered claim.
Canonical: fleet-runner bump-version updates service.yaml, any
const Version = "..." in main.go/version.go, prepends a
CHANGELOG.md entry (see "Per-repo CHANGELOG.md" above), creates the
git tag, and (with --push) pushes commit + tag together:
# Local bump (writes files, prints next steps for review)
ssh root@0docker.com 'pct exec 108 -- /usr/local/bin/fleet-runner bump-version go_<repo> patch'
# Atomic bump + commit + tag + push (one-shot)
ssh root@0docker.com 'pct exec 108 -- /usr/local/bin/fleet-runner bump-version go_<repo> patch --push'
# Write a richer changelog entry instead of the auto-derived commit list:
ssh root@0docker.com 'pct exec 108 -- /usr/local/bin/fleet-runner bump-version go_<repo> minor --changelog "### Added
- new /foo endpoint
### Fixed
- BREAKING: renamed ?q= to ?url=" --push'
# Variants: minor / major / --set 2.0.0After the bump lands, the container is still running the OLD
version until you deploy. Pair with fleet-runner deploy <repo>.
Manual fallback:
cd /path/to/<repo>
# 1. service.yaml (preserve quoting — quoted stays quoted)
sed -i.bak 's/^version: "1.2.3"/version: "1.2.4"/' service.yaml && rm service.yaml.bak
# 2. main.go / version.go const, if present
grep -l 'const Version' *.go
sed -i.bak 's/const Version = "1.2.3"/const Version = "1.2.4"/' main.go && rm main.go.bak
# 3. Commit, tag, push, push tag — ALL FOUR (Gemini forgot step 4)
git add -A && git commit -m "chore: bump version to 1.2.4"
git tag 1.2.4 # NO leading v
git push
git push origin 1.2.4 # tags don't ride `git push` by defaultTag after the commit, push both.
Canonical (only one right answer):
fleet-runner deploy go_<repo>No extra flags needed on Builder LXC 108 — both historical workarounds are resolved as of 2026-08-23:
- Registry cache race —
raw.githubusercontent.com/.../services.jsonis still Fastly-cached withmax-age=300, butfleet-runnernow resolves the registry source itself: when--servicesisn't passed and/root/workspace/services-registry/services.jsonexists, it prefers that local working copy over the CDN automatically (ResolveServicesSource, shipped 2026-05-21). An explicit--services <path>still works if you ever need to force a different source. - Cosign signing — restored fleet-wide; the vault's
cosign-signing-key,cosign-signing-key-password, andcosign-public-keyare all populated again (verified directly againstgo-fleet-secretson 2026-08-23). A plaindeploysigns on push and verifies on pull with no flag.--skip-cosignstill exists as an emergency-only bypass (vault unreachable, key mid-rotation) — don't pass it by default; it prints a loud bypass warning and lands an audit-log row precisely so it's never silently routine.
Post-deploy: re-render the vhost. The embedded nginx-render
step inside deploy often misses the new vhost due to the same
registry-fetch race. Until fleet-runner deploy folds in the local
fetch (separate fix in flight), follow every new-service deploy with:
fleet-runner nginx-render --filter <slug> --push --reloadfleet-runner deploy is idempotent end-to-end. The pipeline is built
to fail fast BEFORE touching prod, and to refuse to declare success
unless every check confirms the new image actually carries the new
code AND the cross-service-call gate is green:
- DNS / vhost / cert — idempotent shape checks.
- Drift detection — reads
service.yamlversionfromorigin/mainviagit show(NOT the long-lived workspace tree, which is shared and routinely stale) and compares to the liveGET /version. If they match, the deploy is an idempotent no-op. - Pre-flight (Go repos) — in a fresh
git worktree add origin/mainon Builder LXC 108, rungo build ./...andgo test ./.... Failure aborts here; prod is not touched. - Build + push —
docker buildx build --platform linux/amd64 --provenance=false --pushtagging both:<version>and:latest. - Pull + digest assertion —
docker compose pullon dockerhost, thendocker inspectthe new:latestdigest. If it equals the previously-running digest, the deploy fails: the new manifest didn't propagate (GHCR auth scope, tag cache, platform mismatch), and rolling forward would be a no-op masquerading as a deploy. - Roll —
docker compose up -drecreates the container. - Health-wait — poll
docker inspect …{{.State.Health.Status}}until "healthy" (up to 90s). "unhealthy" fails immediately; "starting" keeps polling; empty = no HEALTHCHECK directive, trust the container and proceed. - Smoke gate — three probes:
GET /healthmust be 200,GET /selftestmust be 200 (or 404 = "service didn't implement it, skip"); 503 is the codified "internal sources errored" signal and fails the gate.GET /versionmust match the version we just pushed — catches "container restarted but the image didn't roll". - Rollback on smoke fail — captures the previous image digest
before the roll; on smoke failure, retags that digest as
:lateston the dockerhost (no GHCR roundtrip) andcompose up -d, waits for HEALTHCHECK, re-smokes. The new image stays in GHCR but is NOT kept running.
Flags: --force-build (rebuild + roll even when versions match),
--skip-build (assume the image is already in GHCR), --skip-smoke
(offline DR), --no-rollback (leave the new image in place on smoke
fail — for fix-forward scenarios).
Manual fallback (when LXC 108 is unreachable):
If you must deploy manually, do all of these in order — do not skip any:
# 1. Build on an AMD64 host (NOT on an ARM Mac — binary won't run)
docker buildx build --platform linux/amd64 --provenance=false \
-t ghcr.io/baditaflorin/go_<repo>:<version> --push .
# 2. Roll the container forward on the dockerhost
ssh -J root@0docker.com ubuntu_vm@10.10.10.20 '
cd /opt/services/go_<repo>/src
git pull origin main
sudo docker compose pull && sudo docker compose up -d
'
# 3. Update the gateway-served deployment metadata (catalog UI reads it)
ssh -J root@0docker.com florin@10.10.10.10 '
echo "{\"sha\":\"$(git rev-parse HEAD)\",\"version\":\"<version>\",\"deployed_at\":\"$(date -u +%FT%TZ)\"}" \
| sudo tee /etc/nginx/deploy-meta/<slug>.0exec.com.json
sudo nginx -s reload
'
# 4. Smoke test
curl -sSf https://<slug>.<mesh>.com/healthIf step 3 or 4 fails, the deploy is incomplete even though the container is running. Don't declare done until both succeed.
Three commands. Run all three. If anything in the category you touched
is flagged, fix it before stopping. Do not pass --services to any
of these three — converge, audit, and state snapshot don't accept
that flag (confirmed on v0.7.11: flag provided but not defined: -services). They already read the registry correctly without it — see
"Recipe — Deploying a service" above for why the local-copy race no
longer needs a flag at all:
fleet-runner converge
fleet-runner audit --all
fleet-runner state snapshotWhen a change to a shared thing affects N services — a go-common
bump, a changed client signature, a new env contract — don't hand-grep
and hope. fleet-runner rollout finds EVERY affected service, propagates
the change, and proves each still builds.
# 1. PLAN (read-only, the default): see the complete affected set.
# Run on Builder LXC 108 so the code-grep sees the full workspace.
fleet-runner rollout \
--dep github.com/baditaflorin/go-common@<LATEST> \
--grep 'client\.JSProxy(DOM)?\(|GetRendered\(|FetchNetwork\(|RenderJS' \
--graph-callers-of go-js-proxy,go-js-proxy-network,infrastructure-fetch-cache
# 2. APPLY: bump + build/test each consumer, land one auto-merge PR per repo,
# emit a fleet-state/sweeps manifest (revertable via sweep-rollback).
fleet-runner rollout --dep …@<LATEST> --grep '…' --clone --apply --prDiscovery is the union of three sources, because each has blind spots:
--grep (code signature — catches cold consumers the runtime graph never
saw), --graph-callers-of (go-fleet-graph inbound callers), --depends-on
(services.json declared edges). The union is intersected with
kind=container,language=go + excludes, and the backend slugs themselves
are dropped.
Three non-obvious rules:
- The runtime graph is blind to intra-mesh calls. Sibling-service calls
(fetch-cache, js-proxy) use a plain
net/http.Client, notsafehttp, so go-fleet-graph never records the edge.fleet-runner deps <slug>will show ZERO callers for an intra-mesh backend even when 48 services depend on it. For intra-mesh deps the code-grep is authoritative — that's why rollout unions grep with the graph instead of trusting the graph alone. --clonefirst (or run on LXC 108 where the workspace is kept complete): the grep is only as complete as the local checkouts, and the registry has ~110 repos that may not be cloned on a given host. rollout reports a "code-grep is PARTIAL" warning when repos are missing.- Always pass the actual LATEST dep version, never a stale literal —
go get dep@vOldon a repo already ahead silently downgrades it.
When you notice during a real engagement that a fleet service is missing a capability (didn't return needed signal, doesn't handle a class of input, response is silent on a real failure), don't lose the discovery. Capture it, then route through the right tool:
real engagement
│
v
gap noticed ───> add a record to a findings JSON
│ (kind, repo, request, expected, suggested_fix,
│ [optional] auto_apply + patch_unified_diff)
v
bin/autofix.py findings.json [--apply]
│
├── auto_apply: true + fleet repo + clean patch ──> CLONE + APPLY
│ + TEST + PUSH + DEPLOY + /selftest [+ rollback on fail]
│
└── anything else ────────────────────────────────> bin/disclose.py
files the gap as an issue on the right repo (either fleet
repo for fleet_gap; merchant repo for external_leak)
Both tools live in
baditaflorin/go-pentest-leak-bounty-policy/bin/.
Run from any workstation with gh auth status green.
bin/autofix.py lands mechanical fixes unattended. Hard safety
rails — every gate is a clearly-named abort, never silent:
skip_not_fleet— repo not underbaditaflorin/. Third-party repos can never be auto-modified, period.skip_no_auto_apply— finding must opt in withauto_apply: true.skip_no_patch— must include apatch_unified_diff.git apply --check— patch applies cleanly to a fresh clone before any state mutation.- Test gate —
go test ./...(ornpm test) must pass. /selftestgate — post-deploy, the service's/selftestmust return 200./healthonly proves the binary booted;/selftestexercises the patched code path.- Auto-rollback — on
/selftestfail, force-push origin back to the pre-fix SHA AND redeploy the previous image.
bin/disclose.py files an issue when the fix isn't mechanical
(or the repo is third-party). external_leak redacts the token to
the prefix shape only — never re-publicizes the secret.
Findings file shape (single JSON, findings: [...]). One record:
{
"id": "gap-7",
"kind": "fleet_gap",
"repo": "baditaflorin/go-pentest-<svc>",
"auto_apply": true,
"gap_summary": "...",
"patch_unified_diff": "--- a/file\n+++ b/file\n@@ ...\n",
"suggested_fix": "...",
"session_context": "..."
}See
bin/findings-fleet-gaps-example.json
for every dispatch path's shape.
Why this matters for THIS repo: every service that exposes a
public API should ship a /selftest endpoint that exercises its
real dependencies (resolver, upstream API, embedded data). Without
it, the autofix /selftest gate has nothing to verify against and
defaults to a weaker /health probe — which only proves the binary
booted, not that the patched code path works. See
go-pentest-takeover-checker and go-pentest-subfinder (v0.2+)
for the canonical pattern.
-
"Port 8313 is taken, I'll pick 8500 and edit
service.yaml." Usefleet-runner allocate-portand register the squatter. See "Allocating a port" above. And never hand-pick a port at all — deploy resolves the dockerhost compose dir by host_port, so a collision silently clobbers + evicts the live service that owns it (the 2026-05-31 fleet-pipe/svg-icon-inventory incident). Always take the number fromallocate-port. -
"I bumped the version in
service.yamland pushed." Did you tag git AND push the tag AND update the docker image tag? Usefleet-runner bump-version --push. -
"All repos are pushed to origin/main." Pushing code ≠ deploying. The container is on the old image until
fleet-runner deploy(or the manual fallback) runs. -
"I edited
service.yamlport from 8313 to 8500 to avoid conflict." Silent multi-file drift.fleet-runner audit port-matches-registrycatches this. Don't. -
"I ran
git tag X.Y.Z." Did yougit push origin X.Y.Z? Tags don't ridegit pushby default. -
"
fleet-runnerisn't working for me, I'll use a different deploy path." Stop. Either report the exact command + error to the user, or use the manual fallback recipe above and say so in your summary so the user can verify the catalog-meta step landed. -
"
fleet-runner audit <check> --jsonsays unknown check--json." Go'sflagpackage stops parsing at the first non-flag argument, soaudit feature-without-bump --jsonreads--jsonas a positional check name. Put--jsonbefore the check name:fleet-runner audit --json feature-without-bump. Same rule for--severity,--filter. Subcommands with their own FlagSet (audit compose-drift,audit compose-image-drift,audit vhost-drift) accept flags in any order because they parse their own args list. -
bump-versionagainst a stale workspace. Bug fixed in 2026-05-17 afternoon:fleet-runner bump-versionnowgit fetch origin --tagsfirst and refuses to bump if origin/main is ahead. The historical failure shape was: bump created a tag pointing at the wrong SHA,git pushrejected with "would clobber existing tag", recovery required manualgit tag -d <ver>. If you hit that on an older binary, the recovery is still:git -C /root/workspace/<repo> tag -d <ver>then re-run bump-version. -
./binary &smoke tests on Builder LXC 108. Don't. The builder is a build host andfleet-runnerhost — it is not a service host. When an agent runsgo build && ./binary &inside/root/workspace/<repo>/to "quickly check the handler responds," the binary backgrounds, the agent's SSH session ends, and the process becomes an orphan (PPid=1) listening on whichever port it bound. Twelve such orphans were found on 2026-05-21 — three of them silently squatting registered host_ports (e.g. 18107/18224), poised to confuse the next agent who deploys the canonical service to dockerhost and runsss -tlnpto debug. Canonical smoke is thefleet-runner deploysmoke gate against the dockerhost-side container, not an in-process binary on the builder. If you genuinely need to exercise the handler beforedeploy, rungo test ./...(which useshttptest) — no port binding, no orphan risk. If you must bind a port, use127.0.0.1:0(ephemeral) andkillit on the same line that started it.
Symptom. fleet-runner deploy <enricher> (or any container deploy)
rolls back at the smoke gate because the new image's GET /selftest
times out — even though /health is green and the code is fine. The
real cause is somewhere else entirely: a different service is piling
goroutines on a slow shared upstream and starving the whole dockerhost's
scheduler, so every service's /selftest probe stalls past its 8 s
deadline. One service's overload silently fails everyone's deploys.
The canonical instance: go_infrastructure_fetch_cache (port 18205)
proxies render requests to go-js-proxy / go-html-proxy, each held up
to 95 s. A backfill fan-out fires thousands of concurrent render=*
requests; ~8.3k goroutines pile on the saturated renderer (host loadavg
56/20 cores); the scheduler thrashes; downstream enrichers' /selftest
probes time out; healthy deploys fleet-wide roll back. Any service that
proxies to a saturatable sibling (a renderer, a headless browser, a
rate-limited upstream) can produce the same blast radius.
Diagnosis. Confirm it's goroutine pile-up, not the deploying service:
# Goroutine count on the suspected proxy (fetch-cache here): a healthy
# box sits in the low hundreds; a pile-up reads thousands.
curl -s http://<dockerhost>:18205/metrics | grep '^go_goroutines'
# The human-readable stats page surfaces the shed + upstream-error
# counters: a climbing render_shed means the gate is actively shedding
# (good — it's protecting the box); a high upstream_errors means the
# renderer itself is failing/slow (the root cause).
curl -s http://<dockerhost>:18205/ | grep -E 'render_shed|upstream_errors'For the fleet-standard signal, loadshed_shed_total{service,gate} is
emitted on /metrics by every service using go-common/loadshed — a
sustained nonzero rate is the "this box is shedding load" alert.
Live mitigation (no rebuild). The render cap is env-tunable. Lower
MAX_RENDER_INFLIGHT on the host to shed sooner and shrink the pile-up,
then bounce the container — no image rebuild, no fleet-runner deploy:
# On the dockerhost, in the service's compose dir:
cd /opt/services/go_infrastructure_fetch_cache
sed -i 's/^MAX_RENDER_INFLIGHT=.*/MAX_RENDER_INFLIGHT=24/' .env # or add it
sudo docker compose up -d # picks up the new env; no pull/rebuildOnce the pile-up drains, the stalled /selftest probes pass and deploys
stop rolling back. Re-run the blocked deploy. Set the cap back up (or
remove it for the default 64) when the upstream recovers. The durable
fix for the root-cause backfill is to rate-limit the fan-out producer,
not just shed at the cache.
Building a new proxy-to-slow-upstream service? Don't hand-roll the
shed semaphore — use go-common/loadshed (loadshed.New(name, limit) +
TryAcquire/WriteShed, or gate.Guard middleware). It gives you the
fast-503 + Retry-After path and the loadshed_shed_total metric for
free. See go_infrastructure_fetch_cache's renderGate for the
canonical in-line (gate only the expensive sub-path) usage.
modernc.org/sqlite uses real fcntl syscalls for WAL file locking.
Go creates one OS thread per goroutine blocking on a real syscall — unlike
HTTP or postgres, which go through Go's epoll/kqueue poller and do NOT
spawn threads. Under high /verify traffic, an uncancellable goroutine per
request × 5s busy_timeout = 1388 OS threads = 48 GB RAM exhausted on the
host (go-apikey-service incident 2026-05-26).
Any service that opens a modernc.org/sqlite DB MUST follow all three:
// 1. Serialise writes at the Go pool level (channel wait, not fcntl wait).
// This prevents thread explosion: Go queues at the pool, not the syscall.
db.SetMaxOpenConns(1)
db.SetMaxIdleConns(1)
// 2. Cap busy_timeout — belt-and-suspenders, protects against pool leaks.
// Use ≤ 500ms.
dsn := path + "?_pragma=journal_mode(WAL)&_pragma=busy_timeout(500)&..."
// 3. Context-bound ALL goroutines that write to the DB.
// Never: go db.Exec(...)
// Always:
go func(k string, ts int64) {
ctx, cancel := context.WithTimeout(context.Background(), 500*time.Millisecond)
defer cancel()
_, _ = db.ExecContext(ctx, `UPDATE ...`, ts, k)
}(key, now)Audit one-liner — run this in any Go fleet repo workspace to find violations before they cause an incident:
# Find SQLite services missing safety rules
for dir in /opt/services/*/; do
go_mod="$dir/go.mod"
main="$dir/main.go"
[ -f "$go_mod" ] || continue
grep -q "modernc.org/sqlite" "$go_mod" || continue
echo "=== $dir ==="
grep -q "SetMaxOpenConns" "$main" || echo " MISSING: SetMaxOpenConns"
grep -q "busy_timeout" "$main" && \
grep -oP 'busy_timeout\(\K[0-9]+' "$main" | \
awk '{if($1>500) print " HIGH busy_timeout: "$1"ms (reduce to ≤500)"}'
grep -n "go db\." "$main" 2>/dev/null | grep -v '//' | \
sed 's/^/ FIRE-AND-FORGET goroutine: /'
doneCurrent fleet status (as of 2026-05-26):
go-apikey-service: fixed (was the incident service)go-fleet-persona: clean (hadSetMaxOpenConns(1)+ context calls already)- All other services: use postgres/redis (no fcntl risk)
The cardinal rule when you'd otherwise touch every service: modify
the library and bump the dep. A go-common patch plus
fleet-runner update-dep --push github.com/baditaflorin/go-common@vX.Y.Z
beats 130 PRs. To bump only a subset (stragglers, one mesh, one
category) add --repos a,b or --filter mesh=…,category=…,ids=a;b —
update-dep is no longer fleet-wide-only.
go-common ≥ v0.55.0 — /selftest bypasses the fetch cache. The
selftest.Suite now runs every check with
safehttp.WithoutFetchCacheContext(ctx), so /selftest validates the
service's REAL outbound path (DNS + TLS + origin) instead of routing
live probes through a cold fleet cache. Before this, live-probe
selftests (e.g. a detector fetching vercel/netlify/fly) routed through
the cold cache on the freshly-built image and blew past
fleet-runner deploy's 8 s smoke /selftest timeout — false-failing
otherwise-healthy deploys and rolling them back. If you see a deploy
roll back on a /selftest timeout while /health is green, the fix is
to bump the service to go-common ≥ v0.55.0 and redeploy; do NOT reach
for --skip-smoke. Need a single client to skip the cache outside
selftest? safehttp.WithoutFetchCache() (per-client) or
WithoutFetchCacheContext(ctx) (per-request).
Every container repo keeps a CHANGELOG.md at its root, newest
entry on top, in the same shape go-common uses:
## <version> — <YYYY-MM-DD>
### Added | Changed | Fixed | Removed
- one bullet per meaningful change; lead breaking changes with **BREAKING:**Why: with ~220 services moving independently, "what shipped in this
version and could it have broken X?" is otherwise un-answerable without
trawling git. A per-version, human-and-AI-written entry makes
regressions diff-able at a glance and feeds the fleet-wide
fleet-runner changelog digest with intent (not just PR titles).
You (the agent) write the entry — you just made the change, so you
know what happened and why. fleet-runner bump-version scaffolds it
for you: it prepends a ## <newver> — <date> block to CHANGELOG.md
(creating the file if absent), auto-populating it with the commit
subjects since the previous tag. Pass --changelog "### Added\n- …"
to write a richer entry instead of the auto-derived one. The changelog
edit rides the same atomic bump commit as service.yaml + the tag, so
a version never lands without a changelog line. This is opt-out only by
omission — don't bump a service's version without leaving a changelog
entry.
- Local workspace root:
/Users/live/Documents/Codex/2026-05-08/. Sibling repos sit next to this one — read them directly when you need to understand a dependency. - CI: self-hosted Woodpecker CI is live at
https://ci.0exec.com on Builder LXC 108 (see
ADR-0035) —
server + agent at
/opt/woodpecker/on the LXC, capped at 2 concurrent workflows so it doesn't contend withfleet-runner deploy-allbatches on the same box. GitHub webhooks are wired (OAuth App + nginx vhost reusing thewildcard.0exec.comcert); pushes and PRs triggergo build ./... && go test ./...on every activated repo — the same gatefleet-runner deploy's pre-flight already runs, now push-triggered instead of deploy-time-only. A daily cron (docker builder prune -af --filter unused-for=24h, 3:15am) keeps the LXC's build-cache disk usage bounded. Fleet-wide rollout is in progress (batches of.woodpecker.yml+ repo activation via the API, tracked viagit logon this file / repo PRs — no separate rollout ledger). Still don't scaffold GitHub Actions build workflows — that's the billing model this exists to avoid. Repos without their owngo.mod(composite-pattern services built on a sharedgo_composite_runnerbase image) don't get a pipeline — there's no local Go source forgo buildto act on. - Supply chain: prefer npm packages ≥ 3 days old over
@latest— accept known CVEs over zero-day supply-chain injection. - Supply chain: prefer npm packages ≥ 3 days old over
@latest— accept known CVEs over zero-day supply-chain injection.