This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Before planning, editing or committing, Claude Code must read and follow
docs/developer/agent-rules.md. It is the
canonical project-wide rule set. The Claude-specific commands, architecture
orientation and tool instructions below supplement it and never replace it.
Local-first EMS (Energy Management System) controller for Zendure SolarFlow
battery/inverter systems. No YAML automation stack, no cloud dependency for
control decisions. It reads grid-meter load (Shelly / everHome EcoTracker /
Tasmota) and Zendure device telemetry, calculates a power target, and writes
outputLimit (and related state) back to Zendure devices via local HTTP API.
This software writes to real power hardware. Be conservative with changes to control logic, write gates, and safety reconciliation — see "Safety Model" below.
Install deps:
python3 -m venv .venv && source .venv/bin/activate
python -m pip install -r requirements.txt -r requirements-dev.txtCompile check (run after any change to entry script / ems/ / emsctl.py / scripts/check_log_events.py):
python3 -m py_compile ems-solarflow-api-control.py ems/*.py emsctl.py scripts/check_log_events.pySelf-test:
python3 -B ems-solarflow-api-control.py --self-testSimulation (no hardware/network required):
python3 -B ems-solarflow-api-control.py --simulate --max-cycles 1Replay a captured trace:
python3 -B ems-solarflow-api-control.py --replay /path/to/trace.jsonl --onceFull suite:
pytestSingle file / single test:
pytest tests/test_pv_first_charge_balance.py
pytest tests/test_pv_first_charge_balance.py::test_some_case -vOffline power-control regression tests (the required CI check, deterministic, no hardware/network):
pytest tests/ -m "simulation and power_control"Targeted tiers — do not run the full suite for a localized change
(see docs/developer/testing.md):
./scripts/test-fast.sh # unit + contract, no Docker/browser/slow
./scripts/test-admin.sh authority # one Admin functional area
./scripts/test-mqtt.sh
./scripts/test-pr.sh core # one pull-request group
./scripts/test-pr.sh appliance # the appliance group
./scripts/test-rc.sh --list # the release-candidate gatesMarkers (pytest.ini, --strict-markers is on) have two independent
dimensions. Execution level, exactly one per module: unit, contract,
integration, e2e; plus docker, browser, slow. Functional areas, any
number: admin, setup, maintenance, workflow, authority, config,
mqtt, power_control, backup_restore, system_build, appliance. simulation,
regression and mqtt_release remain for the existing gates. Classify a new
module with a module-level pytestmark; tests/test_test_classification.py
enforces the rules.
python3 scripts/check_log_events.py /tmp/ems-sim.log \
--require startup \
--require target_calculationpython3 emsctl.py status
python3 emsctl.py diagnose
python3 emsctl.py diagnose --json
python3 emsctl.py diagnose --control
python3 emsctl.py diagnose --control-quality --sample-seconds 60Operating model is intentionally minimal: one start script, one static config.
python3 ems-solarflow-api-control.pyconfig.json— static installation config (versioned template:config.template.json).data/runtime-state.json— mutable runtime/operator state, created and updated by the EMS and byemsctl.py. The Admin console also mirrors the whitelisted overlapping keys it changed into it on maintenance apply (config → runtime convergence, one-directional, via the samedashboard/runtime_write.pywhitelist). Not a second static config.
The entry script (ems-solarflow-api-control.py) only does bootstrap/coordination:
CLI parsing, config loading, logging setup, client construction, runtime-state
construction, controller startup, main loop. All real implementation lives in ems/.
config.py— config loading, safe parsing, runtime mode helperslogging_utils.py— structuredevent=...logging setupmodels.py— telemetry/capability dataclassesclients.py— HTTP, Zendure, Shelly, Home Assistant clientsruntime_state.py— mutable runtime-state read/writeruntime_intents.py— runtime AC mode intent (ac_output/ac_input) reconciliationtarget_control.py— capability detection and target/output calculation (core control math)controller.py— main EMS control loop, ties everything together (largest module)state_store.py— SQLite store backing battery full-charge assist (and dashboard stats)simulation.py— simulation, replay, preflight, self-test helpersdiagnostics.py— read-onlydiagnoseservice layer (versioned contract); imported by bothemsctl.pyand the dashboardpaths.py— shared project-path resolvers (BASE_DIR,resolve_*_path); import-side-effect-free
Edit the smallest relevant module rather than the entry script or controller.py
when the change is localized (e.g. target math → target_control.py, runtime
state shape → runtime_state.py).
A second product in this repository, and the largest new subsystem: a Debian
package plus systemd units that manage a Raspberry Pi host running the EMS in
Docker. It is not part of the EMS control loop and must not import from ems/.
- Two processes, one privilege boundary: an unprivileged web service
(
web.py,static/app.js) talks over a unix socket to a root agent (agent.py,commands.py) that executes an allowlisted set of operations. The allowlist is enforced on the agent side (protocol.py,operation_schema.py); reaching that socket as an allowed uid is an appliance-takeover capability. - The operating system is patched in place by
apt. There is no second copy and nothing rolls back by itself, so a failed OS upgrade is recovered by re-flashing and restoring a backup — which is whypackages.py's blockers and the backup path matter more here than they look. - Release trust (
release_trust.py,release_attestation.py,artifact_trust.py) is fail-closed by design. Do not weaken a verification path to make a gate pass. - The Appliance Manager updates itself as a signed
.deb(manager_update.py,manager_releases.py,manager_retention.py,manager_install.py,manager_verify.py), never on a timer and always on an operator's button, with an older package installable as readily as a newer one. Three properties are not negotiable:dpkgruns from its own systemd unit rather than the agent's cgroup, every refusal happens before it runs, and the reverter is a copy taken out of the outgoing package. Doing nothing commits an install here, so the deadline inmanager_verify.pyis the only thing standing in for a way back, and it is software rather than firmware. Readdocs/appliance/adr/manager-self-update.mdbefore touching any of it. - One image, three boards:
image_shape.pyandrpi_image_gen.HARDWARE_PROFILESare the one table.grow-root.shis the only thing in this project that repartitions anything, it runs once on a freshly imaged card, and it is gated onems-appliance image-check. persistent_state.RETIRED_SCHEMASmay be added to only by removing a state format, never by adding one. Dropping an axis makes the next package uninstallable on every appliance that recorded it.- Not confirmed on physical hardware — no image has booted on a board, and no
appliance has installed a manager package over HTTPS.
docs/appliance/hardware-validation.mdis the authority on what has and has not been proven; never upgrade a claim there without the evidence it names.
Compile check and tests:
python3 -m py_compile appliance/*.py && node --check appliance/static/app.js
pytest tests/ -k appliance -m "not docker and not browser and not slow"
npx playwright test --config=playwright.appliance.config.tsemsctl.py is a separate large CLI for safe runtime-state edits and diagnostics
(status, device ..., system ..., winter, ha, diagnose ...,
dashboard set-password). The diagnose service layer
(run_install_diagnosis, run_deep_diagnosis, run_hardware_diagnosis,
run_control_diagnosis, run_control_quality_diagnosis — all read-only, reuse
the same data path as the CLI) lives in ems/diagnostics.py; emsctl.py keeps
the thin CLI wrappers (handle_diagnose_command, arg parsing) and re-exports the
service functions. ems/diagnostics.py is import-side-effect-free and must never
import emsctl (so the dashboard can import it directly).
- Reload
runtime-state.jsonif changed - Optionally sync HA helper values
- Read Shelly/EcoTracker/Tasmota house load
- Read Zendure telemetry
- Detect runtime capabilities
- Run state reconciliation when due
- Detect strict night/minSoc idle
- Stabilize total target (
commanded_total_w+ filtered load, deadbands, ramps, filtering) - Allocate target across devices (PV-first, pv_priority_factor)
- Apply device ramp/limits,
min_output_limit, deadband - Write
outputLimitonly behind safety gates
Each device gets a runtime AC role: ac_output (acMode=2, normal output
regulation) or ac_input (acMode=1, excluded from output regulation, may
carry ac_charge_power_w reconciled as inputLimit). acMode/inputLimit
writes for this are owned exclusively by the runtime intent reconciler in
runtime_intents.py — don't add a second writer. Legacy role names
(normal_output, ac_input_charge, reserved) are accepted defensively and
mapped to ac_output/ac_input.
Optional controller lifecycle feature using ems/state_store.py (own SQLite
DB, configured via battery_full_charge_assist.state_database_path,
independent of the dashboard DB). Tracks socLimit == 1 Max-SoC events;
completion requires firmware-reported socLimit == 1 exactly (SOC % and
configured max_soc are not completion thresholds). Assist/restore reuse the
normal safe write helpers and the runtime AC intent reconciler.
Runtime outputLimit writes share the precondition dry_run=false,
simulation_mode=false, not replay, then require the named gate for the device's
transport: API/local-HTTP → allow_hardware_writes=true; local MQTT broker →
allow_mqtt_local_control_writes=true; Zendure cloud MQTT →
allow_mqtt_zendure_control_writes=true. All three gates default on
(RELEASE_WRITE_GATE_DEFAULTS): the template, config upgrade and normal
runtime loading resolve missing gate keys to the same release defaults, while
the simulation/replay safe config and template placeholder safety force every
gate off. Whether a transport actually writes is decided by configuration
presence: per-device write_output_limit opt-in, broker host, API key. Some
hardware generations are not yet validated on physical hardware (see
docs/user/supported-setups.md). Writes dispatch through dev.write_output_limit() and
cfg.control_writes_allowed(dev.control_gate); MQTT control lives in
ems/zendure_mqtt/ (control.py, device_client.py, control_runtime.py,
write_topics.py).
State reconciliation writes (minSoc, socSet, smartMode, gridOffMode,
winter inputLimit, full-charge-assist socSet/acMode/inputLimit)
additionally require allow_state_reconciliation_writes=true and are API-only
(MQTT control devices are output-only, supports_state_reconciliation=False).
The EMS must not run in parallel with another controller writing Zendure
outputLimit.
emsctl.py diagnose --json is a versioned public contract
(schema_version, diagnosis.{status,sections,metrics,root_causes,warnings,errors}).
Root causes use a stable shape: {code, severity, title, message, suggested_next_check}
with severity in info|warning|error and lowercase underscore-separated code.
Support bundle (diagnose --support-bundle) has a fixed file list
(diagnosis.json, control-diagnostics.json, control-quality.json,
redacted-config.json, runtime-state.json, bundle-metadata.json, plus
.txt variants). Bump schema_version/bundle_version and add/update
contract tests for any incompatible change to these.
When touching dashboard UI, read docs/developer/dashboard-style-guide.md first.
State whether the change uses:
- Aggregate / Device style
- Control / Energy stage style
Do not introduce a new dashboard visual system.
Docs are split by audience under docs/ (see docs/README.md for the full
map): user docs in docs/user/, technical reference in docs/technical/,
developer docs in docs/developer/. Notably technical/architecture.md,
technical/control-logic.md, technical/control-flow.md,
technical/runtime-state.md, technical/configuration.md, user/safety.md,
winter-mode.md, dashboard.md, cli.md, developer/development.md,
developer/developer.md. Admin docs: user/admin-console.md (overview),
user/admin-setup.md (new-system flow), user/admin-maintenance.md
(manage-existing-system flow), user/admin-backup-restore.md (Admin
backup/restore workflow), technical/admin-architecture.md (architecture
rules), and technical/admin-discovery.md (full Admin reference).
Update the relevant doc when changing behavior described there.
This project is indexed by GitNexus as ems-solarflow-api-control (34272 symbols, 81411 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
Index stale? Run
node .gitnexus/run.cjs analyzefrom the project root — it auto-selects an available runner. No.gitnexus/run.cjsyet?npx gitnexus analyze(npm 11 crash →npm i -g gitnexus; #1939).
- MUST run impact analysis before editing any symbol. Before modifying a function, class, or method, run
impact({target: "symbolName", direction: "upstream"})and report the blast radius (direct callers, affected processes, risk level) to the user. - MUST run
detect_changes()before committing to verify your changes only affect expected symbols and execution flows. For regression review, compare against the default branch:detect_changes({scope: "compare", base_ref: "main"}). - MUST warn the user if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use
query({search_query: "concept"})to find execution flows instead of grepping. It returns process-grouped results ranked by relevance. - When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use
context({name: "symbolName"}). - For security review,
explain({target: "fileOrSymbol"})lists taint findings (source→sink flows; needsanalyze --pdg).
- NEVER edit a function, class, or method without first running
impacton it. - NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use
renamewhich understands the call graph. - NEVER commit changes without running
detect_changes()to check affected scope.
| Resource | Use for |
|---|---|
gitnexus://repo/ems-solarflow-api-control/context |
Codebase overview, check index freshness |
gitnexus://repo/ems-solarflow-api-control/clusters |
All functional areas |
gitnexus://repo/ems-solarflow-api-control/processes |
All execution flows |
gitnexus://repo/ems-solarflow-api-control/process/{name} |
Step-by-step execution trace |
| Task | Read this skill file |
|---|---|
| Understand architecture / "How does X work?" | .claude/skills/gitnexus/gitnexus-exploring/SKILL.md |
| Blast radius / "What breaks if I change X?" | .claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md |
| Trace bugs / "Why is X failing?" | .claude/skills/gitnexus/gitnexus-debugging/SKILL.md |
| Rename / extract / split / refactor | .claude/skills/gitnexus/gitnexus-refactoring/SKILL.md |
| Tools, resources, schema reference | .claude/skills/gitnexus/gitnexus-guide/SKILL.md |
| Index, status, clean, wiki CLI commands | .claude/skills/gitnexus/gitnexus-cli/SKILL.md |
Serena runs a Python language server over this repo. It is registered globally in
~/.claude.json as serena start-mcp-server --context=claude-code-pragmatic --project-from-cwd, so it activates itself from the working directory — no
activate_project call per session, and no project pinned across repos. Its state
lives in .serena/, git-excluded alongside .gitnexus/ via .git/info/exclude.
Serena and GitNexus overlap on symbol lookup but are not interchangeable:
- GitNexus answers "what does this touch" — call graph, execution flows, blast
radius, risk. It reads a snapshot written by the last
analyzerun, so it lags behind edits made in the current session, silently. - Serena answers "what is this right now" — the language server always reflects the working tree, including edits made a minute ago.
| Question | Tool |
|---|---|
| Blast radius before an edit, risk level | GitNexus impact — mandatory, see above |
| Execution flows, call chains, "how does X work" | GitNexus query / context |
| Scope check before committing | GitNexus detect_changes |
| Taint / source→sink findings | GitNexus explain |
| Current body or signature of a symbol | Serena find_symbol with include_body |
| Structure of a file not yet read | Serena get_symbols_overview |
| Callers/usages, resolved through the type system | Serena find_referencing_symbols |
| Replace a whole function, method or class | Serena replace_symbol_body |
| Type/syntax errors after an edit | Serena get_diagnostics_for_file |
| Config, docs, JSON/YAML, a few lines at a known path | built-in Read / Grep / Edit |
- The GitNexus
impact-before-edit mandate stands unchanged. Serena does not replace it — a language server has no notion of blast radius. - After editing a symbol in this session, trust Serena over GitNexus for that symbol's contents until the index has been re-analyzed. When the two disagree about what the code says, the language server is newer.
- Do not read a whole module to find one function.
get_symbols_overview, thenfind_symbol.ems/controller.pyandemsctl.pyare large enough that this is the difference between a page and a wall of context. - Never rename by find-and-replace. Serena
rename_symbolis language-server-exact; GitNexusrenameis call-graph-aware. Runimpactfirst either way. - Both toolsets are read-mostly and safe to use freely. Neither is a substitute for the safety rules in "Safety Model" — no tool output authorizes a write-gate change.