2026-09-01 21:13:00 -04:00
|
|
|
|
# CLAUDE.md — agent operating manual
|
|
|
|
|
|
|
|
|
|
|
|
Design docs live in `docs/` (00–13, Chinese). They are the specification; this
|
|
|
|
|
|
file distills what is **binding** when writing code. When code and docs conflict,
|
|
|
|
|
|
docs win — or write an ADR changing the doc first.
|
|
|
|
|
|
|
|
|
|
|
|
## Current phase
|
|
|
|
|
|
|
M2: skill contracts, Python skill service, L2 eval harness with baseline
- packages/domain: ForecastRequest, BidOptimizationRequest/Result,
ReportRequest, SkillReport (+ golden and invalid fixtures, exported to
contracts/ and regenerated as pydantic models).
- skills-py/vpp_skills: FastAPI service with versioned registry; load/PV/
price forecasts (same-day-type EWM point forecast, conformal residual
quantiles — coverage test as acceptance gate); bid-optimization MILP on
HiGHS (binary block participation, hard ledger energy bounds, exact
Decimal fit of the rounded curve inside the bounds, revenue distribution
over quantile paths); report generator whose every figure is a
{tool_call_id, path} reference, with a verifier. 48 tests incl. hypothesis
property test that bids respect ledger constraints.
- packages/services: LedgerService.dayAheadBounds (the P7 cascade band
handed to the optimizer); Decimal resolved once for CJS/ESM interop.
- packages/evals: L2 metrics (MAPE, nRMSE, coverage, direction accuracy,
naive/hindsight revenue baselines), HTTP skill client, rolling-origin
harness that pushes each bid through the real ledger, CLI with
--check/--write-baseline; committed baseline on the SYNTHETIC dataset
(no historical Hubei data yet — baselines measure the harness, not KPI).
- CI: evals job boots the skill service and fails on baseline digest drift.
- docs/open-questions: A6 (flexibility marginal cost = offer floor); A4/B6
wired as placeholders. README/CLAUDE.md status → M2 done, M3 next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:29:08 -04:00
|
|
|
|
Docs complete; M1 and M2 implemented (schemas + contracts pipeline; ledger,
|
|
|
|
|
|
snapshot, time-series, relational, ingestion services; Python skill service
|
|
|
|
|
|
with forecast/MILP/report skills; L2 eval harness with committed baseline).
|
|
|
|
|
|
Storage is in-memory reference semantics — persistent adapters arrive with M3.
|
|
|
|
|
|
The eval dataset is synthetic until historical data lands. Next is M3. Build
|
|
|
|
|
|
order is ROADMAP.md (M1→M5). A skill change that moves an L2 metric must
|
|
|
|
|
|
update the baseline in the same change (`npm run eval -w @vpp/evals -- --write-baseline`).
|
2026-09-01 21:13:00 -04:00
|
|
|
|
Do not start a milestone's work before its predecessor's acceptance criteria are
|
|
|
|
|
|
testable, and do not build phase-2 items (edge control links, federation,
|
|
|
|
|
|
interaction/load-control agents) unless explicitly asked.
|
|
|
|
|
|
|
|
|
|
|
|
## Hard rules (from docs/01 invariants — treat as review criteria)
|
|
|
|
|
|
|
|
|
|
|
|
1. **LLM never computes numbers.** Any numeric field in a `Proposal` payload must
|
|
|
|
|
|
be a `{toolCallId, path}` reference into recorded tool-call lineage,
|
|
|
|
|
|
dereferenced by the assembler. If you find yourself parsing a number out of
|
|
|
|
|
|
LLM text into a payload, stop — that's the architecture's one forbidden move.
|
|
|
|
|
|
2. **No LLM in any control path.** Nothing under `services/` or the execution
|
|
|
|
|
|
path may import or await an LLM call. Periodic workflows must run with the
|
|
|
|
|
|
LLM backend down (invariant I6 — there is a test for this).
|
|
|
|
|
|
3. **External effects only via `(Proposal, ExecutionPermit)`.** Gateways/adapters
|
|
|
|
|
|
never accept a bare plan. Permits are short-lived and revocable.
|
|
|
|
|
|
4. **AI cannot approve.** No code path may transition a Proposal to
|
|
|
|
|
|
APPROVED/AUTHORIZED without either a matched envelope + passing checks, or a
|
|
|
|
|
|
human `resume()` with an approver identity distinct from the origination chain.
|
|
|
|
|
|
5. **Everything auditable.** State transitions, tool calls, approvals, permits,
|
|
|
|
|
|
and gateway receipts are event-sourced. Snapshots referenced by lineage are
|
|
|
|
|
|
immutable.
|
|
|
|
|
|
6. **Schemas live in `packages/domain` only.** Python models are generated from
|
|
|
|
|
|
`contracts/*.schema.json` — never hand-edit generated files, never define a
|
|
|
|
|
|
business object schema anywhere else.
|
|
|
|
|
|
|
|
|
|
|
|
## Conventions
|
|
|
|
|
|
|
|
|
|
|
|
- **Naming**: use the canonical code names in GLOSSARY.md. Do not invent new
|
|
|
|
|
|
English names for domain terms that already have one.
|
|
|
|
|
|
- **Language**: architecture docs Chinese; code, identifiers, comments, commit
|
|
|
|
|
|
messages, ADRs English.
|
|
|
|
|
|
- **Data representation** (docs/11 §3.3): money/energy/prices as fixed-point
|
|
|
|
|
|
decimal **strings**; units in field names (`power_mw`, `price_yuan_per_mwh`);
|
|
|
|
|
|
timestamps ISO8601 UTC; market intervals as `{date, interval_index}` with
|
|
|
|
|
|
`interval_minutes` explicit; all IDs strings; enums UPPER_SNAKE string literals.
|
|
|
|
|
|
- **Mastra**: do not trust memorized APIs — check `node_modules/@mastra/*/dist/docs/`
|
|
|
|
|
|
(or the mastra skill) against the installed version before writing framework code.
|
|
|
|
|
|
- **Testing floor**: every invariant above has at least one automated test;
|
|
|
|
|
|
schema changes run golden-fixture validation on both TS and Python sides;
|
|
|
|
|
|
policy pack changes require their rule tests green.
|
|
|
|
|
|
|
|
|
|
|
|
## Do not
|
|
|
|
|
|
|
|
|
|
|
|
- Do not invent business parameters (market deadlines, envelope bounds, loss
|
|
|
|
|
|
budgets, buffer coefficients). Check `docs/open-questions.md`; if a needed
|
|
|
|
|
|
parameter is listed there, wire it as named config with a placeholder value
|
|
|
|
|
|
and a `// OPEN-QUESTION:` comment, and say so in your summary.
|
|
|
|
|
|
- Do not relitigate decisions recorded in `docs/adr/`. If a decision must
|
|
|
|
|
|
change, propose a superseding ADR.
|
|
|
|
|
|
- Do not create root-level summary docs that duplicate `docs/` content
|
|
|
|
|
|
(no architecture.md / tech-stack.md — README points to the sources of truth).
|
|
|
|
|
|
- Do not add a sixth agent, merge the safety chain into business workflows, or
|
|
|
|
|
|
bypass the proposal lifecycle for "internal" effects — these were considered
|
|
|
|
|
|
and rejected (see ADRs 0005–0007).
|
|
|
|
|
|
|
|
|
|
|
|
## Reading order for common tasks
|
|
|
|
|
|
|
|
|
|
|
|
| Task | Read first |
|
|
|
|
|
|
|---|---|
|
|
|
|
|
|
| Domain schemas / contracts | docs/11, docs/00 §4, GLOSSARY.md |
|
|
|
|
|
|
| Runtime, workflows, agents | docs/09, docs/02 |
|
|
|
|
|
|
| Safety chain, approvals, permits | docs/03, docs/01 invariants |
|
|
|
|
|
|
| Skill services (Python) | docs/05 §1–2, docs/11 §3 |
|
|
|
|
|
|
| Anything touching money or bids | docs/07 (scenario), docs/13 §1 |
|
|
|
|
|
|
| Eval harness | docs/12 |
|