Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
5.6 KiB
CLAUDE.md — agent operating manual
Design docs live in docs/ (00–13, Chinese). They are the specification; this
file distills what is binding when writing code. When code and docs conflict,
docs win — or write an ADR changing the doc first.
Current phase
Docs complete; M1–M5 implemented: schemas + contracts pipeline; deterministic
services (ledger, stores, policy engine incl. dispatch rules, envelope with
deviation-streak suspension, authority/permits, file-export + shadow + simulation
gateways, events, lineage assemblers, revenue + power-balance simulation,
review service, Case Desk, kill-switch hierarchy L0–L4 with loss breaker and
abnormal-day protocol, shadow clearing/scoring, KPI computation); Python skill
service (forecasts, bid MILP, report, potential assessment, dispatch
optimization); L2 eval harness with committed baseline (synthetic data); Mastra
runtime with the proposal-lifecycle safety chain (durable suspend/resume),
situation / bid / award-decomposition / review / envelope-review / shadow-close
workflows, trigger service (incl. shadow schedule and live-data ingestion), LLM
port with LLM-down mode, Case Desk API with insight cards, KPI dashboard and
breaker endpoints, shadow replay CLI (npm run shadow -w @vpp/runtime).
Storage is file-backed reference semantics (ledger, time series, counters and
repositories all persist under the data dir) — Postgres/Timescale adapters
later. Phase 1 is feature-complete; what remains is running the shadow period on
real Hubei data (ROADMAP.md M5 acceptance: 20+ consecutive days) and the
governance steps that need business answers (docs/open-questions.md). Build
order is ROADMAP.md (M1→M5). A skill change that moves an L2 metric
must update the baseline in the same change (npm run eval -w @vpp/evals -- --write-baseline).
Workflows are per-runtime factories (createProposalLifecycle(ctx) etc.) — never
module singletons, or a second Mastra instance steals their context.
Do not build phase-2 items (edge control links, federation,
interaction/load-control agents, programmatic bid submission, acting on
envelope-widening recommendations without the docs/03 envelope-review approval)
unless explicitly asked.
Hard rules (from docs/01 invariants — treat as review criteria)
- LLM never computes numbers. Any numeric field in a
Proposalpayload must be a{toolCallId, path}reference into recorded tool-call lineage, dereferenced by the assembler. If you find yourself parsing a number out of LLM text into a payload, stop — that's the architecture's one forbidden move. - No LLM in any control path. Nothing under
services/or the execution path may import or await an LLM call. Periodic workflows must run with the LLM backend down (invariant I6 — there is a test for this). - External effects only via
(Proposal, ExecutionPermit). Gateways/adapters never accept a bare plan. Permits are short-lived and revocable. - AI cannot approve. No code path may transition a Proposal to
APPROVED/AUTHORIZED without either a matched envelope + passing checks, or a
human
resume()with an approver identity distinct from the origination chain. - Everything auditable. State transitions, tool calls, approvals, permits, and gateway receipts are event-sourced. Snapshots referenced by lineage are immutable.
- Schemas live in
packages/domainonly. Python models are generated fromcontracts/*.schema.json— never hand-edit generated files, never define a business object schema anywhere else.
Conventions
- Naming: use the canonical code names in GLOSSARY.md. Do not invent new English names for domain terms that already have one.
- Language: architecture docs Chinese; code, identifiers, comments, commit messages, ADRs English.
- Data representation (docs/11 §3.3): money/energy/prices as fixed-point
decimal strings; units in field names (
power_mw,price_yuan_per_mwh); timestamps ISO8601 UTC; market intervals as{date, interval_index}withinterval_minutesexplicit; all IDs strings; enums UPPER_SNAKE string literals. - Mastra: do not trust memorized APIs — check
node_modules/@mastra/*/dist/docs/(or the mastra skill) against the installed version before writing framework code. - Testing floor: every invariant above has at least one automated test; schema changes run golden-fixture validation on both TS and Python sides; policy pack changes require their rule tests green.
Do not
- Do not invent business parameters (market deadlines, envelope bounds, loss
budgets, buffer coefficients). Check
docs/open-questions.md; if a needed parameter is listed there, wire it as named config with a placeholder value and a// OPEN-QUESTION:comment, and say so in your summary. - Do not relitigate decisions recorded in
docs/adr/. If a decision must change, propose a superseding ADR. - Do not create root-level summary docs that duplicate
docs/content (no architecture.md / tech-stack.md — README points to the sources of truth). - Do not add a sixth agent, merge the safety chain into business workflows, or bypass the proposal lifecycle for "internal" effects — these were considered and rejected (see ADRs 0005–0007).
Reading order for common tasks
| Task | Read first |
|---|---|
| Domain schemas / contracts | docs/11, docs/00 §4, GLOSSARY.md |
| Runtime, workflows, agents | docs/09, docs/02 |
| Safety chain, approvals, permits | docs/03, docs/01 invariants |
| Skill services (Python) | docs/05 §1–2, docs/11 §3 |
| Anything touching money or bids | docs/07 (scenario), docs/13 §1 |
| Eval harness | docs/12 |