vpp-ai-platform/CLAUDE.md
Thomas Bayes 381b6d3521
Some checks failed
ci / typescript (push) Has been cancelled
ci / python (push) Has been cancelled
ci / evals (push) Has been cancelled
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).

Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.

Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.

Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.

Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.

Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.

Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00

97 lines
5.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# CLAUDE.md — agent operating manual
Design docs live in `docs/` (00–13, Chinese). They are the specification; this
file distills what is **binding** when writing code. When code and docs conflict,
docs win — or write an ADR changing the doc first.
## Current phase
Docs complete; M1–M5 implemented: schemas + contracts pipeline; deterministic
services (ledger, stores, policy engine incl. dispatch rules, envelope with
deviation-streak suspension, authority/permits, file-export + shadow + simulation
gateways, events, lineage assemblers, revenue + power-balance simulation,
review service, Case Desk, kill-switch hierarchy L0–L4 with loss breaker and
abnormal-day protocol, shadow clearing/scoring, KPI computation); Python skill
service (forecasts, bid MILP, report, potential assessment, dispatch
optimization); L2 eval harness with committed baseline (synthetic data); Mastra
runtime with the proposal-lifecycle safety chain (durable suspend/resume),
situation / bid / award-decomposition / review / envelope-review / shadow-close
workflows, trigger service (incl. shadow schedule and live-data ingestion), LLM
port with LLM-down mode, Case Desk API with insight cards, KPI dashboard and
breaker endpoints, shadow replay CLI (`npm run shadow -w @vpp/runtime`).
Storage is file-backed reference semantics (ledger, time series, counters and
repositories all persist under the data dir) — Postgres/Timescale adapters
later. Phase 1 is feature-complete; what remains is running the shadow period on
real Hubei data (`ROADMAP.md` M5 acceptance: 20+ consecutive days) and the
governance steps that need business answers (`docs/open-questions.md`). Build
order is ROADMAP.md (M1→M5). A skill change that moves an L2 metric
must update the baseline in the same change (`npm run eval -w @vpp/evals -- --write-baseline`).
Workflows are per-runtime factories (`createProposalLifecycle(ctx)` etc.) — never
module singletons, or a second Mastra instance steals their context.
Do not build phase-2 items (edge control links, federation,
interaction/load-control agents, programmatic bid submission, acting on
envelope-widening recommendations without the docs/03 envelope-review approval)
unless explicitly asked.
## Hard rules (from docs/01 invariants — treat as review criteria)
1. **LLM never computes numbers.** Any numeric field in a `Proposal` payload must
be a `{toolCallId, path}` reference into recorded tool-call lineage,
dereferenced by the assembler. If you find yourself parsing a number out of
LLM text into a payload, stop — that's the architecture's one forbidden move.
2. **No LLM in any control path.** Nothing under `services/` or the execution
path may import or await an LLM call. Periodic workflows must run with the
LLM backend down (invariant I6 — there is a test for this).
3. **External effects only via `(Proposal, ExecutionPermit)`.** Gateways/adapters
never accept a bare plan. Permits are short-lived and revocable.
4. **AI cannot approve.** No code path may transition a Proposal to
APPROVED/AUTHORIZED without either a matched envelope + passing checks, or a
human `resume()` with an approver identity distinct from the origination chain.
5. **Everything auditable.** State transitions, tool calls, approvals, permits,
and gateway receipts are event-sourced. Snapshots referenced by lineage are
immutable.
6. **Schemas live in `packages/domain` only.** Python models are generated from
`contracts/*.schema.json` — never hand-edit generated files, never define a
business object schema anywhere else.
## Conventions
- **Naming**: use the canonical code names in GLOSSARY.md. Do not invent new
English names for domain terms that already have one.
- **Language**: architecture docs Chinese; code, identifiers, comments, commit
messages, ADRs English.
- **Data representation** (docs/11 §3.3): money/energy/prices as fixed-point
decimal **strings**; units in field names (`power_mw`, `price_yuan_per_mwh`);
timestamps ISO8601 UTC; market intervals as `{date, interval_index}` with
`interval_minutes` explicit; all IDs strings; enums UPPER_SNAKE string literals.
- **Mastra**: do not trust memorized APIs — check `node_modules/@mastra/*/dist/docs/`
(or the mastra skill) against the installed version before writing framework code.
- **Testing floor**: every invariant above has at least one automated test;
schema changes run golden-fixture validation on both TS and Python sides;
policy pack changes require their rule tests green.
## Do not
- Do not invent business parameters (market deadlines, envelope bounds, loss
budgets, buffer coefficients). Check `docs/open-questions.md`; if a needed
parameter is listed there, wire it as named config with a placeholder value
and a `// OPEN-QUESTION:` comment, and say so in your summary.
- Do not relitigate decisions recorded in `docs/adr/`. If a decision must
change, propose a superseding ADR.
- Do not create root-level summary docs that duplicate `docs/` content
(no architecture.md / tech-stack.md — README points to the sources of truth).
- Do not add a sixth agent, merge the safety chain into business workflows, or
bypass the proposal lifecycle for "internal" effects — these were considered
and rejected (see ADRs 0005–0007).
## Reading order for common tasks
| Task | Read first |
|---|---|
| Domain schemas / contracts | docs/11, docs/00 §4, GLOSSARY.md |
| Runtime, workflows, agents | docs/09, docs/02 |
| Safety chain, approvals, permits | docs/03, docs/01 invariants |
| Skill services (Python) | docs/05 §1–2, docs/11 §3 |
| Anything touching money or bids | docs/07 (scenario), docs/13 §1 |
| Eval harness | docs/12 |