vpp-ai-platform/CLAUDE.md
Thomas Bayes 8796faca63 M2: skill contracts, Python skill service, L2 eval harness with baseline
- packages/domain: ForecastRequest, BidOptimizationRequest/Result,
  ReportRequest, SkillReport (+ golden and invalid fixtures, exported to
  contracts/ and regenerated as pydantic models).
- skills-py/vpp_skills: FastAPI service with versioned registry; load/PV/
  price forecasts (same-day-type EWM point forecast, conformal residual
  quantiles — coverage test as acceptance gate); bid-optimization MILP on
  HiGHS (binary block participation, hard ledger energy bounds, exact
  Decimal fit of the rounded curve inside the bounds, revenue distribution
  over quantile paths); report generator whose every figure is a
  {tool_call_id, path} reference, with a verifier. 48 tests incl. hypothesis
  property test that bids respect ledger constraints.
- packages/services: LedgerService.dayAheadBounds (the P7 cascade band
  handed to the optimizer); Decimal resolved once for CJS/ESM interop.
- packages/evals: L2 metrics (MAPE, nRMSE, coverage, direction accuracy,
  naive/hindsight revenue baselines), HTTP skill client, rolling-origin
  harness that pushes each bid through the real ledger, CLI with
  --check/--write-baseline; committed baseline on the SYNTHETIC dataset
  (no historical Hubei data yet — baselines measure the harness, not KPI).
- CI: evals job boots the skill service and fails on baseline digest drift.
- docs/open-questions: A6 (flexibility marginal cost = offer floor); A4/B6
  wired as placeholders. README/CLAUDE.md status → M2 done, M3 next.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:29:08 -04:00

81 lines
4.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# CLAUDE.md — agent operating manual
Design docs live in `docs/` (00–13, Chinese). They are the specification; this
file distills what is **binding** when writing code. When code and docs conflict,
docs win — or write an ADR changing the doc first.
## Current phase
Docs complete; M1 and M2 implemented (schemas + contracts pipeline; ledger,
snapshot, time-series, relational, ingestion services; Python skill service
with forecast/MILP/report skills; L2 eval harness with committed baseline).
Storage is in-memory reference semantics — persistent adapters arrive with M3.
The eval dataset is synthetic until historical data lands. Next is M3. Build
order is ROADMAP.md (M1→M5). A skill change that moves an L2 metric must
update the baseline in the same change (`npm run eval -w @vpp/evals -- --write-baseline`).
Do not start a milestone's work before its predecessor's acceptance criteria are
testable, and do not build phase-2 items (edge control links, federation,
interaction/load-control agents) unless explicitly asked.
## Hard rules (from docs/01 invariants — treat as review criteria)
1. **LLM never computes numbers.** Any numeric field in a `Proposal` payload must
be a `{toolCallId, path}` reference into recorded tool-call lineage,
dereferenced by the assembler. If you find yourself parsing a number out of
LLM text into a payload, stop — that's the architecture's one forbidden move.
2. **No LLM in any control path.** Nothing under `services/` or the execution
path may import or await an LLM call. Periodic workflows must run with the
LLM backend down (invariant I6 — there is a test for this).
3. **External effects only via `(Proposal, ExecutionPermit)`.** Gateways/adapters
never accept a bare plan. Permits are short-lived and revocable.
4. **AI cannot approve.** No code path may transition a Proposal to
APPROVED/AUTHORIZED without either a matched envelope + passing checks, or a
human `resume()` with an approver identity distinct from the origination chain.
5. **Everything auditable.** State transitions, tool calls, approvals, permits,
and gateway receipts are event-sourced. Snapshots referenced by lineage are
immutable.
6. **Schemas live in `packages/domain` only.** Python models are generated from
`contracts/*.schema.json` — never hand-edit generated files, never define a
business object schema anywhere else.
## Conventions
- **Naming**: use the canonical code names in GLOSSARY.md. Do not invent new
English names for domain terms that already have one.
- **Language**: architecture docs Chinese; code, identifiers, comments, commit
messages, ADRs English.
- **Data representation** (docs/11 §3.3): money/energy/prices as fixed-point
decimal **strings**; units in field names (`power_mw`, `price_yuan_per_mwh`);
timestamps ISO8601 UTC; market intervals as `{date, interval_index}` with
`interval_minutes` explicit; all IDs strings; enums UPPER_SNAKE string literals.
- **Mastra**: do not trust memorized APIs — check `node_modules/@mastra/*/dist/docs/`
(or the mastra skill) against the installed version before writing framework code.
- **Testing floor**: every invariant above has at least one automated test;
schema changes run golden-fixture validation on both TS and Python sides;
policy pack changes require their rule tests green.
## Do not
- Do not invent business parameters (market deadlines, envelope bounds, loss
budgets, buffer coefficients). Check `docs/open-questions.md`; if a needed
parameter is listed there, wire it as named config with a placeholder value
and a `// OPEN-QUESTION:` comment, and say so in your summary.
- Do not relitigate decisions recorded in `docs/adr/`. If a decision must
change, propose a superseding ADR.
- Do not create root-level summary docs that duplicate `docs/` content
(no architecture.md / tech-stack.md — README points to the sources of truth).
- Do not add a sixth agent, merge the safety chain into business workflows, or
bypass the proposal lifecycle for "internal" effects — these were considered
and rejected (see ADRs 0005–0007).
## Reading order for common tasks
| Task | Read first |
|---|---|
| Domain schemas / contracts | docs/11, docs/00 §4, GLOSSARY.md |
| Runtime, workflows, agents | docs/09, docs/02 |
| Safety chain, approvals, permits | docs/03, docs/01 invariants |
| Skill services (Python) | docs/05 §1–2, docs/11 §3 |
| Anything touching money or bids | docs/07 (scenario), docs/13 §1 |
| Eval harness | docs/12 |