vpp-ai-platform/CLAUDE.md
Thomas Bayes 23fdf49ff2
Some checks are pending
ci / typescript (push) Waiting to run
ci / python (push) Waiting to run
ci / evals (push) Blocked by required conditions
M4: resource agent, envelopes live, review loop, insight cards
- packages/domain: AwardNotice, DISPATCH_PLAN proposal payload, ExecutionReport,
  MeteringRecord, potential-assessment and dispatch-optimization contracts,
  ReviewFinding (+ typed writebacks), SemanticMemoryEntry,
  EnvelopeChangeRequest, InsightCard — exported with fixtures on both sides.
- skills-py: potential-assessment (certified × rolling fulfilment, evidence
  days) and dispatch-optimization (per-interval LP on HiGHS, shortfall
  reported) skills + routes + tests.
- packages/services: dispatch rules in the policy pack (over-allocation,
  award anchor, lineage integrity for allocations); PowerBalanceSimulator;
  SimulationGateway (permit-only, idempotent, seeded execute → ExecutionReports);
  envelope deviation-streak suspension + apply(); ReviewService (attribution,
  reliability EWMA writeback, semantic memory, envelope recommendations as
  change requests); dispatch assembler; skill client methods.
- packages/runtime: resource agent; award-decomposition, review and
  envelope-review workflows; lifecycle selects simulator/gateway by proposal
  type; trigger hooks for awards, execution reports, metering; decide()
  resumes either lifecycle or envelope-review runs; insight cards API.
- Tests: docs/07 D-1 16:00 and D+1 end to end; reliability score 0.9 → 0.880
  and the next assessment de-rates capacity; envelope suspension on a seeded
  3-day streak; WIDEN request applied only by a human. 184 TS + 80 Python.
- docs/open-questions: B10 (reliability/potential parameters). README and
  CLAUDE.md status → M4 done, M5 next.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 19:31:12 -04:00

5.0 KiB
Raw Blame History

CLAUDE.md — agent operating manual

Design docs live in docs/ (00–13, Chinese). They are the specification; this file distills what is binding when writing code. When code and docs conflict, docs win — or write an ADR changing the doc first.

Current phase

Docs complete; M1–M4 implemented: schemas + contracts pipeline; deterministic services (ledger, stores, policy engine incl. dispatch rules, envelope with deviation-streak suspension, authority/permits, file-export + simulation gateways, events, lineage assemblers, revenue + power-balance simulation, review service, Case Desk); Python skill service (forecasts, bid MILP, report, potential assessment, dispatch optimization); L2 eval harness with committed baseline (synthetic data); Mastra runtime with the proposal-lifecycle safety chain (durable suspend/resume), situation / bid / award-decomposition / review / envelope-review workflows, trigger service, LLM port with LLM-down mode, Case Desk API with insight cards. Storage is file-backed reference semantics — Postgres/Timescale adapters later. Next is M5 (shadow run). Build order is ROADMAP.md (M1→M5). A skill change that moves an L2 metric must update the baseline in the same change (npm run eval -w @vpp/evals -- --write-baseline). Workflows are per-runtime factories (createProposalLifecycle(ctx) etc.) — never module singletons, or a second Mastra instance steals their context. Do not start a milestone's work before its predecessor's acceptance criteria are testable, and do not build phase-2 items (edge control links, federation, interaction/load-control agents) unless explicitly asked.

Hard rules (from docs/01 invariants — treat as review criteria)

  1. LLM never computes numbers. Any numeric field in a Proposal payload must be a {toolCallId, path} reference into recorded tool-call lineage, dereferenced by the assembler. If you find yourself parsing a number out of LLM text into a payload, stop — that's the architecture's one forbidden move.
  2. No LLM in any control path. Nothing under services/ or the execution path may import or await an LLM call. Periodic workflows must run with the LLM backend down (invariant I6 — there is a test for this).
  3. External effects only via (Proposal, ExecutionPermit). Gateways/adapters never accept a bare plan. Permits are short-lived and revocable.
  4. AI cannot approve. No code path may transition a Proposal to APPROVED/AUTHORIZED without either a matched envelope + passing checks, or a human resume() with an approver identity distinct from the origination chain.
  5. Everything auditable. State transitions, tool calls, approvals, permits, and gateway receipts are event-sourced. Snapshots referenced by lineage are immutable.
  6. Schemas live in packages/domain only. Python models are generated from contracts/*.schema.json — never hand-edit generated files, never define a business object schema anywhere else.

Conventions

  • Naming: use the canonical code names in GLOSSARY.md. Do not invent new English names for domain terms that already have one.
  • Language: architecture docs Chinese; code, identifiers, comments, commit messages, ADRs English.
  • Data representation (docs/11 §3.3): money/energy/prices as fixed-point decimal strings; units in field names (power_mw, price_yuan_per_mwh); timestamps ISO8601 UTC; market intervals as {date, interval_index} with interval_minutes explicit; all IDs strings; enums UPPER_SNAKE string literals.
  • Mastra: do not trust memorized APIs — check node_modules/@mastra/*/dist/docs/ (or the mastra skill) against the installed version before writing framework code.
  • Testing floor: every invariant above has at least one automated test; schema changes run golden-fixture validation on both TS and Python sides; policy pack changes require their rule tests green.

Do not

  • Do not invent business parameters (market deadlines, envelope bounds, loss budgets, buffer coefficients). Check docs/open-questions.md; if a needed parameter is listed there, wire it as named config with a placeholder value and a // OPEN-QUESTION: comment, and say so in your summary.
  • Do not relitigate decisions recorded in docs/adr/. If a decision must change, propose a superseding ADR.
  • Do not create root-level summary docs that duplicate docs/ content (no architecture.md / tech-stack.md — README points to the sources of truth).
  • Do not add a sixth agent, merge the safety chain into business workflows, or bypass the proposal lifecycle for "internal" effects — these were considered and rejected (see ADRs 0005–0007).

Reading order for common tasks

Task Read first
Domain schemas / contracts docs/11, docs/00 §4, GLOSSARY.md
Runtime, workflows, agents docs/09, docs/02
Safety chain, approvals, permits docs/03, docs/01 invariants
Skill services (Python) docs/05 §1–2, docs/11 §3
Anything touching money or bids docs/07 (scenario), docs/13 §1
Eval harness docs/12