vpp-ai-platform/CLAUDE.md
Thomas Bayes 23fdf49ff2
Some checks are pending
ci / typescript (push) Waiting to run
ci / python (push) Waiting to run
ci / evals (push) Blocked by required conditions
M4: resource agent, envelopes live, review loop, insight cards
- packages/domain: AwardNotice, DISPATCH_PLAN proposal payload, ExecutionReport,
  MeteringRecord, potential-assessment and dispatch-optimization contracts,
  ReviewFinding (+ typed writebacks), SemanticMemoryEntry,
  EnvelopeChangeRequest, InsightCard — exported with fixtures on both sides.
- skills-py: potential-assessment (certified × rolling fulfilment, evidence
  days) and dispatch-optimization (per-interval LP on HiGHS, shortfall
  reported) skills + routes + tests.
- packages/services: dispatch rules in the policy pack (over-allocation,
  award anchor, lineage integrity for allocations); PowerBalanceSimulator;
  SimulationGateway (permit-only, idempotent, seeded execute → ExecutionReports);
  envelope deviation-streak suspension + apply(); ReviewService (attribution,
  reliability EWMA writeback, semantic memory, envelope recommendations as
  change requests); dispatch assembler; skill client methods.
- packages/runtime: resource agent; award-decomposition, review and
  envelope-review workflows; lifecycle selects simulator/gateway by proposal
  type; trigger hooks for awards, execution reports, metering; decide()
  resumes either lifecycle or envelope-review runs; insight cards API.
- Tests: docs/07 D-1 16:00 and D+1 end to end; reliability score 0.9 → 0.880
  and the next assessment de-rates capacity; envelope suspension on a seeded
  3-day streak; WIDEN request applied only by a human. 184 TS + 80 Python.
- docs/open-questions: B10 (reliability/potential parameters). README and
  CLAUDE.md status → M4 done, M5 next.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 19:31:12 -04:00

88 lines
5.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# CLAUDE.md — agent operating manual
Design docs live in `docs/` (00–13, Chinese). They are the specification; this
file distills what is **binding** when writing code. When code and docs conflict,
docs win — or write an ADR changing the doc first.
## Current phase
Docs complete; M1–M4 implemented: schemas + contracts pipeline; deterministic
services (ledger, stores, policy engine incl. dispatch rules, envelope with
deviation-streak suspension, authority/permits, file-export + simulation
gateways, events, lineage assemblers, revenue + power-balance simulation,
review service, Case Desk); Python skill service (forecasts, bid MILP, report,
potential assessment, dispatch optimization); L2 eval harness with committed
baseline (synthetic data); Mastra runtime with the proposal-lifecycle safety
chain (durable suspend/resume), situation / bid / award-decomposition / review /
envelope-review workflows, trigger service, LLM port with LLM-down mode, Case
Desk API with insight cards. Storage is file-backed reference semantics —
Postgres/Timescale adapters later. Next is M5 (shadow run). Build order is ROADMAP.md (M1→M5). A skill change that moves an L2 metric
must update the baseline in the same change (`npm run eval -w @vpp/evals -- --write-baseline`).
Workflows are per-runtime factories (`createProposalLifecycle(ctx)` etc.) — never
module singletons, or a second Mastra instance steals their context.
Do not start a milestone's work before its predecessor's acceptance criteria are
testable, and do not build phase-2 items (edge control links, federation,
interaction/load-control agents) unless explicitly asked.
## Hard rules (from docs/01 invariants — treat as review criteria)
1. **LLM never computes numbers.** Any numeric field in a `Proposal` payload must
be a `{toolCallId, path}` reference into recorded tool-call lineage,
dereferenced by the assembler. If you find yourself parsing a number out of
LLM text into a payload, stop — that's the architecture's one forbidden move.
2. **No LLM in any control path.** Nothing under `services/` or the execution
path may import or await an LLM call. Periodic workflows must run with the
LLM backend down (invariant I6 — there is a test for this).
3. **External effects only via `(Proposal, ExecutionPermit)`.** Gateways/adapters
never accept a bare plan. Permits are short-lived and revocable.
4. **AI cannot approve.** No code path may transition a Proposal to
APPROVED/AUTHORIZED without either a matched envelope + passing checks, or a
human `resume()` with an approver identity distinct from the origination chain.
5. **Everything auditable.** State transitions, tool calls, approvals, permits,
and gateway receipts are event-sourced. Snapshots referenced by lineage are
immutable.
6. **Schemas live in `packages/domain` only.** Python models are generated from
`contracts/*.schema.json` — never hand-edit generated files, never define a
business object schema anywhere else.
## Conventions
- **Naming**: use the canonical code names in GLOSSARY.md. Do not invent new
English names for domain terms that already have one.
- **Language**: architecture docs Chinese; code, identifiers, comments, commit
messages, ADRs English.
- **Data representation** (docs/11 §3.3): money/energy/prices as fixed-point
decimal **strings**; units in field names (`power_mw`, `price_yuan_per_mwh`);
timestamps ISO8601 UTC; market intervals as `{date, interval_index}` with
`interval_minutes` explicit; all IDs strings; enums UPPER_SNAKE string literals.
- **Mastra**: do not trust memorized APIs — check `node_modules/@mastra/*/dist/docs/`
(or the mastra skill) against the installed version before writing framework code.
- **Testing floor**: every invariant above has at least one automated test;
schema changes run golden-fixture validation on both TS and Python sides;
policy pack changes require their rule tests green.
## Do not
- Do not invent business parameters (market deadlines, envelope bounds, loss
budgets, buffer coefficients). Check `docs/open-questions.md`; if a needed
parameter is listed there, wire it as named config with a placeholder value
and a `// OPEN-QUESTION:` comment, and say so in your summary.
- Do not relitigate decisions recorded in `docs/adr/`. If a decision must
change, propose a superseding ADR.
- Do not create root-level summary docs that duplicate `docs/` content
(no architecture.md / tech-stack.md — README points to the sources of truth).
- Do not add a sixth agent, merge the safety chain into business workflows, or
bypass the proposal lifecycle for "internal" effects — these were considered
and rejected (see ADRs 0005–0007).
## Reading order for common tasks
| Task | Read first |
|---|---|
| Domain schemas / contracts | docs/11, docs/00 §4, GLOSSARY.md |
| Runtime, workflows, agents | docs/09, docs/02 |
| Safety chain, approvals, permits | docs/03, docs/01 invariants |
| Skill services (Python) | docs/05 §1–2, docs/11 §3 |
| Anything touching money or bids | docs/07 (scenario), docs/13 §1 |
| Eval harness | docs/12 |