vpp-ai-platform/README.md
Thomas Bayes 381b6d3521
Some checks failed
ci / typescript (push) Has been cancelled
ci / python (push) Has been cancelled
ci / evals (push) Has been cancelled
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).

Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.

Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.

Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.

Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.

Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.

Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00

118 lines
7.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# VPP AI Platform · 虚拟电厂多时空协同智能运营平台
AI-assisted operations platform for a virtual power plant (VPP) in Hubei, China:
five LLM agents propose market bids, resource dispatch, and load-control plans;
a deterministic safety chain (rule check → simulation → envelope/human approval →
execution permit) governs everything before any external effect. **The LLM never
computes numbers and never touches the second-level control loop.**
**Status: M1–M5 implemented (phase 1 feature-complete).** Design docs 00–14 are complete;
implementation follows [ROADMAP.md](ROADMAP.md). Present today:
- `packages/domain` — zod schemas for every business and safety-chain object, exported to
`contracts/` and regenerated as pydantic models (`skills-py/vpp_contracts`).
- `packages/services` — deterministic services: ledger (with the P7 cascade), snapshot,
time-series, relational and ingestion stores; policy engine + `hubei-spot-bidding` pack;
envelope gate; authority (fresh check, permits, revocation); file-export gateway;
event store/bus; lineage recorder + P2 assembler; revenue-scenario simulation; Case Desk.
- `skills-py/vpp_skills` — Python skill service: load/PV/price forecasts with calibrated
quantiles, bid-optimization MILP, report generator.
- `packages/evals` — L2 eval harness with a committed baseline (on a **synthetic** dataset).
- `packages/runtime` — Mastra runtime: `proposal-lifecycle` (the safety chain, suspend/resume
for human approval, durable across restarts), `day-ahead-situation`, `day-ahead-bid`;
`award-decomposition` (resource agent: potential assessment → dispatch optimization →
DISPATCH_PLAN through the same chain to a simulation gateway), `review` (D+1
plan-vs-actual attribution → ReviewFinding → semantic memory + reliability writebacks +
envelope change requests), `envelope-review` (the only path that changes an envelope,
human-approved); envelope deviation-streak suspension; trigger service (scheduled / event /
manual via router agent); provider-abstracted LLM port with an LLM-down mode; Case Desk
HTTP API incl. AI insight cards. Bid release is a file export for manual upload (degraded
channel by design) or, in shadow mode, a recorded-never-submitted receipt.
- **Shadow run (M5, the phase-1 acceptance form)**: `VPP_MODE=SHADOW` runs the full loop on
live data with every external effect simulated — bids recorded, never submitted; the
market's answer simulated from the actual clearing price; dispatch to the simulation
gateway; D+1 review — and scores every day shadow-vs-human-vs-hindsight (`ShadowDayRecord`)
with lineage-completeness audit, plus an auto-generated docs/12 §4 KPI report. Kill-switch
hierarchy L0–L4 (`BreakerService`: permit revocation, envelope suspension, loss breaker,
channel breaker, AI-off) with the abnormal-day protocol and an auditable drill
(`POST /breakers/drill`). Replay driver: `npm run shadow -w @vpp/runtime -- --days 21`.
All eight docs/01 invariants have automated tests (`packages/services/test/chain.test.ts`,
`packages/runtime/test/lifecycle.test.ts`); the docs/07 D-1 16:00 and D+1 sections run in
`packages/runtime/test/m4.test.ts`; the ROADMAP M5 acceptance (21 consecutive shadow days with
complete lineage, KPI report, one widening recommendation not acted on) and the kill-switch
drill run in `packages/runtime/test/m5.test.ts`. Storage is file-backed reference semantics;
Postgres/Timescale adapters and the programmatic trading-platform channel are later work.
Next: run the shadow period on real Hubei data and close the business parameters in
`docs/open-questions.md`.
## Development
Requires Node 24 and Python 3.11 (`skills-py/.python-version`; the generated
pydantic models use `StrEnum` and PEP 604 unions).
```sh
npm ci
npm run check # typecheck, re-export contracts, run TS tests
npm run eval -w @vpp/evals -- --check # L2 harness vs baseline (needs the skill service below)
npm run start -w @vpp/runtime # runtime + Case Desk API on :4100 (VPP_MODE=SHADOW default; VPP_LLM_MODEL unset = LLM-down mode)
npm run shadow -w @vpp/runtime -- --days 21 # replay 21 shadow days from the dataset through the live skill service
cd skills-py
uv venv --python 3.11 .venv && uv pip install -r requirements.txt # or python3.11 -m venv
source .venv/bin/activate
bash scripts/generate_models.sh # regenerate pydantic models (committed, never hand-edited)
python -m pytest -q
python -m uvicorn vpp_skills.app:app --port 8000 # skill service for the eval harness
```
CI (`.github/workflows/ci.yml`) runs both sides and fails if `contracts/` or
`skills-py/vpp_contracts` are not regenerated after a schema change.
## Orientation
| You are… | Start with |
|---|---|
| A coding agent about to implement | [CLAUDE.md](CLAUDE.md), then docs/00, 01, 09, 11 |
| New to the project | [docs/00-overview.md](docs/00-overview.md) → [docs/07-scenario-walkthrough.md](docs/07-scenario-walkthrough.md) (the end-to-end reference scenario) |
| Reviewing the business case | [proposal.md](proposal.md) (申报材料, source of requirements) |
| Looking for a settled decision | [docs/adr/](docs/adr/) |
| Wondering what's still undecided | [docs/open-questions.md](docs/open-questions.md) |
## Document map (docs are in Chinese; implementation-facing files in English)
| Doc | Content |
|---|---|
| [00-overview](docs/00-overview.md) | System context, two-plane architecture, business objects |
| [01-principles](docs/01-principles.md) | 9 principles + 8 hard invariants (binding for all code) |
| [02-cognitive-plane](docs/02-cognitive-plane.md) | Five agents, Runtime, memory, Case Desk |
| [03-safety-chain](docs/03-safety-chain.md) | Proposal state machine, envelopes, permits, staleness |
| [04-control-plane](docs/04-control-plane.md) | Execution engine, edge autonomy, time/space cascades |
| [05-skills-and-data](docs/05-skills-and-data.md) | Skill contracts, five-store data layer, policy packs |
| [06-integration](docs/06-integration.md) | External system boundaries and degraded channels |
| [14-shadow-run-runbook](docs/14-shadow-run-runbook.md) | Shadow-run operations, kill-switch levels (trigger / authority / recovery), KPI definitions as implemented |
| [07-scenario-walkthrough](docs/07-scenario-walkthrough.md) | Day-ahead spot bidding, D-1 → D → D+1 |
| [08-implementation](docs/08-implementation.md) | Stack, LLM abstraction, deployment, milestones |
| [09-runtime-implementation](docs/09-runtime-implementation.md) | Runtime on Mastra: workflows, suspend/resume, lineage |
| [10-federation](docs/10-federation.md) | Cross-province boundary: signed artifacts only |
| [11-contracts](docs/11-contracts.md) | Ports, canonical objects, TS↔Python contract pipeline |
| [12-evaluation](docs/12-evaluation.md) | Four-layer evals, change gates, KPI definitions |
| [13-risks-failure-modes](docs/13-risks-failure-modes.md) | FMEA, top-5 risks, kill-switch hierarchy |
Supporting: [brainstorming.md](brainstorming.md) is an independent peer review
whose findings were integrated (see docs/01 invariants note);
[GLOSSARY.md](GLOSSARY.md) maps Chinese domain terms to canonical code names.
## Target repository layout (from docs/09 §7)
```
packages/
├── domain/ # zod schemas — single source of truth for all business objects
├── runtime/ # Mastra instance, workflows, agents, tool registry, triggers
├── services/ # deterministic services: policy engine, envelopes, ledger, authority
├── adapters/ # anti-corruption layers: trading platform, dispatch, metering
├── evals/ # eval harness, datasets, judges (docs/12 §5)
└── skills-py/ # Python skill services (forecasting, MILP optimization, simulation)
contracts/ # generated JSON Schema + golden fixtures (cross-language contract)
```