Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
118 lines
7.7 KiB
Markdown
118 lines
7.7 KiB
Markdown
# VPP AI Platform · 虚拟电厂多时空协同智能运营平台
|
||
|
||
AI-assisted operations platform for a virtual power plant (VPP) in Hubei, China:
|
||
five LLM agents propose market bids, resource dispatch, and load-control plans;
|
||
a deterministic safety chain (rule check → simulation → envelope/human approval →
|
||
execution permit) governs everything before any external effect. **The LLM never
|
||
computes numbers and never touches the second-level control loop.**
|
||
|
||
**Status: M1–M5 implemented (phase 1 feature-complete).** Design docs 00–14 are complete;
|
||
implementation follows [ROADMAP.md](ROADMAP.md). Present today:
|
||
|
||
- `packages/domain` — zod schemas for every business and safety-chain object, exported to
|
||
`contracts/` and regenerated as pydantic models (`skills-py/vpp_contracts`).
|
||
- `packages/services` — deterministic services: ledger (with the P7 cascade), snapshot,
|
||
time-series, relational and ingestion stores; policy engine + `hubei-spot-bidding` pack;
|
||
envelope gate; authority (fresh check, permits, revocation); file-export gateway;
|
||
event store/bus; lineage recorder + P2 assembler; revenue-scenario simulation; Case Desk.
|
||
- `skills-py/vpp_skills` — Python skill service: load/PV/price forecasts with calibrated
|
||
quantiles, bid-optimization MILP, report generator.
|
||
- `packages/evals` — L2 eval harness with a committed baseline (on a **synthetic** dataset).
|
||
- `packages/runtime` — Mastra runtime: `proposal-lifecycle` (the safety chain, suspend/resume
|
||
for human approval, durable across restarts), `day-ahead-situation`, `day-ahead-bid`;
|
||
`award-decomposition` (resource agent: potential assessment → dispatch optimization →
|
||
DISPATCH_PLAN through the same chain to a simulation gateway), `review` (D+1
|
||
plan-vs-actual attribution → ReviewFinding → semantic memory + reliability writebacks +
|
||
envelope change requests), `envelope-review` (the only path that changes an envelope,
|
||
human-approved); envelope deviation-streak suspension; trigger service (scheduled / event /
|
||
manual via router agent); provider-abstracted LLM port with an LLM-down mode; Case Desk
|
||
HTTP API incl. AI insight cards. Bid release is a file export for manual upload (degraded
|
||
channel by design) or, in shadow mode, a recorded-never-submitted receipt.
|
||
- **Shadow run (M5, the phase-1 acceptance form)**: `VPP_MODE=SHADOW` runs the full loop on
|
||
live data with every external effect simulated — bids recorded, never submitted; the
|
||
market's answer simulated from the actual clearing price; dispatch to the simulation
|
||
gateway; D+1 review — and scores every day shadow-vs-human-vs-hindsight (`ShadowDayRecord`)
|
||
with lineage-completeness audit, plus an auto-generated docs/12 §4 KPI report. Kill-switch
|
||
hierarchy L0–L4 (`BreakerService`: permit revocation, envelope suspension, loss breaker,
|
||
channel breaker, AI-off) with the abnormal-day protocol and an auditable drill
|
||
(`POST /breakers/drill`). Replay driver: `npm run shadow -w @vpp/runtime -- --days 21`.
|
||
|
||
All eight docs/01 invariants have automated tests (`packages/services/test/chain.test.ts`,
|
||
`packages/runtime/test/lifecycle.test.ts`); the docs/07 D-1 16:00 and D+1 sections run in
|
||
`packages/runtime/test/m4.test.ts`; the ROADMAP M5 acceptance (21 consecutive shadow days with
|
||
complete lineage, KPI report, one widening recommendation not acted on) and the kill-switch
|
||
drill run in `packages/runtime/test/m5.test.ts`. Storage is file-backed reference semantics;
|
||
Postgres/Timescale adapters and the programmatic trading-platform channel are later work.
|
||
Next: run the shadow period on real Hubei data and close the business parameters in
|
||
`docs/open-questions.md`.
|
||
|
||
## Development
|
||
|
||
Requires Node 24 and Python 3.11 (`skills-py/.python-version`; the generated
|
||
pydantic models use `StrEnum` and PEP 604 unions).
|
||
|
||
```sh
|
||
npm ci
|
||
npm run check # typecheck, re-export contracts, run TS tests
|
||
npm run eval -w @vpp/evals -- --check # L2 harness vs baseline (needs the skill service below)
|
||
npm run start -w @vpp/runtime # runtime + Case Desk API on :4100 (VPP_MODE=SHADOW default; VPP_LLM_MODEL unset = LLM-down mode)
|
||
npm run shadow -w @vpp/runtime -- --days 21 # replay 21 shadow days from the dataset through the live skill service
|
||
|
||
cd skills-py
|
||
uv venv --python 3.11 .venv && uv pip install -r requirements.txt # or python3.11 -m venv
|
||
source .venv/bin/activate
|
||
bash scripts/generate_models.sh # regenerate pydantic models (committed, never hand-edited)
|
||
python -m pytest -q
|
||
python -m uvicorn vpp_skills.app:app --port 8000 # skill service for the eval harness
|
||
```
|
||
|
||
CI (`.github/workflows/ci.yml`) runs both sides and fails if `contracts/` or
|
||
`skills-py/vpp_contracts` are not regenerated after a schema change.
|
||
|
||
## Orientation
|
||
|
||
| You are… | Start with |
|
||
|---|---|
|
||
| A coding agent about to implement | [CLAUDE.md](CLAUDE.md), then docs/00, 01, 09, 11 |
|
||
| New to the project | [docs/00-overview.md](docs/00-overview.md) → [docs/07-scenario-walkthrough.md](docs/07-scenario-walkthrough.md) (the end-to-end reference scenario) |
|
||
| Reviewing the business case | [proposal.md](proposal.md) (申报材料, source of requirements) |
|
||
| Looking for a settled decision | [docs/adr/](docs/adr/) |
|
||
| Wondering what's still undecided | [docs/open-questions.md](docs/open-questions.md) |
|
||
|
||
## Document map (docs are in Chinese; implementation-facing files in English)
|
||
|
||
| Doc | Content |
|
||
|---|---|
|
||
| [00-overview](docs/00-overview.md) | System context, two-plane architecture, business objects |
|
||
| [01-principles](docs/01-principles.md) | 9 principles + 8 hard invariants (binding for all code) |
|
||
| [02-cognitive-plane](docs/02-cognitive-plane.md) | Five agents, Runtime, memory, Case Desk |
|
||
| [03-safety-chain](docs/03-safety-chain.md) | Proposal state machine, envelopes, permits, staleness |
|
||
| [04-control-plane](docs/04-control-plane.md) | Execution engine, edge autonomy, time/space cascades |
|
||
| [05-skills-and-data](docs/05-skills-and-data.md) | Skill contracts, five-store data layer, policy packs |
|
||
| [06-integration](docs/06-integration.md) | External system boundaries and degraded channels |
|
||
| [14-shadow-run-runbook](docs/14-shadow-run-runbook.md) | Shadow-run operations, kill-switch levels (trigger / authority / recovery), KPI definitions as implemented |
|
||
| [07-scenario-walkthrough](docs/07-scenario-walkthrough.md) | Day-ahead spot bidding, D-1 → D → D+1 |
|
||
| [08-implementation](docs/08-implementation.md) | Stack, LLM abstraction, deployment, milestones |
|
||
| [09-runtime-implementation](docs/09-runtime-implementation.md) | Runtime on Mastra: workflows, suspend/resume, lineage |
|
||
| [10-federation](docs/10-federation.md) | Cross-province boundary: signed artifacts only |
|
||
| [11-contracts](docs/11-contracts.md) | Ports, canonical objects, TS↔Python contract pipeline |
|
||
| [12-evaluation](docs/12-evaluation.md) | Four-layer evals, change gates, KPI definitions |
|
||
| [13-risks-failure-modes](docs/13-risks-failure-modes.md) | FMEA, top-5 risks, kill-switch hierarchy |
|
||
|
||
Supporting: [brainstorming.md](brainstorming.md) is an independent peer review
|
||
whose findings were integrated (see docs/01 invariants note);
|
||
[GLOSSARY.md](GLOSSARY.md) maps Chinese domain terms to canonical code names.
|
||
|
||
## Target repository layout (from docs/09 §7)
|
||
|
||
```
|
||
packages/
|
||
├── domain/ # zod schemas — single source of truth for all business objects
|
||
├── runtime/ # Mastra instance, workflows, agents, tool registry, triggers
|
||
├── services/ # deterministic services: policy engine, envelopes, ledger, authority
|
||
├── adapters/ # anti-corruption layers: trading platform, dispatch, metering
|
||
├── evals/ # eval harness, datasets, judges (docs/12 §5)
|
||
└── skills-py/ # Python skill services (forecasting, MILP optimization, simulation)
|
||
contracts/ # generated JSON Schema + golden fixtures (cross-language contract)
|
||
```
|