Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
|
||
|---|---|---|
| .github/workflows | ||
| contracts | ||
| docs | ||
| packages | ||
| proposal-assets | ||
| skills-py | ||
| .gitignore | ||
| brainstorming.md | ||
| CLAUDE.md | ||
| GLOSSARY.md | ||
| package-lock.json | ||
| package.json | ||
| proposal.md | ||
| README.md | ||
| ROADMAP.md | ||
| tsconfig.base.json | ||
VPP AI Platform · 虚拟电厂多时空协同智能运营平台
AI-assisted operations platform for a virtual power plant (VPP) in Hubei, China: five LLM agents propose market bids, resource dispatch, and load-control plans; a deterministic safety chain (rule check → simulation → envelope/human approval → execution permit) governs everything before any external effect. The LLM never computes numbers and never touches the second-level control loop.
Status: M1–M5 implemented (phase 1 feature-complete). Design docs 00–14 are complete; implementation follows ROADMAP.md. Present today:
packages/domain— zod schemas for every business and safety-chain object, exported tocontracts/and regenerated as pydantic models (skills-py/vpp_contracts).packages/services— deterministic services: ledger (with the P7 cascade), snapshot, time-series, relational and ingestion stores; policy engine +hubei-spot-biddingpack; envelope gate; authority (fresh check, permits, revocation); file-export gateway; event store/bus; lineage recorder + P2 assembler; revenue-scenario simulation; Case Desk.skills-py/vpp_skills— Python skill service: load/PV/price forecasts with calibrated quantiles, bid-optimization MILP, report generator.packages/evals— L2 eval harness with a committed baseline (on a synthetic dataset).packages/runtime— Mastra runtime:proposal-lifecycle(the safety chain, suspend/resume for human approval, durable across restarts),day-ahead-situation,day-ahead-bid;award-decomposition(resource agent: potential assessment → dispatch optimization → DISPATCH_PLAN through the same chain to a simulation gateway),review(D+1 plan-vs-actual attribution → ReviewFinding → semantic memory + reliability writebacks + envelope change requests),envelope-review(the only path that changes an envelope, human-approved); envelope deviation-streak suspension; trigger service (scheduled / event / manual via router agent); provider-abstracted LLM port with an LLM-down mode; Case Desk HTTP API incl. AI insight cards. Bid release is a file export for manual upload (degraded channel by design) or, in shadow mode, a recorded-never-submitted receipt.- Shadow run (M5, the phase-1 acceptance form):
VPP_MODE=SHADOWruns the full loop on live data with every external effect simulated — bids recorded, never submitted; the market's answer simulated from the actual clearing price; dispatch to the simulation gateway; D+1 review — and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord) with lineage-completeness audit, plus an auto-generated docs/12 §4 KPI report. Kill-switch hierarchy L0–L4 (BreakerService: permit revocation, envelope suspension, loss breaker, channel breaker, AI-off) with the abnormal-day protocol and an auditable drill (POST /breakers/drill). Replay driver:npm run shadow -w @vpp/runtime -- --days 21.
All eight docs/01 invariants have automated tests (packages/services/test/chain.test.ts,
packages/runtime/test/lifecycle.test.ts); the docs/07 D-1 16:00 and D+1 sections run in
packages/runtime/test/m4.test.ts; the ROADMAP M5 acceptance (21 consecutive shadow days with
complete lineage, KPI report, one widening recommendation not acted on) and the kill-switch
drill run in packages/runtime/test/m5.test.ts. Storage is file-backed reference semantics;
Postgres/Timescale adapters and the programmatic trading-platform channel are later work.
Next: run the shadow period on real Hubei data and close the business parameters in
docs/open-questions.md.
Development
Requires Node 24 and Python 3.11 (skills-py/.python-version; the generated
pydantic models use StrEnum and PEP 604 unions).
npm ci
npm run check # typecheck, re-export contracts, run TS tests
npm run eval -w @vpp/evals -- --check # L2 harness vs baseline (needs the skill service below)
npm run start -w @vpp/runtime # runtime + Case Desk API on :4100 (VPP_MODE=SHADOW default; VPP_LLM_MODEL unset = LLM-down mode)
npm run shadow -w @vpp/runtime -- --days 21 # replay 21 shadow days from the dataset through the live skill service
cd skills-py
uv venv --python 3.11 .venv && uv pip install -r requirements.txt # or python3.11 -m venv
source .venv/bin/activate
bash scripts/generate_models.sh # regenerate pydantic models (committed, never hand-edited)
python -m pytest -q
python -m uvicorn vpp_skills.app:app --port 8000 # skill service for the eval harness
CI (.github/workflows/ci.yml) runs both sides and fails if contracts/ or
skills-py/vpp_contracts are not regenerated after a schema change.
Orientation
| You are… | Start with |
|---|---|
| A coding agent about to implement | CLAUDE.md, then docs/00, 01, 09, 11 |
| New to the project | docs/00-overview.md → docs/07-scenario-walkthrough.md (the end-to-end reference scenario) |
| Reviewing the business case | proposal.md (申报材料, source of requirements) |
| Looking for a settled decision | docs/adr/ |
| Wondering what's still undecided | docs/open-questions.md |
Document map (docs are in Chinese; implementation-facing files in English)
| Doc | Content |
|---|---|
| 00-overview | System context, two-plane architecture, business objects |
| 01-principles | 9 principles + 8 hard invariants (binding for all code) |
| 02-cognitive-plane | Five agents, Runtime, memory, Case Desk |
| 03-safety-chain | Proposal state machine, envelopes, permits, staleness |
| 04-control-plane | Execution engine, edge autonomy, time/space cascades |
| 05-skills-and-data | Skill contracts, five-store data layer, policy packs |
| 06-integration | External system boundaries and degraded channels |
| 14-shadow-run-runbook | Shadow-run operations, kill-switch levels (trigger / authority / recovery), KPI definitions as implemented |
| 07-scenario-walkthrough | Day-ahead spot bidding, D-1 → D → D+1 |
| 08-implementation | Stack, LLM abstraction, deployment, milestones |
| 09-runtime-implementation | Runtime on Mastra: workflows, suspend/resume, lineage |
| 10-federation | Cross-province boundary: signed artifacts only |
| 11-contracts | Ports, canonical objects, TS↔Python contract pipeline |
| 12-evaluation | Four-layer evals, change gates, KPI definitions |
| 13-risks-failure-modes | FMEA, top-5 risks, kill-switch hierarchy |
Supporting: brainstorming.md is an independent peer review whose findings were integrated (see docs/01 invariants note); GLOSSARY.md maps Chinese domain terms to canonical code names.
Target repository layout (from docs/09 §7)
packages/
├── domain/ # zod schemas — single source of truth for all business objects
├── runtime/ # Mastra instance, workflows, agents, tool registry, triggers
├── services/ # deterministic services: policy engine, envelopes, ledger, authority
├── adapters/ # anti-corruption layers: trading platform, dispatch, metering
├── evals/ # eval harness, datasets, judges (docs/12 §5)
└── skills-py/ # Python skill services (forecasting, MILP optimization, simulation)
contracts/ # generated JSON Schema + golden fixtures (cross-language contract)