vpp-ai-platform/README.md
Thomas Bayes 381b6d3521
Some checks failed
ci / typescript (push) Has been cancelled
ci / python (push) Has been cancelled
ci / evals (push) Has been cancelled
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).

Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.

Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.

Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.

Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.

Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.

Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00

7.7 KiB
Raw Permalink Blame History

VPP AI Platform · 虚拟电厂多时空协同智能运营平台

AI-assisted operations platform for a virtual power plant (VPP) in Hubei, China: five LLM agents propose market bids, resource dispatch, and load-control plans; a deterministic safety chain (rule check → simulation → envelope/human approval → execution permit) governs everything before any external effect. The LLM never computes numbers and never touches the second-level control loop.

Status: M1–M5 implemented (phase 1 feature-complete). Design docs 00–14 are complete; implementation follows ROADMAP.md. Present today:

  • packages/domain — zod schemas for every business and safety-chain object, exported to contracts/ and regenerated as pydantic models (skills-py/vpp_contracts).
  • packages/services — deterministic services: ledger (with the P7 cascade), snapshot, time-series, relational and ingestion stores; policy engine + hubei-spot-bidding pack; envelope gate; authority (fresh check, permits, revocation); file-export gateway; event store/bus; lineage recorder + P2 assembler; revenue-scenario simulation; Case Desk.
  • skills-py/vpp_skills — Python skill service: load/PV/price forecasts with calibrated quantiles, bid-optimization MILP, report generator.
  • packages/evals — L2 eval harness with a committed baseline (on a synthetic dataset).
  • packages/runtime — Mastra runtime: proposal-lifecycle (the safety chain, suspend/resume for human approval, durable across restarts), day-ahead-situation, day-ahead-bid; award-decomposition (resource agent: potential assessment → dispatch optimization → DISPATCH_PLAN through the same chain to a simulation gateway), review (D+1 plan-vs-actual attribution → ReviewFinding → semantic memory + reliability writebacks + envelope change requests), envelope-review (the only path that changes an envelope, human-approved); envelope deviation-streak suspension; trigger service (scheduled / event / manual via router agent); provider-abstracted LLM port with an LLM-down mode; Case Desk HTTP API incl. AI insight cards. Bid release is a file export for manual upload (degraded channel by design) or, in shadow mode, a recorded-never-submitted receipt.
  • Shadow run (M5, the phase-1 acceptance form): VPP_MODE=SHADOW runs the full loop on live data with every external effect simulated — bids recorded, never submitted; the market's answer simulated from the actual clearing price; dispatch to the simulation gateway; D+1 review — and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord) with lineage-completeness audit, plus an auto-generated docs/12 §4 KPI report. Kill-switch hierarchy L0–L4 (BreakerService: permit revocation, envelope suspension, loss breaker, channel breaker, AI-off) with the abnormal-day protocol and an auditable drill (POST /breakers/drill). Replay driver: npm run shadow -w @vpp/runtime -- --days 21.

All eight docs/01 invariants have automated tests (packages/services/test/chain.test.ts, packages/runtime/test/lifecycle.test.ts); the docs/07 D-1 16:00 and D+1 sections run in packages/runtime/test/m4.test.ts; the ROADMAP M5 acceptance (21 consecutive shadow days with complete lineage, KPI report, one widening recommendation not acted on) and the kill-switch drill run in packages/runtime/test/m5.test.ts. Storage is file-backed reference semantics; Postgres/Timescale adapters and the programmatic trading-platform channel are later work. Next: run the shadow period on real Hubei data and close the business parameters in docs/open-questions.md.

Development

Requires Node 24 and Python 3.11 (skills-py/.python-version; the generated pydantic models use StrEnum and PEP 604 unions).

npm ci
npm run check            # typecheck, re-export contracts, run TS tests
npm run eval -w @vpp/evals -- --check   # L2 harness vs baseline (needs the skill service below)
npm run start -w @vpp/runtime            # runtime + Case Desk API on :4100 (VPP_MODE=SHADOW default; VPP_LLM_MODEL unset = LLM-down mode)
npm run shadow -w @vpp/runtime -- --days 21   # replay 21 shadow days from the dataset through the live skill service

cd skills-py
uv venv --python 3.11 .venv && uv pip install -r requirements.txt   # or python3.11 -m venv
source .venv/bin/activate
bash scripts/generate_models.sh   # regenerate pydantic models (committed, never hand-edited)
python -m pytest -q
python -m uvicorn vpp_skills.app:app --port 8000   # skill service for the eval harness

CI (.github/workflows/ci.yml) runs both sides and fails if contracts/ or skills-py/vpp_contracts are not regenerated after a schema change.

Orientation

You are… Start with
A coding agent about to implement CLAUDE.md, then docs/00, 01, 09, 11
New to the project docs/00-overview.md → docs/07-scenario-walkthrough.md (the end-to-end reference scenario)
Reviewing the business case proposal.md (申报材料, source of requirements)
Looking for a settled decision docs/adr/
Wondering what's still undecided docs/open-questions.md

Document map (docs are in Chinese; implementation-facing files in English)

Doc Content
00-overview System context, two-plane architecture, business objects
01-principles 9 principles + 8 hard invariants (binding for all code)
02-cognitive-plane Five agents, Runtime, memory, Case Desk
03-safety-chain Proposal state machine, envelopes, permits, staleness
04-control-plane Execution engine, edge autonomy, time/space cascades
05-skills-and-data Skill contracts, five-store data layer, policy packs
06-integration External system boundaries and degraded channels
14-shadow-run-runbook Shadow-run operations, kill-switch levels (trigger / authority / recovery), KPI definitions as implemented
07-scenario-walkthrough Day-ahead spot bidding, D-1 → D → D+1
08-implementation Stack, LLM abstraction, deployment, milestones
09-runtime-implementation Runtime on Mastra: workflows, suspend/resume, lineage
10-federation Cross-province boundary: signed artifacts only
11-contracts Ports, canonical objects, TS↔Python contract pipeline
12-evaluation Four-layer evals, change gates, KPI definitions
13-risks-failure-modes FMEA, top-5 risks, kill-switch hierarchy

Supporting: brainstorming.md is an independent peer review whose findings were integrated (see docs/01 invariants note); GLOSSARY.md maps Chinese domain terms to canonical code names.

Target repository layout (from docs/09 §7)

packages/
├── domain/       # zod schemas — single source of truth for all business objects
├── runtime/      # Mastra instance, workflows, agents, tool registry, triggers
├── services/     # deterministic services: policy engine, envelopes, ledger, authority
├── adapters/     # anti-corruption layers: trading platform, dispatch, metering
├── evals/        # eval harness, datasets, judges (docs/12 §5)
└── skills-py/    # Python skill services (forecasting, MILP optimization, simulation)
contracts/        # generated JSON Schema + golden fixtures (cross-language contract)