vpp-ai-platform/README.md
Thomas Bayes 23fdf49ff2
Some checks are pending
ci / typescript (push) Waiting to run
ci / python (push) Waiting to run
ci / evals (push) Blocked by required conditions
M4: resource agent, envelopes live, review loop, insight cards
- packages/domain: AwardNotice, DISPATCH_PLAN proposal payload, ExecutionReport,
  MeteringRecord, potential-assessment and dispatch-optimization contracts,
  ReviewFinding (+ typed writebacks), SemanticMemoryEntry,
  EnvelopeChangeRequest, InsightCard — exported with fixtures on both sides.
- skills-py: potential-assessment (certified × rolling fulfilment, evidence
  days) and dispatch-optimization (per-interval LP on HiGHS, shortfall
  reported) skills + routes + tests.
- packages/services: dispatch rules in the policy pack (over-allocation,
  award anchor, lineage integrity for allocations); PowerBalanceSimulator;
  SimulationGateway (permit-only, idempotent, seeded execute → ExecutionReports);
  envelope deviation-streak suspension + apply(); ReviewService (attribution,
  reliability EWMA writeback, semantic memory, envelope recommendations as
  change requests); dispatch assembler; skill client methods.
- packages/runtime: resource agent; award-decomposition, review and
  envelope-review workflows; lifecycle selects simulator/gateway by proposal
  type; trigger hooks for awards, execution reports, metering; decide()
  resumes either lifecycle or envelope-review runs; insight cards API.
- Tests: docs/07 D-1 16:00 and D+1 end to end; reliability score 0.9 → 0.880
  and the next assessment de-rates capacity; envelope suspension on a seeded
  3-day streak; WIDEN request applied only by a human. 184 TS + 80 Python.
- docs/open-questions: B10 (reliability/potential parameters). README and
  CLAUDE.md status → M4 done, M5 next.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 19:31:12 -04:00

6.3 KiB
Raw Blame History

VPP AI Platform · 虚拟电厂多时空协同智能运营平台

AI-assisted operations platform for a virtual power plant (VPP) in Hubei, China: five LLM agents propose market bids, resource dispatch, and load-control plans; a deterministic safety chain (rule check → simulation → envelope/human approval → execution permit) governs everything before any external effect. The LLM never computes numbers and never touches the second-level control loop.

Status: M1–M4 implemented. Design docs 00–13 are complete; implementation follows ROADMAP.md. Present today:

  • packages/domain — zod schemas for every business and safety-chain object, exported to contracts/ and regenerated as pydantic models (skills-py/vpp_contracts).
  • packages/services — deterministic services: ledger (with the P7 cascade), snapshot, time-series, relational and ingestion stores; policy engine + hubei-spot-bidding pack; envelope gate; authority (fresh check, permits, revocation); file-export gateway; event store/bus; lineage recorder + P2 assembler; revenue-scenario simulation; Case Desk.
  • skills-py/vpp_skills — Python skill service: load/PV/price forecasts with calibrated quantiles, bid-optimization MILP, report generator.
  • packages/evals — L2 eval harness with a committed baseline (on a synthetic dataset).
  • packages/runtime — Mastra runtime: proposal-lifecycle (the safety chain, suspend/resume for human approval, durable across restarts), day-ahead-situation, day-ahead-bid; award-decomposition (resource agent: potential assessment → dispatch optimization → DISPATCH_PLAN through the same chain to a simulation gateway), review (D+1 plan-vs-actual attribution → ReviewFinding → semantic memory + reliability writebacks + envelope change requests), envelope-review (the only path that changes an envelope, human-approved); envelope deviation-streak suspension; trigger service (scheduled / event / manual via router agent); provider-abstracted LLM port with an LLM-down mode; Case Desk HTTP API incl. AI insight cards. Bid release is a file export for manual upload (degraded channel by design).

All eight docs/01 invariants have automated tests (packages/services/test/chain.test.ts, packages/runtime/test/lifecycle.test.ts); the docs/07 D-1 16:00 and D+1 sections run in packages/runtime/test/m4.test.ts. Storage is file-backed reference semantics; Postgres/Timescale adapters and the programmatic trading-platform channel are later work. M5 (shadow run) is next.

Development

Requires Node 24 and Python 3.11 (skills-py/.python-version; the generated pydantic models use StrEnum and PEP 604 unions).

npm ci
npm run check            # typecheck, re-export contracts, run TS tests
npm run eval -w @vpp/evals -- --check   # L2 harness vs baseline (needs the skill service below)
npm run start -w @vpp/runtime            # runtime + Case Desk API on :4100 (VPP_LLM_MODEL unset = LLM-down mode)

cd skills-py
uv venv --python 3.11 .venv && uv pip install -r requirements.txt   # or python3.11 -m venv
source .venv/bin/activate
bash scripts/generate_models.sh   # regenerate pydantic models (committed, never hand-edited)
python -m pytest -q
python -m uvicorn vpp_skills.app:app --port 8000   # skill service for the eval harness

CI (.github/workflows/ci.yml) runs both sides and fails if contracts/ or skills-py/vpp_contracts are not regenerated after a schema change.

Orientation

You are… Start with
A coding agent about to implement CLAUDE.md, then docs/00, 01, 09, 11
New to the project docs/00-overview.md → docs/07-scenario-walkthrough.md (the end-to-end reference scenario)
Reviewing the business case proposal.md (申报材料, source of requirements)
Looking for a settled decision docs/adr/
Wondering what's still undecided docs/open-questions.md

Document map (docs are in Chinese; implementation-facing files in English)

Doc Content
00-overview System context, two-plane architecture, business objects
01-principles 9 principles + 8 hard invariants (binding for all code)
02-cognitive-plane Five agents, Runtime, memory, Case Desk
03-safety-chain Proposal state machine, envelopes, permits, staleness
04-control-plane Execution engine, edge autonomy, time/space cascades
05-skills-and-data Skill contracts, five-store data layer, policy packs
06-integration External system boundaries and degraded channels
07-scenario-walkthrough Day-ahead spot bidding, D-1 → D → D+1
08-implementation Stack, LLM abstraction, deployment, milestones
09-runtime-implementation Runtime on Mastra: workflows, suspend/resume, lineage
10-federation Cross-province boundary: signed artifacts only
11-contracts Ports, canonical objects, TS↔Python contract pipeline
12-evaluation Four-layer evals, change gates, KPI definitions
13-risks-failure-modes FMEA, top-5 risks, kill-switch hierarchy

Supporting: brainstorming.md is an independent peer review whose findings were integrated (see docs/01 invariants note); GLOSSARY.md maps Chinese domain terms to canonical code names.

Target repository layout (from docs/09 §7)

packages/
├── domain/       # zod schemas — single source of truth for all business objects
├── runtime/      # Mastra instance, workflows, agents, tool registry, triggers
├── services/     # deterministic services: policy engine, envelopes, ledger, authority
├── adapters/     # anti-corruption layers: trading platform, dispatch, metering
├── evals/        # eval harness, datasets, judges (docs/12 §5)
└── skills-py/    # Python skill services (forecasting, MILP optimization, simulation)
contracts/        # generated JSON Schema + golden fixtures (cross-language contract)