vpp-ai-platform/README.md
Thomas Bayes 23fdf49ff2
Some checks are pending
ci / typescript (push) Waiting to run
ci / python (push) Waiting to run
ci / evals (push) Blocked by required conditions
M4: resource agent, envelopes live, review loop, insight cards
- packages/domain: AwardNotice, DISPATCH_PLAN proposal payload, ExecutionReport,
  MeteringRecord, potential-assessment and dispatch-optimization contracts,
  ReviewFinding (+ typed writebacks), SemanticMemoryEntry,
  EnvelopeChangeRequest, InsightCard — exported with fixtures on both sides.
- skills-py: potential-assessment (certified × rolling fulfilment, evidence
  days) and dispatch-optimization (per-interval LP on HiGHS, shortfall
  reported) skills + routes + tests.
- packages/services: dispatch rules in the policy pack (over-allocation,
  award anchor, lineage integrity for allocations); PowerBalanceSimulator;
  SimulationGateway (permit-only, idempotent, seeded execute → ExecutionReports);
  envelope deviation-streak suspension + apply(); ReviewService (attribution,
  reliability EWMA writeback, semantic memory, envelope recommendations as
  change requests); dispatch assembler; skill client methods.
- packages/runtime: resource agent; award-decomposition, review and
  envelope-review workflows; lifecycle selects simulator/gateway by proposal
  type; trigger hooks for awards, execution reports, metering; decide()
  resumes either lifecycle or envelope-review runs; insight cards API.
- Tests: docs/07 D-1 16:00 and D+1 end to end; reliability score 0.9 → 0.880
  and the next assessment de-rates capacity; envelope suspension on a seeded
  3-day streak; WIDEN request applied only by a human. 184 TS + 80 Python.
- docs/open-questions: B10 (reliability/potential parameters). README and
  CLAUDE.md status → M4 done, M5 next.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 19:31:12 -04:00

105 lines
6.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# VPP AI Platform · 虚拟电厂多时空协同智能运营平台
AI-assisted operations platform for a virtual power plant (VPP) in Hubei, China:
five LLM agents propose market bids, resource dispatch, and load-control plans;
a deterministic safety chain (rule check → simulation → envelope/human approval →
execution permit) governs everything before any external effect. **The LLM never
computes numbers and never touches the second-level control loop.**
**Status: M1–M4 implemented.** Design docs 00–13 are complete; implementation
follows [ROADMAP.md](ROADMAP.md). Present today:
- `packages/domain` — zod schemas for every business and safety-chain object, exported to
`contracts/` and regenerated as pydantic models (`skills-py/vpp_contracts`).
- `packages/services` — deterministic services: ledger (with the P7 cascade), snapshot,
time-series, relational and ingestion stores; policy engine + `hubei-spot-bidding` pack;
envelope gate; authority (fresh check, permits, revocation); file-export gateway;
event store/bus; lineage recorder + P2 assembler; revenue-scenario simulation; Case Desk.
- `skills-py/vpp_skills` — Python skill service: load/PV/price forecasts with calibrated
quantiles, bid-optimization MILP, report generator.
- `packages/evals` — L2 eval harness with a committed baseline (on a **synthetic** dataset).
- `packages/runtime` — Mastra runtime: `proposal-lifecycle` (the safety chain, suspend/resume
for human approval, durable across restarts), `day-ahead-situation`, `day-ahead-bid`;
`award-decomposition` (resource agent: potential assessment → dispatch optimization →
DISPATCH_PLAN through the same chain to a simulation gateway), `review` (D+1
plan-vs-actual attribution → ReviewFinding → semantic memory + reliability writebacks +
envelope change requests), `envelope-review` (the only path that changes an envelope,
human-approved); envelope deviation-streak suspension; trigger service (scheduled / event /
manual via router agent); provider-abstracted LLM port with an LLM-down mode; Case Desk
HTTP API incl. AI insight cards. Bid release is a file export for manual upload (degraded
channel by design).
All eight docs/01 invariants have automated tests (`packages/services/test/chain.test.ts`,
`packages/runtime/test/lifecycle.test.ts`); the docs/07 D-1 16:00 and D+1 sections run in
`packages/runtime/test/m4.test.ts`. Storage is file-backed reference semantics;
Postgres/Timescale adapters and the programmatic trading-platform channel are later work.
M5 (shadow run) is next.
## Development
Requires Node 24 and Python 3.11 (`skills-py/.python-version`; the generated
pydantic models use `StrEnum` and PEP 604 unions).
```sh
npm ci
npm run check # typecheck, re-export contracts, run TS tests
npm run eval -w @vpp/evals -- --check # L2 harness vs baseline (needs the skill service below)
npm run start -w @vpp/runtime # runtime + Case Desk API on :4100 (VPP_LLM_MODEL unset = LLM-down mode)
cd skills-py
uv venv --python 3.11 .venv && uv pip install -r requirements.txt # or python3.11 -m venv
source .venv/bin/activate
bash scripts/generate_models.sh # regenerate pydantic models (committed, never hand-edited)
python -m pytest -q
python -m uvicorn vpp_skills.app:app --port 8000 # skill service for the eval harness
```
CI (`.github/workflows/ci.yml`) runs both sides and fails if `contracts/` or
`skills-py/vpp_contracts` are not regenerated after a schema change.
## Orientation
| You are… | Start with |
|---|---|
| A coding agent about to implement | [CLAUDE.md](CLAUDE.md), then docs/00, 01, 09, 11 |
| New to the project | [docs/00-overview.md](docs/00-overview.md) → [docs/07-scenario-walkthrough.md](docs/07-scenario-walkthrough.md) (the end-to-end reference scenario) |
| Reviewing the business case | [proposal.md](proposal.md) (申报材料, source of requirements) |
| Looking for a settled decision | [docs/adr/](docs/adr/) |
| Wondering what's still undecided | [docs/open-questions.md](docs/open-questions.md) |
## Document map (docs are in Chinese; implementation-facing files in English)
| Doc | Content |
|---|---|
| [00-overview](docs/00-overview.md) | System context, two-plane architecture, business objects |
| [01-principles](docs/01-principles.md) | 9 principles + 8 hard invariants (binding for all code) |
| [02-cognitive-plane](docs/02-cognitive-plane.md) | Five agents, Runtime, memory, Case Desk |
| [03-safety-chain](docs/03-safety-chain.md) | Proposal state machine, envelopes, permits, staleness |
| [04-control-plane](docs/04-control-plane.md) | Execution engine, edge autonomy, time/space cascades |
| [05-skills-and-data](docs/05-skills-and-data.md) | Skill contracts, five-store data layer, policy packs |
| [06-integration](docs/06-integration.md) | External system boundaries and degraded channels |
| [07-scenario-walkthrough](docs/07-scenario-walkthrough.md) | Day-ahead spot bidding, D-1 → D → D+1 |
| [08-implementation](docs/08-implementation.md) | Stack, LLM abstraction, deployment, milestones |
| [09-runtime-implementation](docs/09-runtime-implementation.md) | Runtime on Mastra: workflows, suspend/resume, lineage |
| [10-federation](docs/10-federation.md) | Cross-province boundary: signed artifacts only |
| [11-contracts](docs/11-contracts.md) | Ports, canonical objects, TS↔Python contract pipeline |
| [12-evaluation](docs/12-evaluation.md) | Four-layer evals, change gates, KPI definitions |
| [13-risks-failure-modes](docs/13-risks-failure-modes.md) | FMEA, top-5 risks, kill-switch hierarchy |
Supporting: [brainstorming.md](brainstorming.md) is an independent peer review
whose findings were integrated (see docs/01 invariants note);
[GLOSSARY.md](GLOSSARY.md) maps Chinese domain terms to canonical code names.
## Target repository layout (from docs/09 §7)
```
packages/
├── domain/ # zod schemas — single source of truth for all business objects
├── runtime/ # Mastra instance, workflows, agents, tool registry, triggers
├── services/ # deterministic services: policy engine, envelopes, ledger, authority
├── adapters/ # anti-corruption layers: trading platform, dispatch, metering
├── evals/ # eval harness, datasets, judges (docs/12 §5)
└── skills-py/ # Python skill services (forecasting, MILP optimization, simulation)
contracts/ # generated JSON Schema + golden fixtures (cross-language contract)
```