2026-09-01 21:13:00 -04:00
# VPP AI Platform · 虚拟电厂多时空协同智能运营平台
AI-assisted operations platform for a virtual power plant (VPP) in Hubei, China:
five LLM agents propose market bids, resource dispatch, and load-control plans;
a deterministic safety chain (rule check → simulation → envelope/human approval →
execution permit) governs everything before any external effect. **The LLM never
computes numbers and never touches the second-level control loop.**
M2: skill contracts, Python skill service, L2 eval harness with baseline
- packages/domain: ForecastRequest, BidOptimizationRequest/Result,
ReportRequest, SkillReport (+ golden and invalid fixtures, exported to
contracts/ and regenerated as pydantic models).
- skills-py/vpp_skills: FastAPI service with versioned registry; load/PV/
price forecasts (same-day-type EWM point forecast, conformal residual
quantiles — coverage test as acceptance gate); bid-optimization MILP on
HiGHS (binary block participation, hard ledger energy bounds, exact
Decimal fit of the rounded curve inside the bounds, revenue distribution
over quantile paths); report generator whose every figure is a
{tool_call_id, path} reference, with a verifier. 48 tests incl. hypothesis
property test that bids respect ledger constraints.
- packages/services: LedgerService.dayAheadBounds (the P7 cascade band
handed to the optimizer); Decimal resolved once for CJS/ESM interop.
- packages/evals: L2 metrics (MAPE, nRMSE, coverage, direction accuracy,
naive/hindsight revenue baselines), HTTP skill client, rolling-origin
harness that pushes each bid through the real ledger, CLI with
--check/--write-baseline; committed baseline on the SYNTHETIC dataset
(no historical Hubei data yet — baselines measure the harness, not KPI).
- CI: evals job boots the skill service and fails on baseline digest drift.
- docs/open-questions: A6 (flexibility marginal cost = offer floor); A4/B6
wired as placeholders. README/CLAUDE.md status → M2 done, M3 next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:29:08 -04:00
**Status: M1 and M2 implemented.** Design docs 00– 13 are complete;
implementation follows [ROADMAP.md ](ROADMAP.md ). Present today: domain schemas
(`packages/domain`), the TS↔Python contract pipeline (`contracts/`,
`skills-py/vpp_contracts` ), ledger/snapshot/time-series/relational/ingestion
services (`packages/services`), the Python skill service — load/PV/price
forecasts with calibrated quantiles, bid-optimization MILP, report generator
(`skills-py/vpp_skills`) — and the L2 eval harness with a committed baseline
(`packages/evals`). The eval dataset is **synthetic** (no historical Hubei data
yet); baselines on it measure the harness, not the KPI. M3 (runtime, agents,
safety chain) is next.
M1: close gaps — CI, time-series/relational stores, ingestion skeleton, Python 3.11 pin
- .github/workflows/ci.yml: TS job (typecheck, re-export contracts, drift
check, vitest) and Python job (3.11, regenerate pydantic models, drift
check, pytest) — the dual-side contract test now actually runs in CI.
- packages/services: MemoryTimeSeriesStore (append-only daily-curve
revisions), MemoryRepository + ResourceRegistry (schema-validated,
optimistic versioning), IngestionPipeline (snapshot → quality gate →
store or quarantine). In-memory reference semantics; persistent adapters
arrive with M3 like the ledger.
- skills-py: .python-version + pyproject requires-python >=3.11; codegen
script asserts interpreter version and runs the generator as a module.
- typecheck scripts per package and in root `check`; README status and dev
setup; CLAUDE.md current-phase note.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-01 21:58:49 -04:00
## Development
Requires Node 24 and Python 3.11 (`skills-py/.python-version`; the generated
pydantic models use `StrEnum` and PEP 604 unions).
```sh
npm ci
npm run check # typecheck, re-export contracts, run TS tests
M2: skill contracts, Python skill service, L2 eval harness with baseline
- packages/domain: ForecastRequest, BidOptimizationRequest/Result,
ReportRequest, SkillReport (+ golden and invalid fixtures, exported to
contracts/ and regenerated as pydantic models).
- skills-py/vpp_skills: FastAPI service with versioned registry; load/PV/
price forecasts (same-day-type EWM point forecast, conformal residual
quantiles — coverage test as acceptance gate); bid-optimization MILP on
HiGHS (binary block participation, hard ledger energy bounds, exact
Decimal fit of the rounded curve inside the bounds, revenue distribution
over quantile paths); report generator whose every figure is a
{tool_call_id, path} reference, with a verifier. 48 tests incl. hypothesis
property test that bids respect ledger constraints.
- packages/services: LedgerService.dayAheadBounds (the P7 cascade band
handed to the optimizer); Decimal resolved once for CJS/ESM interop.
- packages/evals: L2 metrics (MAPE, nRMSE, coverage, direction accuracy,
naive/hindsight revenue baselines), HTTP skill client, rolling-origin
harness that pushes each bid through the real ledger, CLI with
--check/--write-baseline; committed baseline on the SYNTHETIC dataset
(no historical Hubei data yet — baselines measure the harness, not KPI).
- CI: evals job boots the skill service and fails on baseline digest drift.
- docs/open-questions: A6 (flexibility marginal cost = offer floor); A4/B6
wired as placeholders. README/CLAUDE.md status → M2 done, M3 next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:29:08 -04:00
npm run eval -w @vpp/evals -- --check # L2 harness vs baseline (needs the skill service below)
M1: close gaps — CI, time-series/relational stores, ingestion skeleton, Python 3.11 pin
- .github/workflows/ci.yml: TS job (typecheck, re-export contracts, drift
check, vitest) and Python job (3.11, regenerate pydantic models, drift
check, pytest) — the dual-side contract test now actually runs in CI.
- packages/services: MemoryTimeSeriesStore (append-only daily-curve
revisions), MemoryRepository + ResourceRegistry (schema-validated,
optimistic versioning), IngestionPipeline (snapshot → quality gate →
store or quarantine). In-memory reference semantics; persistent adapters
arrive with M3 like the ledger.
- skills-py: .python-version + pyproject requires-python >=3.11; codegen
script asserts interpreter version and runs the generator as a module.
- typecheck scripts per package and in root `check`; README status and dev
setup; CLAUDE.md current-phase note.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-01 21:58:49 -04:00
cd skills-py
uv venv --python 3.11 .venv & & uv pip install -r requirements.txt # or python3.11 -m venv
source .venv/bin/activate
bash scripts/generate_models.sh # regenerate pydantic models (committed, never hand-edited)
python -m pytest -q
M2: skill contracts, Python skill service, L2 eval harness with baseline
- packages/domain: ForecastRequest, BidOptimizationRequest/Result,
ReportRequest, SkillReport (+ golden and invalid fixtures, exported to
contracts/ and regenerated as pydantic models).
- skills-py/vpp_skills: FastAPI service with versioned registry; load/PV/
price forecasts (same-day-type EWM point forecast, conformal residual
quantiles — coverage test as acceptance gate); bid-optimization MILP on
HiGHS (binary block participation, hard ledger energy bounds, exact
Decimal fit of the rounded curve inside the bounds, revenue distribution
over quantile paths); report generator whose every figure is a
{tool_call_id, path} reference, with a verifier. 48 tests incl. hypothesis
property test that bids respect ledger constraints.
- packages/services: LedgerService.dayAheadBounds (the P7 cascade band
handed to the optimizer); Decimal resolved once for CJS/ESM interop.
- packages/evals: L2 metrics (MAPE, nRMSE, coverage, direction accuracy,
naive/hindsight revenue baselines), HTTP skill client, rolling-origin
harness that pushes each bid through the real ledger, CLI with
--check/--write-baseline; committed baseline on the SYNTHETIC dataset
(no historical Hubei data yet — baselines measure the harness, not KPI).
- CI: evals job boots the skill service and fails on baseline digest drift.
- docs/open-questions: A6 (flexibility marginal cost = offer floor); A4/B6
wired as placeholders. README/CLAUDE.md status → M2 done, M3 next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:29:08 -04:00
python -m uvicorn vpp_skills.app:app --port 8000 # skill service for the eval harness
M1: close gaps — CI, time-series/relational stores, ingestion skeleton, Python 3.11 pin
- .github/workflows/ci.yml: TS job (typecheck, re-export contracts, drift
check, vitest) and Python job (3.11, regenerate pydantic models, drift
check, pytest) — the dual-side contract test now actually runs in CI.
- packages/services: MemoryTimeSeriesStore (append-only daily-curve
revisions), MemoryRepository + ResourceRegistry (schema-validated,
optimistic versioning), IngestionPipeline (snapshot → quality gate →
store or quarantine). In-memory reference semantics; persistent adapters
arrive with M3 like the ledger.
- skills-py: .python-version + pyproject requires-python >=3.11; codegen
script asserts interpreter version and runs the generator as a module.
- typecheck scripts per package and in root `check`; README status and dev
setup; CLAUDE.md current-phase note.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-01 21:58:49 -04:00
```
CI (`.github/workflows/ci.yml`) runs both sides and fails if `contracts/` or
`skills-py/vpp_contracts` are not regenerated after a schema change.
2026-09-01 21:13:00 -04:00
## Orientation
| You are… | Start with |
|---|---|
| A coding agent about to implement | [CLAUDE.md ](CLAUDE.md ), then docs/00, 01, 09, 11 |
| New to the project | [docs/00-overview.md ](docs/00-overview.md ) → [docs/07-scenario-walkthrough.md ](docs/07-scenario-walkthrough.md ) (the end-to-end reference scenario) |
| Reviewing the business case | [proposal.md ](proposal.md ) (申报材料, source of requirements) |
| Looking for a settled decision | [docs/adr/ ](docs/adr/ ) |
| Wondering what's still undecided | [docs/open-questions.md ](docs/open-questions.md ) |
## Document map (docs are in Chinese; implementation-facing files in English)
| Doc | Content |
|---|---|
| [00-overview ](docs/00-overview.md ) | System context, two-plane architecture, business objects |
| [01-principles ](docs/01-principles.md ) | 9 principles + 8 hard invariants (binding for all code) |
| [02-cognitive-plane ](docs/02-cognitive-plane.md ) | Five agents, Runtime, memory, Case Desk |
| [03-safety-chain ](docs/03-safety-chain.md ) | Proposal state machine, envelopes, permits, staleness |
| [04-control-plane ](docs/04-control-plane.md ) | Execution engine, edge autonomy, time/space cascades |
| [05-skills-and-data ](docs/05-skills-and-data.md ) | Skill contracts, five-store data layer, policy packs |
| [06-integration ](docs/06-integration.md ) | External system boundaries and degraded channels |
| [07-scenario-walkthrough ](docs/07-scenario-walkthrough.md ) | Day-ahead spot bidding, D-1 → D → D+1 |
| [08-implementation ](docs/08-implementation.md ) | Stack, LLM abstraction, deployment, milestones |
| [09-runtime-implementation ](docs/09-runtime-implementation.md ) | Runtime on Mastra: workflows, suspend/resume, lineage |
| [10-federation ](docs/10-federation.md ) | Cross-province boundary: signed artifacts only |
| [11-contracts ](docs/11-contracts.md ) | Ports, canonical objects, TS↔Python contract pipeline |
| [12-evaluation ](docs/12-evaluation.md ) | Four-layer evals, change gates, KPI definitions |
| [13-risks-failure-modes ](docs/13-risks-failure-modes.md ) | FMEA, top-5 risks, kill-switch hierarchy |
Supporting: [brainstorming.md ](brainstorming.md ) is an independent peer review
whose findings were integrated (see docs/01 invariants note);
[GLOSSARY.md ](GLOSSARY.md ) maps Chinese domain terms to canonical code names.
## Target repository layout (from docs/09 §7)
```
packages/
├── domain/ # zod schemas — single source of truth for all business objects
├── runtime/ # Mastra instance, workflows, agents, tool registry, triggers
├── services/ # deterministic services: policy engine, envelopes, ledger, authority
├── adapters/ # anti-corruption layers: trading platform, dispatch, metering
├── evals/ # eval harness, datasets, judges (docs/12 §5)
└── skills-py/ # Python skill services (forecasting, MILP optimization, simulation)
contracts/ # generated JSON Schema + golden fixtures (cross-language contract)
```