Go to file
Thomas Bayes 8796faca63 M2: skill contracts, Python skill service, L2 eval harness with baseline
- packages/domain: ForecastRequest, BidOptimizationRequest/Result,
  ReportRequest, SkillReport (+ golden and invalid fixtures, exported to
  contracts/ and regenerated as pydantic models).
- skills-py/vpp_skills: FastAPI service with versioned registry; load/PV/
  price forecasts (same-day-type EWM point forecast, conformal residual
  quantiles — coverage test as acceptance gate); bid-optimization MILP on
  HiGHS (binary block participation, hard ledger energy bounds, exact
  Decimal fit of the rounded curve inside the bounds, revenue distribution
  over quantile paths); report generator whose every figure is a
  {tool_call_id, path} reference, with a verifier. 48 tests incl. hypothesis
  property test that bids respect ledger constraints.
- packages/services: LedgerService.dayAheadBounds (the P7 cascade band
  handed to the optimizer); Decimal resolved once for CJS/ESM interop.
- packages/evals: L2 metrics (MAPE, nRMSE, coverage, direction accuracy,
  naive/hindsight revenue baselines), HTTP skill client, rolling-origin
  harness that pushes each bid through the real ledger, CLI with
  --check/--write-baseline; committed baseline on the SYNTHETIC dataset
  (no historical Hubei data yet — baselines measure the harness, not KPI).
- CI: evals job boots the skill service and fails on baseline digest drift.
- docs/open-questions: A6 (flexibility marginal cost = offer floor); A4/B6
  wired as placeholders. README/CLAUDE.md status → M2 done, M3 next.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:29:08 -04:00
.github/workflows M2: skill contracts, Python skill service, L2 eval harness with baseline 2026-09-02 06:29:08 -04:00
contracts M2: skill contracts, Python skill service, L2 eval harness with baseline 2026-09-02 06:29:08 -04:00
docs M2: skill contracts, Python skill service, L2 eval harness with baseline 2026-09-02 06:29:08 -04:00
packages M2: skill contracts, Python skill service, L2 eval harness with baseline 2026-09-02 06:29:08 -04:00
proposal-assets init check in 2026-09-01 19:46:59 -04:00
skills-py M2: skill contracts, Python skill service, L2 eval harness with baseline 2026-09-02 06:29:08 -04:00
.gitignore Scaffold repo for implementation handoff 2026-09-01 21:13:00 -04:00
brainstorming.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
CLAUDE.md M2: skill contracts, Python skill service, L2 eval harness with baseline 2026-09-02 06:29:08 -04:00
GLOSSARY.md Scaffold repo for implementation handoff 2026-09-01 21:13:00 -04:00
package-lock.json M2: skill contracts, Python skill service, L2 eval harness with baseline 2026-09-02 06:29:08 -04:00
package.json M1: close gaps — CI, time-series/relational stores, ingestion skeleton, Python 3.11 pin 2026-09-01 21:58:49 -04:00
proposal.md init check in 2026-09-01 19:46:59 -04:00
README.md M2: skill contracts, Python skill service, L2 eval harness with baseline 2026-09-02 06:29:08 -04:00
ROADMAP.md Scaffold repo for implementation handoff 2026-09-01 21:13:00 -04:00
tsconfig.base.json M1: domain schemas, contracts pipeline, ledger and snapshot services 2026-09-01 21:50:22 -04:00

VPP AI Platform · 虚拟电厂多时空协同智能运营平台

AI-assisted operations platform for a virtual power plant (VPP) in Hubei, China: five LLM agents propose market bids, resource dispatch, and load-control plans; a deterministic safety chain (rule check → simulation → envelope/human approval → execution permit) governs everything before any external effect. The LLM never computes numbers and never touches the second-level control loop.

Status: M1 and M2 implemented. Design docs 00–13 are complete; implementation follows ROADMAP.md. Present today: domain schemas (packages/domain), the TS↔Python contract pipeline (contracts/, skills-py/vpp_contracts), ledger/snapshot/time-series/relational/ingestion services (packages/services), the Python skill service — load/PV/price forecasts with calibrated quantiles, bid-optimization MILP, report generator (skills-py/vpp_skills) — and the L2 eval harness with a committed baseline (packages/evals). The eval dataset is synthetic (no historical Hubei data yet); baselines on it measure the harness, not the KPI. M3 (runtime, agents, safety chain) is next.

Development

Requires Node 24 and Python 3.11 (skills-py/.python-version; the generated pydantic models use StrEnum and PEP 604 unions).

npm ci
npm run check            # typecheck, re-export contracts, run TS tests
npm run eval -w @vpp/evals -- --check   # L2 harness vs baseline (needs the skill service below)

cd skills-py
uv venv --python 3.11 .venv && uv pip install -r requirements.txt   # or python3.11 -m venv
source .venv/bin/activate
bash scripts/generate_models.sh   # regenerate pydantic models (committed, never hand-edited)
python -m pytest -q
python -m uvicorn vpp_skills.app:app --port 8000   # skill service for the eval harness

CI (.github/workflows/ci.yml) runs both sides and fails if contracts/ or skills-py/vpp_contracts are not regenerated after a schema change.

Orientation

You are… Start with
A coding agent about to implement CLAUDE.md, then docs/00, 01, 09, 11
New to the project docs/00-overview.md → docs/07-scenario-walkthrough.md (the end-to-end reference scenario)
Reviewing the business case proposal.md (申报材料, source of requirements)
Looking for a settled decision docs/adr/
Wondering what's still undecided docs/open-questions.md

Document map (docs are in Chinese; implementation-facing files in English)

Doc Content
00-overview System context, two-plane architecture, business objects
01-principles 9 principles + 8 hard invariants (binding for all code)
02-cognitive-plane Five agents, Runtime, memory, Case Desk
03-safety-chain Proposal state machine, envelopes, permits, staleness
04-control-plane Execution engine, edge autonomy, time/space cascades
05-skills-and-data Skill contracts, five-store data layer, policy packs
06-integration External system boundaries and degraded channels
07-scenario-walkthrough Day-ahead spot bidding, D-1 → D → D+1
08-implementation Stack, LLM abstraction, deployment, milestones
09-runtime-implementation Runtime on Mastra: workflows, suspend/resume, lineage
10-federation Cross-province boundary: signed artifacts only
11-contracts Ports, canonical objects, TS↔Python contract pipeline
12-evaluation Four-layer evals, change gates, KPI definitions
13-risks-failure-modes FMEA, top-5 risks, kill-switch hierarchy

Supporting: brainstorming.md is an independent peer review whose findings were integrated (see docs/01 invariants note); GLOSSARY.md maps Chinese domain terms to canonical code names.

Target repository layout (from docs/09 §7)

packages/
├── domain/       # zod schemas — single source of truth for all business objects
├── runtime/      # Mastra instance, workflows, agents, tool registry, triggers
├── services/     # deterministic services: policy engine, envelopes, ledger, authority
├── adapters/     # anti-corruption layers: trading platform, dispatch, metering
├── evals/        # eval harness, datasets, judges (docs/12 §5)
└── skills-py/    # Python skill services (forecasting, MILP optimization, simulation)
contracts/        # generated JSON Schema + golden fixtures (cross-language contract)