vpp-ai-platform/docs
Thomas Bayes 8796faca63 M2: skill contracts, Python skill service, L2 eval harness with baseline
- packages/domain: ForecastRequest, BidOptimizationRequest/Result,
  ReportRequest, SkillReport (+ golden and invalid fixtures, exported to
  contracts/ and regenerated as pydantic models).
- skills-py/vpp_skills: FastAPI service with versioned registry; load/PV/
  price forecasts (same-day-type EWM point forecast, conformal residual
  quantiles — coverage test as acceptance gate); bid-optimization MILP on
  HiGHS (binary block participation, hard ledger energy bounds, exact
  Decimal fit of the rounded curve inside the bounds, revenue distribution
  over quantile paths); report generator whose every figure is a
  {tool_call_id, path} reference, with a verifier. 48 tests incl. hypothesis
  property test that bids respect ledger constraints.
- packages/services: LedgerService.dayAheadBounds (the P7 cascade band
  handed to the optimizer); Decimal resolved once for CJS/ESM interop.
- packages/evals: L2 metrics (MAPE, nRMSE, coverage, direction accuracy,
  naive/hindsight revenue baselines), HTTP skill client, rolling-origin
  harness that pushes each bid through the real ledger, CLI with
  --check/--write-baseline; committed baseline on the SYNTHETIC dataset
  (no historical Hubei data yet — baselines measure the harness, not KPI).
- CI: evals job boots the skill service and fails on baseline digest drift.
- docs/open-questions: A6 (flexibility marginal cost = offer floor); A4/B6
  wired as placeholders. README/CLAUDE.md status → M2 done, M3 next.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:29:08 -04:00
..
adr Scaffold repo for implementation handoff 2026-09-01 21:13:00 -04:00
00-overview.md Add risk assessment and failure-mode analysis (doc 13) 2026-09-01 21:06:55 -04:00
01-principles.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
02-cognitive-plane.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
03-safety-chain.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
04-control-plane.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
05-skills-and-data.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
06-integration.md init check in 2026-09-01 19:46:59 -04:00
07-scenario-walkthrough.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
08-implementation.md Add evaluation architecture (doc 12) 2026-09-01 21:03:25 -04:00
09-runtime-implementation.md Add Runtime implementation design (Mastra mapping) 2026-09-01 20:24:32 -04:00
10-federation.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
11-contracts.md Add interfaces and contracts reference (doc 11) 2026-09-01 20:52:14 -04:00
12-evaluation.md Add evaluation architecture (doc 12) 2026-09-01 21:03:25 -04:00
13-risks-failure-modes.md Add risk assessment and failure-mode analysis (doc 13) 2026-09-01 21:06:55 -04:00
open-questions.md M2: skill contracts, Python skill service, L2 eval harness with baseline 2026-09-02 06:29:08 -04:00