M2: skill contracts, Python skill service, L2 eval harness with baseline

- packages/domain: ForecastRequest, BidOptimizationRequest/Result,
  ReportRequest, SkillReport (+ golden and invalid fixtures, exported to
  contracts/ and regenerated as pydantic models).
- skills-py/vpp_skills: FastAPI service with versioned registry; load/PV/
  price forecasts (same-day-type EWM point forecast, conformal residual
  quantiles — coverage test as acceptance gate); bid-optimization MILP on
  HiGHS (binary block participation, hard ledger energy bounds, exact
  Decimal fit of the rounded curve inside the bounds, revenue distribution
  over quantile paths); report generator whose every figure is a
  {tool_call_id, path} reference, with a verifier. 48 tests incl. hypothesis
  property test that bids respect ledger constraints.
- packages/services: LedgerService.dayAheadBounds (the P7 cascade band
  handed to the optimizer); Decimal resolved once for CJS/ESM interop.
- packages/evals: L2 metrics (MAPE, nRMSE, coverage, direction accuracy,
  naive/hindsight revenue baselines), HTTP skill client, rolling-origin
  harness that pushes each bid through the real ledger, CLI with
  --check/--write-baseline; committed baseline on the SYNTHETIC dataset
  (no historical Hubei data yet — baselines measure the harness, not KPI).
- CI: evals job boots the skill service and fails on baseline digest drift.
- docs/open-questions: A6 (flexibility marginal cost = offer floor); A4/B6
  wired as placeholders. README/CLAUDE.md status → M2 done, M3 next.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
This commit is contained in:
Thomas Bayes 2026-09-02 06:29:08 -04:00
parent e7196bc88a
commit 8796faca63
56 changed files with 52156 additions and 24 deletions

View File

@ -44,3 +44,30 @@ jobs:
- run: bash scripts/generate_models.sh
- run: git diff --exit-code -- vpp_contracts/
- run: python -m pytest -q
# L2 eval harness end-to-end (docs/12): TS harness ↔ live Python skill
# service on the committed synthetic dataset. The run must reproduce the
# committed baseline digest exactly — a skill change that moves a metric
# must come with a reviewed baseline update (docs/12 §3 gate).
evals:
runs-on: ubuntu-latest
needs: [typescript, python]
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 24
cache: npm
- uses: actions/setup-python@v5
with:
python-version-file: skills-py/.python-version
cache: pip
cache-dependency-path: skills-py/requirements.txt
- run: npm ci
- run: pip install -r skills-py/requirements.txt
- name: start skill service
working-directory: skills-py
run: |
python -m uvicorn vpp_skills.app:app --port 8000 --log-level warning &
for i in $(seq 1 40); do curl -sf http://127.0.0.1:8000/health && break; sleep 0.5; done
- run: npm run eval -w @vpp/evals -- --check

View File

@ -6,10 +6,13 @@ docs win — or write an ADR changing the doc first.
## Current phase
Docs complete; M1 implemented (schemas, contracts pipeline, ledger, snapshot,
time-series, relational and ingestion services, dual-side CI). Storage is
in-memory reference semantics — persistent adapters arrive with M3. Next is
M2. Build order is ROADMAP.md (M1→M5).
Docs complete; M1 and M2 implemented (schemas + contracts pipeline; ledger,
snapshot, time-series, relational, ingestion services; Python skill service
with forecast/MILP/report skills; L2 eval harness with committed baseline).
Storage is in-memory reference semantics — persistent adapters arrive with M3.
The eval dataset is synthetic until historical data lands. Next is M3. Build
order is ROADMAP.md (M1→M5). A skill change that moves an L2 metric must
update the baseline in the same change (`npm run eval -w @vpp/evals -- --write-baseline`).
Do not start a milestone's work before its predecessor's acceptance criteria are
testable, and do not build phase-2 items (edge control links, federation,
interaction/load-control agents) unless explicitly asked.

View File

@ -6,12 +6,16 @@ a deterministic safety chain (rule check → simulation → envelope/human appro
execution permit) governs everything before any external effect. **The LLM never
computes numbers and never touches the second-level control loop.**
**Status: M1 (data foundation & contracts) implemented.** Design docs 00–13 are
complete; implementation follows [ROADMAP.md](ROADMAP.md). Present today:
domain schemas (`packages/domain`), the TS↔Python contract pipeline
(`contracts/`, `skills-py/vpp_contracts`), and the ledger, snapshot,
time-series, relational and ingestion services (`packages/services`). M2
(Python skills + eval baseline) is next.
**Status: M1 and M2 implemented.** Design docs 00–13 are complete;
implementation follows [ROADMAP.md](ROADMAP.md). Present today: domain schemas
(`packages/domain`), the TS↔Python contract pipeline (`contracts/`,
`skills-py/vpp_contracts`), ledger/snapshot/time-series/relational/ingestion
services (`packages/services`), the Python skill service — load/PV/price
forecasts with calibrated quantiles, bid-optimization MILP, report generator
(`skills-py/vpp_skills`) — and the L2 eval harness with a committed baseline
(`packages/evals`). The eval dataset is **synthetic** (no historical Hubei data
yet); baselines on it measure the harness, not the KPI. M3 (runtime, agents,
safety chain) is next.
## Development
@ -21,12 +25,14 @@ pydantic models use `StrEnum` and PEP 604 unions).
```sh
npm ci
npm run check # typecheck, re-export contracts, run TS tests
npm run eval -w @vpp/evals -- --check # L2 harness vs baseline (needs the skill service below)
cd skills-py
uv venv --python 3.11 .venv && uv pip install -r requirements.txt # or python3.11 -m venv
source .venv/bin/activate
bash scripts/generate_models.sh # regenerate pydantic models (committed, never hand-edited)
python -m pytest -q
python -m uvicorn vpp_skills.app:app --port 8000 # skill service for the eval harness
```
CI (`.github/workflows/ci.yml`) runs both sides and fails if `contracts/` or

View File

@ -0,0 +1,230 @@
{
"market_date": "2026-03-15",
"prices_yuan_per_mwh": {
"interval_minutes": 15,
"date": "2026-03-15",
"values": [
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00"
]
},
"quantities_mwh": {
"interval_minutes": 15,
"date": "2026-03-15",
"values": [
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5"
]
},
"daily_energy_mwh": "1200.0",
"expected_revenue_yuan": "510600.00",
"revenue_distribution_yuan": {
"p10": "432000.00",
"p50": "510600.00",
"p90": "588000.00"
},
"position_bounds": {
"ledger_version": 42,
"daily_energy_min_mwh": "1140.0",
"daily_energy_max_mwh": "1260.0"
},
"solver": {
"name": "highs",
"version": "1.7.0",
"status": "SOLVED",
"objective_value": "487020.00",
"wall_time_ms": 12
},
"binding_constraints": [
"daily_energy_max"
],
"skill_version": "1.0.0"
}

View File

@ -0,0 +1,22 @@
{
"id": "rep-001",
"kind": "DAY_AHEAD_BID_SUMMARY",
"market_date": "2026-03-15",
"sections": [
{
"title": "Bid summary",
"metrics": [
{
"name": "expected_revenue",
"value": "510600.00",
"unit": "yuan"
}
],
"notes": [
"all figures reference solver output tc-001"
]
}
],
"skill_version": "1.0.0",
"generated_at": "2026-03-14T08:30:00Z"
}

View File

@ -0,0 +1,436 @@
{
"market_date": "2026-03-15",
"price_forecast": {
"id": "fc-price-001",
"kind": "PRICE",
"market_date": "2026-03-15",
"unit": "yuan_per_mwh",
"quantiles": {
"p10": {
"interval_minutes": 15,
"date": "2026-03-15",
"values": [
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00",
"360.00"
]
},
"p50": {
"interval_minutes": 15,
"date": "2026-03-15",
"values": [
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50",
"425.50"
]
},
"p90": {
"interval_minutes": 15,
"date": "2026-03-15",
"values": [
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00",
"490.00"
]
}
},
"model": {
"name": "price-forecast",
"version": "1.0.0"
},
"features_snapshot_ref": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
"generated_at": "2026-03-14T06:00:00Z"
},
"adjustable_capacity_mw": {
"interval_minutes": 15,
"date": "2026-03-15",
"values": [
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0",
"60.0"
]
},
"position_bounds": {
"ledger_version": 42,
"daily_energy_min_mwh": "1140.0",
"daily_energy_max_mwh": "1260.0"
},
"risk": {
"risk_aversion": "0.3",
"commitment_buffer_k": "0.9",
"min_block_mwh": "1.0",
"marginal_cost_yuan_per_mwh": "0"
}
}

View File

@ -0,0 +1,230 @@
{
"market_date": "2026-03-15",
"prices_yuan_per_mwh": {
"interval_minutes": 15,
"date": "2026-03-15",
"values": [
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00",
"0.00"
]
},
"quantities_mwh": {
"interval_minutes": 15,
"date": "2026-03-15",
"values": [
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5",
"12.5"
]
},
"daily_energy_mwh": "1200.0",
"expected_revenue_yuan": "510600.00",
"revenue_distribution_yuan": {
"p10": "432000.00",
"p50": "510600.00",
"p90": "588000.00"
},
"position_bounds": {
"ledger_version": 42,
"daily_energy_min_mwh": "1140.0",
"daily_energy_max_mwh": "1260.0"
},
"solver": {
"name": "highs",
"version": "1.7.0",
"status": "OPTIMAL",
"objective_value": "487020.00",
"wall_time_ms": 12
},
"binding_constraints": [
"daily_energy_max"
],
"skill_version": "1.0.0"
}

View File

@ -0,0 +1,213 @@
{
"kind": "PRICE",
"market_date": "2026-03-15",
"unit": "yuan_per_mwh",
"history": [
{
"interval_minutes": 15,
"date": "2026-03-13",
"values": [
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20",
"398.20"
]
},
{
"interval_minutes": 15,
"date": "2026-03-14",
"values": [
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75",
"410.75"
]
}
],
"exogenous": {},
"features_snapshot_ref": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
}

View File

@ -0,0 +1,15 @@
{
"kind": "DAY_AHEAD_BID_SUMMARY",
"market_date": "2026-03-15",
"sources": [
{
"tool_call_id": "tc-001",
"tool": "bid-optimization-milp",
"version": "1.0.0",
"output": {
"expected_revenue_yuan": "510600.00",
"daily_energy_mwh": "1200.0"
}
}
]
}

View File

@ -0,0 +1,26 @@
{
"id": "rep-001",
"kind": "DAY_AHEAD_BID_SUMMARY",
"market_date": "2026-03-15",
"sections": [
{
"title": "Bid summary",
"metrics": [
{
"name": "expected_revenue",
"value": "510600.00",
"unit": "yuan",
"ref": {
"tool_call_id": "tc-001",
"path": "expected_revenue_yuan"
}
}
],
"notes": [
"all figures reference solver output tc-001"
]
}
],
"skill_version": "1.0.0",
"generated_at": "2026-03-14T08:30:00Z"
}

View File

@ -0,0 +1,261 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"market_date": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"price_forecast": {
"type": "object",
"properties": {
"id": {
"type": "string",
"minLength": 1
},
"kind": {
"type": "string",
"enum": [
"LOAD",
"PV",
"PRICE"
]
},
"market_date": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"unit": {
"type": "string",
"enum": [
"mw",
"yuan_per_mwh"
]
},
"quantiles": {
"type": "object",
"properties": {
"p10": {
"type": "object",
"properties": {
"interval_minutes": {
"type": "number",
"const": 15
},
"date": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"values": {
"minItems": 96,
"maxItems": 96,
"type": "array",
"items": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
}
}
},
"required": [
"interval_minutes",
"date",
"values"
],
"additionalProperties": false
},
"p50": {
"type": "object",
"properties": {
"interval_minutes": {
"type": "number",
"const": 15
},
"date": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"values": {
"minItems": 96,
"maxItems": 96,
"type": "array",
"items": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
}
}
},
"required": [
"interval_minutes",
"date",
"values"
],
"additionalProperties": false
},
"p90": {
"type": "object",
"properties": {
"interval_minutes": {
"type": "number",
"const": 15
},
"date": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"values": {
"minItems": 96,
"maxItems": 96,
"type": "array",
"items": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
}
}
},
"required": [
"interval_minutes",
"date",
"values"
],
"additionalProperties": false
}
},
"required": [
"p10",
"p50",
"p90"
],
"additionalProperties": false
},
"model": {
"type": "object",
"properties": {
"name": {
"type": "string",
"minLength": 1
},
"version": {
"type": "string",
"minLength": 1
}
},
"required": [
"name",
"version"
],
"additionalProperties": false
},
"features_snapshot_ref": {
"type": "string",
"pattern": "^[0-9a-f]{64}$"
},
"generated_at": {
"type": "string",
"format": "date-time",
"pattern": "^(?:(?:\\d\\d[2468][048]|\\d\\d[13579][26]|\\d\\d0[48]|[02468][048]00|[13579][26]00)-02-29|\\d{4}-(?:(?:0[13578]|1[02])-(?:0[1-9]|[12]\\d|3[01])|(?:0[469]|11)-(?:0[1-9]|[12]\\d|30)|(?:02)-(?:0[1-9]|1\\d|2[0-8])))T(?:(?:[01]\\d|2[0-3]):[0-5]\\d:[0-5]\\d(?:\\.\\d+)?(?:Z))$"
}
},
"required": [
"id",
"kind",
"market_date",
"unit",
"quantiles",
"model",
"features_snapshot_ref",
"generated_at"
],
"additionalProperties": false
},
"adjustable_capacity_mw": {
"type": "object",
"properties": {
"interval_minutes": {
"type": "number",
"const": 15
},
"date": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"values": {
"minItems": 96,
"maxItems": 96,
"type": "array",
"items": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
}
}
},
"required": [
"interval_minutes",
"date",
"values"
],
"additionalProperties": false
},
"position_bounds": {
"type": "object",
"properties": {
"ledger_version": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"daily_energy_min_mwh": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
},
"daily_energy_max_mwh": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
}
},
"required": [
"ledger_version",
"daily_energy_min_mwh",
"daily_energy_max_mwh"
],
"additionalProperties": false
},
"risk": {
"type": "object",
"properties": {
"risk_aversion": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
},
"commitment_buffer_k": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
},
"min_block_mwh": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
},
"marginal_cost_yuan_per_mwh": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
}
},
"required": [
"risk_aversion",
"commitment_buffer_k",
"min_block_mwh",
"marginal_cost_yuan_per_mwh"
],
"additionalProperties": false
}
},
"required": [
"market_date",
"price_forecast",
"adjustable_capacity_mw",
"position_bounds",
"risk"
],
"additionalProperties": false,
"$id": "https://vpp-ai-platform/contracts/bid_optimization_request.json",
"title": "BidOptimizationRequest"
}

View File

@ -0,0 +1,192 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"market_date": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"prices_yuan_per_mwh": {
"type": "object",
"properties": {
"interval_minutes": {
"type": "number",
"const": 15
},
"date": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"values": {
"minItems": 96,
"maxItems": 96,
"type": "array",
"items": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
}
}
},
"required": [
"interval_minutes",
"date",
"values"
],
"additionalProperties": false
},
"quantities_mwh": {
"type": "object",
"properties": {
"interval_minutes": {
"type": "number",
"const": 15
},
"date": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"values": {
"minItems": 96,
"maxItems": 96,
"type": "array",
"items": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
}
}
},
"required": [
"interval_minutes",
"date",
"values"
],
"additionalProperties": false
},
"daily_energy_mwh": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
},
"expected_revenue_yuan": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
},
"revenue_distribution_yuan": {
"type": "object",
"properties": {
"p10": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
},
"p50": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
},
"p90": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
}
},
"required": [
"p10",
"p50",
"p90"
],
"additionalProperties": false
},
"position_bounds": {
"type": "object",
"properties": {
"ledger_version": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"daily_energy_min_mwh": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
},
"daily_energy_max_mwh": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
}
},
"required": [
"ledger_version",
"daily_energy_min_mwh",
"daily_energy_max_mwh"
],
"additionalProperties": false
},
"solver": {
"type": "object",
"properties": {
"name": {
"type": "string",
"minLength": 1
},
"version": {
"type": "string",
"minLength": 1
},
"status": {
"type": "string",
"enum": [
"OPTIMAL",
"INFEASIBLE",
"TIME_LIMIT",
"ERROR"
]
},
"objective_value": {
"anyOf": [
{
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
},
{
"type": "null"
}
]
},
"wall_time_ms": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
}
},
"required": [
"name",
"version",
"status",
"objective_value",
"wall_time_ms"
],
"additionalProperties": false
},
"binding_constraints": {
"type": "array",
"items": {
"type": "string"
}
},
"skill_version": {
"type": "string",
"pattern": "^\\d+\\.\\d+\\.\\d+$"
}
},
"required": [
"market_date",
"prices_yuan_per_mwh",
"quantities_mwh",
"daily_energy_mwh",
"expected_revenue_yuan",
"revenue_distribution_yuan",
"position_bounds",
"solver",
"binding_constraints",
"skill_version"
],
"additionalProperties": false,
"$id": "https://vpp-ai-platform/contracts/bid_optimization_result.json",
"title": "BidOptimizationResult"
}

View File

@ -0,0 +1,106 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"kind": {
"type": "string",
"enum": [
"LOAD",
"PV",
"PRICE"
]
},
"market_date": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"unit": {
"type": "string",
"enum": [
"mw",
"yuan_per_mwh"
]
},
"history": {
"minItems": 1,
"type": "array",
"items": {
"type": "object",
"properties": {
"interval_minutes": {
"type": "number",
"const": 15
},
"date": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"values": {
"minItems": 96,
"maxItems": 96,
"type": "array",
"items": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
}
}
},
"required": [
"interval_minutes",
"date",
"values"
],
"additionalProperties": false
}
},
"exogenous": {
"type": "object",
"propertyNames": {
"type": "string"
},
"additionalProperties": {
"type": "object",
"properties": {
"interval_minutes": {
"type": "number",
"const": 15
},
"date": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"values": {
"minItems": 96,
"maxItems": 96,
"type": "array",
"items": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
}
}
},
"required": [
"interval_minutes",
"date",
"values"
],
"additionalProperties": false
}
},
"features_snapshot_ref": {
"type": "string",
"pattern": "^[0-9a-f]{64}$"
}
},
"required": [
"kind",
"market_date",
"unit",
"history",
"exogenous",
"features_snapshot_ref"
],
"additionalProperties": false,
"$id": "https://vpp-ai-platform/contracts/forecast_request.json",
"title": "ForecastRequest"
}

View File

@ -0,0 +1,55 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"kind": {
"type": "string",
"enum": [
"DAY_AHEAD_BID_SUMMARY",
"FORECAST_EVAL",
"BID_BACKTEST"
]
},
"market_date": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"sources": {
"minItems": 1,
"type": "array",
"items": {
"type": "object",
"properties": {
"tool_call_id": {
"type": "string",
"minLength": 1
},
"tool": {
"type": "string",
"minLength": 1
},
"version": {
"type": "string",
"pattern": "^\\d+\\.\\d+\\.\\d+$"
},
"output": {}
},
"required": [
"tool_call_id",
"tool",
"version",
"output"
],
"additionalProperties": false
}
}
},
"required": [
"kind",
"market_date",
"sources"
],
"additionalProperties": false,
"$id": "https://vpp-ai-platform/contracts/report_request.json",
"title": "ReportRequest"
}

View File

@ -0,0 +1,111 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"id": {
"type": "string",
"minLength": 1
},
"kind": {
"type": "string",
"enum": [
"DAY_AHEAD_BID_SUMMARY",
"FORECAST_EVAL",
"BID_BACKTEST"
]
},
"market_date": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"sections": {
"type": "array",
"items": {
"type": "object",
"properties": {
"title": {
"type": "string",
"minLength": 1
},
"metrics": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": {
"type": "string",
"minLength": 1
},
"value": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
},
"unit": {
"type": "string",
"minLength": 1
},
"ref": {
"type": "object",
"properties": {
"tool_call_id": {
"type": "string",
"minLength": 1
},
"path": {
"type": "string",
"minLength": 1
}
},
"required": [
"tool_call_id",
"path"
],
"additionalProperties": false
}
},
"required": [
"name",
"value",
"unit",
"ref"
],
"additionalProperties": false
}
},
"notes": {
"type": "array",
"items": {
"type": "string"
}
}
},
"required": [
"title",
"metrics",
"notes"
],
"additionalProperties": false
}
},
"skill_version": {
"type": "string",
"pattern": "^\\d+\\.\\d+\\.\\d+$"
},
"generated_at": {
"type": "string",
"format": "date-time",
"pattern": "^(?:(?:\\d\\d[2468][048]|\\d\\d[13579][26]|\\d\\d0[48]|[02468][048]00|[13579][26]00)-02-29|\\d{4}-(?:(?:0[13578]|1[02])-(?:0[1-9]|[12]\\d|3[01])|(?:0[469]|11)-(?:0[1-9]|[12]\\d|30)|(?:02)-(?:0[1-9]|1\\d|2[0-8])))T(?:(?:[01]\\d|2[0-3]):[0-5]\\d:[0-5]\\d(?:\\.\\d+)?(?:Z))$"
}
},
"required": [
"id",
"kind",
"market_date",
"sections",
"skill_version",
"generated_at"
],
"additionalProperties": false,
"$id": "https://vpp-ai-platform/contracts/skill_report.json",
"title": "SkillReport"
}

View File

@ -14,6 +14,7 @@ values for items on this list** — wire named config with placeholder + an
| A3 | 日内市场机制(是否开、频次、截止) | `market.intraday` |
| A4 | 偏差考核规则:偏差带、考核价格机制(→ MILP 目标函数与风险口径) | `market.deviation` |
| A5 | 中长期持仓对日前申报的约束形式(分解曲线偏差带) | `ledger.da_bounds` |
| A6 | 灵活性资源边际成本口径(用户补偿、设备损耗)——报价优化的报价下限;未定前按 0(价格接受者)申报 | `bidding.marginal_cost` |
## B. 包络与风控参数(owner: 运营团队 · blocks M4 envelopes + M5 breakers)

20
package-lock.json generated
View File

@ -979,6 +979,10 @@
"resolved": "packages/domain",
"link": true
},
"node_modules/@vpp/evals": {
"resolved": "packages/evals",
"link": true
},
"node_modules/@vpp/services": {
"resolved": "packages/services",
"link": true
@ -1246,6 +1250,7 @@
"integrity": "sha512-qcJu88Q2IWqJsDD529JKMdwGm/dvInW4HvQnRwiH9JtihJvzGOscDtHE3x1pBKeUOTysQ8kVmLnJ2kJu7yhcGA==",
"dev": true,
"license": "MIT",
"peer": true,
"engines": {
"node": ">=12"
},
@ -1479,6 +1484,7 @@
"integrity": "sha512-4XP60spRGjSZFf1qYH+dJIkK2znL3zQfl9KkOV9MkkRR/3Dls0dxaBsQPTloEc5BLXWPL9vsOxopxyKoMmDueg==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"esbuild": "^0.27.0 || ^0.28.0",
"fdir": "^6.5.0",
@ -1683,6 +1689,20 @@
"vitest": "^3.0.0"
}
},
"packages/evals": {
"name": "@vpp/evals",
"version": "0.1.0",
"dependencies": {
"@vpp/domain": "*",
"@vpp/services": "*"
},
"devDependencies": {
"@types/node": "^26.4.1",
"tsx": "^4.19.0",
"typescript": "^5.6.0",
"vitest": "^3.0.0"
}
},
"packages/services": {
"name": "@vpp/services",
"version": "0.1.0",

View File

@ -150,6 +150,104 @@ const valid: Record<string, Record<string, unknown>> = {
generated_at: T0,
},
},
forecast_request: {
'price-da': {
kind: 'PRICE',
market_date: DATE,
unit: 'yuan_per_mwh',
history: [curve('398.20', '2026-03-13'), curve('410.75', '2026-03-14')],
exogenous: {},
features_snapshot_ref: REF_A,
},
},
bid_optimization_request: {
'da-basic': {
market_date: DATE,
price_forecast: {
id: 'fc-price-001',
kind: 'PRICE',
market_date: DATE,
unit: 'yuan_per_mwh',
quantiles: { p10: curve('360.00'), p50: curve('425.50'), p90: curve('490.00') },
model: { name: 'price-forecast', version: '1.0.0' },
features_snapshot_ref: REF_A,
generated_at: T0,
},
adjustable_capacity_mw: curve('60.0'),
position_bounds: {
ledger_version: 42,
daily_energy_min_mwh: '1140.0',
daily_energy_max_mwh: '1260.0',
},
risk: {
risk_aversion: '0.3',
commitment_buffer_k: '0.9',
min_block_mwh: '1.0',
marginal_cost_yuan_per_mwh: '0',
},
},
},
bid_optimization_result: {
'da-basic': {
market_date: DATE,
prices_yuan_per_mwh: curve('0.00'),
quantities_mwh: curve('12.5'),
daily_energy_mwh: '1200.0',
expected_revenue_yuan: '510600.00',
revenue_distribution_yuan: { p10: '432000.00', p50: '510600.00', p90: '588000.00' },
position_bounds: {
ledger_version: 42,
daily_energy_min_mwh: '1140.0',
daily_energy_max_mwh: '1260.0',
},
solver: {
name: 'highs',
version: '1.7.0',
status: 'OPTIMAL',
objective_value: '487020.00',
wall_time_ms: 12,
},
binding_constraints: ['daily_energy_max'],
skill_version: '1.0.0',
},
},
report_request: {
'da-bid-summary': {
kind: 'DAY_AHEAD_BID_SUMMARY',
market_date: DATE,
sources: [
{
tool_call_id: 'tc-001',
tool: 'bid-optimization-milp',
version: '1.0.0',
output: { expected_revenue_yuan: '510600.00', daily_energy_mwh: '1200.0' },
},
],
},
},
skill_report: {
'da-bid-summary': {
id: 'rep-001',
kind: 'DAY_AHEAD_BID_SUMMARY',
market_date: DATE,
sections: [
{
title: 'Bid summary',
metrics: [
{
name: 'expected_revenue',
value: '510600.00',
unit: 'yuan',
ref: { tool_call_id: 'tc-001', path: 'expected_revenue_yuan' },
},
],
notes: ['all figures reference solver output tc-001'],
},
],
skill_version: '1.0.0',
generated_at: T1,
},
},
situation_report: {
'normal-day': {
id: 'sit-001',
@ -219,6 +317,22 @@ const invalid: Record<string, Record<string, unknown>> = {
// digest not sha256 hex
'bad-digest': { ...(valid['approval']!['approve'] as object), proposal_digest: 'not-a-digest' },
},
skill_report: {
// a metric without a lineage reference — the one thing a report may never contain (I1)
'metric-without-ref': (() => {
const o = structuredClone(valid['skill_report']!['da-bid-summary']) as any
delete o.sections[0].metrics[0].ref
return o
})(),
},
bid_optimization_result: {
// solver status outside the enum
'bad-solver-status': (() => {
const o = structuredClone(valid['bid_optimization_result']!['da-basic']) as any
o.solver.status = 'SOLVED'
return o
})(),
},
position_update: {
// missing optimistic-concurrency field
'missing-version': (() => {

View File

@ -5,9 +5,12 @@ import { Curve96, Id, IsoUtc, MarketDate, SnapshotRef } from './common.js'
* Forecast output contract (docs/05 §2.1): quantile intervals are mandatory —
* bid risk assessment depends on calibrated bands, not point values.
*/
export const ForecastKind = z.enum(['LOAD', 'PV', 'PRICE'])
export type ForecastKind = z.infer<typeof ForecastKind>
export const ForecastBundle = z.object({
id: Id,
kind: z.enum(['LOAD', 'PV', 'PRICE']),
kind: ForecastKind,
market_date: MarketDate,
unit: z.enum(['mw', 'yuan_per_mwh']),
quantiles: z.object({

View File

@ -8,6 +8,13 @@ import { LedgerView, PositionUpdate } from './ledger.js'
import { Proposal } from './proposal.js'
import { ResourceProfile } from './resource.js'
import { SituationReport } from './situation.js'
import {
BidOptimizationRequest,
BidOptimizationResult,
ForecastRequest,
ReportRequest,
SkillReport,
} from './skill.js'
export * from './common.js'
export * from './proposal.js'
@ -19,6 +26,7 @@ export * from './situation.js'
export * from './resource.js'
export * from './case.js'
export * from './event.js'
export * from './skill.js'
/**
* Registry driving the contracts pipeline: keys become schema/fixture/module
@ -37,4 +45,9 @@ export const schemaRegistry: Record<string, z.ZodType> = {
resource_profile: ResourceProfile,
decision_case: DecisionCase,
event_envelope: EventEnvelope,
forecast_request: ForecastRequest,
bid_optimization_request: BidOptimizationRequest,
bid_optimization_result: BidOptimizationResult,
report_request: ReportRequest,
skill_report: SkillReport,
}

View File

@ -0,0 +1,139 @@
import { z } from 'zod'
import { Curve96, DecimalString, Id, IsoUtc, MarketDate, SnapshotRef } from './common.js'
import { ForecastBundle, ForecastKind } from './forecast.js'
/**
* Skill I/O contracts (docs/05 §2, docs/11 §3). These cross the TS↔Python
* boundary: the runtime's registerSkill wrapper validates them outbound, the
* FastAPI service validates them inbound with the generated pydantic models.
* Skills are stateless — everything they need is in the request; everything
* lineage needs is in the response.
*/
export const SkillVersion = z.string().regex(/^\d+\.\d+\.\d+$/)
/** Forecast skills (load / PV / price): history in, quantile bundle out. */
export const ForecastRequest = z.object({
kind: ForecastKind,
market_date: MarketDate,
unit: z.enum(['mw', 'yuan_per_mwh']),
/** Actuals for past market dates, ascending, all strictly before market_date. */
history: z.array(Curve96).min(1),
/** Exogenous forecasts for market_date keyed by name (e.g. irradiance_w_per_m2). */
exogenous: z.record(z.string(), Curve96),
/** Caller's snapshot of the assembled feature set; echoed into ForecastBundle. */
features_snapshot_ref: SnapshotRef,
})
export type ForecastRequest = z.infer<typeof ForecastRequest>
/**
* Daily energy bounds derived from the position ledger by the TS side
* (LedgerService.dayAheadBounds — the P7 cascade lives in one place). The
* optimizer treats them as hard constraints and echoes them so the rule check
* can verify the bid against the same ledger version.
*/
export const PositionBounds = z.object({
ledger_version: z.int().nonnegative(),
daily_energy_min_mwh: DecimalString,
daily_energy_max_mwh: DecimalString,
})
export type PositionBounds = z.infer<typeof PositionBounds>
export const BidRiskParams = z.object({
/** 0 = maximise P50 revenue, 1 = maximise P10 (floor) revenue. */
risk_aversion: DecimalString,
/** Commitment buffer k: sellable share of certified capacity. OPEN-QUESTION B6. */
commitment_buffer_k: DecimalString,
/** Smallest non-zero interval quantity the market accepts. */
min_block_mwh: DecimalString,
/**
* Marginal cost of delivering flexibility (user compensation, degradation),
* used as the offer-price floor: a seller in a uniform-price market offers
* at cost and lets the forecast allocate quantity. OPEN-QUESTION A6.
*/
marginal_cost_yuan_per_mwh: DecimalString,
})
export type BidRiskParams = z.infer<typeof BidRiskParams>
export const BidOptimizationRequest = z.object({
market_date: MarketDate,
price_forecast: ForecastBundle,
/** Certified adjustable (sellable) capacity per interval, MW. */
adjustable_capacity_mw: Curve96,
position_bounds: PositionBounds,
risk: BidRiskParams,
})
export type BidOptimizationRequest = z.infer<typeof BidOptimizationRequest>
export const SolverStatus = z.enum(['OPTIMAL', 'INFEASIBLE', 'TIME_LIMIT', 'ERROR'])
export const BidOptimizationResult = z.object({
market_date: MarketDate,
prices_yuan_per_mwh: Curve96,
quantities_mwh: Curve96,
daily_energy_mwh: DecimalString,
/** Revenue at the P50 price path, given the bid clears where offer ≤ price. */
expected_revenue_yuan: DecimalString,
revenue_distribution_yuan: z.object({
p10: DecimalString,
p50: DecimalString,
p90: DecimalString,
}),
position_bounds: PositionBounds,
solver: z.object({
name: z.string().min(1),
version: z.string().min(1),
status: SolverStatus,
objective_value: DecimalString.nullable(),
wall_time_ms: z.int().nonnegative(),
}),
binding_constraints: z.array(z.string()),
skill_version: SkillVersion,
})
export type BidOptimizationResult = z.infer<typeof BidOptimizationResult>
/** A number in a report is never typed in: it is a {tool_call_id, path} reference (I1/P2). */
export const MetricRef = z.object({
tool_call_id: Id,
/** JSON-pointer-like dotted path into that tool call's output. */
path: z.string().min(1),
})
export const ReportSource = z.object({
tool_call_id: Id,
tool: Id,
version: SkillVersion,
output: z.unknown(),
})
export const ReportKind = z.enum(['DAY_AHEAD_BID_SUMMARY', 'FORECAST_EVAL', 'BID_BACKTEST'])
export const ReportRequest = z.object({
kind: ReportKind,
market_date: MarketDate,
sources: z.array(ReportSource).min(1),
})
export type ReportRequest = z.infer<typeof ReportRequest>
export const SkillReport = z.object({
id: Id,
kind: ReportKind,
market_date: MarketDate,
sections: z.array(
z.object({
title: z.string().min(1),
metrics: z.array(
z.object({
name: z.string().min(1),
value: DecimalString,
unit: z.string().min(1),
ref: MetricRef,
}),
),
notes: z.array(z.string()),
}),
),
skill_version: SkillVersion,
generated_at: IsoUtc,
})
export type SkillReport = z.infer<typeof SkillReport>

File diff suppressed because it is too large Load Diff

View File

@ -0,0 +1,24 @@
{
"name": "@vpp/evals",
"version": "0.1.0",
"private": true,
"type": "module",
"exports": {
".": "./src/index.ts"
},
"scripts": {
"typecheck": "tsc -p tsconfig.json",
"test": "vitest run",
"eval": "tsx src/cli.ts"
},
"dependencies": {
"@vpp/domain": "*",
"@vpp/services": "*"
},
"devDependencies": {
"@types/node": "^26.4.1",
"tsx": "^4.19.0",
"typescript": "^5.6.0",
"vitest": "^3.0.0"
}
}

View File

@ -0,0 +1,62 @@
{
"layer": "L2",
"dataset": {
"name": "synthetic-hubei-v0-2026-01-01-120d",
"sha256": "3aa36eb46906233b1b311ff03a33e7d663f30ae921fc9a1f3f813ab5205c270a",
"synthetic": true
},
"skill_versions": {
"load-forecast": "1.0.0",
"pv-forecast": "1.0.0",
"price-forecast": "1.0.0",
"bid-optimization-milp": "1.0.0",
"report-generator": "1.0.0"
},
"config": {
"window": 28,
"holdoutFrom": 60,
"holdoutDays": 0,
"risk": {
"risk_aversion": "0.3",
"commitment_buffer_k": "0.9",
"min_block_mwh": "0.5",
"marginal_cost_yuan_per_mwh": "0"
},
"contractShareOfSellable": "0.6",
"daMonthlyDeviationBand": "0.05"
},
"metrics": {
"load": {
"days": 60,
"mape": 0.041,
"nrmse": null,
"coverage_p10_p90": 0.821007,
"direction_accuracy": null
},
"pv": {
"days": 60,
"mape": null,
"nrmse": 0.040938,
"coverage_p10_p90": 0.772109,
"direction_accuracy": null
},
"price": {
"days": 60,
"mape": 0.044013,
"nrmse": null,
"coverage_p10_p90": 0.81059,
"direction_accuracy": 0.962142
},
"bid": {
"days": 60,
"optimal_days": 60,
"ledger_accepted_days": 60,
"revenue_skill_yuan": 7146184.79,
"revenue_naive_yuan": 6012053.56,
"revenue_hindsight_yuan": 7154899.25,
"capture_ratio": 0.998782,
"uplift_vs_naive": 1.188643
}
},
"digest": "4018e32ab695422c613e211aba1392164f3bb9c30076986fdfda220fc9bba047"
}

78
packages/evals/src/cli.ts Normal file
View File

@ -0,0 +1,78 @@
/**
* `npm run eval -w @vpp/evals -- [--base-url URL] [--dataset PATH] [--check] [--write-baseline]`
*
* Runs the L2 harness against a live skill service, archives the run under
* reports/runs/ (gitignored — docs/12 §5 says archived runs live in object
* storage), and either compares it to the committed baseline (--check: exit 1
* on any metric change) or rewrites the baseline (--write-baseline, a
* reviewed change — docs/12 §3 gates a Skill upgrade on its L2 metrics).
*/
import { mkdirSync, readFileSync, writeFileSync, existsSync } from 'node:fs'
import { fileURLToPath } from 'node:url'
import { HttpSkillClient } from './client.js'
import { loadDataset } from './dataset.js'
import { DEFAULT_CONFIG, runL2, stableJson } from './harness.js'
import type { EvalRun } from './harness.js'
const here = fileURLToPath(new URL('.', import.meta.url))
const args = process.argv.slice(2)
const flag = (name: string) => args.includes(name)
const opt = (name: string, dflt: string) => {
const i = args.indexOf(name)
return i >= 0 && args[i + 1] ? args[i + 1]! : dflt
}
const baseUrl = opt('--base-url', process.env['SKILLS_BASE_URL'] ?? 'http://127.0.0.1:8000')
const datasetPath = opt('--dataset', `${here}../datasets/synthetic-hubei-v0.json`)
const baselinePath = `${here}../reports/baselines/l2-synthetic-hubei-v0.json`
const runsDir = `${here}../reports/runs/`
const { dataset, sha256 } = loadDataset(datasetPath)
const run = await runL2(dataset, sha256, new HttpSkillClient(baseUrl), DEFAULT_CONFIG)
mkdirSync(runsDir, { recursive: true })
writeFileSync(`${runsDir}${run.id}.json`, JSON.stringify(run, null, 2) + '\n')
const m = run.metrics
console.log(`L2 eval ${run.id} on ${run.dataset.name}${run.dataset.synthetic ? ' (SYNTHETIC)' : ''}`)
console.log(` load MAPE ${m.load.mape} coverage ${m.load.coverage_p10_p90}`)
console.log(` pv nRMSE ${m.pv.nrmse} coverage(daylight) ${m.pv.coverage_p10_p90}`)
console.log(` price MAPE ${m.price.mape} coverage ${m.price.coverage_p10_p90} direction ${m.price.direction_accuracy}`)
console.log(
` bid optimal ${m.bid.optimal_days}/${m.bid.days} ledger-accepted ${m.bid.ledger_accepted_days}/${m.bid.days}` +
` capture ${m.bid.capture_ratio} uplift-vs-naive ${m.bid.uplift_vs_naive}`,
)
console.log(` digest ${run.digest}`)
/** Baseline = the run minus its per-run identity. */
const baselineOf = (r: EvalRun) => {
const { id: _id, run_at: _at, per_day: _pd, ...rest } = r
return rest
}
if (flag('--write-baseline')) {
mkdirSync(`${here}../reports/baselines/`, { recursive: true })
writeFileSync(baselinePath, JSON.stringify(baselineOf(run), null, 2) + '\n')
console.log(`baseline written: ${baselinePath}`)
}
if (flag('--check')) {
if (!existsSync(baselinePath)) {
console.error(`no baseline at ${baselinePath}; run with --write-baseline first`)
process.exit(2)
}
const baseline = JSON.parse(readFileSync(baselinePath, 'utf8')) as ReturnType<typeof baselineOf>
if (baseline.digest === run.digest) {
console.log('baseline check: identical digest — reproducible')
} else {
console.error('baseline check FAILED: run differs from committed baseline')
const a = stableJson(baseline.metrics)
const b = stableJson(run.metrics)
console.error(` baseline metrics: ${a}`)
console.error(` this run metrics: ${b}`)
for (const key of ['dataset', 'skill_versions', 'config'] as const) {
if (stableJson(baseline[key]) !== stableJson(run[key])) console.error(` ${key} differs`)
}
process.exit(1)
}
}

View File

@ -0,0 +1,75 @@
import {
BidOptimizationRequest,
BidOptimizationResult,
ForecastBundle,
ForecastRequest,
ReportRequest,
SkillReport,
} from '@vpp/domain'
import type { ForecastKind } from '@vpp/domain'
/**
* Typed access to the Python skill service. Both directions are validated
* against the domain schemas — the same discipline registerSkill applies in
* the runtime (docs/09 §4), minus the lineage recording the harness does not
* need. The interface exists so tests can substitute a stub.
*/
export interface SkillClient {
skills(): Promise<Array<{ id: string; version: string; endpoint: string }>>
forecast(kind: ForecastKind, req: ForecastRequest): Promise<ForecastBundle>
optimizeBid(req: BidOptimizationRequest): Promise<BidOptimizationResult>
report(req: ReportRequest): Promise<SkillReport>
}
export class SkillHttpError extends Error {
constructor(
readonly status: number,
readonly url: string,
body: string,
) {
super(`skill call ${url} failed with HTTP ${status}: ${body.slice(0, 500)}`)
this.name = 'SkillHttpError'
}
}
const FORECAST_PATH: Record<ForecastKind, string> = {
LOAD: '/v1/forecast/load',
PV: '/v1/forecast/pv',
PRICE: '/v1/forecast/price',
}
export class HttpSkillClient implements SkillClient {
constructor(private readonly baseUrl: string) {}
private async post<T>(path: string, body: unknown, parse: (x: unknown) => T): Promise<T> {
const url = `${this.baseUrl}${path}`
const res = await fetch(url, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify(body),
})
const text = await res.text()
if (!res.ok) throw new SkillHttpError(res.status, url, text)
return parse(JSON.parse(text))
}
async skills() {
const res = await fetch(`${this.baseUrl}/v1/skills`)
if (!res.ok) throw new SkillHttpError(res.status, `${this.baseUrl}/v1/skills`, await res.text())
return (await res.json()) as Array<{ id: string; version: string; endpoint: string }>
}
forecast(kind: ForecastKind, req: ForecastRequest) {
return this.post(FORECAST_PATH[kind], ForecastRequest.parse(req), (x) => ForecastBundle.parse(x))
}
optimizeBid(req: BidOptimizationRequest) {
return this.post('/v1/optimize/bid', BidOptimizationRequest.parse(req), (x) =>
BidOptimizationResult.parse(x),
)
}
report(req: ReportRequest) {
return this.post('/v1/report', ReportRequest.parse(req), (x) => SkillReport.parse(x))
}
}

View File

@ -0,0 +1,52 @@
import { createHash } from 'node:crypto'
import { readFileSync } from 'node:fs'
import type { Curve96 } from '@vpp/domain'
/**
* Replay dataset shape (docs/12 §2 历史重放集). Today the only instance is the
* SYNTHETIC placeholder written by `python -m vpp_skills.synthetic`; a real
* replay set produced from the event log will have the same shape and a
* `meta.synthetic: false`.
*/
export interface DatasetDay {
date: string
weekend: boolean
load_mw: string[]
pv_mw: string[]
price_yuan_per_mwh: string[]
adjustable_capacity_mw: string[]
}
export interface Dataset {
meta: {
name: string
synthetic: boolean
note: string
generator: string
generator_version: string
seed: number
interval_minutes: 15
peak_load_mw: string
pv_capacity_mw: string
adjustable_capacity_mw: string
}
days: DatasetDay[]
}
export interface LoadedDataset {
dataset: Dataset
sha256: string
}
export function loadDataset(path: string): LoadedDataset {
const raw = readFileSync(path, 'utf8')
return { dataset: JSON.parse(raw) as Dataset, sha256: createHash('sha256').update(raw).digest('hex') }
}
export const toCurve = (values: string[], date: string): Curve96 => ({
interval_minutes: 15,
date,
values,
})
export const nums = (values: string[]): number[] => values.map(Number)

View File

@ -0,0 +1,307 @@
import { createHash } from 'node:crypto'
import type { BidOptimizationResult, BidRiskParams, ForecastBundle, ForecastKind } from '@vpp/domain'
import { Decimal, LedgerService } from '@vpp/services'
import type { SkillClient } from './client.js'
import type { Dataset, DatasetDay } from './dataset.js'
import { nums, toCurve } from './dataset.js'
import {
coverage,
directionAccuracy,
hindsightBid,
mape,
mean,
naiveBid,
nrmse,
realisedRevenue,
} from './metrics.js'
/**
* L2 harness (docs/12 §1 L2, §5 harness/): rolling-origin backtest over a
* replay dataset. For each held-out day it asks the skill service for the three
* forecasts and a bid, scores them against actuals, and pushes the bid through
* the real LedgerService so "MILP output respects ledger constraints" is
* checked by the component that enforces it — not by a re-implementation.
*
* Every run produces an EvalRun whose `digest` covers config, dataset hash,
* skill versions and metrics. Two runs on the same inputs must digest equal
* (reproducibility acceptance); the CLI's --check compares against the
* committed baseline.
*/
export interface HarnessConfig {
/** History days handed to each forecast call. */
window: number
/** Index of the first held-out day (needs ≥ window history before it). */
holdoutFrom: number
/** Number of held-out days; 0 = to the end of the dataset. */
holdoutDays: number
risk: BidRiskParams
/**
* Share of sellable energy (k·cap·0.25·96) used as the daily contract
* position when seeding the ledger. Synthetic-set convenience, not a market
* parameter.
*/
contractShareOfSellable: string
/** Passed to LedgerService; OPEN-QUESTION A5 — placeholder until confirmed. */
daMonthlyDeviationBand: string
}
export const DEFAULT_CONFIG: HarnessConfig = {
window: 28,
holdoutFrom: 60,
holdoutDays: 0,
// OPEN-QUESTION B6 (commitment_buffer_k) — placeholder for harness use only.
// OPEN-QUESTION A6 (marginal_cost) — placeholder 0: offer as a price-taker.
risk: { risk_aversion: '0.3', commitment_buffer_k: '0.9', min_block_mwh: '0.5', marginal_cost_yuan_per_mwh: '0' },
contractShareOfSellable: '0.6',
// OPEN-QUESTION A5 — placeholder, same value the services ledger tests use.
daMonthlyDeviationBand: '0.05',
}
export interface ForecastMetrics {
days: number
mape: number | null
nrmse: number | null
coverage_p10_p90: number
direction_accuracy: number | null
}
export interface BidMetrics {
days: number
optimal_days: number
ledger_accepted_days: number
revenue_skill_yuan: number
revenue_naive_yuan: number
revenue_hindsight_yuan: number
/** skill / hindsight — 1.0 would be perfect foresight. */
capture_ratio: number
/** skill / naive — > 1.0 means the optimiser beats a price-taker. */
uplift_vs_naive: number
}
export interface DayResult {
date: string
load: { mape: number; coverage: number }
pv: { nrmse: number; coverage: number }
price: { mape: number; coverage: number; direction: number }
bid: {
status: string
ledger_accepted: boolean
energy_mwh: string
revenue_skill: number
revenue_naive: number
revenue_hindsight: number
}
}
export interface EvalRun {
id: string
layer: 'L2'
run_at: string
dataset: { name: string; sha256: string; synthetic: boolean }
skill_versions: Record<string, string>
config: HarnessConfig
metrics: { load: ForecastMetrics; pv: ForecastMetrics; price: ForecastMetrics; bid: BidMetrics }
per_day: DayResult[]
/** sha256 over everything above except id/run_at — the reproducibility key. */
digest: string
}
const REF = 'f'.repeat(64)
/** Key-sorted JSON. (services' canonicalJson rejects floats by design; metrics are floats.) */
export const stableJson = (v: unknown): string =>
JSON.stringify(v, (_k, val) =>
val !== null && typeof val === 'object' && !Array.isArray(val)
? Object.fromEntries(Object.entries(val as Record<string, unknown>).sort(([a], [b]) => (a < b ? -1 : 1)))
: val,
)
const FIELD: Record<ForecastKind, keyof DatasetDay> = {
LOAD: 'load_mw',
PV: 'pv_mw',
PRICE: 'price_yuan_per_mwh',
}
const q = (x: number, dp: number) => Number(x.toFixed(dp))
function bandOf(bundle: ForecastBundle) {
return {
p10: nums(bundle.quantiles.p10.values),
p50: nums(bundle.quantiles.p50.values),
p90: nums(bundle.quantiles.p90.values),
}
}
function seedLedger(dataset: Dataset, cfg: HarnessConfig, days: DatasetDay[]): LedgerService {
const ledger = new LedgerService({ daMonthlyDeviationBand: cfg.daMonthlyDeviationBand, clock: () => '2026-01-01T00:00:00Z' })
const k = new Decimal(cfg.risk.commitment_buffer_k)
const share = new Decimal(cfg.contractShareOfSellable)
const months = [...new Set(days.map((d) => d.date.slice(0, 7)))]
let version = 0
for (const month of months) {
const sample = days.find((d) => d.date.startsWith(month))!
const sellable = sample.adjustable_capacity_mw
.reduce((s, v) => s.add(new Decimal(v)), new Decimal(0))
.mul(k)
.mul('0.25')
const daysInMonth = new Date(Date.UTC(Number(month.slice(0, 4)), Number(month.slice(5, 7)), 0)).getUTCDate()
ledger.append({
id: `contract-${month}`,
timescale: 'MONTHLY',
period: month,
kind: 'CONTRACT',
energy_mwh: sellable.mul(share).mul(daysInMonth).toFixed(3),
curve: null,
source_ref: `synthetic-contract-${month}`,
expected_version: version++,
})
}
void dataset
return ledger
}
export async function runL2(
dataset: Dataset,
datasetSha256: string,
client: SkillClient,
cfg: HarnessConfig = DEFAULT_CONFIG,
clock: () => string = () => new Date().toISOString(),
): Promise<EvalRun> {
if (cfg.holdoutFrom < cfg.window) throw new Error('holdoutFrom must be ≥ window')
const end = cfg.holdoutDays > 0 ? Math.min(dataset.days.length, cfg.holdoutFrom + cfg.holdoutDays) : dataset.days.length
const holdout = dataset.days.slice(cfg.holdoutFrom, end)
const ledger = seedLedger(dataset, cfg, holdout)
const skillVersions = Object.fromEntries((await client.skills()).map((s) => [s.id, s.version]))
const perDay: DayResult[] = []
for (let idx = cfg.holdoutFrom; idx < end; idx++) {
const day = dataset.days[idx]!
const history = dataset.days.slice(idx - cfg.window, idx)
const forecasts = {} as Record<ForecastKind, ForecastBundle>
for (const kind of ['LOAD', 'PV', 'PRICE'] as const) {
forecasts[kind] = await client.forecast(kind, {
kind,
market_date: day.date,
unit: kind === 'PRICE' ? 'yuan_per_mwh' : 'mw',
history: history.map((h) => toCurve(h[FIELD[kind]] as string[], h.date)),
exogenous: {},
features_snapshot_ref: REF,
})
}
const bounds = ledger.dayAheadBounds(day.date)
const bid: BidOptimizationResult = await client.optimizeBid({
market_date: day.date,
price_forecast: forecasts.PRICE,
adjustable_capacity_mw: toCurve(day.adjustable_capacity_mw, day.date),
position_bounds: bounds,
risk: cfg.risk,
})
let ledgerAccepted = false
if (bid.solver.status === 'OPTIMAL') {
try {
ledger.append({
id: `bid-${day.date}`,
timescale: 'DAY_AHEAD',
period: day.date,
kind: 'BID_SUBMITTED',
energy_mwh: bid.daily_energy_mwh,
curve: bid.quantities_mwh,
source_ref: `eval-${day.date}`,
expected_version: ledger.read().version,
})
ledgerAccepted = true
} catch {
ledgerAccepted = false
}
}
const actualLoad = nums(day.load_mw)
const actualPv = nums(day.pv_mw)
const actualPrice = nums(day.price_yuan_per_mwh)
const load = bandOf(forecasts.LOAD)
const pv = bandOf(forecasts.PV)
const price = bandOf(forecasts.PRICE)
const capMwh = nums(day.adjustable_capacity_mw).map((c) => c * Number(cfg.risk.commitment_buffer_k) * 0.25)
const eMax = Number(bounds.daily_energy_max_mwh)
const daylight = actualPv.map((v, i) => [v, i] as const).filter(([v]) => v > 0).map(([, i]) => i)
perDay.push({
date: day.date,
load: { mape: q(mape(actualLoad, load.p50), 6), coverage: q(coverage(actualLoad, load.p10, load.p90), 6) },
pv: {
nrmse: q(nrmse(actualPv, pv.p50, Number(dataset.meta.pv_capacity_mw)), 6),
coverage: q(
coverage(
daylight.map((i) => actualPv[i]!),
daylight.map((i) => pv.p10[i]!),
daylight.map((i) => pv.p90[i]!),
),
6,
),
},
price: {
mape: q(mape(actualPrice, price.p50), 6),
coverage: q(coverage(actualPrice, price.p10, price.p90), 6),
direction: q(directionAccuracy(actualPrice, price.p50), 6),
},
bid: {
status: bid.solver.status,
ledger_accepted: ledgerAccepted,
energy_mwh: bid.daily_energy_mwh,
revenue_skill: q(
realisedRevenue({ offers: nums(bid.prices_yuan_per_mwh.values), quantities: nums(bid.quantities_mwh.values) }, actualPrice),
2,
),
revenue_naive: q(realisedRevenue(naiveBid(capMwh, eMax), actualPrice), 2),
revenue_hindsight: q(
realisedRevenue(hindsightBid(actualPrice, capMwh, eMax, Number(cfg.risk.min_block_mwh)), actualPrice),
2,
),
},
})
}
const fm = (pick: (d: DayResult) => { mape?: number; nrmse?: number; coverage: number; direction?: number }): ForecastMetrics => {
const rows = perDay.map(pick)
const has = (k: 'mape' | 'nrmse' | 'direction') => rows.every((r) => r[k] !== undefined)
return {
days: rows.length,
mape: has('mape') ? q(mean(rows.map((r) => r.mape!)), 6) : null,
nrmse: has('nrmse') ? q(mean(rows.map((r) => r.nrmse!)), 6) : null,
coverage_p10_p90: q(mean(rows.map((r) => r.coverage)), 6),
direction_accuracy: has('direction') ? q(mean(rows.map((r) => r.direction!)), 6) : null,
}
}
const sum = (f: (d: DayResult) => number) => q(perDay.reduce((s, d) => s + f(d), 0), 2)
const skill = sum((d) => d.bid.revenue_skill)
const naive = sum((d) => d.bid.revenue_naive)
const hindsight = sum((d) => d.bid.revenue_hindsight)
const body = {
layer: 'L2' as const,
dataset: { name: dataset.meta.name, sha256: datasetSha256, synthetic: dataset.meta.synthetic },
skill_versions: skillVersions,
config: cfg,
metrics: {
load: fm((d) => d.load),
pv: fm((d) => d.pv),
price: fm((d) => d.price),
bid: {
days: perDay.length,
optimal_days: perDay.filter((d) => d.bid.status === 'OPTIMAL').length,
ledger_accepted_days: perDay.filter((d) => d.bid.ledger_accepted).length,
revenue_skill_yuan: skill,
revenue_naive_yuan: naive,
revenue_hindsight_yuan: hindsight,
capture_ratio: hindsight === 0 ? 0 : q(skill / hindsight, 6),
uplift_vs_naive: naive === 0 ? 0 : q(skill / naive, 6),
},
},
per_day: perDay,
}
const digest = createHash('sha256').update(stableJson(body)).digest('hex')
const runAt = clock()
return { id: `l2-${runAt.replace(/[:.]/g, '-')}-${digest.slice(0, 8)}`, run_at: runAt, ...body, digest }
}

View File

@ -0,0 +1,4 @@
export * from './metrics.js'
export * from './dataset.js'
export * from './client.js'
export * from './harness.js'

View File

@ -0,0 +1,103 @@
/**
* L2 skill metrics (docs/12 §1 L2, §4 口径). Pure functions over plain numbers;
* the harness converts decimal strings at the boundary and quantizes results
* when it writes a report. KPI thresholds (≤ 8% etc.) are judged on real data
* and are OPEN-QUESTION C2 in their exact statistical level — nothing here
* hardcodes a pass/fail number.
*/
export const mean = (xs: number[]): number =>
xs.length === 0 ? Number.NaN : xs.reduce((a, b) => a + b, 0) / xs.length
/** Mean absolute percentage error over intervals where actual ≠ 0. */
export function mape(actual: number[], pred: number[]): number {
const terms: number[] = []
for (let i = 0; i < actual.length; i++) {
const a = actual[i]!
if (a !== 0) terms.push(Math.abs((pred[i]! - a) / a))
}
return mean(terms)
}
/** RMSE normalised by installed capacity — the PV metric docs/12 §4 suggests. */
export function nrmse(actual: number[], pred: number[], capacity: number): number {
const se = actual.map((a, i) => (pred[i]! - a) ** 2)
return Math.sqrt(mean(se)) / capacity
}
/** Share of actuals inside [lower, upper]; nominal for a P10–P90 band is 0.80. */
export function coverage(actual: number[], lower: number[], upper: number[]): number {
let hits = 0
for (let i = 0; i < actual.length; i++) {
if (actual[i]! >= lower[i]! && actual[i]! <= upper[i]!) hits++
}
return hits / actual.length
}
/**
* Direction accuracy (docs/12 L2 电价 方向准确率): share of interval pairs whose
* high/low ordering the forecast gets right. Bid optimisation depends on the
* ranking of intervals far more than on absolute price level.
*/
export function directionAccuracy(actual: number[], pred: number[]): number {
let agree = 0
let pairs = 0
for (let i = 0; i < actual.length; i++) {
for (let j = i + 1; j < actual.length; j++) {
const da = actual[j]! - actual[i]!
const dp = pred[j]! - pred[i]!
if (da === 0 && dp === 0) continue
pairs++
if (Math.sign(da) === Math.sign(dp)) agree++
}
}
return pairs === 0 ? Number.NaN : agree / pairs
}
export interface Bid {
offers: number[]
quantities: number[]
}
/** Uniform-price clearing: an offer clears where it does not exceed the realised price. */
export function realisedRevenue(bid: Bid, actualPrice: number[]): number {
let total = 0
for (let t = 0; t < actualPrice.length; t++) {
if (bid.offers[t]! <= actualPrice[t]!) total += actualPrice[t]! * bid.quantities[t]!
}
return total
}
/**
* Lower-bound baseline: a price-taker offering a flat profile that meets the
* upper energy bound (or as much as capacity allows), pro-rata to capacity.
*/
export function naiveBid(capMwh: number[], energyMax: number): Bid {
const sellable = capMwh.reduce((a, b) => a + b, 0)
const scale = sellable === 0 ? 0 : Math.min(1, energyMax / sellable)
return { offers: capMwh.map(() => 0), quantities: capMwh.map((c) => c * scale) }
}
/**
* Upper-bound baseline: perfect hindsight — fill the highest-priced intervals
* first up to the energy bound, honouring per-interval capacity and block
* size. Offers at zero so everything clears.
*/
export function hindsightBid(
actualPrice: number[],
capMwh: number[],
energyMax: number,
minBlock: number,
): Bid {
const order = actualPrice.map((_, i) => i).sort((a, b) => actualPrice[b]! - actualPrice[a]!)
const quantities = capMwh.map(() => 0)
let room = energyMax
for (const t of order) {
const q = Math.min(capMwh[t]!, room)
if (q < minBlock) continue
quantities[t] = q
room -= q
if (room <= 0) break
}
return { offers: capMwh.map(() => 0), quantities }
}

View File

@ -0,0 +1,104 @@
import { fileURLToPath } from 'node:url'
import { describe, expect, it } from 'vitest'
import type {
BidOptimizationRequest,
BidOptimizationResult,
ForecastBundle,
ForecastKind,
ForecastRequest,
ReportRequest,
SkillReport,
} from '@vpp/domain'
import type { SkillClient } from '../src/client.js'
import { loadDataset } from '../src/dataset.js'
import { DEFAULT_CONFIG, runL2 } from '../src/harness.js'
const datasetPath = fileURLToPath(new URL('../datasets/synthetic-hubei-v0.json', import.meta.url))
/**
* Stub skills with trivially checkable behaviour: forecast = last history
* day ± 10%; bid = flat quantity hitting the lower energy bound at offer 0.
* Lets the harness itself be tested without the Python service running.
*/
class StubClient implements SkillClient {
calls = 0
async skills() {
return [
{ id: 'load-forecast', version: '0.0.1', endpoint: '/stub' },
{ id: 'bid-optimization-milp', version: '0.0.1', endpoint: '/stub' },
]
}
async forecast(kind: ForecastKind, req: ForecastRequest): Promise<ForecastBundle> {
this.calls++
const last = req.history[req.history.length - 1]!
const scaled = (f: number) => ({
interval_minutes: 15 as const,
date: req.market_date,
values: last.values.map((v) => (Number(v) * f).toFixed(3)),
})
return {
id: `stub-${kind}`,
kind,
market_date: req.market_date,
unit: req.unit,
quantiles: { p10: scaled(0.9), p50: scaled(1), p90: scaled(1.1) },
model: { name: 'stub', version: '0.0.1' },
features_snapshot_ref: req.features_snapshot_ref,
generated_at: '2026-01-01T00:00:00Z',
}
}
async optimizeBid(req: BidOptimizationRequest): Promise<BidOptimizationResult> {
this.calls++
const per = (Number(req.position_bounds.daily_energy_min_mwh) / 96).toFixed(3)
const flat = { interval_minutes: 15 as const, date: req.market_date, values: Array(96).fill(per) as string[] }
const energy = (Number(per) * 96).toFixed(3)
return {
market_date: req.market_date,
prices_yuan_per_mwh: { ...flat, values: Array(96).fill('0.00') },
quantities_mwh: flat,
daily_energy_mwh: energy,
expected_revenue_yuan: '0.00',
revenue_distribution_yuan: { p10: '0.00', p50: '0.00', p90: '0.00' },
position_bounds: req.position_bounds,
solver: { name: 'stub', version: '0', status: 'OPTIMAL', objective_value: '0.00', wall_time_ms: 0 },
binding_constraints: ['daily_energy_min'],
skill_version: '0.0.1',
}
}
async report(_req: ReportRequest): Promise<SkillReport> {
throw new Error('not used')
}
}
const cfg = { ...DEFAULT_CONFIG, window: 7, holdoutFrom: 10, holdoutDays: 5 }
describe('L2 harness', () => {
it('scores every held-out day and pushes each bid through the real ledger', async () => {
const { dataset, sha256 } = loadDataset(datasetPath)
const client = new StubClient()
const run = await runL2(dataset, sha256, client, cfg, () => '2026-09-01T00:00:00Z')
expect(run.per_day).toHaveLength(5)
expect(client.calls).toBe(5 * 4)
expect(run.metrics.bid.optimal_days).toBe(5)
expect(run.metrics.bid.ledger_accepted_days).toBe(5) // stub bids sit exactly on the lower bound
expect(run.metrics.load.coverage_p10_p90).toBeGreaterThan(0)
expect(run.metrics.price.direction_accuracy).toBeGreaterThan(0.5) // yesterday's shape is informative
expect(run.metrics.bid.revenue_hindsight_yuan).toBeGreaterThanOrEqual(run.metrics.bid.revenue_skill_yuan)
expect(run.metrics.bid.revenue_hindsight_yuan).toBeGreaterThanOrEqual(run.metrics.bid.revenue_naive_yuan)
expect(run.skill_versions['load-forecast']).toBe('0.0.1')
expect(run.dataset.synthetic).toBe(true)
})
it('is reproducible: same inputs → same digest, independent of run time', async () => {
const { dataset, sha256 } = loadDataset(datasetPath)
const a = await runL2(dataset, sha256, new StubClient(), cfg, () => '2026-09-01T00:00:00Z')
const b = await runL2(dataset, sha256, new StubClient(), cfg, () => '2026-09-02T00:00:00Z')
expect(a.digest).toBe(b.digest)
expect(a.id).not.toBe(b.id)
})
it('rejects a holdout that starts before the window is filled', async () => {
const { dataset, sha256 } = loadDataset(datasetPath)
await expect(runL2(dataset, sha256, new StubClient(), { ...cfg, holdoutFrom: 3 })).rejects.toThrow(/window/)
})
})

View File

@ -0,0 +1,61 @@
import { describe, expect, it } from 'vitest'
import {
coverage,
directionAccuracy,
hindsightBid,
mape,
naiveBid,
nrmse,
realisedRevenue,
} from '../src/metrics.js'
describe('forecast metrics', () => {
it('mape ignores zero actuals and is 0 for a perfect forecast', () => {
expect(mape([10, 0, 20], [10, 5, 20])).toBe(0)
expect(mape([10, 20], [11, 18])).toBeCloseTo((0.1 + 0.1) / 2)
})
it('nrmse normalises by capacity', () => {
expect(nrmse([0, 10], [0, 10], 20)).toBe(0)
expect(nrmse([10, 10], [12, 8], 20)).toBeCloseTo(2 / 20)
})
it('coverage counts inclusive band hits', () => {
expect(coverage([1, 2, 3, 4], [1, 1, 4, 1], [1, 3, 5, 3])).toBe(0.5)
})
it('direction accuracy scores pairwise ordering, not level', () => {
expect(directionAccuracy([1, 2, 3], [10, 20, 30])).toBe(1)
expect(directionAccuracy([1, 2, 3], [30, 20, 10])).toBe(0)
expect(directionAccuracy([1, 2, 3], [1, 3, 2])).toBeCloseTo(2 / 3)
})
})
describe('bid backtest baselines', () => {
const price = [100, 300, 200, 50]
const cap = [2, 2, 2, 2]
it('realised revenue clears only where offer ≤ price', () => {
expect(realisedRevenue({ offers: [150, 150, 150, 150], quantities: [1, 1, 1, 1] }, price)).toBe(500)
})
it('naive bid is flat, price-taking and meets the energy bound', () => {
const bid = naiveBid(cap, 4)
expect(bid.quantities).toEqual([1, 1, 1, 1])
expect(bid.offers.every((o) => o === 0)).toBe(true)
expect(naiveBid(cap, 100).quantities).toEqual(cap) // capped by capacity
})
it('hindsight bid fills highest-priced intervals first and bounds revenue above', () => {
const bid = hindsightBid(price, cap, 3, 0.5)
expect(bid.quantities).toEqual([0, 2, 1, 0])
const best = realisedRevenue(bid, price)
expect(best).toBe(800)
expect(realisedRevenue(naiveBid(cap, 3), price)).toBeLessThan(best)
})
it('hindsight honours the block size', () => {
const bid = hindsightBid(price, cap, 2.2, 0.5)
expect(bid.quantities).toEqual([0, 2, 0, 0]) // leftover 0.2 < block → not placed
})
})

View File

@ -0,0 +1,8 @@
{
"extends": "../../tsconfig.base.json",
"compilerOptions": {
"rootDir": ".",
"noEmit": true
},
"include": ["src", "test"]
}

View File

@ -0,0 +1,12 @@
import decimalPkg from 'decimal.js-light'
/**
* decimal.js-light ships CommonJS (`module.exports = Decimal`, plus
* `Decimal.Decimal = Decimal`). Node ESM exposes only a default export; some
* bundlers/test runners expose the class itself as the default. Resolve once
* here so callers just `import { Decimal } from './decimal.js'`.
*/
type DecimalCtor = typeof decimalPkg.Decimal
const resolved: unknown = (decimalPkg as { Decimal?: unknown }).Decimal ?? decimalPkg
export const Decimal = resolved as DecimalCtor
export type Decimal = InstanceType<DecimalCtor>

View File

@ -4,3 +4,4 @@ export * from './quality.js'
export * from './timeseries.js'
export * from './relational.js'
export * from './ingest.js'
export { Decimal } from './decimal.js'

View File

@ -1,5 +1,5 @@
import { Decimal } from 'decimal.js-light'
import type { LedgerView, PositionEntry, PositionUpdate, Timescale } from '@vpp/domain'
import { Decimal } from './decimal.js'
import type { LedgerView, PositionBounds, PositionEntry, PositionUpdate, Timescale } from '@vpp/domain'
import { PositionUpdate as PositionUpdateSchema } from '@vpp/domain'
export class LedgerConcurrencyError extends Error {
@ -46,6 +46,22 @@ export class LedgerService {
return this.entries.filter((e) => e.timescale === timescale)
}
/**
* Daily energy bounds for a day-ahead bid on `date`, derived from the
* MONTHLY CONTRACT position (the same rule checkCascade enforces). Handed
* to the bid-optimization skill as hard constraints so the optimizer can
* never produce a bid the ledger would then reject. Throws when there is
* no monthly anchor — same policy as append.
*/
dayAheadBounds(date: string): PositionBounds {
const { min, max } = this.cascadeBand(date)
return {
ledger_version: this.version,
daily_energy_min_mwh: min.toString(),
daily_energy_max_mwh: max.toString(),
}
}
append(update: PositionUpdate): LedgerView {
PositionUpdateSchema.parse(update)
if (update.expected_version !== this.version) {
@ -71,8 +87,18 @@ export class LedgerService {
*/
private checkCascade(update: PositionUpdate): void {
if (update.kind !== 'BID_SUBMITTED' || update.timescale !== 'DAY_AHEAD') return
const { min, max, contracted, daysInMonth, band } = this.cascadeBand(update.period)
const bid = new Decimal(update.energy_mwh)
if (bid.lt(min) || bid.gt(max)) {
throw new CascadeViolation(
`day-ahead bid ${bid.toString()} MWh outside monthly cascade band ` +
`[${min.toString()}, ${max.toString()}] (contracted ${contracted.toString()} MWh / ${daysInMonth} days ± ${band.mul(100).toString()}%)`,
)
}
}
const month = update.period.slice(0, 7)
private cascadeBand(date: string) {
const month = date.slice(0, 7)
const monthly = this.entries.filter(
(e) => e.timescale === 'MONTHLY' && e.kind === 'CONTRACT' && e.period === month,
)
@ -81,22 +107,18 @@ export class LedgerService {
`no MONTHLY CONTRACT position recorded for ${month}; day-ahead bid has no cascade anchor`,
)
}
const contracted = monthly.reduce((sum, e) => sum.add(new Decimal(e.energy_mwh)), new Decimal(0))
const daysInMonth = new Date(
Date.UTC(Number(month.slice(0, 4)), Number(month.slice(5, 7)), 0),
).getUTCDate()
const dailyShare = contracted.div(daysInMonth)
const band = new Decimal(this.cfg.daMonthlyDeviationBand)
const min = dailyShare.mul(new Decimal(1).sub(band))
const max = dailyShare.mul(new Decimal(1).add(band))
const bid = new Decimal(update.energy_mwh)
if (bid.lt(min) || bid.gt(max)) {
throw new CascadeViolation(
`day-ahead bid ${bid.toString()} MWh outside monthly cascade band ` +
`[${min.toString()}, ${max.toString()}] (contracted ${contracted.toString()} MWh / ${daysInMonth} days ± ${band.mul(100).toString()}%)`,
)
return {
min: dailyShare.mul(new Decimal(1).sub(band)),
max: dailyShare.mul(new Decimal(1).add(band)),
contracted,
daysInMonth,
band,
}
}
}

View File

@ -75,3 +75,22 @@ describe('views', () => {
expect(ledger.viewByTimescale('DAY_AHEAD')).toHaveLength(1)
})
})
describe('dayAheadBounds: the cascade band handed to the bid optimizer', () => {
it('returns the same band checkCascade enforces, stamped with the ledger version', () => {
const ledger = makeLedger()
ledger.append(monthlyContract)
expect(ledger.dayAheadBounds('2026-03-15')).toEqual({
ledger_version: 1,
daily_energy_min_mwh: '1140',
daily_energy_max_mwh: '1260',
})
// A bid exactly at each bound is accepted by append — bounds are inclusive.
ledger.append(daBid('1140', 1))
ledger.append({ ...daBid('1260', 2), id: 'pos-da-max' })
})
it('throws when there is no monthly anchor', () => {
expect(() => makeLedger().dayAheadBounds('2026-03-15')).toThrow(CascadeViolation)
})
})

View File

@ -1,3 +1,9 @@
pydantic>=2.7
datamodel-code-generator>=0.26
pytest>=8.0
hypothesis>=6.100
numpy>=2.0
scipy>=1.14
fastapi>=0.115
uvicorn>=0.30
httpx>=0.27

View File

View File

@ -0,0 +1,36 @@
from __future__ import annotations
import sys
from pathlib import Path
import pytest
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from vpp_skills.synthetic import generate_dataset # noqa: E402
REF_A = "a" * 64
def curve_of(values: list[str], date: str) -> dict:
return {"interval_minutes": 15, "date": date, "values": values}
@pytest.fixture(scope="session")
def dataset() -> dict:
return generate_dataset(seed=7, n_days=90)
def forecast_request(dataset: dict, kind: str, target_idx: int, window: int = 28) -> dict:
field = {"LOAD": "load_mw", "PV": "pv_mw", "PRICE": "price_yuan_per_mwh"}[kind]
unit = "yuan_per_mwh" if kind == "PRICE" else "mw"
days = dataset["days"]
hist = days[max(0, target_idx - window) : target_idx]
return {
"kind": kind,
"market_date": days[target_idx]["date"],
"unit": unit,
"history": [curve_of(d[field], d["date"]) for d in hist],
"exogenous": {},
"features_snapshot_ref": REF_A,
}

View File

@ -0,0 +1,60 @@
"""HTTP contract: golden request fixtures go in, schema-valid responses come out."""
from __future__ import annotations
import json
from decimal import Decimal
from pathlib import Path
from fastapi.testclient import TestClient
from vpp_contracts.bid_optimization_result import BidOptimizationResult
from vpp_contracts.forecast_bundle import ForecastBundle
from vpp_contracts.skill_report import SkillReport
from vpp_skills.app import app
from .conftest import forecast_request
FIXTURES = Path(__file__).resolve().parents[2] / "contracts" / "fixtures"
client = TestClient(app)
def test_registry_lists_versions():
r = client.get("/v1/skills")
assert r.status_code == 200
ids = {s["id"]: s for s in r.json()}
assert ids["bid-optimization-milp"]["endpoint"] == "/v1/optimize/bid"
assert ids["load-forecast"]["version"] == "1.0.0"
def test_price_forecast_from_golden_fixture():
req = json.loads((FIXTURES / "forecast_request" / "price-da.json").read_text())
r = client.post("/v1/forecast/price", json=req)
assert r.status_code == 200, r.text
ForecastBundle.model_validate(r.json())
def test_kind_endpoint_mismatch_is_422(dataset):
r = client.post("/v1/forecast/pv", json=forecast_request(dataset, "LOAD", 30))
assert r.status_code == 422
def test_bid_optimization_from_golden_fixture():
req = json.loads((FIXTURES / "bid_optimization_request" / "da-basic.json").read_text())
r = client.post("/v1/optimize/bid", json=req)
assert r.status_code == 200, r.text
out = BidOptimizationResult.model_validate(r.json())
assert out.solver.status.value == "OPTIMAL"
assert Decimal("1140") <= Decimal(out.daily_energy_mwh) <= Decimal("1260")
def test_report_from_golden_fixture():
req = json.loads((FIXTURES / "report_request" / "da-bid-summary.json").read_text())
r = client.post("/v1/report", json=req)
assert r.status_code == 200, r.text
SkillReport.model_validate(r.json())
def test_schema_violation_is_422():
r = client.post("/v1/optimize/bid", json={"market_date": "2026-03-15"})
assert r.status_code == 422

View File

@ -0,0 +1,140 @@
"""Bid MILP: property tests that the output respects ledger bounds and capacity
(ROADMAP M2 acceptance), plus revenue-consistency and infeasibility handling."""
from __future__ import annotations
from decimal import Decimal
import numpy as np
from hypothesis import given, settings
from hypothesis import strategies as st
from vpp_contracts.bid_optimization_request import BidOptimizationRequest
from vpp_skills.bid_milp import optimize_bid
from vpp_skills.numeric import quantize
from .conftest import REF_A, curve_of
DATE = "2026-03-15"
def _request(p50: np.ndarray, spread: float, cap_mw: np.ndarray, e_min: float, e_max: float,
lam: str = "0.3", k: str = "0.9", min_block: str = "0.5", cost: str = "0") -> BidOptimizationRequest:
p10, p90 = p50 * (1 - spread), p50 * (1 + spread)
q = lambda arr, s: [quantize(float(v), s) for v in arr] # noqa: E731
return BidOptimizationRequest.model_validate(
{
"market_date": DATE,
"price_forecast": {
"id": "fc-price", "kind": "PRICE", "market_date": DATE, "unit": "yuan_per_mwh",
"quantiles": {"p10": curve_of(q(p10, 2), DATE), "p50": curve_of(q(p50, 2), DATE), "p90": curve_of(q(p90, 2), DATE)},
"model": {"name": "price-forecast", "version": "1.0.0"},
"features_snapshot_ref": REF_A, "generated_at": "2026-03-14T06:00:00Z",
},
"adjustable_capacity_mw": curve_of(q(cap_mw, 3), DATE),
"position_bounds": {"ledger_version": 42, "daily_energy_min_mwh": quantize(e_min, 3), "daily_energy_max_mwh": quantize(e_max, 3)},
"risk": {"risk_aversion": lam, "commitment_buffer_k": k, "min_block_mwh": min_block, "marginal_cost_yuan_per_mwh": cost},
}
)
def _vals(c) -> np.ndarray:
return np.array([float(Decimal(v.root)) for v in c.values])
@settings(max_examples=40, deadline=None)
@given(
seed=st.integers(0, 10_000),
lam=st.sampled_from(["0", "0.25", "0.5", "1"]),
k=st.sampled_from(["0.5", "0.8", "1"]),
band=st.floats(0.0, 0.3),
)
def test_bid_respects_ledger_bounds_and_capacity(seed, lam, k, band):
rng = np.random.default_rng(seed)
p50 = rng.uniform(200, 800, 96)
cap = rng.uniform(0, 40, 96)
sellable = float(k) * cap.sum() * 0.25
centre = rng.uniform(0.2, 0.9) * sellable
e_min, e_max = centre * (1 - band), centre * (1 + band)
req = _request(p50, 0.2, cap, e_min, e_max, lam=lam, k=k)
res = optimize_bid(req)
assert res.solver.status.value == "OPTIMAL", res.solver
qty = _vals(res.quantities_mwh)
total = float(Decimal(res.daily_energy_mwh))
assert abs(total - qty.sum()) < 1e-6 # headline figure derived from the curve
lo, hi = Decimal(req.position_bounds.daily_energy_min_mwh), Decimal(req.position_bounds.daily_energy_max_mwh)
assert lo <= Decimal(res.daily_energy_mwh) <= hi # P7 cascade: hard constraint, exact
assert np.all(qty <= float(k) * cap * 0.25 + 1e-3)
assert np.all((qty == 0) | (qty >= 0.5 - 1e-3)) # min block honoured
assert np.all(qty >= 0)
def test_prefers_high_price_intervals():
p50 = np.full(96, 300.0)
p50[72:80] = 900.0 # 18:00–20:00 evening peak
cap = np.full(96, 40.0)
res = optimize_bid(_request(p50, 0.1, cap, 40.0, 60.0, lam="0"))
qty = _vals(res.quantities_mwh)
assert qty[72:80].sum() > 0.99 * qty.sum()
assert "daily_energy_max" in res.binding_constraints
def test_revenue_distribution_is_exact_and_ordered():
rng = np.random.default_rng(1)
p50 = rng.uniform(300, 600, 96)
req = _request(p50, 0.15, np.full(96, 30.0), 200.0, 400.0, lam="0.5")
res = optimize_bid(req)
d = res.revenue_distribution_yuan
assert Decimal(d.p10) <= Decimal(d.p50) <= Decimal(d.p90)
assert res.expected_revenue_yuan == d.p50
# Price-taker offer (cost 0) clears everywhere: P50 revenue = Σ p50·q exactly, in Decimal.
p50s = [v.root for v in req.price_forecast.quantiles.p50.values]
qty = [v.root for v in res.quantities_mwh.values]
expected = sum(Decimal(p) * Decimal(q) for p, q in zip(p50s, qty))
assert Decimal(res.expected_revenue_yuan) == expected.quantize(Decimal("0.01"))
def test_offer_is_the_marginal_cost_floor():
p50 = np.full(96, 500.0)
free = optimize_bid(_request(p50, 0.2, np.full(96, 30.0), 100.0, 200.0, cost="0"))
costly = optimize_bid(_request(p50, 0.2, np.full(96, 30.0), 100.0, 200.0, cost="450"))
assert set(_vals(free.prices_yuan_per_mwh)) == {0.0}
assert set(_vals(costly.prices_yuan_per_mwh)) == {450.0}
# At cost 450 the offer no longer clears on the P10 path (400): floor revenue is zero.
assert costly.revenue_distribution_yuan.p10 == "0.00"
assert Decimal(free.revenue_distribution_yuan.p10) > 0
def test_risk_aversion_tilts_allocation_toward_narrow_bands():
"""Two intervals, same P50; one has a wide band. Risk-neutral is indifferent
(fills by index order), risk-averse must prefer the narrow band."""
p50 = np.full(96, 100.0)
p50[[10, 20]] = 500.0
cap = np.zeros(96)
cap[[10, 20]] = 40.0
band = np.full(96, 0.1)
band[10] = 0.6 # interval 10: P10 = 200; interval 20: P10 = 450
q = lambda arr, s: [quantize(float(v), s) for v in arr] # noqa: E731
req = _request(p50, 0.1, cap, 5.0, 5.0, lam="1")
data = req.model_dump()
data["price_forecast"]["quantiles"]["p10"]["values"] = q(p50 * (1 - band), 2)
req = BidOptimizationRequest.model_validate(data)
res = optimize_bid(req)
qty = _vals(res.quantities_mwh)
assert qty[20] == 5.0 and qty[10] == 0.0
def test_infeasible_when_position_exceeds_sellable_energy():
res = optimize_bid(_request(np.full(96, 400.0), 0.1, np.full(96, 10.0), 500.0, 600.0, k="1"))
assert res.solver.status.value == "INFEASIBLE"
assert res.daily_energy_mwh == "0.000"
assert res.binding_constraints == ["daily_energy_min exceeds sellable energy"]
def test_rejects_out_of_range_risk_params():
import pytest
with pytest.raises(ValueError, match="risk_aversion"):
optimize_bid(_request(np.full(96, 400.0), 0.1, np.full(96, 10.0), 10.0, 20.0, lam="1.5"))
with pytest.raises(ValueError, match="commitment_buffer_k"):
optimize_bid(_request(np.full(96, 400.0), 0.1, np.full(96, 10.0), 10.0, 20.0, k="0"))

View File

@ -0,0 +1,81 @@
"""Forecast skill: contract shape, determinism, and the calibration test the
roadmap names as the M2 acceptance gate (coverage of the P10–P90 band)."""
from __future__ import annotations
from datetime import datetime, timezone
from decimal import Decimal
import numpy as np
import pytest
from vpp_contracts.forecast_request import ForecastRequest
from vpp_skills.forecast import forecast
from .conftest import forecast_request
CLOCK = lambda: datetime(2026, 3, 14, 6, 0, tzinfo=timezone.utc) # noqa: E731
def _arr(curve) -> np.ndarray:
return np.array([float(Decimal(v.root)) for v in curve.values])
@pytest.mark.parametrize("kind", ["LOAD", "PV", "PRICE"])
def test_bundle_shape_and_ordering(dataset, kind):
req = ForecastRequest.model_validate(forecast_request(dataset, kind, 40))
b = forecast(req, CLOCK)
assert b.kind.value == kind and b.market_date == req.market_date
p10, p50, p90 = (_arr(getattr(b.quantiles, q)) for q in ("p10", "p50", "p90"))
assert np.all(p10 <= p50 + 1e-9) and np.all(p50 <= p90 + 1e-9)
if kind != "PRICE":
assert np.all(p10 >= 0)
assert b.generated_at.isoformat() == "2026-03-14T06:00:00+00:00"
assert b.model.name == f"{kind.lower()}-forecast"
def test_deterministic(dataset):
req = ForecastRequest.model_validate(forecast_request(dataset, "LOAD", 50))
a, b = forecast(req, CLOCK), forecast(req, CLOCK)
assert a.model_dump() == b.model_dump()
def test_pv_night_stays_zero(dataset):
req = ForecastRequest.model_validate(forecast_request(dataset, "PV", 45))
b = forecast(req, CLOCK)
assert _arr(b.quantiles.p90)[:20].sum() == 0.0 # 00:00–05:00
assert _arr(b.quantiles.p50)[44:52].sum() > 0.0 # midday
@pytest.mark.parametrize("kind", ["LOAD", "PV", "PRICE"])
def test_interval_calibration(dataset, kind):
"""Nominal 80% band must cover roughly 80% of held-out actuals. Materially
under-covering (optimistic bands) is the failure mode docs/12 flags as more
dangerous than point error."""
field = {"LOAD": "load_mw", "PV": "pv_mw", "PRICE": "price_yuan_per_mwh"}[kind]
hits = total = 0
for idx in range(35, 90):
req = ForecastRequest.model_validate(forecast_request(dataset, kind, idx))
b = forecast(req, CLOCK)
actual = np.array([float(Decimal(v)) for v in dataset["days"][idx][field]])
p10, p90 = _arr(b.quantiles.p10), _arr(b.quantiles.p90)
if kind == "PV": # night intervals are trivially covered; score daylight only
mask = actual > 0
actual, p10, p90 = actual[mask], p10[mask], p90[mask]
hits += int(np.sum((actual >= p10) & (actual <= p90)))
total += len(actual)
coverage = hits / total
assert 0.70 <= coverage <= 0.92, f"{kind} P10–P90 coverage {coverage:.3f} off nominal 0.80"
def test_rejects_bad_history(dataset):
base = forecast_request(dataset, "LOAD", 40)
unsorted = {**base, "history": list(reversed(base["history"]))}
with pytest.raises(ValueError, match="ascending"):
forecast(ForecastRequest.model_validate(unsorted), CLOCK)
future = {**base, "market_date": base["history"][0]["date"]}
with pytest.raises(ValueError, match="strictly before"):
forecast(ForecastRequest.model_validate(future), CLOCK)
wrong_unit = {**base, "unit": "yuan_per_mwh"}
with pytest.raises(ValueError, match="unit"):
forecast(ForecastRequest.model_validate(wrong_unit), CLOCK)

View File

@ -0,0 +1,63 @@
from __future__ import annotations
from datetime import datetime, timezone
from vpp_contracts.report_request import ReportRequest
from vpp_skills.report import generate_report, resolve, verify_report
CLOCK = lambda: datetime(2026, 3, 14, 8, 30, tzinfo=timezone.utc) # noqa: E731
SOLVER_OUTPUT = {
"market_date": "2026-03-15",
"prices_yuan_per_mwh": {"interval_minutes": 15, "date": "2026-03-15", "values": ["1.0"] * 96},
"daily_energy_mwh": "1200.000",
"expected_revenue_yuan": "510600.00",
"revenue_distribution_yuan": {"p10": "432000.00", "p50": "510600.00", "p90": "588000.00"},
"solver": {"name": "highs", "status": "OPTIMAL", "objective_value": "487020.00", "wall_time_ms": 12},
"binding_constraints": ["daily_energy_max"],
}
def _req(kind="DAY_AHEAD_BID_SUMMARY") -> ReportRequest:
return ReportRequest.model_validate(
{
"kind": kind,
"market_date": "2026-03-15",
"sources": [{"tool_call_id": "tc-001", "tool": "bid-optimization-milp", "version": "1.0.0", "output": SOLVER_OUTPUT}],
}
)
def test_every_metric_is_a_reference_that_resolves_to_its_value():
req = _req()
rep = generate_report(req, CLOCK)
names = {m.name for m in rep.sections[0].metrics}
assert {"expected_revenue_yuan", "daily_energy_mwh", "revenue_distribution_yuan.p10", "solver.objective_value"} <= names
assert "prices_yuan_per_mwh" not in " ".join(names) # curves are not inlined
assert verify_report(rep, req) == []
for m in rep.sections[0].metrics:
assert m.value == resolve(SOLVER_OUTPUT, m.ref.path)
assert m.ref.tool_call_id == "tc-001"
def test_units_inferred_from_field_names():
rep = generate_report(_req(), CLOCK)
units = {m.name: m.unit for m in rep.sections[0].metrics}
assert units["expected_revenue_yuan"] == "yuan"
assert units["daily_energy_mwh"] == "mwh"
assert units["revenue_distribution_yuan.p90"] == "yuan"
def test_notes_carry_solver_status_and_binding_constraints():
rep = generate_report(_req(), CLOCK)
notes = " | ".join(rep.sections[0].notes)
assert "solver status: OPTIMAL" in notes and "daily_energy_max" in notes
def test_deterministic_id_and_verify_catches_tampering():
req = _req()
a, b = generate_report(req, CLOCK), generate_report(req, CLOCK)
assert a.id == b.id and a.id.startswith("rep-day_ahead_bid_summary-2026-03-15-")
tampered = a.model_copy(deep=True)
tampered.sections[0].metrics[0].value = "999999.00"
assert verify_report(tampered, req)

View File

@ -0,0 +1,129 @@
# generated by datamodel-codegen:
# filename: bid_optimization_request.json
from __future__ import annotations
from enum import StrEnum
from typing import Literal
from pydantic import (
AwareDatetime,
BaseModel,
ConfigDict,
Field,
RootModel,
conint,
constr,
)
class Kind(StrEnum):
LOAD = 'LOAD'
PV = 'PV'
PRICE = 'PRICE'
class Unit(StrEnum):
mw = 'mw'
yuan_per_mwh = 'yuan_per_mwh'
class Value(RootModel[constr(pattern=r'^-?\d+(\.\d+)?$')]):
root: constr(pattern=r'^-?\d+(\.\d+)?$')
class P10(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
interval_minutes: Literal[15]
date: constr(pattern=r'^\d{4}-\d{2}-\d{2}$')
values: list[Value] = Field(..., max_length=96, min_length=96)
class P50(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
interval_minutes: Literal[15]
date: constr(pattern=r'^\d{4}-\d{2}-\d{2}$')
values: list[Value] = Field(..., max_length=96, min_length=96)
class P90(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
interval_minutes: Literal[15]
date: constr(pattern=r'^\d{4}-\d{2}-\d{2}$')
values: list[Value] = Field(..., max_length=96, min_length=96)
class Quantiles(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
p10: P10
p50: P50
p90: P90
class Model(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
name: constr(min_length=1)
version: constr(min_length=1)
class PriceForecast(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
id: constr(min_length=1)
kind: Kind
market_date: constr(pattern=r'^\d{4}-\d{2}-\d{2}$')
unit: Unit
quantiles: Quantiles
model: Model
features_snapshot_ref: constr(pattern=r'^[0-9a-f]{64}$')
generated_at: AwareDatetime
class AdjustableCapacityMw(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
interval_minutes: Literal[15]
date: constr(pattern=r'^\d{4}-\d{2}-\d{2}$')
values: list[Value] = Field(..., max_length=96, min_length=96)
class PositionBounds(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
ledger_version: conint(ge=0, le=9007199254740991)
daily_energy_min_mwh: constr(pattern=r'^-?\d+(\.\d+)?$')
daily_energy_max_mwh: constr(pattern=r'^-?\d+(\.\d+)?$')
class Risk(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
risk_aversion: constr(pattern=r'^-?\d+(\.\d+)?$')
commitment_buffer_k: constr(pattern=r'^-?\d+(\.\d+)?$')
min_block_mwh: constr(pattern=r'^-?\d+(\.\d+)?$')
marginal_cost_yuan_per_mwh: constr(pattern=r'^-?\d+(\.\d+)?$')
class BidOptimizationRequest(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
market_date: constr(pattern=r'^\d{4}-\d{2}-\d{2}$')
price_forecast: PriceForecast
adjustable_capacity_mw: AdjustableCapacityMw
position_bounds: PositionBounds
risk: Risk

View File

@ -0,0 +1,83 @@
# generated by datamodel-codegen:
# filename: bid_optimization_result.json
from __future__ import annotations
from enum import StrEnum
from typing import Literal
from pydantic import BaseModel, ConfigDict, Field, RootModel, conint, constr
class Value(RootModel[constr(pattern=r'^-?\d+(\.\d+)?$')]):
root: constr(pattern=r'^-?\d+(\.\d+)?$')
class PricesYuanPerMwh(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
interval_minutes: Literal[15]
date: constr(pattern=r'^\d{4}-\d{2}-\d{2}$')
values: list[Value] = Field(..., max_length=96, min_length=96)
class QuantitiesMwh(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
interval_minutes: Literal[15]
date: constr(pattern=r'^\d{4}-\d{2}-\d{2}$')
values: list[Value] = Field(..., max_length=96, min_length=96)
class RevenueDistributionYuan(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
p10: constr(pattern=r'^-?\d+(\.\d+)?$')
p50: constr(pattern=r'^-?\d+(\.\d+)?$')
p90: constr(pattern=r'^-?\d+(\.\d+)?$')
class PositionBounds(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
ledger_version: conint(ge=0, le=9007199254740991)
daily_energy_min_mwh: constr(pattern=r'^-?\d+(\.\d+)?$')
daily_energy_max_mwh: constr(pattern=r'^-?\d+(\.\d+)?$')
class Status(StrEnum):
OPTIMAL = 'OPTIMAL'
INFEASIBLE = 'INFEASIBLE'
TIME_LIMIT = 'TIME_LIMIT'
ERROR = 'ERROR'
class Solver(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
name: constr(min_length=1)
version: constr(min_length=1)
status: Status
objective_value: constr(pattern=r'^-?\d+(\.\d+)?$') | None
wall_time_ms: conint(ge=0, le=9007199254740991)
class BidOptimizationResult(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
market_date: constr(pattern=r'^\d{4}-\d{2}-\d{2}$')
prices_yuan_per_mwh: PricesYuanPerMwh
quantities_mwh: QuantitiesMwh
daily_energy_mwh: constr(pattern=r'^-?\d+(\.\d+)?$')
expected_revenue_yuan: constr(pattern=r'^-?\d+(\.\d+)?$')
revenue_distribution_yuan: RevenueDistributionYuan
position_bounds: PositionBounds
solver: Solver
binding_constraints: list[str]
skill_version: constr(pattern=r'^\d+\.\d+\.\d+$')

View File

@ -0,0 +1,54 @@
# generated by datamodel-codegen:
# filename: forecast_request.json
from __future__ import annotations
from enum import StrEnum
from typing import Literal
from pydantic import BaseModel, ConfigDict, Field, RootModel, constr
class Kind(StrEnum):
LOAD = 'LOAD'
PV = 'PV'
PRICE = 'PRICE'
class Unit(StrEnum):
mw = 'mw'
yuan_per_mwh = 'yuan_per_mwh'
class Value(RootModel[constr(pattern=r'^-?\d+(\.\d+)?$')]):
root: constr(pattern=r'^-?\d+(\.\d+)?$')
class HistoryItem(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
interval_minutes: Literal[15]
date: constr(pattern=r'^\d{4}-\d{2}-\d{2}$')
values: list[Value] = Field(..., max_length=96, min_length=96)
class Exogenous(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
interval_minutes: Literal[15]
date: constr(pattern=r'^\d{4}-\d{2}-\d{2}$')
values: list[Value] = Field(..., max_length=96, min_length=96)
class ForecastRequest(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
kind: Kind
market_date: constr(pattern=r'^\d{4}-\d{2}-\d{2}$')
unit: Unit
history: list[HistoryItem] = Field(..., min_length=1)
exogenous: dict[str, Exogenous]
features_snapshot_ref: constr(pattern=r'^[0-9a-f]{64}$')

View File

@ -0,0 +1,34 @@
# generated by datamodel-codegen:
# filename: report_request.json
from __future__ import annotations
from enum import StrEnum
from typing import Any
from pydantic import BaseModel, ConfigDict, Field, constr
class Kind(StrEnum):
DAY_AHEAD_BID_SUMMARY = 'DAY_AHEAD_BID_SUMMARY'
FORECAST_EVAL = 'FORECAST_EVAL'
BID_BACKTEST = 'BID_BACKTEST'
class Source(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
tool_call_id: constr(min_length=1)
tool: constr(min_length=1)
version: constr(pattern=r'^\d+\.\d+\.\d+$')
output: Any
class ReportRequest(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
kind: Kind
market_date: constr(pattern=r'^\d{4}-\d{2}-\d{2}$')
sources: list[Source] = Field(..., min_length=1)

View File

@ -0,0 +1,53 @@
# generated by datamodel-codegen:
# filename: skill_report.json
from __future__ import annotations
from enum import StrEnum
from pydantic import AwareDatetime, BaseModel, ConfigDict, constr
class Kind(StrEnum):
DAY_AHEAD_BID_SUMMARY = 'DAY_AHEAD_BID_SUMMARY'
FORECAST_EVAL = 'FORECAST_EVAL'
BID_BACKTEST = 'BID_BACKTEST'
class Ref(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
tool_call_id: constr(min_length=1)
path: constr(min_length=1)
class Metric(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
name: constr(min_length=1)
value: constr(pattern=r'^-?\d+(\.\d+)?$')
unit: constr(min_length=1)
ref: Ref
class Section(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
title: constr(min_length=1)
metrics: list[Metric]
notes: list[str]
class SkillReport(BaseModel):
model_config = ConfigDict(
extra='forbid',
)
id: constr(min_length=1)
kind: Kind
market_date: constr(pattern=r'^\d{4}-\d{2}-\d{2}$')
sections: list[Section]
skill_version: constr(pattern=r'^\d+\.\d+\.\d+$')
generated_at: AwareDatetime

View File

@ -0,0 +1,15 @@
"""Python skill services (docs/05 §1–2, ADR-0001).
Every skill is typed (generated pydantic contracts in ``vpp_contracts``),
versioned (``SKILL_VERSIONS``) and stateless. No module here may import or call
an LLM: skills are deterministic numeric services and sit on the control path
(CLAUDE.md hard rule 2).
"""
SKILL_VERSIONS: dict[str, str] = {
"load-forecast": "1.0.0",
"pv-forecast": "1.0.0",
"price-forecast": "1.0.0",
"bid-optimization-milp": "1.0.0",
"report-generator": "1.0.0",
}

View File

@ -0,0 +1,81 @@
"""HTTP/JSON skill service (docs/11 §3.4, docs/09 §4).
One process, one route per skill, generated pydantic contracts on both request
and response. The TS runtime's registerSkill wrapper is the only intended
caller; it snapshots request/response for lineage. Run locally with
``uvicorn vpp_skills.app:app``.
"""
from __future__ import annotations
from fastapi import FastAPI, HTTPException
from fastapi.responses import JSONResponse
from vpp_contracts.bid_optimization_request import BidOptimizationRequest
from vpp_contracts.bid_optimization_result import BidOptimizationResult
from vpp_contracts.forecast_bundle import ForecastBundle
from vpp_contracts.forecast_request import ForecastRequest
from vpp_contracts.report_request import ReportRequest
from vpp_contracts.skill_report import SkillReport
from . import SKILL_VERSIONS
from .bid_milp import optimize_bid
from .forecast import forecast
from .report import generate_report
app = FastAPI(title="vpp-skills", version="0.1.0")
ROUTES = {
"load-forecast": "/v1/forecast/load",
"pv-forecast": "/v1/forecast/pv",
"price-forecast": "/v1/forecast/price",
"bid-optimization-milp": "/v1/optimize/bid",
"report-generator": "/v1/report",
}
@app.exception_handler(ValueError)
async def _value_error(_, exc: ValueError) -> JSONResponse:
return JSONResponse(status_code=422, content={"detail": str(exc)})
@app.get("/health")
def health() -> dict:
return {"status": "ok"}
@app.get("/v1/skills")
def skills() -> list[dict]:
"""Registry view: what the TS tool registry pins versions against."""
return [{"id": sid, "version": ver, "endpoint": ROUTES[sid]} for sid, ver in SKILL_VERSIONS.items()]
def _forecast_for(kind: str, req: ForecastRequest) -> ForecastBundle:
if req.kind.value != kind:
raise HTTPException(status_code=422, detail=f"this endpoint serves kind={kind}, got {req.kind.value}")
return forecast(req)
@app.post(ROUTES["load-forecast"], response_model=ForecastBundle)
def forecast_load(req: ForecastRequest) -> ForecastBundle:
return _forecast_for("LOAD", req)
@app.post(ROUTES["pv-forecast"], response_model=ForecastBundle)
def forecast_pv(req: ForecastRequest) -> ForecastBundle:
return _forecast_for("PV", req)
@app.post(ROUTES["price-forecast"], response_model=ForecastBundle)
def forecast_price(req: ForecastRequest) -> ForecastBundle:
return _forecast_for("PRICE", req)
@app.post(ROUTES["bid-optimization-milp"], response_model=BidOptimizationResult)
def optimize(req: BidOptimizationRequest) -> BidOptimizationResult:
return optimize_bid(req)
@app.post(ROUTES["report-generator"], response_model=SkillReport)
def report(req: ReportRequest) -> SkillReport:
return generate_report(req)

View File

@ -0,0 +1,188 @@
"""Day-ahead bid optimization MILP (docs/05 §2.2, docs/07 D-1 08:00).
Decision: for each of 96 intervals, an offer quantity q_t (MWh) and an offer
price. Model (v1):
maximise Σ_t (w_t − c − DEVIATION_PENALTY) · q_t
s.t. q_t ≤ k · cap_t · 0.25 · u_t (sellable share of certified capacity)
q_t ≥ min_block · u_t (market block size, u_t ∈ {0,1})
E_min ≤ Σ_t q_t ≤ E_max (position ledger cascade, P7)
with w_t = (1−λ)·P50_t + λ·P10_t the risk-weighted price (λ = risk_aversion)
and c the marginal cost of delivering flexibility. In a uniform-price market a
seller is paid the clearing price, so the *offer price* is the cost floor c
(price-taker when c = 0) and the forecast's job is allocation: which intervals
get the bounded daily energy. Risk aversion tilts allocation away from
intervals whose price band is wide (low P10). The revenue distribution
evaluates the bid against each forecast quantile path with the clearing rule
"cleared where offer ≤ price".
Solver: HiGHS via scipy.optimize.milp — open source, deterministic, adequate
for 96 binaries. Pyomo/commercial solvers (docs/08) are a swap behind the
same contract.
"""
from __future__ import annotations
import time
from decimal import Decimal
import numpy as np
import scipy
from scipy.optimize import Bounds, LinearConstraint, milp
from scipy.sparse import csr_matrix, hstack, identity
from vpp_contracts.bid_optimization_request import BidOptimizationRequest
from vpp_contracts.bid_optimization_result import BidOptimizationResult
from . import SKILL_VERSIONS
from .numeric import INTERVALS, SCALE_MONEY, SCALE_MWH, SCALE_PRICE, curve, dsum, quantize, to_array
SKILL = "bid-optimization-milp"
# OPEN-QUESTION A4: Hubei deviation-assessment rule (band + penalty price
# mechanism) is undecided. Until it is, the objective carries a zero penalty
# per offered MWh so the term exists and can be wired to config, but it does
# not shape the bid. See docs/open-questions.md → market.deviation.
DEVIATION_PENALTY_YUAN_PER_MWH = Decimal("0")
STATUS_BY_SCIPY = {0: "OPTIMAL", 1: "TIME_LIMIT", 2: "INFEASIBLE", 3: "INFEASIBLE", 4: "ERROR"}
TOL = 1e-6
def _revenue(offer: list[str], qty: list[str], path: list[str]) -> Decimal:
total = Decimal(0)
for o, q, p in zip(offer, qty, path, strict=True):
if Decimal(o) <= Decimal(p):
total += Decimal(p) * Decimal(q)
return total
def _fit_to_bounds(qty: np.ndarray, cap_mwh: np.ndarray, min_block: float, bounds) -> list[str]:
"""Quantize to SCALE_MWH without leaving [E_min, E_max].
Rounding 96 values independently can move the sum by up to 96·½ulp, enough
for the ledger to reject a bid the solver found feasible. Floor each value
(sum can only drop), then top up the deficit — if any — on intervals that
still have capacity headroom, all in exact decimal arithmetic.
"""
ulp = Decimal(1).scaleb(-SCALE_MWH)
vals = [(Decimal(repr(float(v))).quantize(ulp, rounding="ROUND_FLOOR")) for v in qty.tolist()]
vals = [abs(v) if v == 0 else v for v in vals]
caps = [Decimal(repr(float(c))).quantize(ulp, rounding="ROUND_FLOOR") for c in cap_mwh.tolist()]
e_min = Decimal(bounds.daily_energy_min_mwh)
total = sum(vals, Decimal(0))
if total < e_min and total > 0:
deficit = e_min - total
for i in sorted(range(len(vals)), key=lambda j: caps[j] - vals[j], reverse=True):
if vals[i] == 0 or deficit <= 0:
continue # never open a new interval below the block size
room = caps[i] - vals[i]
add = min(room, deficit).quantize(ulp, rounding="ROUND_CEILING")
add = min(add, room)
vals[i] += add
deficit -= add
# If the deficit could not be placed the solver bound was tight to
# machine precision; the ledger check will judge the residual.
return [format(v, "f") for v in vals]
def optimize_bid(req: BidOptimizationRequest) -> BidOptimizationResult:
if req.price_forecast.kind.value != "PRICE":
raise ValueError("price_forecast.kind must be PRICE")
if req.price_forecast.market_date != req.market_date or req.adjustable_capacity_mw.date != req.market_date:
raise ValueError("price forecast and capacity curve must be for market_date")
lam = float(Decimal(req.risk.risk_aversion))
k = float(Decimal(req.risk.commitment_buffer_k))
min_block = float(Decimal(req.risk.min_block_mwh))
e_min = float(Decimal(req.position_bounds.daily_energy_min_mwh))
e_max = float(Decimal(req.position_bounds.daily_energy_max_mwh))
if not 0.0 <= lam <= 1.0:
raise ValueError("risk_aversion must be in [0, 1]")
if not 0.0 < k <= 1.0:
raise ValueError("commitment_buffer_k must be in (0, 1]")
cost = float(Decimal(req.risk.marginal_cost_yuan_per_mwh))
if cost < 0:
raise ValueError("marginal_cost_yuan_per_mwh must be non-negative")
if min_block < 0 or e_min < 0 or e_max < e_min:
raise ValueError("invalid block size or position bounds")
q = req.price_forecast.quantiles
p10 = to_array(v.root for v in q.p10.values)
p50 = to_array(v.root for v in q.p50.values)
cap_mwh = k * to_array(v.root for v in req.adjustable_capacity_mw.values) * 0.25
if np.any(cap_mwh < 0):
raise ValueError("adjustable capacity must be non-negative")
w = (1.0 - lam) * p50 + lam * p10
penalty = float(DEVIATION_PENALTY_YUAN_PER_MWH)
n = INTERVALS
# Variables: x = [q_0..q_95, u_0..u_95]; minimise -(w - cost - penalty)·q
c = np.concatenate([-(w - cost - penalty), np.zeros(n)])
integrality = np.concatenate([np.zeros(n), np.ones(n)])
bounds = Bounds(np.zeros(2 * n), np.concatenate([cap_mwh, np.ones(n)]))
eye = identity(n, format="csr")
a_cap = hstack([eye, -csr_matrix(np.diag(cap_mwh))]) # q - cap·u ≤ 0
a_blk = hstack([eye, -min_block * eye]) # q - min_block·u ≥ 0
a_sum = csr_matrix(np.concatenate([np.ones(n), np.zeros(n)])[None, :])
constraints = [
LinearConstraint(a_cap, -np.inf, 0.0),
LinearConstraint(a_blk, 0.0, np.inf),
LinearConstraint(a_sum, e_min, e_max),
]
t0 = time.perf_counter()
res = milp(c, constraints=constraints, integrality=integrality, bounds=bounds)
wall_ms = int(round((time.perf_counter() - t0) * 1000))
status = STATUS_BY_SCIPY.get(int(res.status), "ERROR")
binding: list[str] = []
if res.x is None:
qty = np.zeros(n)
if e_min > cap_mwh.sum() + TOL:
binding.append("daily_energy_min exceeds sellable energy")
objective = None
else:
qty = np.clip(res.x[:n], 0.0, None)
qty[qty < TOL] = 0.0
total = qty.sum()
if abs(total - e_max) < 1e-4:
binding.append("daily_energy_max")
if abs(total - e_min) < 1e-4:
binding.append("daily_energy_min")
if np.any((cap_mwh > 0) & (np.abs(qty - cap_mwh) < 1e-6)):
binding.append("interval_capacity")
objective = quantize(-float(res.fun), SCALE_MONEY)
quantities = curve(_fit_to_bounds(qty, cap_mwh, min_block, req.position_bounds), req.market_date, SCALE_MWH)
offers = curve(np.full(n, cost), req.market_date, SCALE_PRICE)
qv, ov = quantities["values"], offers["values"]
p10s = [v.root for v in q.p10.values]
p50s = [v.root for v in q.p50.values]
p90s = [v.root for v in q.p90.values]
result = {
"market_date": req.market_date,
"prices_yuan_per_mwh": offers,
"quantities_mwh": quantities,
"daily_energy_mwh": quantize(dsum(qv), SCALE_MWH),
"expected_revenue_yuan": quantize(_revenue(ov, qv, p50s), SCALE_MONEY),
"revenue_distribution_yuan": {
"p10": quantize(_revenue(ov, qv, p10s), SCALE_MONEY),
"p50": quantize(_revenue(ov, qv, p50s), SCALE_MONEY),
"p90": quantize(_revenue(ov, qv, p90s), SCALE_MONEY),
},
"position_bounds": req.position_bounds.model_dump(),
"solver": {
"name": "highs",
"version": f"scipy-{scipy.__version__}",
"status": status,
"objective_value": objective,
"wall_time_ms": wall_ms,
},
"binding_constraints": binding,
"skill_version": SKILL_VERSIONS[SKILL],
}
return BidOptimizationResult.model_validate(result)

View File

@ -0,0 +1,125 @@
"""Forecast skills v1: load / PV / price with quantile intervals (docs/05 §2.1).
Baseline method, deliberately simple and fully deterministic:
* point forecast (P50): exponentially-weighted mean of recent history days of
the same day-type (weekday vs weekend), per 15-min interval;
* interval (P10/P90): empirical quantiles of *backtest residuals* — the same
method is re-run on each history day using only the days before it, and the
pooled (actual − predicted) residuals give the band. This is split-conformal
calibration in its simplest form, so coverage on held-out days tracks the
nominal 80% instead of being an optimistic guess (docs/12 L2: 区间校准单独考核).
Better models (statsforecast, gradient boosting) plug in behind the same
contract; the eval harness (packages/evals) is what decides whether they win.
Exogenous inputs are accepted by the contract but unused by v1 — recorded here
so nobody mistakes v1 for weather-aware.
"""
from __future__ import annotations
from datetime import date, datetime, timezone
from typing import Callable
import numpy as np
from vpp_contracts.forecast_bundle import ForecastBundle
from vpp_contracts.forecast_request import ForecastRequest
from . import SKILL_VERSIONS
from .numeric import INTERVALS, SCALE_MW, SCALE_PRICE, curve, to_array
MODEL_NAME = {"LOAD": "load-forecast", "PV": "pv-forecast", "PRICE": "price-forecast"}
UNIT_FOR_KIND = {"LOAD": "mw", "PV": "mw", "PRICE": "yuan_per_mwh"}
LOOKBACK_DAYS = 14
HALF_LIFE_DAYS = 3.0
MIN_BACKTEST_HISTORY = 3
NOMINAL_LOWER, NOMINAL_UPPER = 0.10, 0.90
def _is_weekend(d: str) -> bool:
return date.fromisoformat(d).weekday() >= 5
def _point_forecast(days: list[tuple[str, np.ndarray]], target: str) -> np.ndarray:
"""EWM over the most recent LOOKBACK_DAYS same-day-type days (fallback: all days)."""
same = [(d, v) for d, v in days if _is_weekend(d) == _is_weekend(target)]
pool = (same if len(same) >= 2 else days)[-LOOKBACK_DAYS:]
ages = np.array([(date.fromisoformat(target) - date.fromisoformat(d)).days for d, _ in pool], float)
w = np.power(0.5, ages / HALF_LIFE_DAYS)
stack = np.stack([v for _, v in pool])
return (w[:, None] * stack).sum(axis=0) / w.sum()
def _residual_quantiles(days: list[tuple[str, np.ndarray]], relative: bool) -> tuple[float, float]:
"""Pooled backtest residuals. Relative (actual/pred − 1) for quantities whose
spread scales with level (load, PV — a night-time zero must not shrink the
noon band); additive for prices, which can sit near or below zero."""
pairs: list[tuple[np.ndarray, np.ndarray]] = []
for i in range(MIN_BACKTEST_HISTORY, len(days)):
d, actual = days[i]
pairs.append((actual, _point_forecast(days[:i], d)))
if not pairs:
# Too little history to backtest: fall back to day-to-day differences.
pairs = [(days[i][1], days[i - 1][1]) for i in range(1, len(days))]
actual = np.concatenate([a for a, _ in pairs])
pred = np.concatenate([p for _, p in pairs])
if relative:
floor = 0.05 * max(float(pred.max()), 1e-9)
mask = pred > floor
pooled = actual[mask] / pred[mask] - 1.0 if mask.any() else np.zeros(1)
else:
pooled = actual - pred
lo, hi = np.quantile(pooled, [NOMINAL_LOWER, NOMINAL_UPPER])
return float(lo), float(hi)
def forecast(
req: ForecastRequest,
clock: Callable[[], datetime] = lambda: datetime.now(timezone.utc),
) -> ForecastBundle:
kind = str(req.kind.value)
if req.unit.value != UNIT_FOR_KIND[kind]:
raise ValueError(f"{kind} forecast unit must be {UNIT_FOR_KIND[kind]}, got {req.unit.value}")
days = [(c.date, to_array(v.root for v in c.values)) for c in req.history]
dates = [d for d, _ in days]
if dates != sorted(dates) or len(set(dates)) != len(dates):
raise ValueError("history must be ascending by date with no duplicates")
if dates[-1] >= req.market_date:
raise ValueError("history must be strictly before market_date")
p50 = _point_forecast(days, req.market_date)
relative = kind in ("LOAD", "PV")
lo, hi = _residual_quantiles(days, relative)
if relative:
p10, p90 = p50 * (1.0 + lo), p50 * (1.0 + hi)
else:
p10, p90 = p50 + lo, p50 + hi
if kind in ("LOAD", "PV"):
p10, p50, p90 = (np.clip(x, 0.0, None) for x in (p10, p50, p90))
if kind == "PV":
# Intervals that were zero on every history day (night) stay exactly zero.
night = np.all(np.stack([v for _, v in days]) == 0.0, axis=0)
for x in (p10, p50, p90):
x[night] = 0.0
scale = SCALE_PRICE if kind == "PRICE" else SCALE_MW
assert p50.shape == (INTERVALS,)
bundle = {
"id": f"fc-{kind.lower()}-{req.market_date}",
"kind": kind,
"market_date": req.market_date,
"unit": req.unit.value,
"quantiles": {
"p10": curve(p10, req.market_date, scale),
"p50": curve(p50, req.market_date, scale),
"p90": curve(p90, req.market_date, scale),
},
"model": {"name": MODEL_NAME[kind], "version": SKILL_VERSIONS[MODEL_NAME[kind]]},
"features_snapshot_ref": req.features_snapshot_ref,
"generated_at": clock().isoformat().replace("+00:00", "Z"),
}
return ForecastBundle.model_validate(bundle)

View File

@ -0,0 +1,55 @@
"""Decimal-string discipline (docs/11 §3.3) at the skill boundary.
Contracts carry money/energy/prices as fixed-point decimal *strings*. Inside a
skill we compute with numpy floats, then quantize back to strings with an
explicit scale — and any headline figure (revenue, energy) is re-derived with
``decimal.Decimal`` from the already-quantized curves, so what a report cites
is exactly reproducible from the curves it cites.
"""
from __future__ import annotations
from decimal import ROUND_HALF_EVEN, Decimal
from typing import Iterable
import numpy as np
INTERVALS = 96
INTERVAL_HOURS = Decimal("0.25")
SCALE_MW = 3
SCALE_MWH = 3
SCALE_PRICE = 2
SCALE_MONEY = 2
def to_array(values: Iterable[str]) -> np.ndarray:
return np.array([float(Decimal(v)) for v in values], dtype=float)
def quantize(x: float | Decimal, scale: int) -> str:
q = Decimal(1).scaleb(-scale)
d = x if isinstance(x, Decimal) else Decimal(repr(float(x)))
out = d.quantize(q, rounding=ROUND_HALF_EVEN)
if out == 0:
out = abs(out) # normalise "-0.000"
return format(out, "f")
def curve(values: np.ndarray | list[str], date: str, scale: int) -> dict:
if isinstance(values, np.ndarray):
vals = [quantize(float(v), scale) for v in values.tolist()]
else:
vals = list(values)
if len(vals) != INTERVALS:
raise ValueError(f"curve must have {INTERVALS} values, got {len(vals)}")
return {"interval_minutes": 15, "date": date, "values": vals}
def dsum(values: Iterable[str]) -> Decimal:
return sum((Decimal(v) for v in values), Decimal(0))
def dot(a: Iterable[str], b: Iterable[str]) -> Decimal:
"""Exact Σ a_i·b_i over decimal strings."""
return sum((Decimal(x) * Decimal(y) for x, y in zip(a, b, strict=True)), Decimal(0))

View File

@ -0,0 +1,124 @@
"""Report generator skill v1 (docs/05 §2.4).
Produces the *numeric skeleton* of a report: every figure is a
``{tool_call_id, path}`` reference into a recorded tool output, with the value
copied from that path — never typed, never computed here. The LLM's role
(docs/07 D-1 08:00 step 3) is prose around these references and arrives with
the runtime in M3; this skill is what makes "数字原样引用求解器输出" checkable.
"""
from __future__ import annotations
import hashlib
import json
import re
from datetime import datetime, timezone
from typing import Any, Callable
from vpp_contracts.report_request import ReportRequest
from vpp_contracts.skill_report import SkillReport
from . import SKILL_VERSIONS
SKILL = "report-generator"
DECIMAL_RE = re.compile(r"^-?\d+(\.\d+)?$")
UNIT_SUFFIXES = (
("_yuan_per_mwh", "yuan_per_mwh"),
("_mwh", "mwh"),
("_mw", "mw"),
("_yuan", "yuan"),
("_pct", "pct"),
("_ms", "ms"),
)
SECTION_TITLE = {
"DAY_AHEAD_BID_SUMMARY": "Day-ahead bid",
"FORECAST_EVAL": "Forecast evaluation",
"BID_BACKTEST": "Bid backtest",
}
def resolve(output: Any, path: str) -> Any:
"""Dereference a dotted path (list indices as integers) into a tool output."""
node = output
for part in path.split("."):
if isinstance(node, list):
node = node[int(part)]
elif isinstance(node, dict):
node = node[part]
else:
raise KeyError(path)
return node
def _unit_for(name: str, parent: str) -> str:
for suffix, unit in UNIT_SUFFIXES:
if name.endswith(suffix) or parent.endswith(suffix):
return unit
return "1"
def _numeric_leaves(node: Any, prefix: str = "", parent: str = "") -> list[tuple[str, str, str]]:
"""(path, value, unit) for every decimal-string scalar; curves/lists are skipped
(a 96-point curve is cited by snapshot ref, not inlined into a report)."""
out: list[tuple[str, str, str]] = []
if isinstance(node, dict):
for key, val in node.items():
path = f"{prefix}.{key}" if prefix else key
if isinstance(val, str) and DECIMAL_RE.match(val):
out.append((path, val, _unit_for(key, parent)))
elif isinstance(val, dict) and "values" not in val:
out.extend(_numeric_leaves(val, path, key))
return out
def generate_report(
req: ReportRequest,
clock: Callable[[], datetime] = lambda: datetime.now(timezone.utc),
) -> SkillReport:
kind = str(req.kind.value)
sections = []
for src in req.sources:
metrics = [
{"name": path, "value": value, "unit": unit, "ref": {"tool_call_id": src.tool_call_id, "path": path}}
for path, value, unit in _numeric_leaves(src.output)
]
notes = [f"source: {src.tool} v{src.version} (tool call {src.tool_call_id})"]
if isinstance(src.output, dict):
solver = src.output.get("solver")
if isinstance(solver, dict) and "status" in solver:
notes.append(f"solver status: {solver['status']}")
for key in ("binding_constraints",):
if isinstance(src.output.get(key), list) and src.output[key]:
notes.append(f"{key}: {', '.join(map(str, src.output[key]))}")
sections.append({"title": f"{SECTION_TITLE[kind]} · {src.tool}", "metrics": metrics, "notes": notes})
digest = hashlib.sha256(
json.dumps([s.model_dump() for s in req.sources], sort_keys=True, separators=(",", ":")).encode()
).hexdigest()[:12]
report = {
"id": f"rep-{kind.lower()}-{req.market_date}-{digest}",
"kind": kind,
"market_date": req.market_date,
"sections": sections,
"skill_version": SKILL_VERSIONS[SKILL],
"generated_at": clock().isoformat().replace("+00:00", "Z"),
}
return SkillReport.model_validate(report)
def verify_report(report: SkillReport, req: ReportRequest) -> list[str]:
"""Judge (docs/12 L1 数字一致性): every metric value must equal the referenced output."""
by_id = {s.tool_call_id: s.output for s in req.sources}
problems = []
for section in report.sections:
for m in section.metrics:
try:
actual = resolve(by_id[m.ref.tool_call_id], m.ref.path)
except (KeyError, IndexError, ValueError):
problems.append(f"{m.name}: unresolvable ref {m.ref.tool_call_id}#{m.ref.path}")
continue
if str(actual) != m.value:
problems.append(f"{m.name}: value {m.value} != referenced {actual}")
return problems

View File

@ -0,0 +1,114 @@
"""Deterministic SYNTHETIC dataset for the M2 eval baseline.
There is no historical Hubei data in this repository yet (docs/12 §2 历史重放集
comes from the event log once the system runs; partner data is OPEN-QUESTION
D-class). Until real data lands, the eval harness needs *something* with
realistic structure — daily load shape, PV bell with cloud days, price with an
evening peak and occasional spikes — so metrics, calibration tests and the
reproducibility check exercise real code paths.
Every number here is invented for shape only. Baselines computed on it are
baselines of the *harness*, not of the business KPI (预测误差 ≤ 8% is judged
on real data, docs/12 §4). The dataset file says so in its ``meta``.
"""
from __future__ import annotations
import argparse
import json
from datetime import date, timedelta
from pathlib import Path
import numpy as np
from .numeric import INTERVALS, SCALE_MW, SCALE_PRICE, quantize
GENERATOR = "vpp_skills.synthetic"
GENERATOR_VERSION = "0.1.0"
def _profiles() -> tuple[np.ndarray, np.ndarray, np.ndarray]:
t = np.arange(INTERVALS) / 4.0 # hours
load = 0.55 + 0.25 * np.exp(-((t - 11) ** 2) / 8) + 0.35 * np.exp(-((t - 19.5) ** 2) / 6) - 0.15 * np.exp(-((t - 3) ** 2) / 10)
pv = np.clip(np.sin(np.pi * (t - 6.5) / 12.5), 0.0, None) ** 1.4
pv[(t < 6.5) | (t > 19.0)] = 0.0
price = 0.7 + 0.2 * np.exp(-((t - 11) ** 2) / 6) + 0.5 * np.exp(-((t - 19) ** 2) / 4) - 0.25 * np.exp(-((t - 3.5) ** 2) / 12)
return load, pv, price
def generate_dataset(
seed: int = 20260301,
start: str = "2026-01-01",
n_days: int = 120,
peak_load_mw: float = 80.0,
pv_capacity_mw: float = 30.0,
adjustable_capacity_mw: float = 24.0,
base_price: float = 400.0,
) -> dict:
rng = np.random.default_rng(seed)
load_p, pv_p, price_p = _profiles()
d0 = date.fromisoformat(start)
days = []
load_ar = 0.0
cloud = 0.8
for i in range(n_days):
d = d0 + timedelta(days=i)
weekend = d.weekday() >= 5
load_ar = 0.7 * load_ar + rng.normal(0, 0.04)
day_factor = (0.85 if weekend else 1.0) * (1.0 + load_ar)
load = peak_load_mw * load_p * day_factor * (1.0 + rng.normal(0, 0.02, INTERVALS))
load = np.clip(load, 0.0, None)
cloud = float(np.clip(0.6 * cloud + 0.4 * rng.beta(4, 1.5), 0.05, 1.0))
pv = pv_capacity_mw * pv_p * cloud * (1.0 + rng.normal(0, 0.05, INTERVALS))
pv = np.clip(pv, 0.0, None)
pv[pv_p == 0.0] = 0.0
net = (load - pv) / peak_load_mw
price = base_price * (price_p + 0.35 * (net - net.mean())) * (1.0 + rng.normal(0, 0.03, INTERVALS))
if rng.random() < 0.06: # spike day: evening scarcity
price[72:88] *= rng.uniform(1.6, 2.4)
price = np.clip(price, 0.0, None)
days.append(
{
"date": d.isoformat(),
"weekend": weekend,
"load_mw": [quantize(v, SCALE_MW) for v in load],
"pv_mw": [quantize(v, SCALE_MW) for v in pv],
"price_yuan_per_mwh": [quantize(v, SCALE_PRICE) for v in price],
"adjustable_capacity_mw": [quantize(adjustable_capacity_mw, SCALE_MW)] * INTERVALS,
}
)
return {
"meta": {
"name": f"synthetic-hubei-v0-{start}-{n_days}d",
"synthetic": True,
"note": "SYNTHETIC shape-only placeholder. Replace with historical Hubei replay data; "
"baselines on this set measure the harness, not the business KPI.",
"generator": GENERATOR,
"generator_version": GENERATOR_VERSION,
"seed": seed,
"interval_minutes": 15,
"peak_load_mw": quantize(peak_load_mw, SCALE_MW),
"pv_capacity_mw": quantize(pv_capacity_mw, SCALE_MW),
"adjustable_capacity_mw": quantize(adjustable_capacity_mw, SCALE_MW),
},
"days": days,
}
def main() -> None:
ap = argparse.ArgumentParser(description="write the synthetic eval dataset")
ap.add_argument("--out", required=True, type=Path)
ap.add_argument("--seed", type=int, default=20260301)
ap.add_argument("--days", type=int, default=120)
args = ap.parse_args()
ds = generate_dataset(seed=args.seed, n_days=args.days)
args.out.parent.mkdir(parents=True, exist_ok=True)
args.out.write_text(json.dumps(ds, indent=1) + "\n")
print(f"wrote {args.out} ({len(ds['days'])} days)")
if __name__ == "__main__":
main()