vpp-ai-platform/skills-py/tests/conftest.py
Thomas Bayes 8796faca63 M2: skill contracts, Python skill service, L2 eval harness with baseline
- packages/domain: ForecastRequest, BidOptimizationRequest/Result,
  ReportRequest, SkillReport (+ golden and invalid fixtures, exported to
  contracts/ and regenerated as pydantic models).
- skills-py/vpp_skills: FastAPI service with versioned registry; load/PV/
  price forecasts (same-day-type EWM point forecast, conformal residual
  quantiles — coverage test as acceptance gate); bid-optimization MILP on
  HiGHS (binary block participation, hard ledger energy bounds, exact
  Decimal fit of the rounded curve inside the bounds, revenue distribution
  over quantile paths); report generator whose every figure is a
  {tool_call_id, path} reference, with a verifier. 48 tests incl. hypothesis
  property test that bids respect ledger constraints.
- packages/services: LedgerService.dayAheadBounds (the P7 cascade band
  handed to the optimizer); Decimal resolved once for CJS/ESM interop.
- packages/evals: L2 metrics (MAPE, nRMSE, coverage, direction accuracy,
  naive/hindsight revenue baselines), HTTP skill client, rolling-origin
  harness that pushes each bid through the real ledger, CLI with
  --check/--write-baseline; committed baseline on the SYNTHETIC dataset
  (no historical Hubei data yet — baselines measure the harness, not KPI).
- CI: evals job boots the skill service and fails on baseline digest drift.
- docs/open-questions: A6 (flexibility marginal cost = offer floor); A4/B6
  wired as placeholders. README/CLAUDE.md status → M2 done, M3 next.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:29:08 -04:00

37 lines
1019 B
Python

from __future__ import annotations
import sys
from pathlib import Path
import pytest
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from vpp_skills.synthetic import generate_dataset # noqa: E402
REF_A = "a" * 64
def curve_of(values: list[str], date: str) -> dict:
return {"interval_minutes": 15, "date": date, "values": values}
@pytest.fixture(scope="session")
def dataset() -> dict:
return generate_dataset(seed=7, n_days=90)
def forecast_request(dataset: dict, kind: str, target_idx: int, window: int = 28) -> dict:
field = {"LOAD": "load_mw", "PV": "pv_mw", "PRICE": "price_yuan_per_mwh"}[kind]
unit = "yuan_per_mwh" if kind == "PRICE" else "mw"
days = dataset["days"]
hist = days[max(0, target_idx - window) : target_idx]
return {
"kind": kind,
"market_date": days[target_idx]["date"],
"unit": unit,
"history": [curve_of(d[field], d["date"]) for d in hist],
"exogenous": {},
"features_snapshot_ref": REF_A,
}