vpp-ai-platform/skills-py/tests/conftest.py

37 lines
1019 B
Python
Raw Permalink Normal View History

M2: skill contracts, Python skill service, L2 eval harness with baseline - packages/domain: ForecastRequest, BidOptimizationRequest/Result, ReportRequest, SkillReport (+ golden and invalid fixtures, exported to contracts/ and regenerated as pydantic models). - skills-py/vpp_skills: FastAPI service with versioned registry; load/PV/ price forecasts (same-day-type EWM point forecast, conformal residual quantiles — coverage test as acceptance gate); bid-optimization MILP on HiGHS (binary block participation, hard ledger energy bounds, exact Decimal fit of the rounded curve inside the bounds, revenue distribution over quantile paths); report generator whose every figure is a {tool_call_id, path} reference, with a verifier. 48 tests incl. hypothesis property test that bids respect ledger constraints. - packages/services: LedgerService.dayAheadBounds (the P7 cascade band handed to the optimizer); Decimal resolved once for CJS/ESM interop. - packages/evals: L2 metrics (MAPE, nRMSE, coverage, direction accuracy, naive/hindsight revenue baselines), HTTP skill client, rolling-origin harness that pushes each bid through the real ledger, CLI with --check/--write-baseline; committed baseline on the SYNTHETIC dataset (no historical Hubei data yet — baselines measure the harness, not KPI). - CI: evals job boots the skill service and fails on baseline digest drift. - docs/open-questions: A6 (flexibility marginal cost = offer floor); A4/B6 wired as placeholders. README/CLAUDE.md status → M2 done, M3 next. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:29:08 -04:00
from __future__ import annotations
import sys
from pathlib import Path
import pytest
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from vpp_skills.synthetic import generate_dataset # noqa: E402
REF_A = "a" * 64
def curve_of(values: list[str], date: str) -> dict:
return {"interval_minutes": 15, "date": date, "values": values}
@pytest.fixture(scope="session")
def dataset() -> dict:
return generate_dataset(seed=7, n_days=90)
def forecast_request(dataset: dict, kind: str, target_idx: int, window: int = 28) -> dict:
field = {"LOAD": "load_mw", "PV": "pv_mw", "PRICE": "price_yuan_per_mwh"}[kind]
unit = "yuan_per_mwh" if kind == "PRICE" else "mw"
days = dataset["days"]
hist = days[max(0, target_idx - window) : target_idx]
return {
"kind": kind,
"market_date": days[target_idx]["date"],
"unit": unit,
"history": [curve_of(d[field], d["date"]) for d in hist],
"exogenous": {},
"features_snapshot_ref": REF_A,
}