vpp-ai-platform/.github/workflows/ci.yml
Thomas Bayes 8796faca63 M2: skill contracts, Python skill service, L2 eval harness with baseline
- packages/domain: ForecastRequest, BidOptimizationRequest/Result,
  ReportRequest, SkillReport (+ golden and invalid fixtures, exported to
  contracts/ and regenerated as pydantic models).
- skills-py/vpp_skills: FastAPI service with versioned registry; load/PV/
  price forecasts (same-day-type EWM point forecast, conformal residual
  quantiles — coverage test as acceptance gate); bid-optimization MILP on
  HiGHS (binary block participation, hard ledger energy bounds, exact
  Decimal fit of the rounded curve inside the bounds, revenue distribution
  over quantile paths); report generator whose every figure is a
  {tool_call_id, path} reference, with a verifier. 48 tests incl. hypothesis
  property test that bids respect ledger constraints.
- packages/services: LedgerService.dayAheadBounds (the P7 cascade band
  handed to the optimizer); Decimal resolved once for CJS/ESM interop.
- packages/evals: L2 metrics (MAPE, nRMSE, coverage, direction accuracy,
  naive/hindsight revenue baselines), HTTP skill client, rolling-origin
  harness that pushes each bid through the real ledger, CLI with
  --check/--write-baseline; committed baseline on the SYNTHETIC dataset
  (no historical Hubei data yet — baselines measure the harness, not KPI).
- CI: evals job boots the skill service and fails on baseline digest drift.
- docs/open-questions: A6 (flexibility marginal cost = offer floor); A4/B6
  wired as placeholders. README/CLAUDE.md status → M2 done, M3 next.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:29:08 -04:00

74 lines
2.5 KiB
YAML

name: ci
on:
push:
branches: [master]
pull_request:
jobs:
# Domain schemas + deterministic services (TS side of the contract pipeline).
typescript:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 24
cache: npm
- run: npm ci
- run: npm run typecheck
# ADR-0002: contracts/ is generated from packages/domain and committed.
# A dirty tree here means someone changed a zod schema without
# re-exporting, or hand-edited a generated file.
- run: npm run export:schemas && npm run make:fixtures
- run: git diff --exit-code -- contracts/
- run: npm test
# Generated pydantic models must accept/reject the same golden fixtures the
# TS side does (docs/11 §3.2 dual-side contract test).
python:
runs-on: ubuntu-latest
defaults:
run:
working-directory: skills-py
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version-file: skills-py/.python-version
cache: pip
cache-dependency-path: skills-py/requirements.txt
- run: pip install -r requirements.txt
# Generated models are committed but never hand-edited: regenerate and
# require a clean tree.
- run: bash scripts/generate_models.sh
- run: git diff --exit-code -- vpp_contracts/
- run: python -m pytest -q
# L2 eval harness end-to-end (docs/12): TS harness ↔ live Python skill
# service on the committed synthetic dataset. The run must reproduce the
# committed baseline digest exactly — a skill change that moves a metric
# must come with a reviewed baseline update (docs/12 §3 gate).
evals:
runs-on: ubuntu-latest
needs: [typescript, python]
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 24
cache: npm
- uses: actions/setup-python@v5
with:
python-version-file: skills-py/.python-version
cache: pip
cache-dependency-path: skills-py/requirements.txt
- run: npm ci
- run: pip install -r skills-py/requirements.txt
- name: start skill service
working-directory: skills-py
run: |
python -m uvicorn vpp_skills.app:app --port 8000 --log-level warning &
for i in $(seq 1 40); do curl -sf http://127.0.0.1:8000/health && break; sleep 0.5; done
- run: npm run eval -w @vpp/evals -- --check