M2: skill contracts, Python skill service, L2 eval harness with baseline
- packages/domain: ForecastRequest, BidOptimizationRequest/Result,
ReportRequest, SkillReport (+ golden and invalid fixtures, exported to
contracts/ and regenerated as pydantic models).
- skills-py/vpp_skills: FastAPI service with versioned registry; load/PV/
price forecasts (same-day-type EWM point forecast, conformal residual
quantiles — coverage test as acceptance gate); bid-optimization MILP on
HiGHS (binary block participation, hard ledger energy bounds, exact
Decimal fit of the rounded curve inside the bounds, revenue distribution
over quantile paths); report generator whose every figure is a
{tool_call_id, path} reference, with a verifier. 48 tests incl. hypothesis
property test that bids respect ledger constraints.
- packages/services: LedgerService.dayAheadBounds (the P7 cascade band
handed to the optimizer); Decimal resolved once for CJS/ESM interop.
- packages/evals: L2 metrics (MAPE, nRMSE, coverage, direction accuracy,
naive/hindsight revenue baselines), HTTP skill client, rolling-origin
harness that pushes each bid through the real ledger, CLI with
--check/--write-baseline; committed baseline on the SYNTHETIC dataset
(no historical Hubei data yet — baselines measure the harness, not KPI).
- CI: evals job boots the skill service and fails on baseline digest drift.
- docs/open-questions: A6 (flexibility marginal cost = offer floor); A4/B6
wired as placeholders. README/CLAUDE.md status → M2 done, M3 next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:29:08 -04:00
|
|
|
"""HTTP/JSON skill service (docs/11 §3.4, docs/09 §4).
|
|
|
|
|
|
|
|
|
|
One process, one route per skill, generated pydantic contracts on both request
|
|
|
|
|
and response. The TS runtime's registerSkill wrapper is the only intended
|
|
|
|
|
caller; it snapshots request/response for lineage. Run locally with
|
|
|
|
|
``uvicorn vpp_skills.app:app``.
|
|
|
|
|
"""
|
|
|
|
|
|
|
|
|
|
from __future__ import annotations
|
|
|
|
|
|
|
|
|
|
from fastapi import FastAPI, HTTPException
|
|
|
|
|
from fastapi.responses import JSONResponse
|
|
|
|
|
|
|
|
|
|
from vpp_contracts.bid_optimization_request import BidOptimizationRequest
|
|
|
|
|
from vpp_contracts.bid_optimization_result import BidOptimizationResult
|
M4: resource agent, envelopes live, review loop, insight cards
- packages/domain: AwardNotice, DISPATCH_PLAN proposal payload, ExecutionReport,
MeteringRecord, potential-assessment and dispatch-optimization contracts,
ReviewFinding (+ typed writebacks), SemanticMemoryEntry,
EnvelopeChangeRequest, InsightCard — exported with fixtures on both sides.
- skills-py: potential-assessment (certified × rolling fulfilment, evidence
days) and dispatch-optimization (per-interval LP on HiGHS, shortfall
reported) skills + routes + tests.
- packages/services: dispatch rules in the policy pack (over-allocation,
award anchor, lineage integrity for allocations); PowerBalanceSimulator;
SimulationGateway (permit-only, idempotent, seeded execute → ExecutionReports);
envelope deviation-streak suspension + apply(); ReviewService (attribution,
reliability EWMA writeback, semantic memory, envelope recommendations as
change requests); dispatch assembler; skill client methods.
- packages/runtime: resource agent; award-decomposition, review and
envelope-review workflows; lifecycle selects simulator/gateway by proposal
type; trigger hooks for awards, execution reports, metering; decide()
resumes either lifecycle or envelope-review runs; insight cards API.
- Tests: docs/07 D-1 16:00 and D+1 end to end; reliability score 0.9 → 0.880
and the next assessment de-rates capacity; envelope suspension on a seeded
3-day streak; WIDEN request applied only by a human. 184 TS + 80 Python.
- docs/open-questions: B10 (reliability/potential parameters). README and
CLAUDE.md status → M4 done, M5 next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 19:31:12 -04:00
|
|
|
from vpp_contracts.dispatch_optimization_request import DispatchOptimizationRequest
|
|
|
|
|
from vpp_contracts.dispatch_optimization_result import DispatchOptimizationResult
|
|
|
|
|
from vpp_contracts.potential_assessment_request import PotentialAssessmentRequest
|
|
|
|
|
from vpp_contracts.potential_assessment_result import PotentialAssessmentResult
|
M2: skill contracts, Python skill service, L2 eval harness with baseline
- packages/domain: ForecastRequest, BidOptimizationRequest/Result,
ReportRequest, SkillReport (+ golden and invalid fixtures, exported to
contracts/ and regenerated as pydantic models).
- skills-py/vpp_skills: FastAPI service with versioned registry; load/PV/
price forecasts (same-day-type EWM point forecast, conformal residual
quantiles — coverage test as acceptance gate); bid-optimization MILP on
HiGHS (binary block participation, hard ledger energy bounds, exact
Decimal fit of the rounded curve inside the bounds, revenue distribution
over quantile paths); report generator whose every figure is a
{tool_call_id, path} reference, with a verifier. 48 tests incl. hypothesis
property test that bids respect ledger constraints.
- packages/services: LedgerService.dayAheadBounds (the P7 cascade band
handed to the optimizer); Decimal resolved once for CJS/ESM interop.
- packages/evals: L2 metrics (MAPE, nRMSE, coverage, direction accuracy,
naive/hindsight revenue baselines), HTTP skill client, rolling-origin
harness that pushes each bid through the real ledger, CLI with
--check/--write-baseline; committed baseline on the SYNTHETIC dataset
(no historical Hubei data yet — baselines measure the harness, not KPI).
- CI: evals job boots the skill service and fails on baseline digest drift.
- docs/open-questions: A6 (flexibility marginal cost = offer floor); A4/B6
wired as placeholders. README/CLAUDE.md status → M2 done, M3 next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:29:08 -04:00
|
|
|
from vpp_contracts.forecast_bundle import ForecastBundle
|
|
|
|
|
from vpp_contracts.forecast_request import ForecastRequest
|
|
|
|
|
from vpp_contracts.report_request import ReportRequest
|
|
|
|
|
from vpp_contracts.skill_report import SkillReport
|
|
|
|
|
|
|
|
|
|
from . import SKILL_VERSIONS
|
|
|
|
|
from .bid_milp import optimize_bid
|
M4: resource agent, envelopes live, review loop, insight cards
- packages/domain: AwardNotice, DISPATCH_PLAN proposal payload, ExecutionReport,
MeteringRecord, potential-assessment and dispatch-optimization contracts,
ReviewFinding (+ typed writebacks), SemanticMemoryEntry,
EnvelopeChangeRequest, InsightCard — exported with fixtures on both sides.
- skills-py: potential-assessment (certified × rolling fulfilment, evidence
days) and dispatch-optimization (per-interval LP on HiGHS, shortfall
reported) skills + routes + tests.
- packages/services: dispatch rules in the policy pack (over-allocation,
award anchor, lineage integrity for allocations); PowerBalanceSimulator;
SimulationGateway (permit-only, idempotent, seeded execute → ExecutionReports);
envelope deviation-streak suspension + apply(); ReviewService (attribution,
reliability EWMA writeback, semantic memory, envelope recommendations as
change requests); dispatch assembler; skill client methods.
- packages/runtime: resource agent; award-decomposition, review and
envelope-review workflows; lifecycle selects simulator/gateway by proposal
type; trigger hooks for awards, execution reports, metering; decide()
resumes either lifecycle or envelope-review runs; insight cards API.
- Tests: docs/07 D-1 16:00 and D+1 end to end; reliability score 0.9 → 0.880
and the next assessment de-rates capacity; envelope suspension on a seeded
3-day streak; WIDEN request applied only by a human. 184 TS + 80 Python.
- docs/open-questions: B10 (reliability/potential parameters). README and
CLAUDE.md status → M4 done, M5 next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 19:31:12 -04:00
|
|
|
from .dispatch_opt import optimize_dispatch
|
|
|
|
|
from .potential import assess_potential
|
M2: skill contracts, Python skill service, L2 eval harness with baseline
- packages/domain: ForecastRequest, BidOptimizationRequest/Result,
ReportRequest, SkillReport (+ golden and invalid fixtures, exported to
contracts/ and regenerated as pydantic models).
- skills-py/vpp_skills: FastAPI service with versioned registry; load/PV/
price forecasts (same-day-type EWM point forecast, conformal residual
quantiles — coverage test as acceptance gate); bid-optimization MILP on
HiGHS (binary block participation, hard ledger energy bounds, exact
Decimal fit of the rounded curve inside the bounds, revenue distribution
over quantile paths); report generator whose every figure is a
{tool_call_id, path} reference, with a verifier. 48 tests incl. hypothesis
property test that bids respect ledger constraints.
- packages/services: LedgerService.dayAheadBounds (the P7 cascade band
handed to the optimizer); Decimal resolved once for CJS/ESM interop.
- packages/evals: L2 metrics (MAPE, nRMSE, coverage, direction accuracy,
naive/hindsight revenue baselines), HTTP skill client, rolling-origin
harness that pushes each bid through the real ledger, CLI with
--check/--write-baseline; committed baseline on the SYNTHETIC dataset
(no historical Hubei data yet — baselines measure the harness, not KPI).
- CI: evals job boots the skill service and fails on baseline digest drift.
- docs/open-questions: A6 (flexibility marginal cost = offer floor); A4/B6
wired as placeholders. README/CLAUDE.md status → M2 done, M3 next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:29:08 -04:00
|
|
|
from .forecast import forecast
|
|
|
|
|
from .report import generate_report
|
|
|
|
|
|
|
|
|
|
app = FastAPI(title="vpp-skills", version="0.1.0")
|
|
|
|
|
|
|
|
|
|
ROUTES = {
|
|
|
|
|
"load-forecast": "/v1/forecast/load",
|
|
|
|
|
"pv-forecast": "/v1/forecast/pv",
|
|
|
|
|
"price-forecast": "/v1/forecast/price",
|
|
|
|
|
"bid-optimization-milp": "/v1/optimize/bid",
|
|
|
|
|
"report-generator": "/v1/report",
|
M4: resource agent, envelopes live, review loop, insight cards
- packages/domain: AwardNotice, DISPATCH_PLAN proposal payload, ExecutionReport,
MeteringRecord, potential-assessment and dispatch-optimization contracts,
ReviewFinding (+ typed writebacks), SemanticMemoryEntry,
EnvelopeChangeRequest, InsightCard — exported with fixtures on both sides.
- skills-py: potential-assessment (certified × rolling fulfilment, evidence
days) and dispatch-optimization (per-interval LP on HiGHS, shortfall
reported) skills + routes + tests.
- packages/services: dispatch rules in the policy pack (over-allocation,
award anchor, lineage integrity for allocations); PowerBalanceSimulator;
SimulationGateway (permit-only, idempotent, seeded execute → ExecutionReports);
envelope deviation-streak suspension + apply(); ReviewService (attribution,
reliability EWMA writeback, semantic memory, envelope recommendations as
change requests); dispatch assembler; skill client methods.
- packages/runtime: resource agent; award-decomposition, review and
envelope-review workflows; lifecycle selects simulator/gateway by proposal
type; trigger hooks for awards, execution reports, metering; decide()
resumes either lifecycle or envelope-review runs; insight cards API.
- Tests: docs/07 D-1 16:00 and D+1 end to end; reliability score 0.9 → 0.880
and the next assessment de-rates capacity; envelope suspension on a seeded
3-day streak; WIDEN request applied only by a human. 184 TS + 80 Python.
- docs/open-questions: B10 (reliability/potential parameters). README and
CLAUDE.md status → M4 done, M5 next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 19:31:12 -04:00
|
|
|
"potential-assessment": "/v1/assess/potential",
|
|
|
|
|
"dispatch-optimization": "/v1/optimize/dispatch",
|
M2: skill contracts, Python skill service, L2 eval harness with baseline
- packages/domain: ForecastRequest, BidOptimizationRequest/Result,
ReportRequest, SkillReport (+ golden and invalid fixtures, exported to
contracts/ and regenerated as pydantic models).
- skills-py/vpp_skills: FastAPI service with versioned registry; load/PV/
price forecasts (same-day-type EWM point forecast, conformal residual
quantiles — coverage test as acceptance gate); bid-optimization MILP on
HiGHS (binary block participation, hard ledger energy bounds, exact
Decimal fit of the rounded curve inside the bounds, revenue distribution
over quantile paths); report generator whose every figure is a
{tool_call_id, path} reference, with a verifier. 48 tests incl. hypothesis
property test that bids respect ledger constraints.
- packages/services: LedgerService.dayAheadBounds (the P7 cascade band
handed to the optimizer); Decimal resolved once for CJS/ESM interop.
- packages/evals: L2 metrics (MAPE, nRMSE, coverage, direction accuracy,
naive/hindsight revenue baselines), HTTP skill client, rolling-origin
harness that pushes each bid through the real ledger, CLI with
--check/--write-baseline; committed baseline on the SYNTHETIC dataset
(no historical Hubei data yet — baselines measure the harness, not KPI).
- CI: evals job boots the skill service and fails on baseline digest drift.
- docs/open-questions: A6 (flexibility marginal cost = offer floor); A4/B6
wired as placeholders. README/CLAUDE.md status → M2 done, M3 next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:29:08 -04:00
|
|
|
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@app.exception_handler(ValueError)
|
|
|
|
|
async def _value_error(_, exc: ValueError) -> JSONResponse:
|
|
|
|
|
return JSONResponse(status_code=422, content={"detail": str(exc)})
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@app.get("/health")
|
|
|
|
|
def health() -> dict:
|
|
|
|
|
return {"status": "ok"}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@app.get("/v1/skills")
|
|
|
|
|
def skills() -> list[dict]:
|
|
|
|
|
"""Registry view: what the TS tool registry pins versions against."""
|
|
|
|
|
return [{"id": sid, "version": ver, "endpoint": ROUTES[sid]} for sid, ver in SKILL_VERSIONS.items()]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _forecast_for(kind: str, req: ForecastRequest) -> ForecastBundle:
|
|
|
|
|
if req.kind.value != kind:
|
|
|
|
|
raise HTTPException(status_code=422, detail=f"this endpoint serves kind={kind}, got {req.kind.value}")
|
|
|
|
|
return forecast(req)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@app.post(ROUTES["load-forecast"], response_model=ForecastBundle)
|
|
|
|
|
def forecast_load(req: ForecastRequest) -> ForecastBundle:
|
|
|
|
|
return _forecast_for("LOAD", req)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@app.post(ROUTES["pv-forecast"], response_model=ForecastBundle)
|
|
|
|
|
def forecast_pv(req: ForecastRequest) -> ForecastBundle:
|
|
|
|
|
return _forecast_for("PV", req)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@app.post(ROUTES["price-forecast"], response_model=ForecastBundle)
|
|
|
|
|
def forecast_price(req: ForecastRequest) -> ForecastBundle:
|
|
|
|
|
return _forecast_for("PRICE", req)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@app.post(ROUTES["bid-optimization-milp"], response_model=BidOptimizationResult)
|
|
|
|
|
def optimize(req: BidOptimizationRequest) -> BidOptimizationResult:
|
|
|
|
|
return optimize_bid(req)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@app.post(ROUTES["report-generator"], response_model=SkillReport)
|
|
|
|
|
def report(req: ReportRequest) -> SkillReport:
|
|
|
|
|
return generate_report(req)
|
M4: resource agent, envelopes live, review loop, insight cards
- packages/domain: AwardNotice, DISPATCH_PLAN proposal payload, ExecutionReport,
MeteringRecord, potential-assessment and dispatch-optimization contracts,
ReviewFinding (+ typed writebacks), SemanticMemoryEntry,
EnvelopeChangeRequest, InsightCard — exported with fixtures on both sides.
- skills-py: potential-assessment (certified × rolling fulfilment, evidence
days) and dispatch-optimization (per-interval LP on HiGHS, shortfall
reported) skills + routes + tests.
- packages/services: dispatch rules in the policy pack (over-allocation,
award anchor, lineage integrity for allocations); PowerBalanceSimulator;
SimulationGateway (permit-only, idempotent, seeded execute → ExecutionReports);
envelope deviation-streak suspension + apply(); ReviewService (attribution,
reliability EWMA writeback, semantic memory, envelope recommendations as
change requests); dispatch assembler; skill client methods.
- packages/runtime: resource agent; award-decomposition, review and
envelope-review workflows; lifecycle selects simulator/gateway by proposal
type; trigger hooks for awards, execution reports, metering; decide()
resumes either lifecycle or envelope-review runs; insight cards API.
- Tests: docs/07 D-1 16:00 and D+1 end to end; reliability score 0.9 → 0.880
and the next assessment de-rates capacity; envelope suspension on a seeded
3-day streak; WIDEN request applied only by a human. 184 TS + 80 Python.
- docs/open-questions: B10 (reliability/potential parameters). README and
CLAUDE.md status → M4 done, M5 next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 19:31:12 -04:00
|
|
|
|
|
|
|
|
|
|
|
|
|
@app.post(ROUTES["potential-assessment"], response_model=PotentialAssessmentResult)
|
|
|
|
|
def potential(req: PotentialAssessmentRequest) -> PotentialAssessmentResult:
|
|
|
|
|
return assess_potential(req)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@app.post(ROUTES["dispatch-optimization"], response_model=DispatchOptimizationResult)
|
|
|
|
|
def dispatch(req: DispatchOptimizationRequest) -> DispatchOptimizationResult:
|
|
|
|
|
return optimize_dispatch(req)
|