vpp-ai-platform/contracts/fixtures/kpi_report/window-30d.json

52 lines
1.2 KiB
JSON
Raw Normal View History

M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records bids without submitting them, simulates the market answer from the actual day-ahead clearing price, dispatches to the simulation gateway, runs the D+1 review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord) with a lineage-completeness audit. KpiReport regenerated after every day per docs/12 §4 definitions (C1/C2 placeholders as named config). Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1 envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids), L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED), L4 AI-off (templates run, no Proposal created); abnormal-day protocol on EXTREME situations; per-level authority (B8 placeholder); auditable drill. Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints and insight cards, replay CLI (npm run shadow). Ledger, time series and streak counters are file-backed so a multi-week run survives restarts. Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW receipt channel; contracts, fixtures and pydantic models regenerated. Services: L2 metrics moved from evals so the shadow run and the harness share one implementation. Fixes: envelope/permit validity compared ISO timestamps as strings ('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only, metrics unchanged) — rewritten from the live service. Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/ recovery, KPI definitions as implemented); README and CLAUDE.md status. Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN recommendation produced but not acted on; restart durability; drill; L2/L3/L4 and abnormal-day paths; API. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00
{
"id": "kpi-2026-03-15",
"window": {
"from": "2026-02-14",
"to": "2026-03-15",
"days": 30
},
"kpis": [
{
"id": "FORECAST_LOAD_MAPE",
"value": "0.052",
"unit": "1",
"target": "0.08",
"comparator": "LTE",
"samples": 30,
"status": "MEET",
"definition": "aggregate day-ahead 96-interval load MAPE, rolling window mean"
},
{
"id": "CROSS_REGION_MATCH",
"value": null,
"unit": "1",
"target": "0.85",
"comparator": "GTE",
"samples": 0,
"status": "NOT_APPLICABLE",
"definition": "federation commitments vs delivery confirmations (phase 2)"
}
],
"comparison": {
"shadow_yuan": "15156000.00",
"human_yuan": "13337280.00",
"hindsight_yuan": "16200000.00",
"naive_yuan": "15913800.00",
"capture_ratio": "0.935556",
"uplift_vs_naive": "0.952383",
"days_with_human_baseline": 30
},
"shadow": {
"days": 30,
"complete_days": 30,
"consecutive_complete_days": 30,
"first_date": "2026-02-14",
"last_date": "2026-03-15",
"released_days": 27,
"pending_days": 3,
"widen_recommendations": 1,
"breaker_trips": 0
},
"generated_at": "2026-03-16T03:00:00Z"
}