vpp-ai-platform/contracts/schema/insight_card.json

139 lines
3.1 KiB
JSON
Raw Normal View History

M4: resource agent, envelopes live, review loop, insight cards - packages/domain: AwardNotice, DISPATCH_PLAN proposal payload, ExecutionReport, MeteringRecord, potential-assessment and dispatch-optimization contracts, ReviewFinding (+ typed writebacks), SemanticMemoryEntry, EnvelopeChangeRequest, InsightCard — exported with fixtures on both sides. - skills-py: potential-assessment (certified × rolling fulfilment, evidence days) and dispatch-optimization (per-interval LP on HiGHS, shortfall reported) skills + routes + tests. - packages/services: dispatch rules in the policy pack (over-allocation, award anchor, lineage integrity for allocations); PowerBalanceSimulator; SimulationGateway (permit-only, idempotent, seeded execute → ExecutionReports); envelope deviation-streak suspension + apply(); ReviewService (attribution, reliability EWMA writeback, semantic memory, envelope recommendations as change requests); dispatch assembler; skill client methods. - packages/runtime: resource agent; award-decomposition, review and envelope-review workflows; lifecycle selects simulator/gateway by proposal type; trigger hooks for awards, execution reports, metering; decide() resumes either lifecycle or envelope-review runs; insight cards API. - Tests: docs/07 D-1 16:00 and D+1 end to end; reliability score 0.9 → 0.880 and the next assessment de-rates capacity; envelope suspension on a seeded 3-day streak; WIDEN request applied only by a human. 184 TS + 80 Python. - docs/open-questions: B10 (reliability/potential parameters). README and CLAUDE.md status → M4 done, M5 next. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 19:31:12 -04:00
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"id": {
"type": "string",
"minLength": 1
},
"page": {
"type": "string",
"enum": [
"TRADING_DESK",
"RESOURCE_POOL",
"OPERATIONS_DASHBOARD"
]
},
"market_date": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"title": {
"type": "string",
"minLength": 1
},
"headline": {
"type": "string",
"minLength": 1
},
"metrics": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": {
"type": "string",
"minLength": 1
},
"value": {
"type": "string",
"pattern": "^-?\\d+(\\.\\d+)?$"
},
"unit": {
"type": "string",
"minLength": 1
},
"ref": {
"anyOf": [
{
"type": "object",
"properties": {
"tool_call_id": {
"type": "string",
"minLength": 1
},
"path": {
"type": "string",
"minLength": 1
}
},
"required": [
"tool_call_id",
"path"
],
"additionalProperties": false
},
{
"type": "null"
}
]
}
},
"required": [
"name",
"value",
"unit",
"ref"
],
"additionalProperties": false
}
},
"source": {
"type": "object",
"properties": {
"kind": {
"type": "string",
"enum": [
"SITUATION_REPORT",
"REVIEW_FINDING",
"DECISION_CASE",
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records bids without submitting them, simulates the market answer from the actual day-ahead clearing price, dispatches to the simulation gateway, runs the D+1 review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord) with a lineage-completeness audit. KpiReport regenerated after every day per docs/12 §4 definitions (C1/C2 placeholders as named config). Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1 envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids), L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED), L4 AI-off (templates run, no Proposal created); abnormal-day protocol on EXTREME situations; per-level authority (B8 placeholder); auditable drill. Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints and insight cards, replay CLI (npm run shadow). Ledger, time series and streak counters are file-backed so a multi-week run survives restarts. Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW receipt channel; contracts, fixtures and pydantic models regenerated. Services: L2 metrics moved from evals so the shadow run and the harness share one implementation. Fixes: envelope/permit validity compared ISO timestamps as strings ('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only, metrics unchanged) — rewritten from the live service. Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/ recovery, KPI definitions as implemented); README and CLAUDE.md status. Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN recommendation produced but not acted on; restart durability; drill; L2/L3/L4 and abnormal-day paths; API. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00
"POTENTIAL_ASSESSMENT",
"KPI_REPORT",
"SHADOW_DAY",
"BREAKER"
M4: resource agent, envelopes live, review loop, insight cards - packages/domain: AwardNotice, DISPATCH_PLAN proposal payload, ExecutionReport, MeteringRecord, potential-assessment and dispatch-optimization contracts, ReviewFinding (+ typed writebacks), SemanticMemoryEntry, EnvelopeChangeRequest, InsightCard — exported with fixtures on both sides. - skills-py: potential-assessment (certified × rolling fulfilment, evidence days) and dispatch-optimization (per-interval LP on HiGHS, shortfall reported) skills + routes + tests. - packages/services: dispatch rules in the policy pack (over-allocation, award anchor, lineage integrity for allocations); PowerBalanceSimulator; SimulationGateway (permit-only, idempotent, seeded execute → ExecutionReports); envelope deviation-streak suspension + apply(); ReviewService (attribution, reliability EWMA writeback, semantic memory, envelope recommendations as change requests); dispatch assembler; skill client methods. - packages/runtime: resource agent; award-decomposition, review and envelope-review workflows; lifecycle selects simulator/gateway by proposal type; trigger hooks for awards, execution reports, metering; decide() resumes either lifecycle or envelope-review runs; insight cards API. - Tests: docs/07 D-1 16:00 and D+1 end to end; reliability score 0.9 → 0.880 and the next assessment de-rates capacity; envelope suspension on a seeded 3-day streak; WIDEN request applied only by a human. 184 TS + 80 Python. - docs/open-questions: B10 (reliability/potential parameters). README and CLAUDE.md status → M4 done, M5 next. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 19:31:12 -04:00
]
},
"id": {
"type": "string",
"minLength": 1
},
"ref": {
"anyOf": [
{
"type": "string",
"pattern": "^[0-9a-f]{64}$"
},
{
"type": "null"
}
]
}
},
"required": [
"kind",
"id",
"ref"
],
"additionalProperties": false
},
"generated_at": {
"type": "string",
"format": "date-time",
"pattern": "^(?:(?:\\d\\d[2468][048]|\\d\\d[13579][26]|\\d\\d0[48]|[02468][048]00|[13579][26]00)-02-29|\\d{4}-(?:(?:0[13578]|1[02])-(?:0[1-9]|[12]\\d|3[01])|(?:0[469]|11)-(?:0[1-9]|[12]\\d|30)|(?:02)-(?:0[1-9]|1\\d|2[0-8])))T(?:(?:[01]\\d|2[0-3]):[0-5]\\d:[0-5]\\d(?:\\.\\d+)?(?:Z))$"
}
},
"required": [
"id",
"page",
"market_date",
"title",
"headline",
"metrics",
"source",
"generated_at"
],
"additionalProperties": false,
"$id": "https://vpp-ai-platform/contracts/insight_card.json",
"title": "InsightCard"
}