vpp-ai-platform/docs
Thomas Bayes 381b6d3521
Some checks failed
ci / typescript (push) Has been cancelled
ci / python (push) Has been cancelled
ci / evals (push) Has been cancelled
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).

Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.

Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.

Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.

Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.

Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.

Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00
..
adr Scaffold repo for implementation handoff 2026-09-01 21:13:00 -04:00
external docs: drop edition tagline from architecture brief PDFs 2026-09-02 13:17:49 -04:00
00-overview.md Add risk assessment and failure-mode analysis (doc 13) 2026-09-01 21:06:55 -04:00
01-principles.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
02-cognitive-plane.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
03-safety-chain.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
04-control-plane.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
05-skills-and-data.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
06-integration.md init check in 2026-09-01 19:46:59 -04:00
07-scenario-walkthrough.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
08-implementation.md Add evaluation architecture (doc 12) 2026-09-01 21:03:25 -04:00
09-runtime-implementation.md Add Runtime implementation design (Mastra mapping) 2026-09-01 20:24:32 -04:00
10-federation.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
11-contracts.md Add interfaces and contracts reference (doc 11) 2026-09-01 20:52:14 -04:00
12-evaluation.md Add evaluation architecture (doc 12) 2026-09-01 21:03:25 -04:00
13-risks-failure-modes.md Add risk assessment and failure-mode analysis (doc 13) 2026-09-01 21:06:55 -04:00
14-shadow-run-runbook.md M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI 2026-09-02 22:54:14 -04:00
open-questions.md M4: resource agent, envelopes live, review loop, insight cards 2026-09-02 19:31:12 -04:00