vpp-ai-platform/docs
stewart hu 6c999186c0 Add evaluation architecture (doc 12)
Four-layer eval design (LLM tasks / skills / safety chain / end-to-end
decision quality), datasets built on event-log replay (I7), change
gates mapping each change type to required evals — envelope widening
approvable only on shadow/online L4 data — and measurable definitions
for the proposal's core KPIs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019u5SLNweVio6ozJX7yfxQr

🔮 View transcript: https://logs.lojong.info/s/e8u90k3t33w590r7b5y7yzqh
2026-09-01 21:03:25 -04:00
..
00-overview.md Add evaluation architecture (doc 12) 2026-09-01 21:03:25 -04:00
01-principles.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
02-cognitive-plane.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
03-safety-chain.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
04-control-plane.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
05-skills-and-data.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
06-integration.md init check in 2026-09-01 19:46:59 -04:00
07-scenario-walkthrough.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
08-implementation.md Add evaluation architecture (doc 12) 2026-09-01 21:03:25 -04:00
09-runtime-implementation.md Add Runtime implementation design (Mastra mapping) 2026-09-01 20:24:32 -04:00
10-federation.md Integrate peer review findings into architecture docs 2026-09-01 20:44:12 -04:00
11-contracts.md Add interfaces and contracts reference (doc 11) 2026-09-01 20:52:14 -04:00
12-evaluation.md Add evaluation architecture (doc 12) 2026-09-01 21:03:25 -04:00