Four-layer eval design (LLM tasks / skills / safety chain / end-to-end decision quality), datasets built on event-log replay (I7), change gates mapping each change type to required evals — envelope widening approvable only on shadow/online L4 data — and measurable definitions for the proposal's core KPIs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019u5SLNweVio6ozJX7yfxQr 🔮 View transcript: https://logs.lojong.info/s/e8u90k3t33w590r7b5y7yzqh |
||
|---|---|---|
| .. | ||
| 00-overview.md | ||
| 01-principles.md | ||
| 02-cognitive-plane.md | ||
| 03-safety-chain.md | ||
| 04-control-plane.md | ||
| 05-skills-and-data.md | ||
| 06-integration.md | ||
| 07-scenario-walkthrough.md | ||
| 08-implementation.md | ||
| 09-runtime-implementation.md | ||
| 10-federation.md | ||
| 11-contracts.md | ||
| 12-evaluation.md | ||