Four-layer eval design (LLM tasks / skills / safety chain / end-to-end
decision quality), datasets built on event-log replay (I7), change
gates mapping each change type to required evals — envelope widening
approvable only on shadow/online L4 data — and measurable definitions
for the proposal's core KPIs.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019u5SLNweVio6ozJX7yfxQr🔮 View transcript: https://logs.lojong.info/s/e8u90k3t33w590r7b5y7yzqh
Consolidates plane-boundary ports, service/event contracts, and the
canonical domain-object table (merged with peer review §12), plus the
TS↔Python strategy: zod as schema source, committed JSON Schema
artifacts, generated pydantic models, golden-fixture contract tests
in both CIs, and cross-language data-representation rules.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019u5SLNweVio6ozJX7yfxQr🔮 View transcript: https://logs.lojong.info/s/e8u90k3t33w590r7b5y7yzqh