diff --git a/docs/external/architecture-brief.en.md b/docs/external/architecture-brief.en.md
new file mode 100644
index 0000000..452c5b2
--- /dev/null
+++ b/docs/external/architecture-brief.en.md
@@ -0,0 +1,74 @@
+# VPP AI Platform — Architecture Brief
+
+**Virtual Power Plant Multi-Timescale Intelligent Operations Platform, built on the GuangMing Power LLM**
+
+*Condensed external edition · English · v0.1 · September 2026 · Full version: [system-architecture.en.md](system-architecture.en.md)*
+
+---
+
+The platform is an AI-assisted decision and controlled-execution system for a virtual power plant in Hubei. Five LLM agents propose market bids, dispatch plans and load-control plans. A deterministic safety chain governs everything before any external effect. Three commitments define the design:
+
+- **Two planes, one channel.** A slow cognitive plane (LLM agents, minutes to days) is separated from a fast control plane (deterministic execution, seconds). The safety chain is the only path between them.
+- **The LLM never computes a number.** Prices, volumes and setpoints come only from versioned forecasting, optimisation and simulation tools. Every numeric field in a proposal is a reference to a recorded tool call.
+- **Autonomy through envelopes.** Humans approve envelopes (bounded authority with a validity window), not every action. In-envelope actions auto-pass after rule check and simulation. Everything else escalates to a human. Envelopes start empty and widen only on evidence.
+
+## 1. System context
+
+
+
+The platform sits between the provincial trading, dispatch and metering systems above and the aggregated flexible resources below (31 MW today, at least 80 MW targeted). Operators work through a Case Desk, where every AI output is attached to an accountable case with an objective, an owner and a deadline. The platform overlays the existing 3060 platform: it reuses master data and the permission system, and it converges every bid and control action into one audited channel so nothing can bypass the safety chain.
+
+## 2. Agent runtime
+
+
+
+Workflows are the entry point, not agents. Every business process (day-ahead bidding, review, intraday correction) is a deterministic, durable chain of steps. Agents are invoked at specific steps with a whitelisted tool set. Scheduled and event triggers instantiate workflow templates without touching the LLM. Only manual requests pass through the router agent, whose output is constrained by schema to a registered template ID plus parameters. The LLM decides *what* to do; the engine decides *how*. Every trigger, step, tool call, transition, approval and permit is written to an event-sourced log.
+
+## 3. Context management and memory
+
+
+
+Context for each agent step is assembled by a deterministic upstream step that reads the versioned ledger, immutable snapshots, explicitly declared semantic memory and effective knowledge entries. The agent does not decide what to look up, so every decision is reproducible. Agent output carries numbers only as references into recorded tool-call lineage. The assembler dereferences them, which means LLM text cannot carry a number into a proposal.
+
+Three memory layers serve different lifetimes. Working memory lives for one task. Episodic memory is the event log and time-series store, permanent and queried rather than recalled. Semantic memory holds structured lessons written back by the D+1 review workflow, versioned and injected explicitly by each template. Authoritative operational facts always live in the domain database, never in agent memory.
+
+## 4. Safety chain and state machine
+
+
+
+Every intended action is a `Proposal` with a typed payload, full lineage and an immutable digest. One lifecycle workflow implements the chain for all proposal types:
+
+| Stage | What happens |
+|---|---|
+| Rule check | Deterministic policy pack: compliance, constraints, ledger consistency, lineage completeness, separation of duties |
+| Simulation | Revenue backtest for bids, power balance for control plans. A warning always escalates to a human |
+| Envelope check | Within an approved envelope: auto-approved with a full record. Otherwise the run suspends into the Case Desk inbox |
+| Human approval | A resume by an approver whose identity is separated from the origination chain. Binds to the digest and a validity window |
+| Fresh-state check | Immediately before release: resources online, constraints satisfied, evidence unchanged. Otherwise `STALE` and re-enter |
+| Permit and release | An `ExecutionPermit` is issued: short-lived, revocable, digest-matched. Gateways accept only the `(Proposal, ExecutionPermit)` pair |
+
+Suspended approvals survive process restarts. Permit revocation is the fastest rollback: the gateway refuses further commands and edge terminals fall back to the last approved plan or a safe curve.
+
+## 5. Technology stack
+
+
+
+The runtime, agents, workflows, domain schemas and deterministic services are TypeScript on Mastra, chosen for typed tool contracts and durable suspend/resume. The professional-model skills are Python HTTP services. LLM calls go through a capability-oriented port with a model router, so the backend moves from commercial APIs in development to a local open model in pre-production and the GuangMing Power LLM adapter in production with no business-code change. The runtime also has an LLM-down mode: forecasting, optimisation, checking, approval, bidding and control all run without the LLM. Roughly a third of the platform comes from the framework. The differentiating assets are the deterministic services.
+
+## 6. Interfaces and contracts
+
+
+
+Hand-written zod schemas are the single source of truth. The build exports JSON Schema to a committed contracts directory, from which pydantic models are generated and never hand-edited. Golden fixtures, valid and invalid, are validated by both the TypeScript and Python CI, and any disagreement fails the build. Inside the platform, each trust boundary is a port that the LLM cannot call: proposal submission, policy engine, simulation, envelope, authority, gateway, ledger, evidence and Case Desk. Data representation rules are part of the contract: decimal strings for money and energy, units in field names, UTC timestamps, intervals as date plus index, string IDs, upper-case enums.
+
+## 7. Evaluation-driven development
+
+
+
+Because every decision is replayable from retained evidence, the audit trail is also the evaluation dataset. Four layers localise problems: L1 LLM tasks, L2 professional-model skills, L3 safety-chain correctness including one automated test per invariant, and L4 end-to-end decision quality through shadow runs and historical backtests. Evaluation is a change gate. Prompt changes require the L1 set and an L4 backtest. A backend switch requires the full L1 set and an L4 shadow comparison. Skill upgrades must update their committed L2 baseline in the same change. Policy-pack upgrades require all rule tests green. Envelope widening is approved only on L4 shadow or online evidence. Every review finding becomes a candidate evaluation case, so failures feed the next baseline.
+
+---
+
+**Status.** Milestones M1 to M4 are delivered: schemas and contract pipeline, deterministic services, Python skills with a committed evaluation baseline, the runtime with the full safety chain, resource dispatch and review workflows, and the Case Desk. All eight invariants have passing automated tests, including the LLM-down run and rejection of AI self-approval. M5, a shadow run on live data with all external effects simulated, is next. Business parameters such as market windows, envelope bounds and loss budgets are held as named configuration pending confirmation by the operations team.
+
+*Diagrams are generated from `diagrams/gen.mjs` via `diagrams/render.sh`.*
diff --git a/docs/external/diagrams/gen.mjs b/docs/external/diagrams/gen.mjs
new file mode 100644
index 0000000..b8b2e68
--- /dev/null
+++ b/docs/external/diagrams/gen.mjs
@@ -0,0 +1,286 @@
+import { writeFileSync } from 'node:fs'
+
+const C = {
+ ink: '#1B2331', ink2: '#4A5568', ink3: '#7A8494', line: '#B9C2CE',
+ teal: '#2C6B70', tealSoft: '#E4F0F0', amber: '#B86A14', amberSoft: '#F7ECDD',
+ slate: '#EEF1F5', red: '#A23B3B', redSoft: '#F6E3E3', white: '#FFFFFF',
+}
+const FONT = 'font-family="Helvetica Neue,Helvetica,Arial,sans-serif"'
+const esc = (s) => s.replace(/&/g, '&').replace(//g, '>')
+
+function text(x, y, s, { size = 13, fill = C.ink, anchor = 'middle', weight = 400, mono = false, italic = false } = {}) {
+ const fam = mono ? 'font-family="Menlo,SFMono-Regular,Consolas,monospace"' : ''
+ return `${esc(s)}`
+}
+// title + sub lines centered in box
+function box(x, y, w, h, title, sub = [], o = {}) {
+ const { fill = C.white, stroke = C.line, sw = 1.2, dash = '', tcolor = C.ink, tsize = 14, ssize = 12, scolor = C.ink2, r = 5, tag = null, tagColor = C.amber, mono = false } = o
+ const lines = 1 + sub.length
+ const lh = 16
+ const total = tsize + 4 + sub.length * lh
+ let cy = y + h / 2 - total / 2 + tsize - 2
+ let s = ``
+ s += text(x + w / 2, cy, title, { size: tsize, fill: tcolor, weight: 600, mono })
+ for (const l of sub) { cy += lh; s += text(x + w / 2, cy, l, { size: ssize, fill: scolor }) }
+ if (tag) {
+ const tw = tag.length * 7 + 14
+ s += `` + text(x + w - tw / 2 - 6, y + 4, tag, { size: 10, fill: C.white, weight: 700 })
+ }
+ return s
+}
+function group(x, y, w, h, label, o = {}) {
+ const { stroke = C.line, fill = 'none', color = C.ink3 } = o
+ return `` +
+ text(x + 12, y + 18, label.toUpperCase(), { size: 10.5, fill: color, anchor: 'start', weight: 700 })
+}
+function arrow(x1, y1, x2, y2, label = '', o = {}) {
+ const { dashed = false, both = false, color = C.ink2, lx = 0, ly = -6, path = null, lsize = 11.5, lfill = C.ink2 } = o
+ const d = path || `M${x1},${y1} L${x2},${y2}`
+ let s = ``
+ if (label) {
+ const mx = (x1 + x2) / 2 + lx, my = (y1 + y2) / 2 + ly
+ const tw = label.length * 6.4 + 8
+ s += `` + text(mx, my + 2, label, { size: lsize, fill: lfill })
+ }
+ return s
+}
+function svg(w, h, body) {
+ return `
`
+}
+const legend = (x, y, items) => items.map(([fill, stroke, lbl], i) =>
+ `` + text(x + i * 150 + 20, y + 1, lbl, { size: 11, fill: C.ink2, anchor: 'start' })).join('')
+
+const out = {}
+
+// ---------- 1. system context ----------
+{
+ let b = ''
+ b += group(30, 40, 250, 420, 'External systems')
+ const ext = [['Trading platform', ['bids · clearing · disclosures']], ['Dispatch / load mgmt', ['regulation requests']], ['Metering / settlement', ['meter data · statements']], ['Weather & market data', ['read-only feeds']]]
+ ext.forEach(([t, s], i) => { b += box(50, 70 + i * 95, 210, 70, t, s, { fill: C.slate }) })
+
+ b += group(330, 40, 600, 490, 'VPP AI Platform')
+ b += box(360, 70, 540, 90, 'Cognitive plane', ['five LLM agents · agent runtime · workflows', 'minutes to days'], { fill: C.amberSoft, stroke: C.amber, tag: 'LLM' })
+ b += box(360, 215, 540, 90, 'Safety chain', ['rule check → simulation → envelope / human approval → fresh check → permit', 'deterministic · the only path between planes'], { fill: C.tealSoft, stroke: C.teal, tag: 'NO LLM', tagColor: C.teal })
+ b += box(360, 360, 540, 80, 'Control plane', ['execution engine · edge terminals · seconds'], { fill: C.tealSoft, stroke: C.teal, tag: 'NO LLM', tagColor: C.teal })
+ b += box(360, 465, 540, 45, 'Data foundation', [], { fill: C.slate, tsize: 13 })
+ b += text(630, 500, 'time-series · relational · knowledge/vector · feature store · event log', { size: 11.5, fill: C.ink2 })
+ b += arrow(630, 160, 630, 215, 'Proposal (digest + lineage)', { lx: 110 })
+ b += arrow(630, 305, 630, 360, '(Proposal, ExecutionPermit)', { lx: 110 })
+ b += arrow(280, 115, 360, 115, '', { both: true })
+ b += arrow(360, 260, 280, 100, '', { path: 'M360,260 L300,260 L300,130 L280,130' })
+ b += text(292, 253, 'bid + permit', { size: 10.5, fill: C.ink2, anchor: 'end' })
+ b += arrow(900, 400, 980, 400, '')
+ b += arrow(980, 125, 900, 125, '', { both: true })
+ b += box(980, 90, 190, 70, 'Operators', ['Case Desk · approvals', 'scenario forks · AI cards'], { fill: C.white })
+ b += box(980, 360, 190, 80, 'Flexible resources', ['PV · storage · EV charging', 'HVAC · industrial · microgrids'], { fill: C.white })
+ b += arrow(630, 440, 630, 465)
+ b += arrow(360, 400, 360, 140, '', { path: 'M360,400 L345,400 L345,140 L360,140', dashed: true })
+ b += `execution feedback`
+ b += text(1075, 470, '31 MW today → ≥ 80 MW target', { size: 11, fill: C.ink3 })
+ out['01-system-context'] = svg(1200, 550, b)
+}
+
+// ---------- 2. agent runtime ----------
+{
+ let b = ''
+ b += group(30, 40, 250, 330, 'Task triggers')
+ b += box(50, 70, 210, 70, 'Scheduled', ['day-ahead 06:00 · bid window · D+1 review'], { fill: C.tealSoft, stroke: C.teal })
+ b += box(50, 165, 210, 70, 'Event', ['deviation · resource offline · clearing'], { fill: C.tealSoft, stroke: C.teal })
+ b += box(50, 260, 210, 70, 'Manual', ['operator question or task'], { fill: C.slate })
+ b += box(330, 260, 190, 70, 'Router agent', ['intent → template id + params', 'schema-constrained enum'], { fill: C.amberSoft, stroke: C.amber, tag: 'LLM' })
+ b += arrow(260, 295, 330, 295)
+ b += box(580, 120, 240, 130, 'Workflow engine', ['deterministic step chains', 'durable · suspend/resume', 'one template per business process'], { fill: C.tealSoft, stroke: C.teal, tag: 'NO LLM', tagColor: C.teal })
+ b += arrow(260, 105, 580, 160, 'no LLM in the path', { path: 'M260,105 L420,105 L420,160 L580,160', lx: 110, ly: -37 })
+ b += arrow(260, 200, 580, 200, 'rule-matched', { lx: -20 })
+ b += arrow(520, 295, 580, 240, '', { path: 'M520,295 L560,295 L560,240 L580,240' })
+ b += box(880, 70, 260, 80, 'Agent step', ['one of five agents, invoked at a step', 'tools whitelisted per agent'], { fill: C.amberSoft, stroke: C.amber, tag: 'LLM' })
+ b += arrow(820, 150, 880, 110, '', { path: 'M820,150 L850,150 L850,110 L880,110' })
+ b += box(880, 190, 260, 80, 'Tool registry', ['typed, versioned skill contracts', 'every call snapshotted into lineage'], { fill: C.tealSoft, stroke: C.teal })
+ b += arrow(1010, 150, 1010, 190)
+ b += box(880, 310, 260, 60, 'Python skill services', ['forecast · MILP · simulation'], { fill: C.slate })
+ b += arrow(1010, 270, 1010, 310, 'HTTP / JSON', { lx: 60 })
+ b += box(580, 300, 240, 70, 'Proposal → safety chain', ['numbers dereferenced from lineage'], { fill: C.white, stroke: C.ink2 })
+ b += arrow(700, 250, 700, 300)
+ b += box(330, 400, 810, 46, 'Event log (event-sourced): triggers, steps, tool calls, transitions, approvals, permits, receipts', [], { fill: C.slate, tsize: 12.5, dash: '' })
+ b += arrow(700, 370, 700, 400, '', { dashed: true })
+ b += legend(30, 445, [[C.amberSoft, C.amber, 'LLM involved'], [C.tealSoft, C.teal, 'Deterministic']])
+ b += text(30, 480, 'Workflows are the entry point, not agents. The LLM decides what to do (pick a template); the engine decides how, step by step.', { size: 12, fill: C.ink2, anchor: 'start', italic: true })
+ out['02-agent-runtime'] = svg(1200, 500, b)
+}
+
+// ---------- 3. context & memory ----------
+{
+ let b = ''
+ b += group(30, 40, 1140, 210, 'Context assembly for one agent step')
+ const src = [['Position ledger', 'versioned read'], ['Time-series / snapshots', 'immutable, checksummed'], ['Semantic memory', 'lessons declared by template'], ['Knowledge base (RAG)', 'effective rules only']]
+ src.forEach(([t, s], i) => { b += box(50 + i * 180, 70, 165, 60, t, [s], { fill: C.slate, tsize: 12.5, ssize: 11 }) })
+ b += box(50, 170, 705, 55, 'fetchContext step (deterministic, reproducible)', ['the agent does not decide what to look up'], { fill: C.tealSoft, stroke: C.teal, tsize: 13 })
+ for (let i = 0; i < 4; i++) b += arrow(132 + i * 180, 130, 132 + i * 180, 170)
+ b += box(790, 70, 170, 155, 'Agent step', ['prose + tool calls', 'numbers only as', '{toolCallId, path}', 'references'], { fill: C.amberSoft, stroke: C.amber, tag: 'LLM' })
+ b += arrow(755, 197, 790, 197)
+ b += box(995, 70, 160, 155, 'Assembler', ['dereferences refs', 'from recorded lineage', '→ Proposal', 'LLM text cannot', 'carry a number'], { fill: C.tealSoft, stroke: C.teal })
+ b += arrow(960, 147, 995, 147)
+
+ b += group(30, 280, 1140, 165, 'Three memory layers')
+ b += box(50, 315, 355, 105, 'Working memory', ['current task context, intermediate results', 'runtime memory · task lifetime', 'step-to-step parameters preferred over memory'], { fill: C.white })
+ b += box(423, 315, 355, 105, 'Episodic memory', ['operating states, trading, control, responses', 'event log + time-series · permanent (audit)', 'queried through fetchContext, not recalled'], { fill: C.white })
+ b += box(796, 315, 355, 105, 'Semantic memory', ['structured lessons from D+1 review', 'knowledge + strategy store · versioned', 'written by ReviewFinding, injected explicitly'], { fill: C.white })
+ b += text(30, 475, 'Authoritative facts live in the domain database, never in agent memory (invariant I8). No automatic memory recall feeds production decisions.', { size: 12, fill: C.ink2, anchor: 'start', italic: true })
+ out['03-context-memory'] = svg(1200, 495, b)
+}
+
+// ---------- 4. safety chain state machine ----------
+{
+ let b = ''
+ const Y = 120, H = 56, W = 118, G = 32
+ const stages = [
+ ['DRAFT', 'agent output'], ['RULE_CHECK', 'policy pack'], ['SIMULATION', 'backtest / balance'], ['ENVELOPE_CHECK', 'within bounds?'],
+ ['APPROVED', 'auto or human'], ['FRESH_CHECK', 're-verify state'], ['AUTHORIZED', 'permit issued'], ['RELEASED', 'gateway accepts'],
+ ]
+ const xs = stages.map((_, i) => 30 + i * (W + G))
+ stages.forEach(([t, s], i) => {
+ const hi = i >= 1 && i <= 6
+ b += box(xs[i], Y, W, H, t, [s], { fill: hi ? C.tealSoft : C.white, stroke: hi ? C.teal : C.line, tsize: 12, ssize: 10.5, mono: true })
+ if (i < stages.length - 1) b += arrow(xs[i] + W, Y + H / 2, xs[i + 1], Y + H / 2)
+ })
+ b += group(xs[1] - 12, 52, xs[6] + W + 12 - xs[1] + 12, 148, 'Safety chain · deterministic · shared by every proposal type', { stroke: C.teal, color: C.teal })
+ // human branch
+ b += box(xs[3] + 20, 250, 200, 56, 'PENDING_HUMAN', ['suspended run · Case Desk inbox'], { fill: C.amberSoft, stroke: C.amber, tsize: 12, ssize: 10.5, mono: true })
+ b += arrow(xs[3] + W / 2, Y + H, xs[3] + 60, 250, 'outside envelope', { path: `M${xs[3] + W / 2},${Y + H} L${xs[3] + W / 2},250`, lx: -70, ly: 8 })
+ b += arrow(xs[2] + W / 2, Y + H, xs[3] + 20, 278, 'sim warning', { path: `M${xs[2] + W / 2},${Y + H} L${xs[2] + W / 2},278 L${xs[3] + 20},278`, lx: -45, ly: 40 })
+ b += arrow(xs[3] + 220, 278, xs[4] + W / 2, Y + H, 'human resume', { path: `M${xs[3] + 220},278 L${xs[4] + W / 2},278 L${xs[4] + W / 2},${Y + H}`, lx: 40, ly: 40 })
+ b += text(xs[4] + W / 2 + 16, 300, 'approver ≠ originator (I1)', { size: 10.5, fill: C.ink3, anchor: 'start' })
+ b += arrow(xs[3] + W / 2 + 30, Y, xs[4] + W / 2 - 30, Y, 'within envelope → AUTO_APPROVED', { path: `M${xs[3] + W - 10},${Y} L${xs[3] + W - 10},${Y - 24} L${xs[4] + 10},${Y - 24} L${xs[4] + 10},${Y}`, ly: -26, lx: 0 })
+ // rejected
+ b += box(xs[1], 350, W, 44, 'REJECTED', [], { fill: C.redSoft, stroke: C.red, tsize: 12, mono: true })
+ b += arrow(xs[1] + W / 2, Y + H, xs[1] + W / 2, 350, 'rule failure', { lx: 44 })
+ b += arrow(xs[3] + 60, 306, xs[1] + W, 372, 'human rejects', { path: `M${xs[3] + 60},306 L${xs[3] + 60},372 L${xs[1] + W},372`, lx: -30, ly: 45 })
+ // stale
+ b += box(xs[5], 350, W, 44, 'STALE', [], { fill: C.redSoft, stroke: C.red, tsize: 12, mono: true })
+ b += arrow(xs[5] + W / 2, Y + H, xs[5] + W / 2, 350, 'evidence changed', { lx: 60 })
+ b += arrow(xs[5] + W / 2, 350, xs[1] + W - 15, Y + H, 're-enter chain', { path: `M${xs[5] + W / 2 - 30},350 L${xs[5] + W / 2 - 30},330 L${xs[1] + W - 15},330 L${xs[1] + W - 15},${Y + H + 2}`, dashed: true, lx: -172, ly: 59 })
+ // after release
+ b += text(xs[7] + W / 2, 210, 'EXECUTING → COMPLETED', { size: 11, fill: C.ink2, mono: true })
+ b += text(xs[7] + W / 2, 226, 'or ROLLED_BACK', { size: 11, fill: C.ink2, mono: true })
+ b += text(xs[7] + W / 2, 244, '→ ExecutionReport → review', { size: 10.5, fill: C.ink3 })
+ // notes
+ b += text(30, 430, 'Approval binds to the proposal digest and a validity window (I3). Gateways accept only (Proposal, ExecutionPermit): short-lived, revocable, digest-matched (I4).', { size: 12, fill: C.ink2, anchor: 'start' })
+ b += text(30, 452, 'Simulation warnings always escalate to a human. No code path lets the AI reach APPROVED. Envelopes start empty and widen only on shadow-run evidence.', { size: 12, fill: C.ink2, anchor: 'start' })
+ out['04-safety-chain'] = svg(1290, 475, b)
+}
+
+// ---------- 5. tech stack ----------
+{
+ let b = ''
+ const row = (y, h, label, items, o = {}) => {
+ b += text(30, y + h / 2 + 4, label, { size: 11, fill: C.ink3, anchor: 'start', weight: 700 })
+ const x0 = 190, W = 980, n = items.length, g = 12, w = (W - g * (n - 1)) / n
+ items.forEach(([t, s], i) => { b += box(x0 + i * (w + g), y, w, h, t, s, { fill: o.fill || C.white, stroke: o.stroke || C.line, tsize: 13, ssize: 11, tag: o.tag && i === 0 ? o.tag : null, tagColor: o.tagColor }) })
+ }
+ row(40, 64, 'OPERATOR UI', [['Case Desk', ['approval inbox · case view · lineage expansion']], ['AI insight cards', ['read-only projections in existing business pages']]])
+ row(120, 80, 'RUNTIME', [['TypeScript + Mastra', ['workflows (durable suspend/resume)', 'agents as controlled steps · tool registry']], ['Trigger service', ['cron · event bus consumers', 'router agent for manual requests']], ['LLM port', ['capability interface + model router', 'LLM-down mode for degraded runs']]], { fill: C.amberSoft, stroke: C.amber })
+ row(226, 80, 'SERVICES', [['Policy engine', ['versioned, tested policy packs']], ['Envelope + authority', ['envelope match · fresh check · permits']], ['Ledger + lineage', ['position cascade · snapshots · assembler']], ['Gateways + Case Desk', ['file export · simulation · cases']]], { fill: C.tealSoft, stroke: C.teal, tag: 'NO LLM', tagColor: C.teal })
+ row(326, 64, 'CONTRACTS', [['Domain schemas (zod)', ['single source of truth']], ['JSON Schema', ['exported, committed, versioned']], ['Golden fixtures', ['validated by both CIs']]])
+ row(406, 72, 'SKILLS', [['Python services (FastAPI)', ['load / PV / price forecasts with quantiles', 'bid MILP · dispatch MILP · potential · report']], ['Generated pydantic models', ['never hand-edited']]], { fill: C.slate })
+ row(494, 72, 'DATA', [['PostgreSQL + pgvector', ['relational · knowledge · outbox events']], ['TimescaleDB / IoTDB', ['telemetry curves']], ['Object storage', ['snapshots · archives']], ['Kafka (phase 2)', ['outbox relay, same interface']]])
+ row(582, 72, 'LLM BACKENDS', [['Commercial API', ['development']], ['Local open model', ['pre-production']], ['GuangMing Power LLM', ['production, on-prem adapter']]], { fill: C.amberSoft, stroke: C.amber })
+ b += text(30, 688, 'Backend swaps are gated by the L1 evaluation set and an L4 shadow comparison, not judgment. About a third of the platform is framework; the differentiating parts are the deterministic services.', { size: 12, fill: C.ink2, anchor: 'start', italic: true })
+ out['05-tech-stack'] = svg(1200, 706, b)
+}
+
+// ---------- 6. interfaces & contracts ----------
+{
+ let b = ''
+ b += group(30, 40, 560, 260, 'Contract pipeline (one source of truth)')
+ b += box(55, 75, 220, 60, 'zod schemas', ['packages/domain · hand-written'], { fill: C.tealSoft, stroke: C.teal, mono: true, tsize: 13 })
+ b += box(55, 165, 220, 60, 'JSON Schema', ['contracts/ · committed · versioned'], { fill: C.slate, mono: true, tsize: 13 })
+ b += arrow(165, 135, 165, 165, 'build export', { lx: 60 })
+ b += box(340, 165, 225, 60, 'pydantic models', ['generated · never hand-edited'], { fill: C.slate, mono: true, tsize: 13 })
+ b += arrow(275, 195, 340, 195, 'codegen', { ly: -8 })
+ b += box(340, 75, 225, 60, 'Golden fixtures', ['valid + invalid samples per object'], { fill: C.white })
+ b += arrow(275, 105, 340, 105, '', {})
+ b += arrow(452, 135, 452, 165, '', { dashed: true })
+ b += box(55, 245, 510, 40, 'TS CI and Python CI both validate every fixture; any disagreement fails the build', [], { fill: C.white, stroke: C.ink2, tsize: 12 })
+ b += arrow(165, 225, 165, 245, '', { dashed: true })
+ b += arrow(452, 225, 452, 245, '', { dashed: true })
+
+ b += group(620, 40, 550, 260, 'Ports = trust boundaries (deterministic, not exposed to the LLM)')
+ const ports = [
+ ['ProposalSubmissionPort', 'the only way into the safety chain'],
+ ['PolicyEnginePort · SimulationPort', 'rule check and simulation'],
+ ['EnvelopePort · AuthorityPort', 'envelope match · fresh check · permits · revoke'],
+ ['GatewayPort', 'accepts (Proposal, ExecutionPermit) only'],
+ ['LedgerPort · EvidencePort', 'versioned ledger · event-sourced evidence'],
+ ['CaseDeskPort', 'open · read · apply command · watch'],
+ ]
+ ports.forEach(([p, d], i) => {
+ const y = 78 + i * 36
+ b += text(640, y + 12, p, { size: 12, fill: C.teal, anchor: 'start', mono: true, weight: 600 })
+ b += text(1150, y + 12, d, { size: 11.5, fill: C.ink2, anchor: 'end' })
+ if (i < ports.length - 1) b += ``
+ })
+
+ b += group(30, 330, 1140, 150, 'Cross-process interfaces and data representation rules')
+ const ifs = [['TS runtime → Python skills', 'HTTP/JSON, schema-validated both ends'], ['Case Desk / front end → runtime', 'HTTP/JSON approvals and case API'], ['External systems', 'anti-corruption adapters, raw + translated logged'], ['Edge control link', 'grid protocol stack, outside the JSON domain']]
+ ifs.forEach(([t, s], i) => { b += box(50 + i * 280, 360, 265, 56, t, [s], { fill: C.white, tsize: 12.5, ssize: 11 }) })
+ b += text(50, 448, 'Money, energy and prices as fixed-point decimal strings · units in field names (power_mw, price_yuan_per_mwh) · ISO 8601 UTC · intervals as {date, interval_index}', { size: 11.5, fill: C.ink2, anchor: 'start' })
+ b += text(50, 466, 'all IDs strings · enums UPPER_SNAKE · breaking schema changes bump the major version and ship a migration note', { size: 11.5, fill: C.ink2, anchor: 'start' })
+ out['06-contracts'] = svg(1200, 500, b)
+}
+
+// ---------- 7. eval-driven development ----------
+{
+ let b = ''
+ b += group(30, 40, 470, 345, 'Four evaluation layers')
+ const layers = [
+ ['L4', 'End-to-end decision quality', 'shadow run · historical backtest · counterfactual attribution'],
+ ['L3', 'Safety chain correctness', 'policy pack tests · invariant tests I1–I8 · red-team cases'],
+ ['L2', 'Professional-model skills', 'MAPE · interval coverage · solver quality vs. bounds'],
+ ['L1', 'LLM tasks', 'routing accuracy · schema compliance · numeric consistency'],
+ ]
+ layers.forEach(([l, t, s], i) => {
+ const y = 70 + i * 72
+ b += ``
+ b += text(72, y + 36, l, { size: 16, fill: C.teal, weight: 700, mono: true })
+ b += text(105, y + 26, t, { size: 13, fill: C.ink, anchor: 'start', weight: 600 })
+ b += text(105, y + 45, s, { size: 11, fill: C.ink2, anchor: 'start' })
+ })
+ b += text(50, 370, 'A problem at L4 is localised to the layer that caused it.', { size: 11.5, fill: C.ink3, anchor: 'start', italic: true })
+
+ b += group(530, 40, 640, 345, 'Change gates: what must pass before a change ships')
+ const gates = [
+ ['Prompt / agent instructions', 'L1 task set + L4 fixed-day backtest, no regression'],
+ ['LLM backend switch', 'full L1 set + L4 shadow comparison vs. baseline'],
+ ['Skill model upgrade', 'all L2 metrics for that skill + downstream L4 backtest'],
+ ['Policy pack upgrade', 'L3 pack tests green + impact analysis on pending proposals'],
+ ['Contract schema change', 'golden fixtures on both sides + compatibility check'],
+ ['Envelope widening', 'L4 shadow / online evidence only — no data, no widening'],
+ ]
+ gates.forEach(([c, g], i) => {
+ const y = 78 + i * 46
+ b += text(550, y + 14, c, { size: 12.5, fill: C.ink, anchor: 'start', weight: 600 })
+ b += text(550, y + 31, g, { size: 11.5, fill: i === 5 ? C.amber : C.ink2, anchor: 'start' })
+ if (i < gates.length - 1) b += ``
+ })
+
+ // loop
+ b += group(30, 410, 1140, 130, 'The audit trail is the evaluation dataset')
+ const loop = [['Production run', ['event log + immutable snapshots']], ['Replay datasets', ['representative days, frozen']], ['Eval run (archived)', ['committed baseline, CI --check']], ['ReviewFinding', ['D+1 attribution']], ['New eval cases', ['every finding → candidate case']]]
+ loop.forEach(([t, s], i) => {
+ b += box(50 + i * 228, 445, 200, 60, t, s, { fill: i === 2 ? C.tealSoft : C.white, stroke: i === 2 ? C.teal : C.line, tsize: 12.5, ssize: 11 })
+ if (i < loop.length - 1) b += arrow(250 + i * 228, 475, 278 + i * 228, 475)
+ })
+ b += arrow(1062, 505, 150, 505, '', { path: 'M1062,505 L1062,525 L150,525 L150,506', dashed: true })
+ out['07-eval-driven'] = svg(1200, 560, b)
+}
+
+for (const [name, html] of Object.entries(out)) writeFileSync(`${process.argv[2]}/${name}.html`, html)
+console.log(Object.keys(out).join('\n'))
diff --git a/docs/external/diagrams/render.sh b/docs/external/diagrams/render.sh
new file mode 100755
index 0000000..cb0f678
--- /dev/null
+++ b/docs/external/diagrams/render.sh
@@ -0,0 +1,16 @@
+#!/usr/bin/env bash
+# Regenerates docs/external/img/*.png from gen.mjs (hand-authored SVG) using headless Chrome.
+set -euo pipefail
+cd "$(dirname "$0")"
+CHROME="${CHROME:-/Applications/Google Chrome.app/Contents/MacOS/Google Chrome}"
+tmp=$(mktemp -d)
+node gen.mjs "$tmp" >/dev/null
+for f in "$tmp"/*.html; do
+ n=$(basename "${f%.html}")
+ dims=$(grep -o 'width="[0-9]*" height="[0-9]*" font' "$f" | head -1 | grep -o '[0-9]*' | tr '\n' ' ')
+ set -- $dims
+ "$CHROME" --headless=new --disable-gpu --hide-scrollbars --force-device-scale-factor=2 \
+ --window-size="$1,$2" --screenshot="../img/$n.png" "$f" >/dev/null 2>&1
+done
+rm -rf "$tmp"
+ls ../img
diff --git a/docs/external/img/01-system-context.png b/docs/external/img/01-system-context.png
new file mode 100644
index 0000000..fd9b197
Binary files /dev/null and b/docs/external/img/01-system-context.png differ
diff --git a/docs/external/img/02-agent-runtime.png b/docs/external/img/02-agent-runtime.png
new file mode 100644
index 0000000..c859ee7
Binary files /dev/null and b/docs/external/img/02-agent-runtime.png differ
diff --git a/docs/external/img/03-context-memory.png b/docs/external/img/03-context-memory.png
new file mode 100644
index 0000000..95ac7dd
Binary files /dev/null and b/docs/external/img/03-context-memory.png differ
diff --git a/docs/external/img/04-safety-chain.png b/docs/external/img/04-safety-chain.png
new file mode 100644
index 0000000..7cbf7e6
Binary files /dev/null and b/docs/external/img/04-safety-chain.png differ
diff --git a/docs/external/img/05-tech-stack.png b/docs/external/img/05-tech-stack.png
new file mode 100644
index 0000000..2e2aa1c
Binary files /dev/null and b/docs/external/img/05-tech-stack.png differ
diff --git a/docs/external/img/06-contracts.png b/docs/external/img/06-contracts.png
new file mode 100644
index 0000000..a65d7db
Binary files /dev/null and b/docs/external/img/06-contracts.png differ
diff --git a/docs/external/img/07-eval-driven.png b/docs/external/img/07-eval-driven.png
new file mode 100644
index 0000000..8b0daf0
Binary files /dev/null and b/docs/external/img/07-eval-driven.png differ