M3: Mastra runtime, safety chain, two agents, Case Desk v1
- packages/domain: safety-chain objects (ValidationResult, SimulationResult,
EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft,
RouterDecision, BidExportFile) + fixtures on both sides.
- packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack
(digest-valid, bid-format, price-limits, quantity-non-negative,
ledger-consistency, lineage-integrity, originator-permission — each with
pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits,
revoke, gateway validate); FileExportGateway (idempotent receipts);
Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario
simulator; CaseDeskService; FsRepository; skill HTTP client moved here.
- packages/runtime: createRuntime (LibSQL storage, per-runtime workflow
factories), proposal-lifecycle (rule check → simulation → envelope gate
with suspend/resume → fresh check + permit → release), day-ahead-situation,
day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort
(Mastra/Scripted/Null), Case Desk HTTP API, dev entry point.
- Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with
LLM down, restart survival of a suspended approval, permit expiry and
revocation, replay of a released proposal, trigger scheduling. 141 TS +
60 Python tests.
- Known gaps: ledger not yet persisted (replayed on restart); STALE ends the
run instead of looping to rule check; synthetic data stands in for
historical replay.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:55:55 -04:00
|
|
|
{
|
|
|
|
|
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
|
|
|
|
"type": "object",
|
|
|
|
|
"properties": {
|
|
|
|
|
"receipt_id": {
|
|
|
|
|
"type": "string",
|
|
|
|
|
"minLength": 1
|
|
|
|
|
},
|
|
|
|
|
"proposal_digest": {
|
|
|
|
|
"type": "string",
|
|
|
|
|
"pattern": "^[0-9a-f]{64}$"
|
|
|
|
|
},
|
|
|
|
|
"permit_id": {
|
|
|
|
|
"type": "string",
|
|
|
|
|
"minLength": 1
|
|
|
|
|
},
|
|
|
|
|
"channel": {
|
|
|
|
|
"type": "string",
|
|
|
|
|
"enum": [
|
|
|
|
|
"FILE_EXPORT",
|
|
|
|
|
"TRADING_PLATFORM_API",
|
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00
|
|
|
"SIMULATION_GATEWAY",
|
|
|
|
|
"SHADOW"
|
M3: Mastra runtime, safety chain, two agents, Case Desk v1
- packages/domain: safety-chain objects (ValidationResult, SimulationResult,
EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft,
RouterDecision, BidExportFile) + fixtures on both sides.
- packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack
(digest-valid, bid-format, price-limits, quantity-non-negative,
ledger-consistency, lineage-integrity, originator-permission — each with
pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits,
revoke, gateway validate); FileExportGateway (idempotent receipts);
Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario
simulator; CaseDeskService; FsRepository; skill HTTP client moved here.
- packages/runtime: createRuntime (LibSQL storage, per-runtime workflow
factories), proposal-lifecycle (rule check → simulation → envelope gate
with suspend/resume → fresh check + permit → release), day-ahead-situation,
day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort
(Mastra/Scripted/Null), Case Desk HTTP API, dev entry point.
- Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with
LLM down, restart survival of a suspended approval, permit expiry and
revocation, replay of a released proposal, trigger scheduling. 141 TS +
60 Python tests.
- Known gaps: ledger not yet persisted (replayed on restart); STALE ends the
run instead of looping to rule check; synthetic data stands in for
historical replay.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:55:55 -04:00
|
|
|
]
|
|
|
|
|
},
|
|
|
|
|
"idempotency_key": {
|
|
|
|
|
"type": "string",
|
|
|
|
|
"minLength": 1
|
|
|
|
|
},
|
|
|
|
|
"artifact_ref": {
|
|
|
|
|
"type": "string",
|
|
|
|
|
"minLength": 1
|
|
|
|
|
},
|
|
|
|
|
"accepted_at": {
|
|
|
|
|
"type": "string",
|
|
|
|
|
"format": "date-time",
|
|
|
|
|
"pattern": "^(?:(?:\\d\\d[2468][048]|\\d\\d[13579][26]|\\d\\d0[48]|[02468][048]00|[13579][26]00)-02-29|\\d{4}-(?:(?:0[13578]|1[02])-(?:0[1-9]|[12]\\d|3[01])|(?:0[469]|11)-(?:0[1-9]|[12]\\d|30)|(?:02)-(?:0[1-9]|1\\d|2[0-8])))T(?:(?:[01]\\d|2[0-3]):[0-5]\\d:[0-5]\\d(?:\\.\\d+)?(?:Z))$"
|
|
|
|
|
}
|
|
|
|
|
},
|
|
|
|
|
"required": [
|
|
|
|
|
"receipt_id",
|
|
|
|
|
"proposal_digest",
|
|
|
|
|
"permit_id",
|
|
|
|
|
"channel",
|
|
|
|
|
"idempotency_key",
|
|
|
|
|
"artifact_ref",
|
|
|
|
|
"accepted_at"
|
|
|
|
|
],
|
|
|
|
|
"additionalProperties": false,
|
|
|
|
|
"$id": "https://vpp-ai-platform/contracts/execution_receipt.json",
|
|
|
|
|
"title": "ExecutionReceipt"
|
|
|
|
|
}
|