M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00
import { z } from 'zod'
import { AwardNotice , Curve96 , ExecutionReport , HumanBidRecord , MarketDate , MeteringRecord , RouterDecision } from '@vpp/domain'
import type { EventEnvelope , KpiReport , ShadowDayRecord } from '@vpp/domain'
import { shadowClearing } from '@vpp/services'
M4: resource agent, envelopes live, review loop, insight cards
- packages/domain: AwardNotice, DISPATCH_PLAN proposal payload, ExecutionReport,
MeteringRecord, potential-assessment and dispatch-optimization contracts,
ReviewFinding (+ typed writebacks), SemanticMemoryEntry,
EnvelopeChangeRequest, InsightCard — exported with fixtures on both sides.
- skills-py: potential-assessment (certified × rolling fulfilment, evidence
days) and dispatch-optimization (per-interval LP on HiGHS, shortfall
reported) skills + routes + tests.
- packages/services: dispatch rules in the policy pack (over-allocation,
award anchor, lineage integrity for allocations); PowerBalanceSimulator;
SimulationGateway (permit-only, idempotent, seeded execute → ExecutionReports);
envelope deviation-streak suspension + apply(); ReviewService (attribution,
reliability EWMA writeback, semantic memory, envelope recommendations as
change requests); dispatch assembler; skill client methods.
- packages/runtime: resource agent; award-decomposition, review and
envelope-review workflows; lifecycle selects simulator/gateway by proposal
type; trigger hooks for awards, execution reports, metering; decide()
resumes either lifecycle or envelope-review runs; insight cards API.
- Tests: docs/07 D-1 16:00 and D+1 end to end; reliability score 0.9 → 0.880
and the next assessment de-rates capacity; envelope suspension on a seeded
3-day streak; WIDEN request applied only by a human. 184 TS + 80 Python.
- docs/open-questions: B10 (reliability/potential parameters). README and
CLAUDE.md status → M4 done, M5 next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 19:31:12 -04:00
import { ENVELOPE_REVIEW_STEP_ID } from './workflows/envelope-review.js'
M3: Mastra runtime, safety chain, two agents, Case Desk v1
- packages/domain: safety-chain objects (ValidationResult, SimulationResult,
EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft,
RouterDecision, BidExportFile) + fixtures on both sides.
- packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack
(digest-valid, bid-format, price-limits, quantity-non-negative,
ledger-consistency, lineage-integrity, originator-permission — each with
pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits,
revoke, gateway validate); FileExportGateway (idempotent receipts);
Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario
simulator; CaseDeskService; FsRepository; skill HTTP client moved here.
- packages/runtime: createRuntime (LibSQL storage, per-runtime workflow
factories), proposal-lifecycle (rule check → simulation → envelope gate
with suspend/resume → fresh check + permit → release), day-ahead-situation,
day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort
(Mastra/Scripted/Null), Case Desk HTTP API, dev entry point.
- Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with
LLM down, restart survival of a suspended approval, permit expiry and
revocation, replay of a released proposal, trigger scheduling. 141 TS +
60 Python tests.
- Known gaps: ledger not yet persisted (replayed on restart); STALE ends the
run instead of looping to rule check; synthetic data stands in for
historical replay.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:55:55 -04:00
import { LlmUnavailable } from './llm.js'
import type { Runtime } from './runtime.js'
import { ENVELOPE_GATE_STEP_ID } from './workflows/proposal-lifecycle.js'
import type { ResumeDecision } from './workflows/proposal-lifecycle.js'
/ * *
* Trigger service ( docs / 02 § 1 , docs / 09 § 1 ) : three sources , one shape —
* "instantiate template X with params Y" .
* scheduled : cron entries ( D - 1 06 :00 situation , 08 :00 bid ) — no LLM in the path
* event : bus consumers with deterministic rules
* manual : the router agent turns free text into a RouterDecision
* ( a template id + params , never a new flow ) ; on LLM outage the
* operator picks the template explicitly .
* Approval is a resume on the lifecycle run ( docs / 09 § 2 ) .
* /
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00
export type ScheduledTemplate = 'dayAheadSituation' | 'dayAheadBid' | 'shadowClearing' | 'shadowClose'
M3: Mastra runtime, safety chain, two agents, Case Desk v1
- packages/domain: safety-chain objects (ValidationResult, SimulationResult,
EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft,
RouterDecision, BidExportFile) + fixtures on both sides.
- packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack
(digest-valid, bid-format, price-limits, quantity-non-negative,
ledger-consistency, lineage-integrity, originator-permission — each with
pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits,
revoke, gateway validate); FileExportGateway (idempotent receipts);
Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario
simulator; CaseDeskService; FsRepository; skill HTTP client moved here.
- packages/runtime: createRuntime (LibSQL storage, per-runtime workflow
factories), proposal-lifecycle (rule check → simulation → envelope gate
with suspend/resume → fresh check + permit → release), day-ahead-situation,
day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort
(Mastra/Scripted/Null), Case Desk HTTP API, dev entry point.
- Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with
LLM down, restart survival of a suspended approval, permit expiry and
revocation, replay of a released proposal, trigger scheduling. 141 TS +
60 Python tests.
- Known gaps: ledger not yet persisted (replayed on restart); STALE ends the
run instead of looping to rule check; synthetic data stands in for
historical replay.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:55:55 -04:00
export interface ScheduleEntry {
id : string
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00
/** 'HH:MM' Asia/Shanghai */
M3: Mastra runtime, safety chain, two agents, Case Desk v1
- packages/domain: safety-chain objects (ValidationResult, SimulationResult,
EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft,
RouterDecision, BidExportFile) + fixtures on both sides.
- packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack
(digest-valid, bid-format, price-limits, quantity-non-negative,
ledger-consistency, lineage-integrity, originator-permission — each with
pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits,
revoke, gateway validate); FileExportGateway (idempotent receipts);
Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario
simulator; CaseDeskService; FsRepository; skill HTTP client moved here.
- packages/runtime: createRuntime (LibSQL storage, per-runtime workflow
factories), proposal-lifecycle (rule check → simulation → envelope gate
with suspend/resume → fresh check + permit → release), day-ahead-situation,
day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort
(Mastra/Scripted/Null), Case Desk HTTP API, dev entry point.
- Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with
LLM down, restart survival of a suspended approval, permit expiry and
revocation, replay of a released proposal, trigger scheduling. 141 TS +
60 Python tests.
- Known gaps: ledger not yet persisted (replayed on restart); STALE ends the
run instead of looping to rule check; synthetic data stands in for
historical replay.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:55:55 -04:00
localTime : string
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00
workflow : ScheduledTemplate
/** Market date relative to the local day the entry fires on: +1 = tomorrow (D-1 stages), -1 = yesterday (D+1 close). */
marketDateOffsetDays? : 1 | - 1
M3: Mastra runtime, safety chain, two agents, Case Desk v1
- packages/domain: safety-chain objects (ValidationResult, SimulationResult,
EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft,
RouterDecision, BidExportFile) + fixtures on both sides.
- packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack
(digest-valid, bid-format, price-limits, quantity-non-negative,
ledger-consistency, lineage-integrity, originator-permission — each with
pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits,
revoke, gateway validate); FileExportGateway (idempotent receipts);
Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario
simulator; CaseDeskService; FsRepository; skill HTTP client moved here.
- packages/runtime: createRuntime (LibSQL storage, per-runtime workflow
factories), proposal-lifecycle (rule check → simulation → envelope gate
with suspend/resume → fresh check + permit → release), day-ahead-situation,
day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort
(Mastra/Scripted/Null), Case Desk HTTP API, dev entry point.
- Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with
LLM down, restart survival of a suspended approval, permit expiry and
revocation, replay of a released proposal, trigger scheduling. 141 TS +
60 Python tests.
- Known gaps: ledger not yet persisted (replayed on restart); STALE ends the
run instead of looping to rule check; synthetic data stands in for
historical replay.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:55:55 -04:00
}
export const DEFAULT_SCHEDULE : ScheduleEntry [ ] = [
{ id : 'situation-0600' , localTime : '06:00' , workflow : 'dayAheadSituation' } , // docs/07 D-1 06:00
{ id : 'bid-0800' , localTime : '08:00' , workflow : 'dayAheadBid' } , // docs/07 D-1 08:00; window is OPEN-QUESTION A1
]
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00
/** Shadow mode adds the simulated market answer (docs/07 D-1 16:00) and the D+1 close (execution → metering → review → scoring). */
export const SHADOW_SCHEDULE : ScheduleEntry [ ] = [
. . . DEFAULT_SCHEDULE ,
{ id : 'shadow-clearing-1600' , localTime : '16:00' , workflow : 'shadowClearing' } , // docs/07 D-1 16:00; publication time is OPEN-QUESTION A1
{ id : 'shadow-close-0200' , localTime : '02:00' , workflow : 'shadowClose' , marketDateOffsetDays : - 1 } , // D+1 02:00 for D
]
export const marketDateFor = ( nowIso : string , offsetDays : number ) : string = > {
M3: Mastra runtime, safety chain, two agents, Case Desk v1
- packages/domain: safety-chain objects (ValidationResult, SimulationResult,
EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft,
RouterDecision, BidExportFile) + fixtures on both sides.
- packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack
(digest-valid, bid-format, price-limits, quantity-non-negative,
ledger-consistency, lineage-integrity, originator-permission — each with
pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits,
revoke, gateway validate); FileExportGateway (idempotent receipts);
Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario
simulator; CaseDeskService; FsRepository; skill HTTP client moved here.
- packages/runtime: createRuntime (LibSQL storage, per-runtime workflow
factories), proposal-lifecycle (rule check → simulation → envelope gate
with suspend/resume → fresh check + permit → release), day-ahead-situation,
day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort
(Mastra/Scripted/Null), Case Desk HTTP API, dev entry point.
- Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with
LLM down, restart survival of a suspended approval, permit expiry and
revocation, replay of a released proposal, trigger scheduling. 141 TS +
60 Python tests.
- Known gaps: ledger not yet persisted (replayed on restart); STALE ends the
run instead of looping to rule check; synthetic data stands in for
historical replay.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:55:55 -04:00
const shanghai = new Date ( new Date ( nowIso ) . getTime ( ) + 8 * 3 _600_000 )
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00
shanghai . setUTCDate ( shanghai . getUTCDate ( ) + offsetDays )
M3: Mastra runtime, safety chain, two agents, Case Desk v1
- packages/domain: safety-chain objects (ValidationResult, SimulationResult,
EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft,
RouterDecision, BidExportFile) + fixtures on both sides.
- packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack
(digest-valid, bid-format, price-limits, quantity-non-negative,
ledger-consistency, lineage-integrity, originator-permission — each with
pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits,
revoke, gateway validate); FileExportGateway (idempotent receipts);
Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario
simulator; CaseDeskService; FsRepository; skill HTTP client moved here.
- packages/runtime: createRuntime (LibSQL storage, per-runtime workflow
factories), proposal-lifecycle (rule check → simulation → envelope gate
with suspend/resume → fresh check + permit → release), day-ahead-situation,
day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort
(Mastra/Scripted/Null), Case Desk HTTP API, dev entry point.
- Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with
LLM down, restart survival of a suspended approval, permit expiry and
revocation, replay of a released proposal, trigger scheduling. 141 TS +
60 Python tests.
- Known gaps: ledger not yet persisted (replayed on restart); STALE ends the
run instead of looping to rule check; synthetic data stands in for
historical replay.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:55:55 -04:00
return shanghai . toISOString ( ) . slice ( 0 , 10 )
}
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00
export const nextMarketDate = ( nowIso : string ) : string = > marketDateFor ( nowIso , 1 )
/** Live-data feed (docs/06 §3): actual curves for a market date; each goes through the quality gate. */
export const MarketDataInput = z . object ( {
market_date : MarketDate ,
load_mw : z.array ( z . string ( ) ) . length ( 96 ) . optional ( ) ,
pv_mw : z.array ( z . string ( ) ) . length ( 96 ) . optional ( ) ,
price_yuan_per_mwh : z.array ( z . string ( ) ) . length ( 96 ) . optional ( ) ,
source : z.string ( ) . min ( 1 ) . optional ( ) ,
} )
export type MarketDataInput = z . infer < typeof MarketDataInput >
M3: Mastra runtime, safety chain, two agents, Case Desk v1
- packages/domain: safety-chain objects (ValidationResult, SimulationResult,
EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft,
RouterDecision, BidExportFile) + fixtures on both sides.
- packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack
(digest-valid, bid-format, price-limits, quantity-non-negative,
ledger-consistency, lineage-integrity, originator-permission — each with
pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits,
revoke, gateway validate); FileExportGateway (idempotent receipts);
Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario
simulator; CaseDeskService; FsRepository; skill HTTP client moved here.
- packages/runtime: createRuntime (LibSQL storage, per-runtime workflow
factories), proposal-lifecycle (rule check → simulation → envelope gate
with suspend/resume → fresh check + permit → release), day-ahead-situation,
day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort
(Mastra/Scripted/Null), Case Desk HTTP API, dev entry point.
- Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with
LLM down, restart survival of a suspended approval, permit expiry and
revocation, replay of a released proposal, trigger scheduling. 141 TS +
60 Python tests.
- Known gaps: ledger not yet persisted (replayed on restart); STALE ends the
run instead of looping to rule check; synthetic data stands in for
historical replay.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:55:55 -04:00
export class TriggerService {
private timer : NodeJS.Timeout | null = null
private readonly firedToday = new Set < string > ( )
private readonly unsubscribe : Array < ( ) = > void > = [ ]
constructor (
private readonly rt : Runtime ,
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00
private readonly schedule : ScheduleEntry [ ] = rt . ctx . config . mode === 'SHADOW' ? SHADOW_SCHEDULE : DEFAULT_SCHEDULE ,
M3: Mastra runtime, safety chain, two agents, Case Desk v1
- packages/domain: safety-chain objects (ValidationResult, SimulationResult,
EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft,
RouterDecision, BidExportFile) + fixtures on both sides.
- packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack
(digest-valid, bid-format, price-limits, quantity-non-negative,
ledger-consistency, lineage-integrity, originator-permission — each with
pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits,
revoke, gateway validate); FileExportGateway (idempotent receipts);
Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario
simulator; CaseDeskService; FsRepository; skill HTTP client moved here.
- packages/runtime: createRuntime (LibSQL storage, per-runtime workflow
factories), proposal-lifecycle (rule check → simulation → envelope gate
with suspend/resume → fresh check + permit → release), day-ahead-situation,
day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort
(Mastra/Scripted/Null), Case Desk HTTP API, dev entry point.
- Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with
LLM down, restart survival of a suspended approval, permit expiry and
revocation, replay of a released proposal, trigger scheduling. 141 TS +
60 Python tests.
- Known gaps: ledger not yet persisted (replayed on restart); STALE ends the
run instead of looping to rule check; synthetic data stands in for
historical replay.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:55:55 -04:00
) { }
// ---- scheduled ---------------------------------------------------------
/** Polls the clock; fires each entry once per local day. Deterministic, LLM-free. */
startScheduler ( pollMs = 30 _000 ) : void {
if ( this . timer ) return
this . timer = setInterval ( ( ) = > void this . tick ( ) , pollMs )
this . timer . unref ( )
}
stop ( ) : void {
if ( this . timer ) clearInterval ( this . timer )
this . timer = null
for ( const u of this . unsubscribe ) u ( )
}
async tick ( nowIso = this . rt . ctx . clock ( ) ) : Promise < string [ ] > {
const local = new Date ( new Date ( nowIso ) . getTime ( ) + 8 * 3 _600_000 )
const hhmm = local . toISOString ( ) . slice ( 11 , 16 )
const day = local . toISOString ( ) . slice ( 0 , 10 )
const fired : string [ ] = [ ]
for ( const entry of this . schedule ) {
const key = ` ${ day } : ${ entry . id } `
if ( hhmm >= entry . localTime && ! this . firedToday . has ( key ) ) {
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00
const marketDate = marketDateFor ( nowIso , entry . marketDateOffsetDays ? ? 1 )
try {
await this . fire ( entry , marketDate )
this . firedToday . add ( key )
fired . push ( entry . id )
} catch ( e ) {
// Late data (e.g. no clearing price yet) is normal in a live shadow run: record it and retry next tick.
this . rt . ctx . events . append ( { event_type : 'ScheduledTriggerFailed' , payload : { entry : entry.id , workflow : entry.workflow , market_date : marketDate , error : ( e as Error ) . message } , correlation_id : ` schedule- ${ entry . id } ` } )
}
M3: Mastra runtime, safety chain, two agents, Case Desk v1
- packages/domain: safety-chain objects (ValidationResult, SimulationResult,
EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft,
RouterDecision, BidExportFile) + fixtures on both sides.
- packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack
(digest-valid, bid-format, price-limits, quantity-non-negative,
ledger-consistency, lineage-integrity, originator-permission — each with
pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits,
revoke, gateway validate); FileExportGateway (idempotent receipts);
Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario
simulator; CaseDeskService; FsRepository; skill HTTP client moved here.
- packages/runtime: createRuntime (LibSQL storage, per-runtime workflow
factories), proposal-lifecycle (rule check → simulation → envelope gate
with suspend/resume → fresh check + permit → release), day-ahead-situation,
day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort
(Mastra/Scripted/Null), Case Desk HTTP API, dev entry point.
- Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with
LLM down, restart survival of a suspended approval, permit expiry and
revocation, replay of a released proposal, trigger scheduling. 141 TS +
60 Python tests.
- Known gaps: ledger not yet persisted (replayed on restart); STALE ends the
run instead of looping to rule check; synthetic data stands in for
historical replay.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:55:55 -04:00
}
}
return fired
}
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00
private async fire ( entry : ScheduleEntry , marketDate : string ) : Promise < void > {
if ( entry . workflow === 'shadowClearing' ) await this . shadowClearing ( marketDate )
else if ( entry . workflow === 'shadowClose' ) await this . shadowClose ( marketDate )
else await this . startTemplate ( entry . workflow , { market_date : marketDate } , 'SCHEDULED' , entry . id )
}
M3: Mastra runtime, safety chain, two agents, Case Desk v1
- packages/domain: safety-chain objects (ValidationResult, SimulationResult,
EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft,
RouterDecision, BidExportFile) + fixtures on both sides.
- packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack
(digest-valid, bid-format, price-limits, quantity-non-negative,
ledger-consistency, lineage-integrity, originator-permission — each with
pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits,
revoke, gateway validate); FileExportGateway (idempotent receipts);
Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario
simulator; CaseDeskService; FsRepository; skill HTTP client moved here.
- packages/runtime: createRuntime (LibSQL storage, per-runtime workflow
factories), proposal-lifecycle (rule check → simulation → envelope gate
with suspend/resume → fresh check + permit → release), day-ahead-situation,
day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort
(Mastra/Scripted/Null), Case Desk HTTP API, dev entry point.
- Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with
LLM down, restart survival of a suspended approval, permit expiry and
revocation, replay of a released proposal, trigger scheduling. 141 TS +
60 Python tests.
- Known gaps: ledger not yet persisted (replayed on restart); STALE ends the
run instead of looping to rule check; synthetic data stands in for
historical replay.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:55:55 -04:00
// ---- event -------------------------------------------------------------
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00
/ * *
* Deterministic event rules . A SituationPublished with EXTREME risk starts the
* abnormal - day protocol ( docs / 13 § 1 ) : envelopes are off for that market date —
* every proposal goes to a human — and a case is opened for the review .
* /
M3: Mastra runtime, safety chain, two agents, Case Desk v1
- packages/domain: safety-chain objects (ValidationResult, SimulationResult,
EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft,
RouterDecision, BidExportFile) + fixtures on both sides.
- packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack
(digest-valid, bid-format, price-limits, quantity-non-negative,
ledger-consistency, lineage-integrity, originator-permission — each with
pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits,
revoke, gateway validate); FileExportGateway (idempotent receipts);
Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario
simulator; CaseDeskService; FsRepository; skill HTTP client moved here.
- packages/runtime: createRuntime (LibSQL storage, per-runtime workflow
factories), proposal-lifecycle (rule check → simulation → envelope gate
with suspend/resume → fresh check + permit → release), day-ahead-situation,
day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort
(Mastra/Scripted/Null), Case Desk HTTP API, dev entry point.
- Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with
LLM down, restart survival of a suspended approval, permit expiry and
revocation, replay of a released proposal, trigger scheduling. 141 TS +
60 Python tests.
- Known gaps: ledger not yet persisted (replayed on restart); STALE ends the
run instead of looping to rule check; synthetic data stands in for
historical replay.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:55:55 -04:00
startEventConsumers ( ) : void {
this . unsubscribe . push (
this . rt . ctx . events . subscribe ( 'SituationPublished' , ( evt : EventEnvelope ) = > {
const payload = evt . payload as { report : { market_date : string ; risk_level : string ; id : string } }
if ( payload . report . risk_level === 'EXTREME' ) {
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00
this . rt . ctx . breakers . declareAbnormalDay ( payload . report . market_date , { id : 'runtime' , role : 'system' } , ` situation ${ payload . report . id } : risk EXTREME (OPEN-QUESTION B7 triggers) ` , this . rt . ctx . clock ( ) )
M3: Mastra runtime, safety chain, two agents, Case Desk v1
- packages/domain: safety-chain objects (ValidationResult, SimulationResult,
EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft,
RouterDecision, BidExportFile) + fixtures on both sides.
- packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack
(digest-valid, bid-format, price-limits, quantity-non-negative,
ledger-consistency, lineage-integrity, originator-permission — each with
pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits,
revoke, gateway validate); FileExportGateway (idempotent receipts);
Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario
simulator; CaseDeskService; FsRepository; skill HTTP client moved here.
- packages/runtime: createRuntime (LibSQL storage, per-runtime workflow
factories), proposal-lifecycle (rule check → simulation → envelope gate
with suspend/resume → fresh check + permit → release), day-ahead-situation,
day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort
(Mastra/Scripted/Null), Case Desk HTTP API, dev entry point.
- Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with
LLM down, restart survival of a suspended approval, permit expiry and
revocation, replay of a released proposal, trigger scheduling. 141 TS +
60 Python tests.
- Known gaps: ledger not yet persisted (replayed on restart); STALE ends the
run instead of looping to rule check; synthetic data stands in for
historical replay.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:55:55 -04:00
this . rt . ctx . caseDesk . open ( {
id : this.rt.ctx.newId ( 'case' ) ,
kind : 'ADHOC_ANALYSIS' ,
objective : ` Abnormal-day protocol review for ${ payload . report . market_date } (situation ${ payload . report . id } ) ` ,
owner : 'user-ops-lead' , // OPEN-QUESTION B8
deadline : ` ${ payload . report . market_date } T00:00:00Z ` ,
} )
}
} ) ,
)
}
M4: resource agent, envelopes live, review loop, insight cards
- packages/domain: AwardNotice, DISPATCH_PLAN proposal payload, ExecutionReport,
MeteringRecord, potential-assessment and dispatch-optimization contracts,
ReviewFinding (+ typed writebacks), SemanticMemoryEntry,
EnvelopeChangeRequest, InsightCard — exported with fixtures on both sides.
- skills-py: potential-assessment (certified × rolling fulfilment, evidence
days) and dispatch-optimization (per-interval LP on HiGHS, shortfall
reported) skills + routes + tests.
- packages/services: dispatch rules in the policy pack (over-allocation,
award anchor, lineage integrity for allocations); PowerBalanceSimulator;
SimulationGateway (permit-only, idempotent, seeded execute → ExecutionReports);
envelope deviation-streak suspension + apply(); ReviewService (attribution,
reliability EWMA writeback, semantic memory, envelope recommendations as
change requests); dispatch assembler; skill client methods.
- packages/runtime: resource agent; award-decomposition, review and
envelope-review workflows; lifecycle selects simulator/gateway by proposal
type; trigger hooks for awards, execution reports, metering; decide()
resumes either lifecycle or envelope-review runs; insight cards API.
- Tests: docs/07 D-1 16:00 and D+1 end to end; reliability score 0.9 → 0.880
and the next assessment de-rates capacity; envelope suspension on a seeded
3-day streak; WIDEN request applied only by a human. 184 TS + 80 Python.
- docs/open-questions: B10 (reliability/potential parameters). README and
CLAUDE.md status → M4 done, M5 next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 19:31:12 -04:00
/** Trading-platform adapter stand-in: an award arrives → ledger + decomposition (docs/07 D-1 16:00). */
async onAward ( award : AwardNotice ) {
const a = AwardNotice . parse ( award )
this . rt . ctx . events . append ( { event_type : 'AwardReceived' , payload : a , correlation_id : a.id } )
const wf = this . rt . mastra . getWorkflow ( 'awardDecomposition' )
const run = await wf . createRun ( )
const result = await run . start ( { inputData : { award : a } } )
if ( result . status === 'failed' ) throw result . error
return { run_id : run.runId , result : result.status === 'success' ? result.result : null , status : result.status }
}
/** Execution engine / edge feedback stand-in: execution reports are evidence for review and potential assessment. */
onExecutionReports ( reports : ExecutionReport [ ] ) {
for ( const r of reports ) {
const parsed = ExecutionReport . parse ( r )
this . rt . ctx . executionReports . put ( parsed )
this . rt . ctx . events . append ( { event_type : 'ExecutionReport' , payload : parsed , correlation_id : parsed.dispatch_proposal_digest } )
}
}
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
2026-09-02 22:54:14 -04:00
/** Live-data feed: actual curves for a market date through the ingestion pipeline (quality gate + evidence snapshot). */
onMarketData ( input : MarketDataInput ) {
const m = MarketDataInput . parse ( input )
const source = m . source ? ? 'market-data-feed'
const series : Array < [ string , string [ ] | undefined ] > = [ [ 'load:aggregate' , m . load_mw ] , [ 'pv:aggregate' , m . pv_mw ] , [ 'price:da' , m . price_yuan_per_mwh ] ]
const accepted : string [ ] = [ ]
const quarantined : Array < { series : string ; issues : string [ ] } > = [ ]
for ( const [ id , values ] of series ) {
if ( ! values ) continue
const curve = Curve96 . parse ( { interval_minutes : 15 , date : m.market_date , values } )
const r = this . rt . ctx . ingestion . ingestCurve ( id , curve , source )
if ( r . accepted ) accepted . push ( id )
else quarantined . push ( { series : id , issues : r.issues } )
}
this . rt . ctx . events . append ( { event_type : 'MarketDataIngested' , payload : { market_date : m.market_date , source , accepted , quarantined } , correlation_id : ` market-data- ${ m . market_date } ` } )
return { market_date : m.market_date , accepted , quarantined }
}
/** The human trader's actual submission for a market date (shadow comparison baseline). */
onHumanBid ( record : HumanBidRecord ) {
const h = HumanBidRecord . parse ( record )
this . rt . ctx . humanBids . put ( h )
this . rt . ctx . events . append ( { event_type : 'HumanBidRecorded' , payload : { id : h.id , market_date : h.market_date , source : h.source } , correlation_id : ` shadow- ${ h . market_date } ` } )
return h
}
/ * *
* Shadow market answer ( docs / 07 D - 1 16 :00 in shadow mode ) : the released
* shadow bid is cleared against the actual day - ahead price and the resulting
* AwardNotice goes down the normal award path . Idempotent per bid digest .
* /
async shadowClearing ( marketDate : string ) {
const ctx = this . rt . ctx
const bid = ctx . caseDesk . proposals
. list ( )
. map ( ( r ) = > r . value )
. filter ( ( p ) = > p . type === 'BID' && p . payload . market_date === marketDate && p . status === 'RELEASED' )
. sort ( ( a , b ) = > ( a . created_at < b . created_at ? - 1 : 1 ) )
. at ( - 1 )
if ( ! bid ) {
ctx . events . append ( { event_type : 'ShadowClearingSkipped' , payload : { market_date : marketDate , reason : 'no RELEASED bid for the day' } , correlation_id : ` shadow- ${ marketDate } ` } )
return null
}
const existing = ctx . awards . list ( ) . find ( ( a ) = > a . value . bid_proposal_digest === bid . digest )
if ( existing ) return { award : existing.value , decomposition : null }
const price = ctx . timeseries . latest ( 'price:da' , marketDate ) ? . curve
if ( ! price ) throw new Error ( ` no clearing price for ${ marketDate } yet — ingest market data first ` )
const award = shadowClearing ( bid , price , { id : ` award-shadow- ${ marketDate } ` , now : ctx.clock ( ) } )
ctx . events . append ( { event_type : 'ShadowCleared' , payload : { market_date : marketDate , award_id : award.id , bid_proposal_digest : bid.digest , awarded_mwh : award.awarded_mwh.values.reduce ( ( s , v ) = > s + Number ( v ) , 0 ) . toFixed ( 3 ) } , correlation_id : ` shadow- ${ marketDate } ` } )
const decomposition = await this . onAward ( award )
return { award , decomposition }
}
/** D+1 close for a shadow day: simulated execution → metering → review → comparison → KPI (shadow-close workflow). */
async shadowClose ( marketDate : string ) : Promise < { run_id : string ; record : ShadowDayRecord ; kpi : KpiReport ; review_ran : boolean ; execution_simulated : boolean } > {
const wf = this . rt . mastra . getWorkflow ( 'shadowClose' )
const run = await wf . createRun ( )
const result = await run . start ( { inputData : { market_date : marketDate } } )
if ( result . status === 'failed' ) throw result . error
if ( result . status !== 'success' ) throw new Error ( ` shadow close for ${ marketDate } ended ${ result . status } ` )
return { run_id : run.runId , . . . result . result }
}
M4: resource agent, envelopes live, review loop, insight cards
- packages/domain: AwardNotice, DISPATCH_PLAN proposal payload, ExecutionReport,
MeteringRecord, potential-assessment and dispatch-optimization contracts,
ReviewFinding (+ typed writebacks), SemanticMemoryEntry,
EnvelopeChangeRequest, InsightCard — exported with fixtures on both sides.
- skills-py: potential-assessment (certified × rolling fulfilment, evidence
days) and dispatch-optimization (per-interval LP on HiGHS, shortfall
reported) skills + routes + tests.
- packages/services: dispatch rules in the policy pack (over-allocation,
award anchor, lineage integrity for allocations); PowerBalanceSimulator;
SimulationGateway (permit-only, idempotent, seeded execute → ExecutionReports);
envelope deviation-streak suspension + apply(); ReviewService (attribution,
reliability EWMA writeback, semantic memory, envelope recommendations as
change requests); dispatch assembler; skill client methods.
- packages/runtime: resource agent; award-decomposition, review and
envelope-review workflows; lifecycle selects simulator/gateway by proposal
type; trigger hooks for awards, execution reports, metering; decide()
resumes either lifecycle or envelope-review runs; insight cards API.
- Tests: docs/07 D-1 16:00 and D+1 end to end; reliability score 0.9 → 0.880
and the next assessment de-rates capacity; envelope suspension on a seeded
3-day streak; WIDEN request applied only by a human. 184 TS + 80 Python.
- docs/open-questions: B10 (reliability/potential parameters). README and
CLAUDE.md status → M4 done, M5 next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 19:31:12 -04:00
/** Metering adapter stand-in: D+1 metering arrives → review workflow (docs/07 D+1). */
async onMetering ( record : MeteringRecord ) {
const m = MeteringRecord . parse ( record )
this . rt . ctx . metering . put ( m )
this . rt . ctx . events . append ( { event_type : 'MeteringArrived' , payload : m , correlation_id : m.id } )
const wf = this . rt . mastra . getWorkflow ( 'review' )
const run = await wf . createRun ( )
const result = await run . start ( { inputData : { market_date : m.market_date } } )
if ( result . status === 'failed' ) throw result . error
return { run_id : run.runId , result : result.status === 'success' ? result.result : null }
}
M3: Mastra runtime, safety chain, two agents, Case Desk v1
- packages/domain: safety-chain objects (ValidationResult, SimulationResult,
EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft,
RouterDecision, BidExportFile) + fixtures on both sides.
- packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack
(digest-valid, bid-format, price-limits, quantity-non-negative,
ledger-consistency, lineage-integrity, originator-permission — each with
pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits,
revoke, gateway validate); FileExportGateway (idempotent receipts);
Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario
simulator; CaseDeskService; FsRepository; skill HTTP client moved here.
- packages/runtime: createRuntime (LibSQL storage, per-runtime workflow
factories), proposal-lifecycle (rule check → simulation → envelope gate
with suspend/resume → fresh check + permit → release), day-ahead-situation,
day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort
(Mastra/Scripted/Null), Case Desk HTTP API, dev entry point.
- Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with
LLM down, restart survival of a suspended approval, permit expiry and
revocation, replay of a released proposal, trigger scheduling. 141 TS +
60 Python tests.
- Known gaps: ledger not yet persisted (replayed on restart); STALE ends the
run instead of looping to rule check; synthetic data stands in for
historical replay.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:55:55 -04:00
// ---- manual ------------------------------------------------------------
async route ( message : string ) : Promise < RouterDecision > {
const prompt = ` Operator request: " ${ message } ". Today (UTC) is ${ this . rt . ctx . clock ( ) } . Default market_date is ${ nextMarketDate ( this . rt . ctx . clock ( ) ) } . `
try {
return await this . rt . ctx . llm . structured ( 'router-agent' , prompt , RouterDecision )
} catch ( e ) {
if ( e instanceof LlmUnavailable ) throw new LlmUnavailable ( 'router needs an LLM; pick a template explicitly via startTemplate' )
throw e
}
}
async manual ( message : string ) : Promise < { decision : RouterDecision ; run_id : string | null } > {
const decision = await this . route ( message )
if ( decision . workflow_id === 'adhoc-analysis' ) return { decision , run_id : null }
const run_id = await this . startTemplate ( decision . workflow_id === 'day-ahead-bid' ? 'dayAheadBid' : 'dayAheadSituation' , decision . params , 'MANUAL' , 'router' )
return { decision , run_id }
}
async startTemplate ( workflow : 'dayAheadSituation' | 'dayAheadBid' , params : { market_date : string } , trigger : 'SCHEDULED' | 'EVENT' | 'MANUAL' , source : string ) : Promise < string > {
const wf = this . rt . mastra . getWorkflow ( workflow )
const run = await wf . createRun ( )
this . rt . ctx . events . append ( { event_type : 'WorkflowTriggered' , payload : { workflow , params , trigger , source , run_id : run.runId } , correlation_id : run.runId } )
const result = await run . start ( { inputData : params } )
this . rt . ctx . events . append ( { event_type : 'WorkflowFinished' , payload : { workflow , run_id : run.runId , status : result.status } , correlation_id : run.runId } )
if ( result . status === 'failed' ) throw result . error
return run . runId
}
// ---- approval = resume -------------------------------------------------
async decide ( runId : string , decision : ResumeDecision ) {
M4: resource agent, envelopes live, review loop, insight cards
- packages/domain: AwardNotice, DISPATCH_PLAN proposal payload, ExecutionReport,
MeteringRecord, potential-assessment and dispatch-optimization contracts,
ReviewFinding (+ typed writebacks), SemanticMemoryEntry,
EnvelopeChangeRequest, InsightCard — exported with fixtures on both sides.
- skills-py: potential-assessment (certified × rolling fulfilment, evidence
days) and dispatch-optimization (per-interval LP on HiGHS, shortfall
reported) skills + routes + tests.
- packages/services: dispatch rules in the policy pack (over-allocation,
award anchor, lineage integrity for allocations); PowerBalanceSimulator;
SimulationGateway (permit-only, idempotent, seeded execute → ExecutionReports);
envelope deviation-streak suspension + apply(); ReviewService (attribution,
reliability EWMA writeback, semantic memory, envelope recommendations as
change requests); dispatch assembler; skill client methods.
- packages/runtime: resource agent; award-decomposition, review and
envelope-review workflows; lifecycle selects simulator/gateway by proposal
type; trigger hooks for awards, execution reports, metering; decide()
resumes either lifecycle or envelope-review runs; insight cards API.
- Tests: docs/07 D-1 16:00 and D+1 end to end; reliability score 0.9 → 0.880
and the next assessment de-rates capacity; envelope suspension on a seeded
3-day streak; WIDEN request applied only by a human. 184 TS + 80 Python.
- docs/open-questions: B10 (reliability/potential parameters). README and
CLAUDE.md status → M4 done, M5 next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 19:31:12 -04:00
const pending = this . rt . ctx . caseDesk . pending . list ( ) . find ( ( p ) = > p . value . run_id === runId ) ? . value
if ( pending ? . workflow === 'envelopeReview' ) {
const run = await this . rt . mastra . getWorkflow ( 'envelopeReview' ) . createRun ( { runId } )
const res = await run . resume ( { step : ENVELOPE_REVIEW_STEP_ID , resumeData : decision } )
return res . status === 'success' ? { status : 'success' as const , result : { outcome : res.result.applied ? ( 'RELEASED' as const ) : ( 'REJECTED' as const ) , . . . res . result } } : res
}
M3: Mastra runtime, safety chain, two agents, Case Desk v1
- packages/domain: safety-chain objects (ValidationResult, SimulationResult,
EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft,
RouterDecision, BidExportFile) + fixtures on both sides.
- packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack
(digest-valid, bid-format, price-limits, quantity-non-negative,
ledger-consistency, lineage-integrity, originator-permission — each with
pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits,
revoke, gateway validate); FileExportGateway (idempotent receipts);
Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario
simulator; CaseDeskService; FsRepository; skill HTTP client moved here.
- packages/runtime: createRuntime (LibSQL storage, per-runtime workflow
factories), proposal-lifecycle (rule check → simulation → envelope gate
with suspend/resume → fresh check + permit → release), day-ahead-situation,
day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort
(Mastra/Scripted/Null), Case Desk HTTP API, dev entry point.
- Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with
LLM down, restart survival of a suspended approval, permit expiry and
revocation, replay of a released proposal, trigger scheduling. 141 TS +
60 Python tests.
- Known gaps: ledger not yet persisted (replayed on restart); STALE ends the
run instead of looping to rule check; synthetic data stands in for
historical replay.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:55:55 -04:00
const wf = this . rt . mastra . getWorkflow ( 'proposalLifecycle' )
const run = await wf . createRun ( { runId } )
return run . resume ( { step : ENVELOPE_GATE_STEP_ID , resumeData : decision } )
}
}