381b6d3521
6 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
381b6d3521 |
M5: shadow run, kill-switch hierarchy L0–L4, KPI dashboard, shadow replay CLI
Shadow run (ROADMAP M5, phase-1 acceptance form): SHADOW runtime mode records
bids without submitting them, simulates the market answer from the actual
day-ahead clearing price, dispatches to the simulation gateway, runs the D+1
review, and scores every day shadow-vs-human-vs-hindsight (ShadowDayRecord)
with a lineage-completeness audit. KpiReport regenerated after every day per
docs/12 §4 definitions (C1/C2 placeholders as named config).
Kill switches (docs/13 §8): BreakerService with L0 permit revocation, L1
envelope suspension, L2 loss breaker (mark-to-market, reduce-only bids),
L3 channel breaker (bids fall back to the file channel, dispatch BLOCKED),
L4 AI-off (templates run, no Proposal created); abnormal-day protocol on
EXTREME situations; per-level authority (B8 placeholder); auditable drill.
Runtime: shadow-close workflow, shadow schedule entries, live-data ingestion
through the quality gate, human-bid ingestion, breaker/KPI/shadow endpoints
and insight cards, replay CLI (npm run shadow). Ledger, time series and
streak counters are file-backed so a multi-week run survives restarts.
Domain: HumanBidRecord, ShadowDayRecord, KpiReport, BreakerRecord, SHADOW
receipt channel; contracts, fixtures and pydantic models regenerated.
Services: L2 metrics moved from evals so the shadow run and the harness
share one implementation.
Fixes: envelope/permit validity compared ISO timestamps as strings
('…00Z' vs '…00.000Z'); L2 baseline was stale since M4 (skill_versions only,
metrics unchanged) — rewritten from the live service.
Docs: docs/14 shadow-run runbook (timeline, breaker trigger/authority/
recovery, KPI definitions as implemented); README and CLAUDE.md status.
Tests: 21 consecutive shadow days with complete lineage, KPI report, WIDEN
recommendation produced but not acted on; restart durability; drill; L2/L3/L4
and abnormal-day paths; API.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017wrZgPL9LoKaD69BpEQU4v
|
||
|
|
23fdf49ff2 |
M4: resource agent, envelopes live, review loop, insight cards
- packages/domain: AwardNotice, DISPATCH_PLAN proposal payload, ExecutionReport, MeteringRecord, potential-assessment and dispatch-optimization contracts, ReviewFinding (+ typed writebacks), SemanticMemoryEntry, EnvelopeChangeRequest, InsightCard — exported with fixtures on both sides. - skills-py: potential-assessment (certified × rolling fulfilment, evidence days) and dispatch-optimization (per-interval LP on HiGHS, shortfall reported) skills + routes + tests. - packages/services: dispatch rules in the policy pack (over-allocation, award anchor, lineage integrity for allocations); PowerBalanceSimulator; SimulationGateway (permit-only, idempotent, seeded execute → ExecutionReports); envelope deviation-streak suspension + apply(); ReviewService (attribution, reliability EWMA writeback, semantic memory, envelope recommendations as change requests); dispatch assembler; skill client methods. - packages/runtime: resource agent; award-decomposition, review and envelope-review workflows; lifecycle selects simulator/gateway by proposal type; trigger hooks for awards, execution reports, metering; decide() resumes either lifecycle or envelope-review runs; insight cards API. - Tests: docs/07 D-1 16:00 and D+1 end to end; reliability score 0.9 → 0.880 and the next assessment de-rates capacity; envelope suspension on a seeded 3-day streak; WIDEN request applied only by a human. 184 TS + 80 Python. - docs/open-questions: B10 (reliability/potential parameters). README and CLAUDE.md status → M4 done, M5 next. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA |
||
|
|
1cc21e0dc6 |
M3: Mastra runtime, safety chain, two agents, Case Desk v1
- packages/domain: safety-chain objects (ValidationResult, SimulationResult, EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft, RouterDecision, BidExportFile) + fixtures on both sides. - packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack (digest-valid, bid-format, price-limits, quantity-non-negative, ledger-consistency, lineage-integrity, originator-permission — each with pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits, revoke, gateway validate); FileExportGateway (idempotent receipts); Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario simulator; CaseDeskService; FsRepository; skill HTTP client moved here. - packages/runtime: createRuntime (LibSQL storage, per-runtime workflow factories), proposal-lifecycle (rule check → simulation → envelope gate with suspend/resume → fresh check + permit → release), day-ahead-situation, day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort (Mastra/Scripted/Null), Case Desk HTTP API, dev entry point. - Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with LLM down, restart survival of a suspended approval, permit expiry and revocation, replay of a released proposal, trigger scheduling. 141 TS + 60 Python tests. - Known gaps: ledger not yet persisted (replayed on restart); STALE ends the run instead of looping to rule check; synthetic data stands in for historical replay. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA |
||
|
|
8796faca63 |
M2: skill contracts, Python skill service, L2 eval harness with baseline
- packages/domain: ForecastRequest, BidOptimizationRequest/Result,
ReportRequest, SkillReport (+ golden and invalid fixtures, exported to
contracts/ and regenerated as pydantic models).
- skills-py/vpp_skills: FastAPI service with versioned registry; load/PV/
price forecasts (same-day-type EWM point forecast, conformal residual
quantiles — coverage test as acceptance gate); bid-optimization MILP on
HiGHS (binary block participation, hard ledger energy bounds, exact
Decimal fit of the rounded curve inside the bounds, revenue distribution
over quantile paths); report generator whose every figure is a
{tool_call_id, path} reference, with a verifier. 48 tests incl. hypothesis
property test that bids respect ledger constraints.
- packages/services: LedgerService.dayAheadBounds (the P7 cascade band
handed to the optimizer); Decimal resolved once for CJS/ESM interop.
- packages/evals: L2 metrics (MAPE, nRMSE, coverage, direction accuracy,
naive/hindsight revenue baselines), HTTP skill client, rolling-origin
harness that pushes each bid through the real ledger, CLI with
--check/--write-baseline; committed baseline on the SYNTHETIC dataset
(no historical Hubei data yet — baselines measure the harness, not KPI).
- CI: evals job boots the skill service and fails on baseline digest drift.
- docs/open-questions: A6 (flexibility marginal cost = offer floor); A4/B6
wired as placeholders. README/CLAUDE.md status → M2 done, M3 next.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
|
||
|
|
e7196bc88a |
M1: close gaps — CI, time-series/relational stores, ingestion skeleton, Python 3.11 pin
- .github/workflows/ci.yml: TS job (typecheck, re-export contracts, drift check, vitest) and Python job (3.11, regenerate pydantic models, drift check, pytest) — the dual-side contract test now actually runs in CI. - packages/services: MemoryTimeSeriesStore (append-only daily-curve revisions), MemoryRepository + ResourceRegistry (schema-validated, optimistic versioning), IngestionPipeline (snapshot → quality gate → store or quarantine). In-memory reference semantics; persistent adapters arrive with M3 like the ledger. - skills-py: .python-version + pyproject requires-python >=3.11; codegen script asserts interpreter version and runs the generator as a module. - typecheck scripts per package and in root `check`; README status and dev setup; CLAUDE.md current-phase note. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA |
||
|
|
f681e134cc |
M1: domain schemas, contracts pipeline, ledger and snapshot services
- packages/domain: zod schemas for the 07-scenario object set with a schemaRegistry driving schema export, fixture generation, and Python module naming; decimal-string/sha256/curve primitives enforce the docs/11 §3.3 representation rules at the type layer - contracts/: 12 JSON Schemas (draft 2020-12), 12 golden fixtures, 3 invalid fixtures crafted to fail on both sides (float money, bad digest, missing concurrency version) - skills-py: generated pydantic models (committed, never hand-edited), regeneration script, mirror pytest using JSON-mode validation — dual-side contract tests agree on all 15 fixtures - packages/services: content-addressed immutable snapshot store (memory + fs, rejects floats), LedgerService v1 with optimistic concurrency and the P7 constraint cascade (monthly position bounds day-ahead bids; no-anchor bids rejected), quality-gate stub - cascade band shape is an OPEN-QUESTION A5 placeholder (pro-rata daily share ±band) — see ledger.ts checkCascade All green: tsc typecheck, 41 TS tests, 15 Python tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019u5SLNweVio6ozJX7yfxQr 🔮 View transcript: https://logs.lojong.info/s/e8u90k3t33w590r7b5y7yzqh |