Commit Graph

2 Commits

Author SHA1 Message Date
Thomas Bayes
1cc21e0dc6 M3: Mastra runtime, safety chain, two agents, Case Desk v1
Some checks are pending
ci / typescript (push) Waiting to run
ci / python (push) Waiting to run
ci / evals (push) Blocked by required conditions
- packages/domain: safety-chain objects (ValidationResult, SimulationResult,
  EnvelopeMatch, StaleDenial, ExecutionReceipt, LineageRef/BidProposalDraft,
  RouterDecision, BidExportFile) + fixtures on both sides.
- packages/services: proposal digest; PolicyEngine + hubei-spot-bidding pack
  (digest-valid, bid-format, price-limits, quantity-non-negative,
  ledger-consistency, lineage-integrity, originator-permission — each with
  pass/fail tests); EnvelopeService; AuthorityService (fresh check, permits,
  revoke, gateway validate); FileExportGateway (idempotent receipts);
  Memory/File EventBus; LineageRecorder + P2 assembler; RevenueScenario
  simulator; CaseDeskService; FsRepository; skill HTTP client moved here.
- packages/runtime: createRuntime (LibSQL storage, per-runtime workflow
  factories), proposal-lifecycle (rule check → simulation → envelope gate
  with suspend/resume → fresh check + permit → release), day-ahead-situation,
  day-ahead-bid, TriggerService (scheduled/event/manual), LlmPort
  (Mastra/Scripted/Null), Case Desk HTTP API, dev entry point.
- Tests: all eight docs/01 invariants, docs/07 06:00→08:30 end to end with
  LLM down, restart survival of a suspended approval, permit expiry and
  revocation, replay of a released proposal, trigger scheduling. 141 TS +
  60 Python tests.
- Known gaps: ledger not yet persisted (replayed on restart); STALE ends the
  run instead of looping to rule check; synthetic data stands in for
  historical replay.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:55:55 -04:00
Thomas Bayes
8796faca63 M2: skill contracts, Python skill service, L2 eval harness with baseline
- packages/domain: ForecastRequest, BidOptimizationRequest/Result,
  ReportRequest, SkillReport (+ golden and invalid fixtures, exported to
  contracts/ and regenerated as pydantic models).
- skills-py/vpp_skills: FastAPI service with versioned registry; load/PV/
  price forecasts (same-day-type EWM point forecast, conformal residual
  quantiles — coverage test as acceptance gate); bid-optimization MILP on
  HiGHS (binary block participation, hard ledger energy bounds, exact
  Decimal fit of the rounded curve inside the bounds, revenue distribution
  over quantile paths); report generator whose every figure is a
  {tool_call_id, path} reference, with a verifier. 48 tests incl. hypothesis
  property test that bids respect ledger constraints.
- packages/services: LedgerService.dayAheadBounds (the P7 cascade band
  handed to the optimizer); Decimal resolved once for CJS/ESM interop.
- packages/evals: L2 metrics (MAPE, nRMSE, coverage, direction accuracy,
  naive/hindsight revenue baselines), HTTP skill client, rolling-origin
  harness that pushes each bid through the real ledger, CLI with
  --check/--write-baseline; committed baseline on the SYNTHETIC dataset
  (no historical Hubei data yet — baselines measure the harness, not KPI).
- CI: evals job boots the skill service and fails on baseline digest drift.
- docs/open-questions: A6 (flexibility marginal cost = offer floor); A4/B6
  wired as placeholders. README/CLAUDE.md status → M2 done, M3 next.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UoYoGYzHkFyv3ALenkRPhA
2026-09-02 06:29:08 -04:00