vpp-ai-platform/ROADMAP.md
stewart hu 80835138e9 Scaffold repo for implementation handoff
- README: orientation, doc map, target package layout (pointers only,
  no duplicated architecture content)
- CLAUDE.md: agent operating manual — invariants as code-review rules,
  conventions, do-not list, task reading order
- ROADMAP: M1-M5 with verifiable acceptance criteria, phase-2 fence
- GLOSSARY: canonical Chinese-term → code-name mapping
- docs/adr/: eight ADRs recording settled decisions and rejected
  alternatives
- docs/open-questions.md: consolidated TODO(业务) tracker by owner and
  blocking milestone
- .gitignore; untrack .DS_Store

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019u5SLNweVio6ozJX7yfxQr

🔮 View transcript: https://logs.lojong.info/s/e8u90k3t33w590r7b5y7yzqh
2026-09-01 21:13:00 -04:00

4.4 KiB
Raw Permalink Blame History

ROADMAP

Phase 1 (系统研发与省内能力落地, 2025.12–2026.5) as five milestones. Each has acceptance criteria a coding agent can verify. Sequencing rationale: M1–M3 need no external-party scheduling; the demo-able bidding loop lands earliest (docs/08 §4). Phase 2 items are listed but not to be built yet.

M1 — Data foundation & contracts

Deliverables:

  • packages/domain: zod schemas for the 07-scenario object set — Proposal (+digest), Approval, ExecutionPermit, Envelope, PositionLedger/PositionUpdate, ForecastBundle, SituationReport, ResourceProfile, DecisionCase, EventEnvelope.
  • contracts/: JSON Schema export pipeline + golden fixtures; pydantic codegen wired for skills-py; dual-side contract tests in CI.
  • Ingestion pipeline skeleton: time-series store + relational store + snapshot store (immutable, checksummed); data-quality gate stub.
  • Ledger service v1: read with version, append, per-timescale views.

Accept when: golden fixtures validate identically in TS and Python CI; a ledger constraint-cascade test passes (monthly position bounds a day-ahead write); snapshots are content-addressed and immutable.

M2 — Skills v1 & eval baseline

Deliverables:

  • skills-py: load forecast, PV forecast, price forecast (all with quantile intervals), bid-optimization MILP, report generator. HTTP/JSON per contracts.
  • packages/evals: harness + L2 baselines (MAPE, interval coverage, backtest revenue vs naive/hindsight bounds) on historical data.

Accept when: forecasts return calibrated quantiles (coverage test); MILP output respects ledger constraints in property tests; eval runs are reproducible and archived; baseline report exists for each skill.

M3 — Runtime, two agents, safety chain, Case Desk

Deliverables:

  • packages/runtime: Mastra instance; trigger service (cron/event/manual); router agent with template-enum output; registerSkill lineage wrapper; outbox event bus.
  • Workflows: day-ahead-situation, day-ahead-bid, proposal-lifecycle (rule check → simulation → envelope gate with suspend/resume → fresh check → permit → release).
  • packages/services: policy engine v1 (+ first Hubei policy pack, from open-questions once confirmed), envelope service, authority service (permits), audit/lineage assembly.
  • Case Desk v1: approval inbox (resume endpoint), case view, lineage expansion.
  • Bid release path = file export for manual upload (degraded channel by design).

Accept when: docs/07 timeline 06:00→08:30 runs end-to-end on historical data; all eight invariant tests pass (incl. LLM-down degraded run and AI-self-approval rejection); a suspended approval survives process restart; permit expiry blocks a late release.

M4 — Resource agent, envelopes live, review loop

Deliverables:

  • Resource dispatch agent + potential-assessment skill + award-decomposition flow.
  • Envelope model end-to-end: bounded auto-approval, deviation-streak suspension, envelope re-approval workflow.
  • Review (复盘) workflow: ledger plan-vs-actual per timescale, attribution, ReviewFinding writebacks (semantic memory, reliability scores, envelope recommendations).
  • AI insight cards embedded in existing business pages (read-only projections).

Accept when: docs/07 D-1 16:00 and D+1 sections run; a ReviewFinding measurably updates a ResourceProfile reliability score; envelope suspension triggers on a seeded deviation streak.

M5 — Shadow run (phase-1 acceptance form)

Deliverables:

  • Full loop on live data, all external effects simulated (bids generated, not submitted; commands to a simulation gateway).
  • Daily shadow-vs-human-vs-hindsight comparison report; KPI dashboard per docs/12 §4 definitions.
  • Loss circuit-breaker, abnormal-day protocol, kill-switch L0–L2 implemented and drilled once.

Accept when: 20+ consecutive shadow days with complete lineage; KPI report auto-generated; one envelope-widening recommendation produced from shadow data (not acted on — that's a phase-2 governance step).

Phase 2 (do not build in phase 1)

Interaction-service & load-control agents; edge control links (real protocols); trading-platform programmatic submission; dispatch/metering live interfaces; federation gateway (build FlexibilityEnvelope schema in M1 — the three no-rework reservations of docs/10 §5 are in scope for phase 1 schemas only); controlled live bidding → gradual envelope widening.