vpp-ai-platform/ROADMAP.md

92 lines
4.4 KiB
Markdown
Raw Normal View History

# ROADMAP
Phase 1 (系统研发与省内能力落地, 2025.12–2026.5) as five milestones. Each has
acceptance criteria a coding agent can verify. Sequencing rationale: M1–M3 need
no external-party scheduling; the demo-able bidding loop lands earliest
(docs/08 §4). Phase 2 items are listed but **not to be built yet**.
## M1 — Data foundation & contracts
Deliverables:
- `packages/domain`: zod schemas for the 07-scenario object set — `Proposal`
(+digest), `Approval`, `ExecutionPermit`, `Envelope`, `PositionLedger`/`PositionUpdate`,
`ForecastBundle`, `SituationReport`, `ResourceProfile`, `DecisionCase`, `EventEnvelope`.
- `contracts/`: JSON Schema export pipeline + golden fixtures; pydantic codegen
wired for `skills-py`; dual-side contract tests in CI.
- Ingestion pipeline skeleton: time-series store + relational store + snapshot
store (immutable, checksummed); data-quality gate stub.
- Ledger service v1: read with version, append, per-timescale views.
Accept when: golden fixtures validate identically in TS and Python CI; a ledger
constraint-cascade test passes (monthly position bounds a day-ahead write);
snapshots are content-addressed and immutable.
## M2 — Skills v1 & eval baseline
Deliverables:
- `skills-py`: load forecast, PV forecast, price forecast (all with quantile
intervals), bid-optimization MILP, report generator. HTTP/JSON per contracts.
- `packages/evals`: harness + L2 baselines (MAPE, interval coverage, backtest
revenue vs naive/hindsight bounds) on historical data.
Accept when: forecasts return calibrated quantiles (coverage test); MILP output
respects ledger constraints in property tests; eval runs are reproducible and
archived; baseline report exists for each skill.
## M3 — Runtime, two agents, safety chain, Case Desk
Deliverables:
- `packages/runtime`: Mastra instance; trigger service (cron/event/manual);
router agent with template-enum output; `registerSkill` lineage wrapper;
outbox event bus.
- Workflows: `day-ahead-situation`, `day-ahead-bid`, `proposal-lifecycle`
(rule check → simulation → envelope gate with suspend/resume → fresh check →
permit → release).
- `packages/services`: policy engine v1 (+ first Hubei policy pack, from
open-questions once confirmed), envelope service, authority service (permits),
audit/lineage assembly.
- Case Desk v1: approval inbox (resume endpoint), case view, lineage expansion.
- Bid release path = file export for manual upload (degraded channel by design).
Accept when: docs/07 timeline 06:00→08:30 runs end-to-end on historical data;
all eight invariant tests pass (incl. LLM-down degraded run and AI-self-approval
rejection); a suspended approval survives process restart; permit expiry blocks
a late release.
## M4 — Resource agent, envelopes live, review loop
Deliverables:
- Resource dispatch agent + potential-assessment skill + award-decomposition flow.
- Envelope model end-to-end: bounded auto-approval, deviation-streak suspension,
envelope re-approval workflow.
- Review (复盘) workflow: ledger plan-vs-actual per timescale, attribution,
`ReviewFinding` writebacks (semantic memory, reliability scores, envelope
recommendations).
- AI insight cards embedded in existing business pages (read-only projections).
Accept when: docs/07 D-1 16:00 and D+1 sections run; a ReviewFinding measurably
updates a ResourceProfile reliability score; envelope suspension triggers on a
seeded deviation streak.
## M5 — Shadow run (phase-1 acceptance form)
Deliverables:
- Full loop on live data, all external effects simulated (bids generated, not
submitted; commands to a simulation gateway).
- Daily shadow-vs-human-vs-hindsight comparison report; KPI dashboard per
docs/12 §4 definitions.
- Loss circuit-breaker, abnormal-day protocol, kill-switch L0–L2 implemented
and drilled once.
Accept when: 20+ consecutive shadow days with complete lineage; KPI report
auto-generated; one envelope-widening recommendation produced from shadow data
(not acted on — that's a phase-2 governance step).
## Phase 2 (do not build in phase 1)
Interaction-service & load-control agents; edge control links (real protocols);
trading-platform programmatic submission; dispatch/metering live interfaces;
federation gateway (build `FlexibilityEnvelope` schema in M1 — the three
no-rework reservations of docs/10 §5 are in scope for phase 1 schemas only);
controlled live bidding → gradual envelope widening.