- README: orientation, doc map, target package layout (pointers only, no duplicated architecture content) - CLAUDE.md: agent operating manual — invariants as code-review rules, conventions, do-not list, task reading order - ROADMAP: M1-M5 with verifiable acceptance criteria, phase-2 fence - GLOSSARY: canonical Chinese-term → code-name mapping - docs/adr/: eight ADRs recording settled decisions and rejected alternatives - docs/open-questions.md: consolidated TODO(业务) tracker by owner and blocking milestone - .gitignore; untrack .DS_Store Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019u5SLNweVio6ozJX7yfxQr 🔮 View transcript: https://logs.lojong.info/s/e8u90k3t33w590r7b5y7yzqh
92 lines
4.4 KiB
Markdown
92 lines
4.4 KiB
Markdown
# ROADMAP
|
||
|
||
Phase 1 (系统研发与省内能力落地, 2025.12–2026.5) as five milestones. Each has
|
||
acceptance criteria a coding agent can verify. Sequencing rationale: M1–M3 need
|
||
no external-party scheduling; the demo-able bidding loop lands earliest
|
||
(docs/08 §4). Phase 2 items are listed but **not to be built yet**.
|
||
|
||
## M1 — Data foundation & contracts
|
||
|
||
Deliverables:
|
||
- `packages/domain`: zod schemas for the 07-scenario object set — `Proposal`
|
||
(+digest), `Approval`, `ExecutionPermit`, `Envelope`, `PositionLedger`/`PositionUpdate`,
|
||
`ForecastBundle`, `SituationReport`, `ResourceProfile`, `DecisionCase`, `EventEnvelope`.
|
||
- `contracts/`: JSON Schema export pipeline + golden fixtures; pydantic codegen
|
||
wired for `skills-py`; dual-side contract tests in CI.
|
||
- Ingestion pipeline skeleton: time-series store + relational store + snapshot
|
||
store (immutable, checksummed); data-quality gate stub.
|
||
- Ledger service v1: read with version, append, per-timescale views.
|
||
|
||
Accept when: golden fixtures validate identically in TS and Python CI; a ledger
|
||
constraint-cascade test passes (monthly position bounds a day-ahead write);
|
||
snapshots are content-addressed and immutable.
|
||
|
||
## M2 — Skills v1 & eval baseline
|
||
|
||
Deliverables:
|
||
- `skills-py`: load forecast, PV forecast, price forecast (all with quantile
|
||
intervals), bid-optimization MILP, report generator. HTTP/JSON per contracts.
|
||
- `packages/evals`: harness + L2 baselines (MAPE, interval coverage, backtest
|
||
revenue vs naive/hindsight bounds) on historical data.
|
||
|
||
Accept when: forecasts return calibrated quantiles (coverage test); MILP output
|
||
respects ledger constraints in property tests; eval runs are reproducible and
|
||
archived; baseline report exists for each skill.
|
||
|
||
## M3 — Runtime, two agents, safety chain, Case Desk
|
||
|
||
Deliverables:
|
||
- `packages/runtime`: Mastra instance; trigger service (cron/event/manual);
|
||
router agent with template-enum output; `registerSkill` lineage wrapper;
|
||
outbox event bus.
|
||
- Workflows: `day-ahead-situation`, `day-ahead-bid`, `proposal-lifecycle`
|
||
(rule check → simulation → envelope gate with suspend/resume → fresh check →
|
||
permit → release).
|
||
- `packages/services`: policy engine v1 (+ first Hubei policy pack, from
|
||
open-questions once confirmed), envelope service, authority service (permits),
|
||
audit/lineage assembly.
|
||
- Case Desk v1: approval inbox (resume endpoint), case view, lineage expansion.
|
||
- Bid release path = file export for manual upload (degraded channel by design).
|
||
|
||
Accept when: docs/07 timeline 06:00→08:30 runs end-to-end on historical data;
|
||
all eight invariant tests pass (incl. LLM-down degraded run and AI-self-approval
|
||
rejection); a suspended approval survives process restart; permit expiry blocks
|
||
a late release.
|
||
|
||
## M4 — Resource agent, envelopes live, review loop
|
||
|
||
Deliverables:
|
||
- Resource dispatch agent + potential-assessment skill + award-decomposition flow.
|
||
- Envelope model end-to-end: bounded auto-approval, deviation-streak suspension,
|
||
envelope re-approval workflow.
|
||
- Review (复盘) workflow: ledger plan-vs-actual per timescale, attribution,
|
||
`ReviewFinding` writebacks (semantic memory, reliability scores, envelope
|
||
recommendations).
|
||
- AI insight cards embedded in existing business pages (read-only projections).
|
||
|
||
Accept when: docs/07 D-1 16:00 and D+1 sections run; a ReviewFinding measurably
|
||
updates a ResourceProfile reliability score; envelope suspension triggers on a
|
||
seeded deviation streak.
|
||
|
||
## M5 — Shadow run (phase-1 acceptance form)
|
||
|
||
Deliverables:
|
||
- Full loop on live data, all external effects simulated (bids generated, not
|
||
submitted; commands to a simulation gateway).
|
||
- Daily shadow-vs-human-vs-hindsight comparison report; KPI dashboard per
|
||
docs/12 §4 definitions.
|
||
- Loss circuit-breaker, abnormal-day protocol, kill-switch L0–L2 implemented
|
||
and drilled once.
|
||
|
||
Accept when: 20+ consecutive shadow days with complete lineage; KPI report
|
||
auto-generated; one envelope-widening recommendation produced from shadow data
|
||
(not acted on — that's a phase-2 governance step).
|
||
|
||
## Phase 2 (do not build in phase 1)
|
||
|
||
Interaction-service & load-control agents; edge control links (real protocols);
|
||
trading-platform programmatic submission; dispatch/metering live interfaces;
|
||
federation gateway (build `FlexibilityEnvelope` schema in M1 — the three
|
||
no-rework reservations of docs/10 §5 are in scope for phase 1 schemas only);
|
||
controlled live bidding → gradual envelope widening.
|