- 01: add eight testable hard invariants (self-approval ban, duty separation, digest-bound approvals, permit-only effects, LLM-less degraded operation) - 03: approval staleness model — proposal digest, FRESH_CHECK/STALE states, APPROVED→AUTHORIZED split with revocable ExecutionPermit; gateways accept permits, never bare plans - 05: rules packaged as versioned, testable Policy Packs pinned in proposal lineage - 02: approval workbench reshaped into a Case Desk (DecisionCase as the operator's accountability unit) - 10 (new): cross-province federation boundary — signed artifacts only, plus three phase-1 no-rework reservations - 08: add M5 shadow-run milestone as phase-1 acceptance form - 00/04/07: object table, doc map, and walkthrough consistency updates - add brainstorming.md (peer review source document) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019u5SLNweVio6ozJX7yfxQr 🔮 View transcript: https://logs.lojong.info/s/e8u90k3t33w590r7b5y7yzqh
598 lines
21 KiB
Markdown
598 lines
21 KiB
Markdown
# VPP AI Platform Architecture Brainstorming
|
|
|
|
## 1. Purpose
|
|
|
|
This document develops the technical architecture proposed for the virtual power
|
|
plant multi-timescale collaborative operations platform. It is intended for
|
|
business architects, platform engineers, algorithm teams, control-system
|
|
engineers, security teams, and project stakeholders.
|
|
|
|
The goal is not to select products or implementation frameworks yet. The
|
|
immediate goal is to agree on:
|
|
|
|
- The system's major trust and responsibility boundaries.
|
|
- The role AI and large language models should play.
|
|
- How decisions become approved external actions.
|
|
- How the platform should divide work across time and spatial scales.
|
|
- Which architecture should be taken into detailed design.
|
|
|
|
## 2. Executive Recommendation
|
|
|
|
The proposal has a strong business vision, but the platform should not be
|
|
implemented literally as five autonomous agents coordinated by a single
|
|
all-purpose runtime.
|
|
|
|
The recommended direction is:
|
|
|
|
1. Use a deterministic, event-driven VPP operating kernel as the operational
|
|
foundation.
|
|
2. Present work to operators through durable operation and decision cases.
|
|
3. Expose a small decision-case lifecycle API to external systems.
|
|
4. Treat the five agents as business-facing roles and replaceable contributors,
|
|
not as trust boundaries or direct controllers.
|
|
5. Allow AI to create, compare, and explain draft decisions, but never to
|
|
authorize or directly execute consequential actions.
|
|
|
|
In short:
|
|
|
|
> The operational domain should remain correct, auditable, and usable when the
|
|
> LLM is unavailable.
|
|
|
|
## 3. What Is Already Strong
|
|
|
|
The proposal establishes several sound principles:
|
|
|
|
- It covers the full operational loop from monitoring and planning to execution
|
|
feedback and review.
|
|
- It explicitly states that AI must not directly control terminal equipment.
|
|
- It includes rule validation, simulation, human approval, rollback, and audit.
|
|
- It builds on existing business processes, algorithms, data, and resource
|
|
access.
|
|
- It evaluates success through operational and financial outcomes rather than
|
|
model capability alone.
|
|
- It recognizes that annual, monthly, day-ahead, intraday, and real-time
|
|
decisions require coordination.
|
|
|
|
These principles should be preserved as architectural invariants.
|
|
|
|
## 4. Architectural Questions in the Current Proposal
|
|
|
|
Several aspects need to be resolved before detailed system design begins.
|
|
|
|
### 4.1 Agent taxonomy is inconsistent
|
|
|
|
The proposal alternates between five agent roles and six business stages. Names
|
|
also change between analysis, resource dispatch or organization, transaction
|
|
gaming or decision-making, transaction service or user interaction, and load
|
|
control or load management.
|
|
|
|
If these names become public APIs or independently deployed services, normal
|
|
changes in the business model will cause expensive coupling. They are better
|
|
treated as product personas or internal capability compositions.
|
|
|
|
### 4.2 The LLM appears too foundational
|
|
|
|
The diagrams place the Guangming power model at the bottom of the architecture.
|
|
This suggests that normal operation depends on LLM availability and behavior.
|
|
|
|
Forecasting, flexibility calculation, optimization, rule enforcement, approval,
|
|
market submission, and device control must remain available without an LLM. The
|
|
LLM should be an advisory and orchestration participant above a deterministic
|
|
operating foundation.
|
|
|
|
### 4.3 The runtime owns too much
|
|
|
|
The proposed Agent Runtime includes workflow orchestration, context
|
|
construction, memory, tools, security, and state synchronization. This risks
|
|
becoming a platform-wide bottleneck and a single failure domain.
|
|
|
|
Workflow durability, operational state, model execution, policy enforcement,
|
|
approval authority, and control execution should have separate responsibilities
|
|
and contracts.
|
|
|
|
### 4.4 Time scales are not separated operationally
|
|
|
|
Second-level equipment control and day-ahead reasoning appear in the same
|
|
logical workflow. They have different availability, latency, determinism, and
|
|
safety requirements.
|
|
|
|
Real-time control should live at the site or edge and use prevalidated policies.
|
|
Cloud or provincial AI workflows should produce bounded plans and control
|
|
envelopes rather than participate in a fast control loop.
|
|
|
|
### 4.5 Knowledge and executable rules are conflated
|
|
|
|
RAG is useful for retrieving regulations, operating procedures, explanations,
|
|
and historical cases. It must not be the authoritative mechanism for enforcing a
|
|
market rule, safety limit, or device constraint.
|
|
|
|
Consequential rules should be stored as versioned, testable, executable policy
|
|
packs. RAG may explain a policy or locate its source, while a deterministic
|
|
policy engine enforces it.
|
|
|
|
### 4.6 Decision provenance is underspecified
|
|
|
|
The architecture needs canonical, versioned representations of forecasts,
|
|
constraints, flexibility, plans, approvals, commands, execution receipts, and
|
|
settlement outcomes. Without these, simulation, approval, replay, and audit
|
|
cannot reliably refer to the same decision.
|
|
|
|
### 4.7 Cross-province operation needs a federation boundary
|
|
|
|
Cross-province collaboration should not imply a single central platform with
|
|
unrestricted access to raw telemetry or terminal control. The exchanged objects
|
|
should normally be signed flexibility envelopes, commitments, bids, awards, and
|
|
settlement facts.
|
|
|
|
## 5. Non-Negotiable Architecture Invariants
|
|
|
|
The following invariants should guide all designs:
|
|
|
|
- AI cannot approve its own proposal.
|
|
- AI cannot directly call a market-submission or device-control endpoint.
|
|
- Every consequential action is bound to an immutable artifact digest and a
|
|
validity window.
|
|
- Approval is invalidated when material evidence, constraints, or the proposed
|
|
effect changes.
|
|
- Operational decisions retain their input data snapshot, rule version, model
|
|
version, solver version, and approvals.
|
|
- External effects are idempotent and produce durable receipts.
|
|
- The platform supports replay from recorded evidence.
|
|
- Degraded operation remains possible without the LLM.
|
|
- Edge control remains safe during provincial-platform or wide-area network
|
|
outages.
|
|
- Operational truth is stored in authoritative domain stores, not in agent
|
|
memory.
|
|
|
|
## 6. Three Architecture Alternatives
|
|
|
|
The following alternatives are intentionally different rather than variations of
|
|
the same agent-runtime design.
|
|
|
|
### 6.1 Alternative A: Minimal Decision Platform
|
|
|
|
This design exposes only three lifecycle operations:
|
|
|
|
```ts
|
|
interface VppDecisionPlatform {
|
|
start(request: DecisionRequest): Promise<DecisionCase>;
|
|
act(action: CaseAction): Promise<DecisionCase>;
|
|
observe(query: CaseQuery): Promise<CaseObservation>;
|
|
}
|
|
```
|
|
|
|
Human requests, scheduled jobs, and operational events all create a versioned
|
|
decision case. The caller never selects an agent, prompt, model, skill, or
|
|
solver.
|
|
|
|
A day-ahead planning case may produce:
|
|
|
|
- Load, renewable-generation, and market-price forecasts.
|
|
- A resource portfolio and flexibility envelope.
|
|
- Candidate market bids.
|
|
- A customer response plan.
|
|
- A dispatch envelope.
|
|
- Rule and simulation results.
|
|
- Separate approval gates for market submission, outreach, and dispatch.
|
|
|
|
An authorized user approves an exact proposal version and digest. Approval does
|
|
not grant the AI direct access to an external system; it allows a guarded
|
|
adapter to process the authorized effect.
|
|
|
|
This design hides orchestration, retries, agent topology, model selection, tool
|
|
invocation, and integration protocols very effectively. Its main risk is that
|
|
`DecisionCase` and the service behind it become generic containers with
|
|
insufficient domain boundaries.
|
|
|
|
### 6.2 Alternative B: Event-Driven VPP Operating Kernel
|
|
|
|
This design makes the deterministic operational domain the foundation and treats
|
|
AI as a replaceable adapter.
|
|
|
|
Representative ports are:
|
|
|
|
```ts
|
|
interface ObservationIngestPort {
|
|
record(observation: Observation): Promise<EvidenceReference>;
|
|
}
|
|
|
|
interface CommandPort {
|
|
execute(command: DomainCommand): Promise<CommandResult>;
|
|
}
|
|
|
|
interface ControlAuthorityPort {
|
|
authorize(
|
|
plan: CandidatePlan,
|
|
simulation: SimulationResult,
|
|
approvals: Approval[],
|
|
): Promise<ExecutionPermit | Denial>;
|
|
}
|
|
|
|
interface ControlGateway {
|
|
dispatch(permit: ExecutionPermit): Promise<ExecutionReceipt>;
|
|
}
|
|
```
|
|
|
|
The core invariant is:
|
|
|
|
> Evidence can trigger decisions; decisions can propose effects; only the
|
|
> authority plane can authorize effects.
|
|
|
|
For a demand-response operation, the system records the request as evidence,
|
|
calculates a forecast and flexibility envelope, creates candidate plans, runs
|
|
safety simulation, collects approval, and requests an execution permit. The
|
|
control gateway accepts a permit, not a raw plan.
|
|
|
|
This alternative offers the strongest replayability, auditability, safety
|
|
isolation, and AI replaceability. Its cost is conceptual complexity: teams must
|
|
handle commands, observations, events, idempotency, revisions, projections, and
|
|
eventual consistency correctly.
|
|
|
|
### 6.3 Alternative C: Operations Case Desk
|
|
|
|
This design optimizes the product around the operator's unit of accountability:
|
|
a bounded operational case.
|
|
|
|
Examples include:
|
|
|
|
- Prepare tomorrow's market bid.
|
|
- Deliver a 20 MW demand-response event.
|
|
- Resolve an execution deviation.
|
|
- Review a completed dispatch.
|
|
|
|
The operator-facing interface is:
|
|
|
|
```ts
|
|
interface OperationsDesk {
|
|
open(kind: CaseKind, input: CaseInput): Promise<CaseId>;
|
|
read(caseId: CaseId): Promise<CaseView>;
|
|
apply(command: CaseCommand): Promise<CommandReceipt>;
|
|
watch(caseId: CaseId): AsyncIterable<CaseEvent>;
|
|
}
|
|
```
|
|
|
|
Each case contains its objective, owner, horizon, evidence, assumptions,
|
|
scenario branches, recommendations, approvals, execution receipts, and outcome.
|
|
An advanced operator can fork a high-price or low-solar scenario without
|
|
directly invoking an agent.
|
|
|
|
This design produces the strongest operator experience and naturally organizes
|
|
audit and collaboration. Its principal risk is turning a case into a permanent
|
|
collection of unrelated activity. Each case therefore needs one accountable
|
|
objective, an explicit deadline, an owner, and a completion contract.
|
|
|
|
## 7. Comparison and Recommended Hybrid
|
|
|
|
The Minimal Decision Platform has the smallest learning and misuse surface for
|
|
external callers. It is a deep interface, but its simplicity can conceal an
|
|
overly centralized implementation.
|
|
|
|
The Event-Driven Operating Kernel is the strongest internal architecture for a
|
|
safety-sensitive and heavily audited system. It makes evidence, authority, and
|
|
external effects explicit. It is less convenient as the operator's primary
|
|
mental model.
|
|
|
|
The Operations Case Desk best matches how business users collaborate and accept
|
|
responsibility for decisions. It is not, by itself, sufficient as the low-level
|
|
operating and control architecture.
|
|
|
|
The recommended hybrid is therefore:
|
|
|
|
- Event-driven deterministic kernel inside.
|
|
- Operations Case Desk for operators.
|
|
- Minimal decision-case API for external systems.
|
|
|
|
The layers complement one another instead of duplicating responsibilities.
|
|
|
|
## 8. Recommended Logical Architecture
|
|
|
|
```mermaid
|
|
flowchart TB
|
|
T[Operators / Schedules / Operational Events] --> C[Operations Case Desk and Decision API]
|
|
C --> P[Durable Process Manager and Decision Ledger]
|
|
|
|
P --> F[Forecasting Services]
|
|
P --> R[Flexibility and Resource Portfolio Services]
|
|
P --> M[Market and Dispatch Optimization]
|
|
P --> U[Customer Response Planning]
|
|
P --> S[Settlement and Review]
|
|
|
|
A[LLM Advisor and Agent Roles] --> P
|
|
A --> F
|
|
A --> R
|
|
A --> M
|
|
A --> U
|
|
|
|
F --> G[Policy Validation]
|
|
R --> G
|
|
M --> G
|
|
U --> G
|
|
|
|
G --> V[Simulation and Fresh-State Verification]
|
|
V --> H[Human Approval]
|
|
H --> E[Execution Authority]
|
|
E --> X[Market / Messaging / Control Gateways]
|
|
X --> D[Site and Edge Controllers]
|
|
D --> O[Execution Observations]
|
|
O --> P
|
|
```
|
|
|
|
### 8.1 Experience plane
|
|
|
|
The experience plane provides the operations desk, dashboards, AI-assisted
|
|
investigation, scenario comparison, approval inbox, and operational reports.
|
|
Pages are projections over cases and the decision ledger rather than independent
|
|
workflow silos.
|
|
|
|
### 8.2 Orchestration plane
|
|
|
|
The orchestration plane owns durable process state, deadlines, retries,
|
|
compensation, correlation, and case progression. It invokes domain capabilities
|
|
but does not perform forecasting, optimization, authorization, or control
|
|
itself.
|
|
|
|
### 8.3 Decision-services plane
|
|
|
|
Domain services provide deterministic or bounded capabilities:
|
|
|
|
- Load, generation, and price forecasting.
|
|
- Resource capability and flexibility assessment.
|
|
- Portfolio aggregation and commitment management.
|
|
- Market bidding and dispatch optimization.
|
|
- Customer segmentation and response planning.
|
|
- Simulation, deviation analysis, settlement, and attribution.
|
|
|
|
These services accept complete, versioned inputs and return typed artifacts with
|
|
provenance.
|
|
|
|
### 8.4 AI advisory plane
|
|
|
|
The AI layer provides intent understanding, task decomposition, retrieval,
|
|
explanation, report generation, tool selection, and candidate strategy
|
|
generation. The five proposed agents can exist here as business-specific
|
|
contributors.
|
|
|
|
There should be no `executeDeviceCommand` or unrestricted `submitMarketBid` tool
|
|
available to an agent.
|
|
|
|
### 8.5 Authority and execution plane
|
|
|
|
This plane enforces executable policies, separation of duties, approval
|
|
requirements, current-state verification, permit expiry, and effect limits. It
|
|
is the only route to market, messaging, and control gateways.
|
|
|
|
### 8.6 Evidence and data plane
|
|
|
|
This plane records raw observations, normalized telemetry, business events,
|
|
reference data, feature values, forecasts, model outputs, decisions, and
|
|
execution outcomes. It supports point-in-time reconstruction and replay.
|
|
|
|
## 9. Decision and Execution Lifecycle
|
|
|
|
A consequential artifact should follow an explicit lifecycle:
|
|
|
|
```text
|
|
DRAFT
|
|
-> VALIDATED
|
|
-> SIMULATED
|
|
-> APPROVED
|
|
-> AUTHORIZED
|
|
-> ISSUED
|
|
-> ACKNOWLEDGED
|
|
-> COMPLETED / FAILED / ROLLED_BACK
|
|
```
|
|
|
|
The LLM may create or explain a `DRAFT`. It cannot advance an artifact to
|
|
`AUTHORIZED`.
|
|
|
|
Approval should bind to:
|
|
|
|
- The exact artifact digest.
|
|
- The approved effect scope.
|
|
- Financial and quantity limits.
|
|
- A validity window.
|
|
- The approver's identity and role.
|
|
- The evidence and policy versions used for evaluation.
|
|
|
|
If material telemetry, constraints, rules, or proposed effects change, the
|
|
approval becomes stale.
|
|
|
|
## 10. Separation by Time Scale
|
|
|
|
### Seconds
|
|
|
|
Site and edge controllers own equipment protection, local interlocks, fast
|
|
feedback control, and prevalidated fallback behavior. There is no LLM dependency
|
|
in this loop.
|
|
|
|
### One to fifteen minutes
|
|
|
|
Streaming services perform telemetry processing, state estimation, deviation
|
|
detection, rolling forecasts, and bounded corrective optimization. Any automatic
|
|
response must be preauthorized and constrained.
|
|
|
|
### Intraday and day-ahead
|
|
|
|
The provincial platform performs market analysis, portfolio planning, scenario
|
|
comparison, operator review, and approval. This is the main operating range for
|
|
AI-assisted decision workflows.
|
|
|
|
### Monthly and annual
|
|
|
|
Planning services support contract strategy, resource acquisition, capacity
|
|
planning, model training, policy analysis, and long-horizon simulation.
|
|
|
|
## 11. Separation by Spatial Scale
|
|
|
|
### Device and site
|
|
|
|
The site retains device protocols, protection constraints, local control, and
|
|
detailed telemetry. It exposes a controlled resource capability interface
|
|
upward.
|
|
|
|
### Aggregation unit
|
|
|
|
The aggregation layer calculates site or resource-group flexibility envelopes,
|
|
manages local commitments, and translates bounded dispatch envelopes into site
|
|
plans.
|
|
|
|
### Provincial platform
|
|
|
|
The provincial layer manages portfolios, markets, operator decisions, approval,
|
|
reporting, and province-wide optimization.
|
|
|
|
### Cross-province federation
|
|
|
|
Federated platforms should exchange signed and versioned business artifacts such
|
|
as:
|
|
|
|
- Flexibility envelopes.
|
|
- Available capacity and reserve.
|
|
- Commitments and constraints.
|
|
- Bids and awards.
|
|
- Delivery and settlement facts.
|
|
|
|
Raw device control authority and unrestricted customer-level telemetry should
|
|
not cross this boundary by default.
|
|
|
|
## 12. Canonical Domain Contracts
|
|
|
|
Before selecting implementation frameworks, the project should define and
|
|
version the following objects:
|
|
|
|
- `Resource`
|
|
- `ResourceCapability`
|
|
- `Portfolio`
|
|
- `Commitment`
|
|
- `FlexibilityEnvelope`
|
|
- `ForecastBundle`
|
|
- `MarketOpportunity`
|
|
- `DecisionCase`
|
|
- `CandidatePlan`
|
|
- `ConstraintSet`
|
|
- `ValidationResult`
|
|
- `SimulationResult`
|
|
- `Approval`
|
|
- `ExecutionPermit`
|
|
- `DispatchOrder`
|
|
- `ExecutionReceipt`
|
|
- `Settlement`
|
|
- `PolicyPack`
|
|
|
|
Every material plan should contain:
|
|
|
|
- Its validity window.
|
|
- The input-data snapshot or evidence references.
|
|
- Market-rule and policy versions.
|
|
- Model and solver versions.
|
|
- Objectives and constraints.
|
|
- Uncertainty and confidence information.
|
|
- Expected physical and financial effects.
|
|
- An immutable digest.
|
|
- Current validation and approval state.
|
|
|
|
## 13. Data, Knowledge, Rules, and Memory
|
|
|
|
These four concepts should remain distinct.
|
|
|
|
### Operational data
|
|
|
|
Telemetry, market facts, commitments, user responses, and execution observations
|
|
are authoritative domain data. They require quality indicators, timestamps,
|
|
lineage, retention, and access controls.
|
|
|
|
### Knowledge
|
|
|
|
Regulations, procedures, manuals, historical reports, and operating experience
|
|
can be indexed for retrieval and explanation. Retrieved content is evidence for
|
|
a user or model, not automatically an executable rule.
|
|
|
|
### Rules
|
|
|
|
Market rules, approval rules, equipment constraints, financial limits, and
|
|
security policies should be versioned, executable, testable, and auditable.
|
|
|
|
### Memory
|
|
|
|
Agent memory should contain convenience context and curated lessons, not
|
|
authoritative operational state. Feedback should enter a controlled evaluation
|
|
and promotion process before it changes a model, policy, or operating template.
|
|
|
|
## 14. Reliability and Security Considerations
|
|
|
|
The detailed design should account for:
|
|
|
|
- Idempotent handling of repeated events, approvals, submissions, and commands.
|
|
- Optimistic concurrency to reject stale approvals and conflicting actions.
|
|
- Late, missing, duplicated, or low-quality telemetry.
|
|
- Model and solver timeouts with deterministic fallbacks.
|
|
- Network partition between provincial and site systems.
|
|
- Partial device acceptance and compensating dispatch.
|
|
- Permit expiry and revocation.
|
|
- Segregation of operator, approver, and administrator duties.
|
|
- Field-level authorization for customer, market, and infrastructure-sensitive
|
|
data.
|
|
- Immutable audit records and tamper evidence.
|
|
- Tenant, portfolio, site, and provincial data boundaries.
|
|
- Observability for every decision and external effect.
|
|
|
|
## 15. Recommended First Vertical Slice
|
|
|
|
The first implementation should be one complete operational flow rather than
|
|
five partially connected agents:
|
|
|
|
> Day-ahead portfolio plan -> forecast -> flexibility assessment -> candidate
|
|
> bids -> simulation -> operator approval -> simulated market submission ->
|
|
> execution replay -> settlement review.
|
|
|
|
The slice should initially run in shadow mode using historical and live data
|
|
without producing real external effects. It should prove:
|
|
|
|
- Canonical domain contracts.
|
|
- Evidence and decision lineage.
|
|
- Durable orchestration and case state.
|
|
- Rule and model versioning.
|
|
- Scenario comparison.
|
|
- Approval bound to immutable artifacts.
|
|
- External-effect simulation and receipts.
|
|
- Replay and outcome attribution.
|
|
- Operation without the LLM.
|
|
|
|
Once this foundation works, the proposed analysis, resource, trading,
|
|
customer-interaction, and load-management agents can be introduced as
|
|
contributors to the same case lifecycle.
|
|
|
|
## 16. Decisions Required Next
|
|
|
|
The following decisions materially affect the detailed architecture:
|
|
|
|
1. Is the new platform an AI overlay on the existing 3060 platform, a gradual
|
|
replacement, or a separate system of engagement?
|
|
2. Which existing system remains the source of truth for resources, customers,
|
|
market positions, dispatch instructions, and settlement?
|
|
3. Is the first release recommendation-only, capable of controlled market
|
|
submission, or capable of controlled dispatch?
|
|
4. Which actions may eventually become automatically authorized within
|
|
predefined limits?
|
|
5. Which functions and data must remain at the site, provincial, or dedicated
|
|
security zone?
|
|
6. What protocols and latency guarantees currently exist between the provincial
|
|
platform, aggregation units, and edge terminals?
|
|
7. Which province-specific rules must become executable policy packs in the
|
|
first release?
|
|
8. What is the acceptable behavior when forecasts, telemetry, the LLM, an
|
|
optimizer, or a downstream system is unavailable?
|
|
|
|
## 17. Suggested Next Architecture Work
|
|
|
|
After the decisions above are answered, the next design artifacts should be:
|
|
|
|
1. A system-context diagram showing existing systems and ownership boundaries.
|
|
2. A source-of-truth matrix for all major domain entities.
|
|
3. A canonical domain vocabulary and schema definitions.
|
|
4. A detailed day-ahead decision-case sequence.
|
|
5. A dispatch authority and execution threat model.
|
|
6. A latency, availability, recovery, and data-retention requirements matrix.
|
|
7. A deployment-zone and cross-province federation design.
|
|
8. An incremental roadmap built from end-to-end vertical slices.
|