vpp-ai-platform/brainstorming.md

598 lines
21 KiB
Markdown
Raw Permalink Normal View History

# VPP AI Platform Architecture Brainstorming
## 1. Purpose
This document develops the technical architecture proposed for the virtual power
plant multi-timescale collaborative operations platform. It is intended for
business architects, platform engineers, algorithm teams, control-system
engineers, security teams, and project stakeholders.
The goal is not to select products or implementation frameworks yet. The
immediate goal is to agree on:
- The system's major trust and responsibility boundaries.
- The role AI and large language models should play.
- How decisions become approved external actions.
- How the platform should divide work across time and spatial scales.
- Which architecture should be taken into detailed design.
## 2. Executive Recommendation
The proposal has a strong business vision, but the platform should not be
implemented literally as five autonomous agents coordinated by a single
all-purpose runtime.
The recommended direction is:
1. Use a deterministic, event-driven VPP operating kernel as the operational
foundation.
2. Present work to operators through durable operation and decision cases.
3. Expose a small decision-case lifecycle API to external systems.
4. Treat the five agents as business-facing roles and replaceable contributors,
not as trust boundaries or direct controllers.
5. Allow AI to create, compare, and explain draft decisions, but never to
authorize or directly execute consequential actions.
In short:
> The operational domain should remain correct, auditable, and usable when the
> LLM is unavailable.
## 3. What Is Already Strong
The proposal establishes several sound principles:
- It covers the full operational loop from monitoring and planning to execution
feedback and review.
- It explicitly states that AI must not directly control terminal equipment.
- It includes rule validation, simulation, human approval, rollback, and audit.
- It builds on existing business processes, algorithms, data, and resource
access.
- It evaluates success through operational and financial outcomes rather than
model capability alone.
- It recognizes that annual, monthly, day-ahead, intraday, and real-time
decisions require coordination.
These principles should be preserved as architectural invariants.
## 4. Architectural Questions in the Current Proposal
Several aspects need to be resolved before detailed system design begins.
### 4.1 Agent taxonomy is inconsistent
The proposal alternates between five agent roles and six business stages. Names
also change between analysis, resource dispatch or organization, transaction
gaming or decision-making, transaction service or user interaction, and load
control or load management.
If these names become public APIs or independently deployed services, normal
changes in the business model will cause expensive coupling. They are better
treated as product personas or internal capability compositions.
### 4.2 The LLM appears too foundational
The diagrams place the Guangming power model at the bottom of the architecture.
This suggests that normal operation depends on LLM availability and behavior.
Forecasting, flexibility calculation, optimization, rule enforcement, approval,
market submission, and device control must remain available without an LLM. The
LLM should be an advisory and orchestration participant above a deterministic
operating foundation.
### 4.3 The runtime owns too much
The proposed Agent Runtime includes workflow orchestration, context
construction, memory, tools, security, and state synchronization. This risks
becoming a platform-wide bottleneck and a single failure domain.
Workflow durability, operational state, model execution, policy enforcement,
approval authority, and control execution should have separate responsibilities
and contracts.
### 4.4 Time scales are not separated operationally
Second-level equipment control and day-ahead reasoning appear in the same
logical workflow. They have different availability, latency, determinism, and
safety requirements.
Real-time control should live at the site or edge and use prevalidated policies.
Cloud or provincial AI workflows should produce bounded plans and control
envelopes rather than participate in a fast control loop.
### 4.5 Knowledge and executable rules are conflated
RAG is useful for retrieving regulations, operating procedures, explanations,
and historical cases. It must not be the authoritative mechanism for enforcing a
market rule, safety limit, or device constraint.
Consequential rules should be stored as versioned, testable, executable policy
packs. RAG may explain a policy or locate its source, while a deterministic
policy engine enforces it.
### 4.6 Decision provenance is underspecified
The architecture needs canonical, versioned representations of forecasts,
constraints, flexibility, plans, approvals, commands, execution receipts, and
settlement outcomes. Without these, simulation, approval, replay, and audit
cannot reliably refer to the same decision.
### 4.7 Cross-province operation needs a federation boundary
Cross-province collaboration should not imply a single central platform with
unrestricted access to raw telemetry or terminal control. The exchanged objects
should normally be signed flexibility envelopes, commitments, bids, awards, and
settlement facts.
## 5. Non-Negotiable Architecture Invariants
The following invariants should guide all designs:
- AI cannot approve its own proposal.
- AI cannot directly call a market-submission or device-control endpoint.
- Every consequential action is bound to an immutable artifact digest and a
validity window.
- Approval is invalidated when material evidence, constraints, or the proposed
effect changes.
- Operational decisions retain their input data snapshot, rule version, model
version, solver version, and approvals.
- External effects are idempotent and produce durable receipts.
- The platform supports replay from recorded evidence.
- Degraded operation remains possible without the LLM.
- Edge control remains safe during provincial-platform or wide-area network
outages.
- Operational truth is stored in authoritative domain stores, not in agent
memory.
## 6. Three Architecture Alternatives
The following alternatives are intentionally different rather than variations of
the same agent-runtime design.
### 6.1 Alternative A: Minimal Decision Platform
This design exposes only three lifecycle operations:
```ts
interface VppDecisionPlatform {
start(request: DecisionRequest): Promise<DecisionCase>;
act(action: CaseAction): Promise<DecisionCase>;
observe(query: CaseQuery): Promise<CaseObservation>;
}
```
Human requests, scheduled jobs, and operational events all create a versioned
decision case. The caller never selects an agent, prompt, model, skill, or
solver.
A day-ahead planning case may produce:
- Load, renewable-generation, and market-price forecasts.
- A resource portfolio and flexibility envelope.
- Candidate market bids.
- A customer response plan.
- A dispatch envelope.
- Rule and simulation results.
- Separate approval gates for market submission, outreach, and dispatch.
An authorized user approves an exact proposal version and digest. Approval does
not grant the AI direct access to an external system; it allows a guarded
adapter to process the authorized effect.
This design hides orchestration, retries, agent topology, model selection, tool
invocation, and integration protocols very effectively. Its main risk is that
`DecisionCase` and the service behind it become generic containers with
insufficient domain boundaries.
### 6.2 Alternative B: Event-Driven VPP Operating Kernel
This design makes the deterministic operational domain the foundation and treats
AI as a replaceable adapter.
Representative ports are:
```ts
interface ObservationIngestPort {
record(observation: Observation): Promise<EvidenceReference>;
}
interface CommandPort {
execute(command: DomainCommand): Promise<CommandResult>;
}
interface ControlAuthorityPort {
authorize(
plan: CandidatePlan,
simulation: SimulationResult,
approvals: Approval[],
): Promise<ExecutionPermit | Denial>;
}
interface ControlGateway {
dispatch(permit: ExecutionPermit): Promise<ExecutionReceipt>;
}
```
The core invariant is:
> Evidence can trigger decisions; decisions can propose effects; only the
> authority plane can authorize effects.
For a demand-response operation, the system records the request as evidence,
calculates a forecast and flexibility envelope, creates candidate plans, runs
safety simulation, collects approval, and requests an execution permit. The
control gateway accepts a permit, not a raw plan.
This alternative offers the strongest replayability, auditability, safety
isolation, and AI replaceability. Its cost is conceptual complexity: teams must
handle commands, observations, events, idempotency, revisions, projections, and
eventual consistency correctly.
### 6.3 Alternative C: Operations Case Desk
This design optimizes the product around the operator's unit of accountability:
a bounded operational case.
Examples include:
- Prepare tomorrow's market bid.
- Deliver a 20 MW demand-response event.
- Resolve an execution deviation.
- Review a completed dispatch.
The operator-facing interface is:
```ts
interface OperationsDesk {
open(kind: CaseKind, input: CaseInput): Promise<CaseId>;
read(caseId: CaseId): Promise<CaseView>;
apply(command: CaseCommand): Promise<CommandReceipt>;
watch(caseId: CaseId): AsyncIterable<CaseEvent>;
}
```
Each case contains its objective, owner, horizon, evidence, assumptions,
scenario branches, recommendations, approvals, execution receipts, and outcome.
An advanced operator can fork a high-price or low-solar scenario without
directly invoking an agent.
This design produces the strongest operator experience and naturally organizes
audit and collaboration. Its principal risk is turning a case into a permanent
collection of unrelated activity. Each case therefore needs one accountable
objective, an explicit deadline, an owner, and a completion contract.
## 7. Comparison and Recommended Hybrid
The Minimal Decision Platform has the smallest learning and misuse surface for
external callers. It is a deep interface, but its simplicity can conceal an
overly centralized implementation.
The Event-Driven Operating Kernel is the strongest internal architecture for a
safety-sensitive and heavily audited system. It makes evidence, authority, and
external effects explicit. It is less convenient as the operator's primary
mental model.
The Operations Case Desk best matches how business users collaborate and accept
responsibility for decisions. It is not, by itself, sufficient as the low-level
operating and control architecture.
The recommended hybrid is therefore:
- Event-driven deterministic kernel inside.
- Operations Case Desk for operators.
- Minimal decision-case API for external systems.
The layers complement one another instead of duplicating responsibilities.
## 8. Recommended Logical Architecture
```mermaid
flowchart TB
T[Operators / Schedules / Operational Events] --> C[Operations Case Desk and Decision API]
C --> P[Durable Process Manager and Decision Ledger]
P --> F[Forecasting Services]
P --> R[Flexibility and Resource Portfolio Services]
P --> M[Market and Dispatch Optimization]
P --> U[Customer Response Planning]
P --> S[Settlement and Review]
A[LLM Advisor and Agent Roles] --> P
A --> F
A --> R
A --> M
A --> U
F --> G[Policy Validation]
R --> G
M --> G
U --> G
G --> V[Simulation and Fresh-State Verification]
V --> H[Human Approval]
H --> E[Execution Authority]
E --> X[Market / Messaging / Control Gateways]
X --> D[Site and Edge Controllers]
D --> O[Execution Observations]
O --> P
```
### 8.1 Experience plane
The experience plane provides the operations desk, dashboards, AI-assisted
investigation, scenario comparison, approval inbox, and operational reports.
Pages are projections over cases and the decision ledger rather than independent
workflow silos.
### 8.2 Orchestration plane
The orchestration plane owns durable process state, deadlines, retries,
compensation, correlation, and case progression. It invokes domain capabilities
but does not perform forecasting, optimization, authorization, or control
itself.
### 8.3 Decision-services plane
Domain services provide deterministic or bounded capabilities:
- Load, generation, and price forecasting.
- Resource capability and flexibility assessment.
- Portfolio aggregation and commitment management.
- Market bidding and dispatch optimization.
- Customer segmentation and response planning.
- Simulation, deviation analysis, settlement, and attribution.
These services accept complete, versioned inputs and return typed artifacts with
provenance.
### 8.4 AI advisory plane
The AI layer provides intent understanding, task decomposition, retrieval,
explanation, report generation, tool selection, and candidate strategy
generation. The five proposed agents can exist here as business-specific
contributors.
There should be no `executeDeviceCommand` or unrestricted `submitMarketBid` tool
available to an agent.
### 8.5 Authority and execution plane
This plane enforces executable policies, separation of duties, approval
requirements, current-state verification, permit expiry, and effect limits. It
is the only route to market, messaging, and control gateways.
### 8.6 Evidence and data plane
This plane records raw observations, normalized telemetry, business events,
reference data, feature values, forecasts, model outputs, decisions, and
execution outcomes. It supports point-in-time reconstruction and replay.
## 9. Decision and Execution Lifecycle
A consequential artifact should follow an explicit lifecycle:
```text
DRAFT
-> VALIDATED
-> SIMULATED
-> APPROVED
-> AUTHORIZED
-> ISSUED
-> ACKNOWLEDGED
-> COMPLETED / FAILED / ROLLED_BACK
```
The LLM may create or explain a `DRAFT`. It cannot advance an artifact to
`AUTHORIZED`.
Approval should bind to:
- The exact artifact digest.
- The approved effect scope.
- Financial and quantity limits.
- A validity window.
- The approver's identity and role.
- The evidence and policy versions used for evaluation.
If material telemetry, constraints, rules, or proposed effects change, the
approval becomes stale.
## 10. Separation by Time Scale
### Seconds
Site and edge controllers own equipment protection, local interlocks, fast
feedback control, and prevalidated fallback behavior. There is no LLM dependency
in this loop.
### One to fifteen minutes
Streaming services perform telemetry processing, state estimation, deviation
detection, rolling forecasts, and bounded corrective optimization. Any automatic
response must be preauthorized and constrained.
### Intraday and day-ahead
The provincial platform performs market analysis, portfolio planning, scenario
comparison, operator review, and approval. This is the main operating range for
AI-assisted decision workflows.
### Monthly and annual
Planning services support contract strategy, resource acquisition, capacity
planning, model training, policy analysis, and long-horizon simulation.
## 11. Separation by Spatial Scale
### Device and site
The site retains device protocols, protection constraints, local control, and
detailed telemetry. It exposes a controlled resource capability interface
upward.
### Aggregation unit
The aggregation layer calculates site or resource-group flexibility envelopes,
manages local commitments, and translates bounded dispatch envelopes into site
plans.
### Provincial platform
The provincial layer manages portfolios, markets, operator decisions, approval,
reporting, and province-wide optimization.
### Cross-province federation
Federated platforms should exchange signed and versioned business artifacts such
as:
- Flexibility envelopes.
- Available capacity and reserve.
- Commitments and constraints.
- Bids and awards.
- Delivery and settlement facts.
Raw device control authority and unrestricted customer-level telemetry should
not cross this boundary by default.
## 12. Canonical Domain Contracts
Before selecting implementation frameworks, the project should define and
version the following objects:
- `Resource`
- `ResourceCapability`
- `Portfolio`
- `Commitment`
- `FlexibilityEnvelope`
- `ForecastBundle`
- `MarketOpportunity`
- `DecisionCase`
- `CandidatePlan`
- `ConstraintSet`
- `ValidationResult`
- `SimulationResult`
- `Approval`
- `ExecutionPermit`
- `DispatchOrder`
- `ExecutionReceipt`
- `Settlement`
- `PolicyPack`
Every material plan should contain:
- Its validity window.
- The input-data snapshot or evidence references.
- Market-rule and policy versions.
- Model and solver versions.
- Objectives and constraints.
- Uncertainty and confidence information.
- Expected physical and financial effects.
- An immutable digest.
- Current validation and approval state.
## 13. Data, Knowledge, Rules, and Memory
These four concepts should remain distinct.
### Operational data
Telemetry, market facts, commitments, user responses, and execution observations
are authoritative domain data. They require quality indicators, timestamps,
lineage, retention, and access controls.
### Knowledge
Regulations, procedures, manuals, historical reports, and operating experience
can be indexed for retrieval and explanation. Retrieved content is evidence for
a user or model, not automatically an executable rule.
### Rules
Market rules, approval rules, equipment constraints, financial limits, and
security policies should be versioned, executable, testable, and auditable.
### Memory
Agent memory should contain convenience context and curated lessons, not
authoritative operational state. Feedback should enter a controlled evaluation
and promotion process before it changes a model, policy, or operating template.
## 14. Reliability and Security Considerations
The detailed design should account for:
- Idempotent handling of repeated events, approvals, submissions, and commands.
- Optimistic concurrency to reject stale approvals and conflicting actions.
- Late, missing, duplicated, or low-quality telemetry.
- Model and solver timeouts with deterministic fallbacks.
- Network partition between provincial and site systems.
- Partial device acceptance and compensating dispatch.
- Permit expiry and revocation.
- Segregation of operator, approver, and administrator duties.
- Field-level authorization for customer, market, and infrastructure-sensitive
data.
- Immutable audit records and tamper evidence.
- Tenant, portfolio, site, and provincial data boundaries.
- Observability for every decision and external effect.
## 15. Recommended First Vertical Slice
The first implementation should be one complete operational flow rather than
five partially connected agents:
> Day-ahead portfolio plan -> forecast -> flexibility assessment -> candidate
> bids -> simulation -> operator approval -> simulated market submission ->
> execution replay -> settlement review.
The slice should initially run in shadow mode using historical and live data
without producing real external effects. It should prove:
- Canonical domain contracts.
- Evidence and decision lineage.
- Durable orchestration and case state.
- Rule and model versioning.
- Scenario comparison.
- Approval bound to immutable artifacts.
- External-effect simulation and receipts.
- Replay and outcome attribution.
- Operation without the LLM.
Once this foundation works, the proposed analysis, resource, trading,
customer-interaction, and load-management agents can be introduced as
contributors to the same case lifecycle.
## 16. Decisions Required Next
The following decisions materially affect the detailed architecture:
1. Is the new platform an AI overlay on the existing 3060 platform, a gradual
replacement, or a separate system of engagement?
2. Which existing system remains the source of truth for resources, customers,
market positions, dispatch instructions, and settlement?
3. Is the first release recommendation-only, capable of controlled market
submission, or capable of controlled dispatch?
4. Which actions may eventually become automatically authorized within
predefined limits?
5. Which functions and data must remain at the site, provincial, or dedicated
security zone?
6. What protocols and latency guarantees currently exist between the provincial
platform, aggregation units, and edge terminals?
7. Which province-specific rules must become executable policy packs in the
first release?
8. What is the acceptable behavior when forecasts, telemetry, the LLM, an
optimizer, or a downstream system is unavailable?
## 17. Suggested Next Architecture Work
After the decisions above are answered, the next design artifacts should be:
1. A system-context diagram showing existing systems and ownership boundaries.
2. A source-of-truth matrix for all major domain entities.
3. A canonical domain vocabulary and schema definitions.
4. A detailed day-ahead decision-case sequence.
5. A dispatch authority and execution threat model.
6. A latency, availability, recovery, and data-retention requirements matrix.
7. A deployment-zone and cross-province federation design.
8. An incremental roadmap built from end-to-end vertical slices.