- 01: add eight testable hard invariants (self-approval ban, duty separation, digest-bound approvals, permit-only effects, LLM-less degraded operation) - 03: approval staleness model — proposal digest, FRESH_CHECK/STALE states, APPROVED→AUTHORIZED split with revocable ExecutionPermit; gateways accept permits, never bare plans - 05: rules packaged as versioned, testable Policy Packs pinned in proposal lineage - 02: approval workbench reshaped into a Case Desk (DecisionCase as the operator's accountability unit) - 10 (new): cross-province federation boundary — signed artifacts only, plus three phase-1 no-rework reservations - 08: add M5 shadow-run milestone as phase-1 acceptance form - 00/04/07: object table, doc map, and walkthrough consistency updates - add brainstorming.md (peer review source document) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019u5SLNweVio6ozJX7yfxQr 🔮 View transcript: https://logs.lojong.info/s/e8u90k3t33w590r7b5y7yzqh
21 KiB
VPP AI Platform Architecture Brainstorming
1. Purpose
This document develops the technical architecture proposed for the virtual power plant multi-timescale collaborative operations platform. It is intended for business architects, platform engineers, algorithm teams, control-system engineers, security teams, and project stakeholders.
The goal is not to select products or implementation frameworks yet. The immediate goal is to agree on:
- The system's major trust and responsibility boundaries.
- The role AI and large language models should play.
- How decisions become approved external actions.
- How the platform should divide work across time and spatial scales.
- Which architecture should be taken into detailed design.
2. Executive Recommendation
The proposal has a strong business vision, but the platform should not be implemented literally as five autonomous agents coordinated by a single all-purpose runtime.
The recommended direction is:
- Use a deterministic, event-driven VPP operating kernel as the operational foundation.
- Present work to operators through durable operation and decision cases.
- Expose a small decision-case lifecycle API to external systems.
- Treat the five agents as business-facing roles and replaceable contributors, not as trust boundaries or direct controllers.
- Allow AI to create, compare, and explain draft decisions, but never to authorize or directly execute consequential actions.
In short:
The operational domain should remain correct, auditable, and usable when the LLM is unavailable.
3. What Is Already Strong
The proposal establishes several sound principles:
- It covers the full operational loop from monitoring and planning to execution feedback and review.
- It explicitly states that AI must not directly control terminal equipment.
- It includes rule validation, simulation, human approval, rollback, and audit.
- It builds on existing business processes, algorithms, data, and resource access.
- It evaluates success through operational and financial outcomes rather than model capability alone.
- It recognizes that annual, monthly, day-ahead, intraday, and real-time decisions require coordination.
These principles should be preserved as architectural invariants.
4. Architectural Questions in the Current Proposal
Several aspects need to be resolved before detailed system design begins.
4.1 Agent taxonomy is inconsistent
The proposal alternates between five agent roles and six business stages. Names also change between analysis, resource dispatch or organization, transaction gaming or decision-making, transaction service or user interaction, and load control or load management.
If these names become public APIs or independently deployed services, normal changes in the business model will cause expensive coupling. They are better treated as product personas or internal capability compositions.
4.2 The LLM appears too foundational
The diagrams place the Guangming power model at the bottom of the architecture. This suggests that normal operation depends on LLM availability and behavior.
Forecasting, flexibility calculation, optimization, rule enforcement, approval, market submission, and device control must remain available without an LLM. The LLM should be an advisory and orchestration participant above a deterministic operating foundation.
4.3 The runtime owns too much
The proposed Agent Runtime includes workflow orchestration, context construction, memory, tools, security, and state synchronization. This risks becoming a platform-wide bottleneck and a single failure domain.
Workflow durability, operational state, model execution, policy enforcement, approval authority, and control execution should have separate responsibilities and contracts.
4.4 Time scales are not separated operationally
Second-level equipment control and day-ahead reasoning appear in the same logical workflow. They have different availability, latency, determinism, and safety requirements.
Real-time control should live at the site or edge and use prevalidated policies. Cloud or provincial AI workflows should produce bounded plans and control envelopes rather than participate in a fast control loop.
4.5 Knowledge and executable rules are conflated
RAG is useful for retrieving regulations, operating procedures, explanations, and historical cases. It must not be the authoritative mechanism for enforcing a market rule, safety limit, or device constraint.
Consequential rules should be stored as versioned, testable, executable policy packs. RAG may explain a policy or locate its source, while a deterministic policy engine enforces it.
4.6 Decision provenance is underspecified
The architecture needs canonical, versioned representations of forecasts, constraints, flexibility, plans, approvals, commands, execution receipts, and settlement outcomes. Without these, simulation, approval, replay, and audit cannot reliably refer to the same decision.
4.7 Cross-province operation needs a federation boundary
Cross-province collaboration should not imply a single central platform with unrestricted access to raw telemetry or terminal control. The exchanged objects should normally be signed flexibility envelopes, commitments, bids, awards, and settlement facts.
5. Non-Negotiable Architecture Invariants
The following invariants should guide all designs:
- AI cannot approve its own proposal.
- AI cannot directly call a market-submission or device-control endpoint.
- Every consequential action is bound to an immutable artifact digest and a validity window.
- Approval is invalidated when material evidence, constraints, or the proposed effect changes.
- Operational decisions retain their input data snapshot, rule version, model version, solver version, and approvals.
- External effects are idempotent and produce durable receipts.
- The platform supports replay from recorded evidence.
- Degraded operation remains possible without the LLM.
- Edge control remains safe during provincial-platform or wide-area network outages.
- Operational truth is stored in authoritative domain stores, not in agent memory.
6. Three Architecture Alternatives
The following alternatives are intentionally different rather than variations of the same agent-runtime design.
6.1 Alternative A: Minimal Decision Platform
This design exposes only three lifecycle operations:
interface VppDecisionPlatform {
start(request: DecisionRequest): Promise<DecisionCase>;
act(action: CaseAction): Promise<DecisionCase>;
observe(query: CaseQuery): Promise<CaseObservation>;
}
Human requests, scheduled jobs, and operational events all create a versioned decision case. The caller never selects an agent, prompt, model, skill, or solver.
A day-ahead planning case may produce:
- Load, renewable-generation, and market-price forecasts.
- A resource portfolio and flexibility envelope.
- Candidate market bids.
- A customer response plan.
- A dispatch envelope.
- Rule and simulation results.
- Separate approval gates for market submission, outreach, and dispatch.
An authorized user approves an exact proposal version and digest. Approval does not grant the AI direct access to an external system; it allows a guarded adapter to process the authorized effect.
This design hides orchestration, retries, agent topology, model selection, tool
invocation, and integration protocols very effectively. Its main risk is that
DecisionCase and the service behind it become generic containers with
insufficient domain boundaries.
6.2 Alternative B: Event-Driven VPP Operating Kernel
This design makes the deterministic operational domain the foundation and treats AI as a replaceable adapter.
Representative ports are:
interface ObservationIngestPort {
record(observation: Observation): Promise<EvidenceReference>;
}
interface CommandPort {
execute(command: DomainCommand): Promise<CommandResult>;
}
interface ControlAuthorityPort {
authorize(
plan: CandidatePlan,
simulation: SimulationResult,
approvals: Approval[],
): Promise<ExecutionPermit | Denial>;
}
interface ControlGateway {
dispatch(permit: ExecutionPermit): Promise<ExecutionReceipt>;
}
The core invariant is:
Evidence can trigger decisions; decisions can propose effects; only the authority plane can authorize effects.
For a demand-response operation, the system records the request as evidence, calculates a forecast and flexibility envelope, creates candidate plans, runs safety simulation, collects approval, and requests an execution permit. The control gateway accepts a permit, not a raw plan.
This alternative offers the strongest replayability, auditability, safety isolation, and AI replaceability. Its cost is conceptual complexity: teams must handle commands, observations, events, idempotency, revisions, projections, and eventual consistency correctly.
6.3 Alternative C: Operations Case Desk
This design optimizes the product around the operator's unit of accountability: a bounded operational case.
Examples include:
- Prepare tomorrow's market bid.
- Deliver a 20 MW demand-response event.
- Resolve an execution deviation.
- Review a completed dispatch.
The operator-facing interface is:
interface OperationsDesk {
open(kind: CaseKind, input: CaseInput): Promise<CaseId>;
read(caseId: CaseId): Promise<CaseView>;
apply(command: CaseCommand): Promise<CommandReceipt>;
watch(caseId: CaseId): AsyncIterable<CaseEvent>;
}
Each case contains its objective, owner, horizon, evidence, assumptions, scenario branches, recommendations, approvals, execution receipts, and outcome. An advanced operator can fork a high-price or low-solar scenario without directly invoking an agent.
This design produces the strongest operator experience and naturally organizes audit and collaboration. Its principal risk is turning a case into a permanent collection of unrelated activity. Each case therefore needs one accountable objective, an explicit deadline, an owner, and a completion contract.
7. Comparison and Recommended Hybrid
The Minimal Decision Platform has the smallest learning and misuse surface for external callers. It is a deep interface, but its simplicity can conceal an overly centralized implementation.
The Event-Driven Operating Kernel is the strongest internal architecture for a safety-sensitive and heavily audited system. It makes evidence, authority, and external effects explicit. It is less convenient as the operator's primary mental model.
The Operations Case Desk best matches how business users collaborate and accept responsibility for decisions. It is not, by itself, sufficient as the low-level operating and control architecture.
The recommended hybrid is therefore:
- Event-driven deterministic kernel inside.
- Operations Case Desk for operators.
- Minimal decision-case API for external systems.
The layers complement one another instead of duplicating responsibilities.
8. Recommended Logical Architecture
flowchart TB
T[Operators / Schedules / Operational Events] --> C[Operations Case Desk and Decision API]
C --> P[Durable Process Manager and Decision Ledger]
P --> F[Forecasting Services]
P --> R[Flexibility and Resource Portfolio Services]
P --> M[Market and Dispatch Optimization]
P --> U[Customer Response Planning]
P --> S[Settlement and Review]
A[LLM Advisor and Agent Roles] --> P
A --> F
A --> R
A --> M
A --> U
F --> G[Policy Validation]
R --> G
M --> G
U --> G
G --> V[Simulation and Fresh-State Verification]
V --> H[Human Approval]
H --> E[Execution Authority]
E --> X[Market / Messaging / Control Gateways]
X --> D[Site and Edge Controllers]
D --> O[Execution Observations]
O --> P
8.1 Experience plane
The experience plane provides the operations desk, dashboards, AI-assisted investigation, scenario comparison, approval inbox, and operational reports. Pages are projections over cases and the decision ledger rather than independent workflow silos.
8.2 Orchestration plane
The orchestration plane owns durable process state, deadlines, retries, compensation, correlation, and case progression. It invokes domain capabilities but does not perform forecasting, optimization, authorization, or control itself.
8.3 Decision-services plane
Domain services provide deterministic or bounded capabilities:
- Load, generation, and price forecasting.
- Resource capability and flexibility assessment.
- Portfolio aggregation and commitment management.
- Market bidding and dispatch optimization.
- Customer segmentation and response planning.
- Simulation, deviation analysis, settlement, and attribution.
These services accept complete, versioned inputs and return typed artifacts with provenance.
8.4 AI advisory plane
The AI layer provides intent understanding, task decomposition, retrieval, explanation, report generation, tool selection, and candidate strategy generation. The five proposed agents can exist here as business-specific contributors.
There should be no executeDeviceCommand or unrestricted submitMarketBid tool
available to an agent.
8.5 Authority and execution plane
This plane enforces executable policies, separation of duties, approval requirements, current-state verification, permit expiry, and effect limits. It is the only route to market, messaging, and control gateways.
8.6 Evidence and data plane
This plane records raw observations, normalized telemetry, business events, reference data, feature values, forecasts, model outputs, decisions, and execution outcomes. It supports point-in-time reconstruction and replay.
9. Decision and Execution Lifecycle
A consequential artifact should follow an explicit lifecycle:
DRAFT
-> VALIDATED
-> SIMULATED
-> APPROVED
-> AUTHORIZED
-> ISSUED
-> ACKNOWLEDGED
-> COMPLETED / FAILED / ROLLED_BACK
The LLM may create or explain a DRAFT. It cannot advance an artifact to
AUTHORIZED.
Approval should bind to:
- The exact artifact digest.
- The approved effect scope.
- Financial and quantity limits.
- A validity window.
- The approver's identity and role.
- The evidence and policy versions used for evaluation.
If material telemetry, constraints, rules, or proposed effects change, the approval becomes stale.
10. Separation by Time Scale
Seconds
Site and edge controllers own equipment protection, local interlocks, fast feedback control, and prevalidated fallback behavior. There is no LLM dependency in this loop.
One to fifteen minutes
Streaming services perform telemetry processing, state estimation, deviation detection, rolling forecasts, and bounded corrective optimization. Any automatic response must be preauthorized and constrained.
Intraday and day-ahead
The provincial platform performs market analysis, portfolio planning, scenario comparison, operator review, and approval. This is the main operating range for AI-assisted decision workflows.
Monthly and annual
Planning services support contract strategy, resource acquisition, capacity planning, model training, policy analysis, and long-horizon simulation.
11. Separation by Spatial Scale
Device and site
The site retains device protocols, protection constraints, local control, and detailed telemetry. It exposes a controlled resource capability interface upward.
Aggregation unit
The aggregation layer calculates site or resource-group flexibility envelopes, manages local commitments, and translates bounded dispatch envelopes into site plans.
Provincial platform
The provincial layer manages portfolios, markets, operator decisions, approval, reporting, and province-wide optimization.
Cross-province federation
Federated platforms should exchange signed and versioned business artifacts such as:
- Flexibility envelopes.
- Available capacity and reserve.
- Commitments and constraints.
- Bids and awards.
- Delivery and settlement facts.
Raw device control authority and unrestricted customer-level telemetry should not cross this boundary by default.
12. Canonical Domain Contracts
Before selecting implementation frameworks, the project should define and version the following objects:
ResourceResourceCapabilityPortfolioCommitmentFlexibilityEnvelopeForecastBundleMarketOpportunityDecisionCaseCandidatePlanConstraintSetValidationResultSimulationResultApprovalExecutionPermitDispatchOrderExecutionReceiptSettlementPolicyPack
Every material plan should contain:
- Its validity window.
- The input-data snapshot or evidence references.
- Market-rule and policy versions.
- Model and solver versions.
- Objectives and constraints.
- Uncertainty and confidence information.
- Expected physical and financial effects.
- An immutable digest.
- Current validation and approval state.
13. Data, Knowledge, Rules, and Memory
These four concepts should remain distinct.
Operational data
Telemetry, market facts, commitments, user responses, and execution observations are authoritative domain data. They require quality indicators, timestamps, lineage, retention, and access controls.
Knowledge
Regulations, procedures, manuals, historical reports, and operating experience can be indexed for retrieval and explanation. Retrieved content is evidence for a user or model, not automatically an executable rule.
Rules
Market rules, approval rules, equipment constraints, financial limits, and security policies should be versioned, executable, testable, and auditable.
Memory
Agent memory should contain convenience context and curated lessons, not authoritative operational state. Feedback should enter a controlled evaluation and promotion process before it changes a model, policy, or operating template.
14. Reliability and Security Considerations
The detailed design should account for:
- Idempotent handling of repeated events, approvals, submissions, and commands.
- Optimistic concurrency to reject stale approvals and conflicting actions.
- Late, missing, duplicated, or low-quality telemetry.
- Model and solver timeouts with deterministic fallbacks.
- Network partition between provincial and site systems.
- Partial device acceptance and compensating dispatch.
- Permit expiry and revocation.
- Segregation of operator, approver, and administrator duties.
- Field-level authorization for customer, market, and infrastructure-sensitive data.
- Immutable audit records and tamper evidence.
- Tenant, portfolio, site, and provincial data boundaries.
- Observability for every decision and external effect.
15. Recommended First Vertical Slice
The first implementation should be one complete operational flow rather than five partially connected agents:
Day-ahead portfolio plan -> forecast -> flexibility assessment -> candidate bids -> simulation -> operator approval -> simulated market submission -> execution replay -> settlement review.
The slice should initially run in shadow mode using historical and live data without producing real external effects. It should prove:
- Canonical domain contracts.
- Evidence and decision lineage.
- Durable orchestration and case state.
- Rule and model versioning.
- Scenario comparison.
- Approval bound to immutable artifacts.
- External-effect simulation and receipts.
- Replay and outcome attribution.
- Operation without the LLM.
Once this foundation works, the proposed analysis, resource, trading, customer-interaction, and load-management agents can be introduced as contributors to the same case lifecycle.
16. Decisions Required Next
The following decisions materially affect the detailed architecture:
- Is the new platform an AI overlay on the existing 3060 platform, a gradual replacement, or a separate system of engagement?
- Which existing system remains the source of truth for resources, customers, market positions, dispatch instructions, and settlement?
- Is the first release recommendation-only, capable of controlled market submission, or capable of controlled dispatch?
- Which actions may eventually become automatically authorized within predefined limits?
- Which functions and data must remain at the site, provincial, or dedicated security zone?
- What protocols and latency guarantees currently exist between the provincial platform, aggregation units, and edge terminals?
- Which province-specific rules must become executable policy packs in the first release?
- What is the acceptable behavior when forecasts, telemetry, the LLM, an optimizer, or a downstream system is unavailable?
17. Suggested Next Architecture Work
After the decisions above are answered, the next design artifacts should be:
- A system-context diagram showing existing systems and ownership boundaries.
- A source-of-truth matrix for all major domain entities.
- A canonical domain vocabulary and schema definitions.
- A detailed day-ahead decision-case sequence.
- A dispatch authority and execution threat model.
- A latency, availability, recovery, and data-retention requirements matrix.
- A deployment-zone and cross-province federation design.
- An incremental roadmap built from end-to-end vertical slices.