Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bg6jx9vHNHzB91qyV64GQ7
48 KiB
VPP AI Platform — System Architecture
Virtual Power Plant Multi-Timescale Intelligent Operations Platform, built on the GuangMing Power LLM
External architecture overview · English edition · v0.1 · September 2026
1. Executive summary
The VPP AI Platform is an AI-assisted decision and controlled-execution system for operating a virtual power plant (VPP) in Hubei Province. It aggregates distributed flexible resources (photovoltaics, battery storage, EV charging, HVAC, industrial loads, microgrids and parks), currently 31 MW with a target of at least 80 MW, and operates them across the annual, monthly, day-ahead, intraday and real-time timescales against the Hubei electricity spot market, the dispatch and load-management systems, and the metering and settlement systems.
Five LLM-based agents support the six-stage business loop of monitoring, resource organisation, trading decisions, customer response, control execution and deviation review. The defining property of the architecture is that the AI proposes and explains, but never computes the numbers and never touches the control loop. Every action with an external effect passes through a deterministic safety chain of rule check, simulation, envelope or human approval, fresh-state check and execution permit before any gateway will accept it.
Three design commitments make this concrete:
- Two planes, one channel. A slow cognitive plane (LLM agents, minutes to days) is strictly separated from a fast control plane (deterministic execution, seconds). The safety chain is the only path between them.
- The LLM never produces a number. Bid prices, dispatch volumes and setpoints come only from versioned forecasting, optimisation and simulation tools. Every numeric field in a proposal is a reference to a recorded tool call. A hallucinated number cannot be written into a proposal by construction.
- Tiered autonomy through envelopes. Humans approve envelopes (policy-level bounds with validity windows) rather than every action. Actions inside an envelope pass automatically after rule check and simulation. Anything outside an envelope, or flagged by simulation, escalates to a human. Envelopes start empty and widen only on evidence.
Everything is event-sourced and replayable, which satisfies regulatory audit requirements (MLPS compliance) and, as a side effect, makes the audit trail the platform's evaluation dataset.
2. System context
flowchart TB
subgraph EXT["External systems"]
TP[Trading platform]
DP[Dispatch / load management]
MS[Metering / settlement]
WX[Weather and market data]
end
subgraph PLAT["VPP AI Platform"]
COG["Cognitive plane<br/>five agents + agent runtime"]
SAFE["Safety chain<br/>rule check → simulation → approval → permit"]
CTRL["Control plane<br/>execution engine + edge terminals"]
DATA[("Data foundation<br/>time-series · relational · knowledge · feature · event log")]
end
RES["Flexible resources<br/>PV · storage · EV charging · HVAC · industrial load · microgrids"]
EXT <--> COG
COG -->|Proposal| SAFE
SAFE -->|approved plan + permit| CTRL
CTRL --> RES
CTRL -->|execution feedback| COG
COG <--> DATA
CTRL --> DATA
Upstream, the platform interfaces with the provincial trading platform (bid submission, clearing results, market disclosures), the dispatch and load-management systems (regulation requests, capability declarations, baselines and assessments) and the metering and settlement systems (meter data, settlement statements).
Downstream, it aggregates customer-side resources through a four-level spatial hierarchy: province, aggregation unit, customer site, device.
Internally, operators work through a Case Desk, where every piece of AI output is attached to an accountable operating case with an objective, an owner and a deadline.
The platform is an overlay on the existing "3060" platform rather than a replacement. It reuses the existing master-data services (customer, asset, device and site registries) and the existing identity and permission system, and it converges every bid or control action into a single audited channel so that no action can bypass the safety chain.
3. Architecture principles
Nine principles govern every design decision. They are the review criteria for any module or change.
| # | Principle | In one sentence |
|---|---|---|
| P1 | Two-plane separation | The LLM never appears on the second-level control path. |
| P2 | The LLM does not compute | All numbers come from the professional model layer and enter proposals untouched. |
| P3 | Envelope autonomy | Humans approve envelopes; in-envelope actions auto-pass; out-of-envelope actions escalate. |
| P4 | Object-based communication | Agents exchange typed, versioned business objects, never free text. |
| P5 | Everything is auditable | Event sourcing by default; every proposal carries full lineage. |
| P6 | Provider abstraction | LLM calls go through a capability-oriented layer; the backend is swappable. |
| P7 | Constraint cascade | Annual to real-time is one cascade of constraints held in a shared position ledger. |
| P8 | Tiered data storage | Five stores with explicit writers and readers, not one database for everything. |
| P9 | Edge autonomy | Edge terminals close the loop locally and keep running through disconnection. |
The principles are backed by eight hard invariants, each of which has an automated test:
| # | Invariant |
|---|---|
| I1 | The AI cannot approve its own proposal. The approver identity is separated from the origination chain. |
| I2 | Operation, approval and administration roles are separated in the permission system. |
| I3 | Approvals bind to an immutable proposal digest and a validity window. Material change in evidence invalidates the approval. |
| I4 | External effects are accepted only against an execution permit that is short-lived and revocable. |
| I5 | External effects are idempotent and produce a durable receipt. |
| I6 | The system degrades gracefully when the LLM is unavailable. Forecasting, optimisation, checking, approval, bidding and control do not depend on the LLM being up. |
| I7 | Any historical decision can be replayed in full from retained evidence. |
| I8 | The authoritative store for operational facts is the domain database, not agent memory. |
These invariants were independently converged on by a separate peer review of the design, which we take as evidence that the architectural direction is sound.
4. The two-plane architecture
| Cognitive plane | Control plane | |
|---|---|---|
| Timescale | Minutes to days (collaborative decision response within 3 minutes) | Seconds |
| Core components | Five LLM agents and the agent runtime | Execution engine and edge control terminals |
| Decision mode | LLM orchestration plus professional-model computation | Deterministic rules, execution within pre-approved envelopes |
| Output | Proposals, plans, operating envelopes | Control commands, execution feedback |
| LLM involvement | Yes: intent understanding, task planning, tool orchestration, explanation | No, never |
The cognitive plane is deployed in the management information zone. The control plane sits in the production control zone. The only crossings between zones are the safety chain's one-way controlled delivery of approved plans and the return path for telemetry and execution feedback. This directly implements the mandate that AI never issues control commands.
5. Cognitive plane
5.1 Agent runtime
The runtime is the operating system of the cognitive plane. The five agents are the business processes that run on it. A key and somewhat counter-intuitive design point: workflows are the entry point, not agents. Every business process (day-ahead bidding, review, intraday correction) is a deterministic chain of steps. Agents are invoked only at specific steps and do not roam freely.
flowchart LR
subgraph TRIG["Task triggers"]
T1[Manual<br/>natural language, ad-hoc analysis]
T2[Event<br/>deviation, PV anomaly, resource offline]
T3[Scheduled<br/>day-ahead forecast, bid window, review]
end
TRIG --> IR["Intent parsing (LLM)<br/>manual triggers only"]
TRIG --> WF
IR --> WF["Workflow engine<br/>deterministic, durable, resumable"]
WF --> AG[Agent steps]
AG --> TL[Tool registry<br/>typed, versioned, lineage-recording]
AG <--> MEM[Memory service]
WF --> OUT[Proposal / business output]
WF -.-> LOG[(Event log)]
The three trigger sources all converge on the same thing: instantiate a workflow template with parameters.
- Scheduled triggers (the daily bid, the D+1 review) instantiate their template directly and never pass through the LLM.
- Event triggers are evaluated by deterministic rules on the event bus and instantiate the matching template.
- Manual triggers are the only entry point that passes through LLM intent parsing. The router agent's output is constrained by schema to "one of the registered template IDs plus parameters". The LLM can choose a wrong template, which is correctable and auditable, but it cannot invent a process.
This compresses non-determinism to the smallest necessary scope: the LLM decides what to do, the workflow engine decides how to do it, step by step.
Two further disciplines are enforced in code. Context for each agent step is assembled deterministically by upstream steps (database queries, ledger reads, knowledge retrieval), so decisions are reproducible. And agent outputs carry numeric fields only as references to recorded tool calls, which the proposal assembler dereferences. The LLM cannot type a number into a proposal.
5.2 The five agents
Each agent is defined by a role, a context-building strategy, a set of workflow templates, a whitelisted tool set, a memory view and the business objects it produces.
| Agent | Responsibility | Produces | Phase |
|---|---|---|---|
| Intelligent analysis | Situational awareness, forecast orchestration, anomaly detection, attribution, risk warning. Leads the D+1 review workflow. | SituationReport |
1 |
| Resource dispatch | Resource onboarding and profiling, flexibility potential assessment, grouping, dispatch plan generation. Capacity verification must reference the latest execution feedback. | ResourceProfile, dispatch-plan proposals |
1 |
| Trading | Market rule adaptation (retrieval over the rule base), price assessment, bid optimisation, revenue and risk estimation, bid generation. Must read the position ledger for upper-timescale constraints before bidding. | Bid proposals, PositionUpdate |
1 |
| Customer interaction | Customer profiling, response willingness, segmentation, invitation strategy, incentive matching, fulfilment tracking. Batch outreach is an external action and defaults to human approval. | Invitation proposals, Commitment |
2 |
| Load control (proposal only) | Control target decomposition, control plan generation, deviation monitoring, correction proposals. Output always stops at a proposal; execution belongs to the control plane. | Control-plan proposals | 2 |
Review (deviation attribution after settlement) is deliberately not a sixth agent. It is a cross-agent workflow led by the analysis agent whose authoritative inputs are the position ledger and the event log.
5.3 Memory
| Layer | Content | Storage | Lifetime |
|---|---|---|---|
| Working memory | Current task context and intermediate results | Runtime memory | Task |
| Episodic memory | Operating states, trading results, control results, customer responses | Event log and time-series store | Permanent (audit) |
| Semantic memory | Structured lessons from review: forecast bias patterns by weather type, customer reliability, strategy effectiveness | Knowledge store and strategy library | Long-term, versioned |
The platform does not feed production decisions from an agent framework's automatic memory recall. Which lessons are injected into a workflow is declared explicitly by the template (for example, "forecast deviation patterns on comparable weather days" for the bidding flow), so it is testable and reviewable.
5.4 Case Desk: the human interface
The operator's unit of work is not "an agent's output" but a DecisionCase: one business matter with a single objective, an owner, a deadline and a completion condition. A case aggregates, by reference, all evidence (situation reports, forecasts, profile snapshots), the related proposals and their states, approvals and permits, scenario branches, and the outcome.
- Cases are opened automatically by the trigger sources (a daily bid case, a deviation case) and can be opened manually for special analyses.
- Approval inboxes, dashboards and reports are projections over cases, the ledger and the event log, not separate feature silos.
- Operators can fork scenarios inside a case ("what if tomorrow has a high-price window?"). Agents rerun the calculation tools to produce comparison branches, but scenario forks never produce new proposals.
- Free text exists only at the human-to-platform boundary. Agent-to-agent communication is always typed objects.
- AI insight cards embedded in existing business pages expand to show lineage: conclusion, referenced tool outputs, data sources.
6. The safety chain
The safety chain is the only channel between the cognitive and control planes and, in the context of MLPS compliance and dispatch regulation, the platform's real product. Reviews and acceptance focus on the completeness and traceability of this chain.
6.1 The Proposal object
Every intended action, whether a market bid, a customer invitation, a control plan or a dispatch plan, is a Proposal. It carries a typed payload, full lineage (the originating agent and trigger, every tool call with version and input/output snapshot references, data snapshot references, the ledger version read), the matched envelope if any, an immutable content digest of payload plus lineage, and a complete audit trail of state transitions.
Hard constraint: any numeric field in the payload must be traceable to a tool call in the lineage. This is the machine-checkable form of "the LLM does not compute", and the rule-check stage enforces it.
6.2 State machine
stateDiagram-v2
[*] --> DRAFT: agent generates
DRAFT --> RULE_CHECK: submit
RULE_CHECK --> REJECTED: rule failure
RULE_CHECK --> SIMULATION: pass
SIMULATION --> PENDING_HUMAN: simulation warning (always escalates)
SIMULATION --> ENVELOPE_CHECK: pass
ENVELOPE_CHECK --> AUTO_APPROVED: within envelope
ENVELOPE_CHECK --> PENDING_HUMAN: outside envelope / no envelope
PENDING_HUMAN --> APPROVED: human approval (may be multi-level)
PENDING_HUMAN --> REJECTED: human rejection
AUTO_APPROVED --> FRESH_CHECK: pre-release check
APPROVED --> FRESH_CHECK: pre-release check
FRESH_CHECK --> AUTHORIZED: pass, ExecutionPermit issued
FRESH_CHECK --> STALE: material state changed
STALE --> RULE_CHECK: re-enter chain
AUTHORIZED --> RELEASED: gateway accepts against permit
RELEASED --> EXECUTING
EXECUTING --> COMPLETED
EXECUTING --> ROLLED_BACK
REJECTED --> [*]
COMPLETED --> [*]: ExecutionReport → review
ROLLED_BACK --> [*]: incident report → review + alert
The stages:
- Rule check is a deterministic policy engine. It validates compliance (trading rules, bid format), constraints (resource limits, contract boundaries, ledger consistency), permission (is this agent allowed to originate this proposal type) and lineage completeness. It also enforces separation of duties: the approver may not appear in the proposal's origination chain, and there is no code path by which the AI reaches an approval transition.
- Simulation is chosen by proposal type: revenue and risk scenario backtesting for bids, power-balance or power-flow simulation for control plans. A simulation warning always escalates to a human, without exception.
- Envelope check matches the proposal against approved envelopes. Within bounds, it is auto-approved with a full record. Otherwise it suspends for human approval.
- Fresh-state check and permit are described below.
The chain is implemented once, as a single durable workflow shared by all proposal types. Approvals can hang for hours, so the workflow persists its suspended state and survives process restarts and multi-replica deployment. Timeouts are configurable, for example an escalation alert if a bid is still unapproved 30 minutes before the submission deadline.
6.3 Approval binding and staleness
An approval is not "approval of this proposal". It is "approval of this digest version of the proposal, valid in this window". The approval records the approver's identity and role from the existing permission system, the approved scope and limits, the validity window, and the versions of evidence relied on (policy pack, ledger version, data snapshots).
An approval becomes invalid when its window expires or when material evidence changes: a referenced resource goes offline, the ledger position moves, the policy pack is upgraded, or telemetry departs significantly from the snapshot at approval time. This is detected by the fresh-state check immediately before release.
6.4 Execution permits
Approval is not permission to execute. Hours can pass between approval and execution, and the physical world changes. Immediately before release, the authority service re-runs the critical checks (resources online, constraints still satisfied, approval still valid) and, if they pass, issues an ExecutionPermit: bound to the approved digest, short-lived (minutes for control plans, until the deadline for bids), with effect limits narrowed from the approval scope, and revocable.
Gateways accept permits, not plans. The trading-platform adapter and the execution engine accept only a (Proposal, ExecutionPermit) pair with a matching digest and an unexpired, unrevoked permit. Even if an approved plan were leaked or replayed, it could produce no external effect without a valid permit.
A typical sequence: a bid is approved at 08:30; at 09:00 a 10 MW storage station goes offline; at 09:30 the adapter requests a permit; the fresh-state check finds the verified capacity is no longer valid; the proposal becomes STALE and re-enters the chain with reduced capacity, where it is either re-approved or lands inside the envelope.
6.5 Envelopes and the autonomy dial
An Envelope is a human-approved, multi-level-signed authorisation with a scope (proposal type, timescales, resource set), type-specific bounds (for bids: price and volume ranges, position deviation; for control: per-customer and total curtailment limits, allowed windows; for invitations: eligible customer groups, incentive caps), a validity window, and escalation conditions under which it automatically suspends (a streak of in-envelope executions with deviations over threshold, or a review finding recommending contraction).
This is the same trust model that grid dispatch already applies to AGC units, which are autonomous within a regulation dead band. It is an extension of an existing model to VPP resources, not a new one.
| Stage | Envelope state | Effect |
|---|---|---|
| Initial launch | Empty | Every proposal goes to a human. Equivalent to the traditional mode. |
| Trial operation | Narrow (for example bid prices within a small tolerance of the forecast baseline) | Routine days are automatic, volatile days are manual. |
| Mature operation | Widened progressively on review evidence | Humans handle only exceptions and envelope renewals. |
Both widening and narrowing are driven by review data and recorded. The growth of the platform's autonomy is itself auditable.
6.6 Rollback
Permit revocation is the highest-priority rollback: the gateway immediately refuses subsequent commands and edge terminals fall back to the previous approved plan or a safe curve. Any anomaly during execution (check failure, deviation over threshold, channel fault) triggers a fallback: control plans revert to the last approved plan or a safe shutdown curve held locally at the edge; bids can be withdrawn and resubmitted before the deadline, or enter deviation management after it. Rollbacks are recorded events and mandatorily generate a review task.
7. Control plane
The control plane is deterministic. It contains no LLM and no "intelligence", only reliable execution within approved plans and envelopes.
7.1 Spatial hierarchy
| Level | Decides | Timescale | Autonomy |
|---|---|---|---|
| Province | Aggregate plans, envelopes, allocation across aggregation units | 15 minutes to a day | — |
| Aggregation unit | Rebalancing within the unit, substitution of failed resources | Minutes | Rebalancing within the unit envelope |
| Customer site | Device-group coordination, local closed loop | Seconds to minutes | Continues through disconnection |
| Device | Setpoint execution | Seconds | Device self-protection |
The province issues only targets plus envelopes at aggregation-unit granularity. Device-level allocation happens locally in the aggregation layer. Aggregation units report their adjustable capability upward as a FlexibilityEnvelope, the same artefact used for inter-provincial exchange. The province does no device-level micromanagement, which reduces communication and latency requirements and matches the zoning requirements of the security regime.
7.2 Execution engine
The engine takes proposals in RELEASED state and performs plan decomposition to aggregation units, delivery over an encrypted, authenticated, idempotent channel, execution tracking against plan, deviation handling and fallback. When a deviation exceeds threshold, the fast path is deterministic: rebalance within the current approved envelope at the aggregation layer, and emit an event to the cognitive plane. The LLM participates only in the next round of plan correction, never on the fast path. The engine outputs an ExecutionReport to the event log for review.
7.3 Edge terminals
Each edge terminal holds a local copy of the current approved plan, envelope and safe shutdown curve, updated atomically on every delivery. It closes the loop locally at second resolution and enforces device protection constraints regardless of upstream commands. On disconnection from the province, it continues executing the last approved plan to the end of its period, then switches to the safe curve. Actions during disconnection are logged locally and back-filled on reconnection. Loss of contact with the provincial platform never means loss of control over resources.
7.4 Timescale cascade
Annual, monthly, day-ahead, intraday and real-time are not five systems but one cascade in which each layer's decisions become the next layer's constraints. Cascade state lives in the shared position ledger. All agents read and write it through the runtime, and no agent may hold private position state. Rule check enforces cascade integrity: any lower-timescale proposal must reference the current ledger version and may not breach upper-timescale constraints. Deviation review is, by construction, a per-timescale comparison of ledger plan values against metered actuals, so the ledger doubles as the settlement reconciliation backbone.
8. Professional model layer and data foundation
8.1 Skills
Every skill is a typed, versioned, stateless service registered in the runtime's tool registry. The LLM can call only registered tools, and every call automatically records input and output snapshot references plus the version into proposal lineage. Internally a skill can be anything: statistical models, deep learning, mixed-integer programming, a simulation engine.
| Category | Skills | Notes |
|---|---|---|
| Forecasting | Short-term load, PV output, spot price | 96-interval curves with quantile intervals (P10/P50/P90). Uncertainty intervals are mandatory because bid risk estimation depends on them. |
| Assessment and optimisation | Flexibility potential assessment, bid optimisation (MILP), aggregation and dispatch optimisation (MILP), command decomposition, incentive optimisation | Take ledger constraints and risk preferences as inputs. |
| Simulation and validation | Revenue scenario backtest, power-balance and power-flow simulation, bid format validation | Used by the safety chain. |
| Analysis and generation | Attribution analysis, anomaly detection, report generation | The LLM contributes prose; every number is a reference to a tool output. |
8.2 Five stores
| Store | Content | Writers | Readers |
|---|---|---|---|
| Time-series | Load, generation, price and weather telemetry | Ingestion pipeline, edge upload | Forecast skills, dashboards |
| Relational | Customer records, contracts, asset registry, market results, position ledger | Business applications, settlement ingestion | All |
| Knowledge and vector | Trading rules, policies, device constraint documents, cases (retrieval-augmented generation) | Knowledge operations | Trading agent, rule-check support, Q&A |
| Feature | Forecast-model features, consistent online and offline | Feature pipeline | Forecast training and inference |
| Event log | All business events, proposal transitions, audit | Runtime (event sourcing) | Review, audit, regulatory evidence |
Governance points: master data management via the existing registries; a data quality gate on everything entering the feature store (missing rate, jumps, timestamp alignment), with data quality propagated through lineage so proposals built on degraded data are flagged and treated more conservatively by the envelope check; immutable, checksummed snapshots for everything a proposal references; and tiered de-identification for customer behaviour data.
8.3 Policy packs: knowledge vs. rules
Knowledge (retrievable) and rules (executable) are strictly separated. Retrieval is used to explain rules, locate sources and support questions. Any consequential check (market rules, safety limits, device constraints, financial caps) is enforced by a policy pack: a versioned, tested, deployable bundle of executable rules, each with positive and negative test cases and references back to the source text in the knowledge base. Proposal lineage records the policy pack version used, so "which rules were applied at the time" is always answerable and replayable. Upgrading a policy pack triggers an impact analysis that scans pending and in-envelope proposals for re-check. Every bid rejected by the trading platform is fed back as a candidate knowledge entry and rule.
9. External integration
All external systems connect through anti-corruption-layer adapters. External message formats are translated into internal business objects inside the adapter, and external field semantics never leak into the core domain model. Every adapter records both the raw message and the translated object. Availability assumptions are pessimistic by design: timeouts, idempotent retries and degraded paths are all explicit.
| Interface | Exchanges |
|---|---|
| Trading platform | Bid submission and modification (outbound, after release against permit); clearing results into the ledger; market disclosures into the time-series store. Degraded path: if the interface fails, the platform exports a standard bid file for manual upload. This path is drilled. |
| Dispatch and load management | Regulation requests as inbound events; adjustable capacity declarations outbound; baselines and assessment results inbound for review. |
| Metering and settlement | Meter data into the time-series store and the ledger's actuals; settlement statements for revenue review. Meter-data arrival is the trigger for the D+1 review workflow. |
| Weather and market data | Read-only sources through the ingestion pipeline with source tagging and quality gating. |
Channel convergence is a key governance step at integration: after launch, even manual bids go through the platform. A human creates the proposal directly, skipping agent generation but not the checks and the audit trail. Otherwise the ledger and the audit have blind spots.
10. Reference scenario: day-ahead spot bidding
One business day exercises every component. Market times below are placeholders pending confirmation of Hubei spot-market rules.
| Time | Event |
|---|---|
| D-1 06:00 | Scheduled trigger instantiates the day-ahead situation workflow with no LLM intent parsing. The analysis agent calls the load, PV and price forecast tools, runs anomaly detection over yesterday's telemetry, attributes significant movements ("high temperature tomorrow, HVAC load expected up"), and publishes a SituationReport to the event bus. |
| D-1 07:00 | The resource dispatch agent wakes on the report. The flexibility assessment tool verifies the next day's adjustable capacity per interval, using the latest execution feedback. A storage station with an abnormal state of charge has its capacity reduced and an alert event is raised. |
| D-1 08:00 | The trading agent retrieves the currently effective bidding rules, reads the position ledger for the monthly contract position and the resulting day-ahead volume bounds, assembles the optimiser input (price forecast intervals, verified capacity, ledger constraints, risk preferences), and calls the bid optimisation MILP. The LLM writes the bid explanation for the approver, referencing the solver's numbers verbatim. A bid proposal is assembled with full lineage. |
| D-1 08:30 | Safety chain. Rule check passes format, price limits, ledger consistency and lineage completeness. Simulation backtests the bid over a large set of price scenarios and finds the tail loss within tolerance. The envelope check finds the curve within the approved bid envelope and auto-approves. Fresh-state check re-verifies resources and ledger, a permit is issued, and the adapter submits the bid against it. If the weather were extreme and the optimiser's result outside the envelope, the trader would instead see the explanation, the expanded lineage and the simulation result in the Case Desk and approve, modify or reject. |
| D-1 16:00 | Clearing results arrive: the awarded schedule and clearing prices are written to the ledger's day-ahead layer. The resource dispatch agent decomposes the award into a dispatch plan proposal, which goes through the chain (with power-balance simulation) and, once approved, is decomposed to aggregation units and delivered to edge terminals with envelope parameters. |
| D | Edge terminals execute the 96-interval plan locally. At 10:42 an industrial customer's load deviates beyond threshold. The execution engine's deterministic fast path rebalances within the aggregation unit's envelope using storage, and emits an event. On the slow path, the load control agent generates an intraday correction proposal, which passes the chain and updates the setpoint sequence. If the correction would breach the envelope, it escalates to a human, and the customer interaction agent drafts a temporary invitation proposal, which defaults to human approval. |
| D+1 | Meter data arrives and triggers the review workflow. Plan versus actual is compared per timescale (bid, clearing, execution, metering). Attribution decomposes the deviation into forecast, response and control components. A ReviewFinding writes back to semantic memory ("HVAC load is systematically under-forecast on hot days"), to the customer's reliability score, and as an envelope recommendation ("26 of 26 in-envelope executions within tolerance this month, consider widening the price tolerance"), which triggers an envelope re-approval workflow that a human must approve. |
11. Implementation
11.1 Technology stack
| Layer | Choice | Rationale |
|---|---|---|
| Agent runtime, agents, workflows, domain schemas, deterministic services | TypeScript on Mastra | Typed tool contracts and business objects, durable workflows with suspend/resume for hours-long approvals |
| Professional-model skills | Python HTTP services (forecasting, MILP optimisation, simulation) | The ML and optimisation ecosystem lives in Python, and it suits the partner research teams |
| Contracts | JSON Schema exported from the TypeScript domain package, pydantic models generated from it, golden fixtures validated by both sides in CI | One source of truth across languages; the contract is the organisational boundary between teams |
| Event bus | PostgreSQL transactional outbox in phase 1, relayed to Kafka in phase 2 behind an unchanged interface | Phase-1 volume is bid-frequency, not telemetry-frequency |
| Storage | PostgreSQL with pgvector, TimescaleDB or IoTDB, object storage for archives | Tiered storage per P8 |
| Front end | AI cards embedded progressively in existing business pages, plus new approval and dashboard views | Overlay, not replacement |
Data representation rules are part of the contract: money, energy and prices as fixed-point decimal strings, units in field names, ISO 8601 UTC timestamps, market intervals as date plus interval index, string identifiers, upper-case string enumerations.
Roughly a third of the platform is provided by the agent framework. The differentiating assets are custom-built: the policy engine, the envelope model and checker, the fresh-state check and permit service, the position ledger, the Case Desk with multi-level approval inbox, the external adapters, the execution engine and edge protocol, and lineage assembly.
11.2 LLM provider abstraction
All LLM calls go through a capability-oriented interface (complete, plan, extract, explain) and a model router that selects a backend by task type and environment: commercial APIs in development, local open-weight models in pre-production, the GuangMing Power LLM adapter in production. Business code does not change when the backend changes.
The abstraction is designed to the lowest common denominator of capabilities. Until the GuangMing interface specification is confirmed, it assumes an OpenAI-compatible protocol, a conservative context length and no native tool calling, with a degraded prompt-plus-JSON-parsing path. Switching a backend is gated by evaluation, not judgment: the full L1 evaluation set plus an L4 shadow comparison against the commercial baseline (see section 13).
11.3 Deployment
┌─ Provincial cloud (management information zone) ───────────────────┐
│ Cognitive plane (runtime + agents) LLM inference (GuangMing / local) │
│ Skill services Data foundation Safety chain Case Desk │
└──────────────┬──────────────────────────────────────────────────────┘
│ one-way controlled interface (approved plans / envelopes)
┌─ Production control zone ──────────────────────────────────────────┐
│ Execution engine Controlled-channel gateway Bypass monitoring │
└──────────────┬──────────────────────────────────────────────────────┘
│ encrypted channel
┌─ Edge ─────────────────────────────────────────────────────────────┐
│ Aggregation-unit edge nodes → local CPS / data gateways → devices │
└────────────────────────────────────────────────────────────────────┘
Forecast-model training uses a cloud GPU cluster. LLM inference runs on a provincial local inference engine. Second-level control depends on no inference service at all.
12. Security and compliance
- Zoning. The cognitive plane resides entirely in the management zone. Only the safety chain's one-way delivery and the telemetry return path cross into the control zone.
- Channel security. Controlled channels are encrypted end to end with mutual authentication. Control commands are signed and verified at the edge before execution.
- Bypass monitoring. Anomalous command patterns on the channel (frequency, magnitude, out-of-range) are monitored independently of business logic. Alerts reach humans directly and can trip the channel, which puts the edge into its disconnection strategy. The system remains safe when the channel is tripped.
- Separation of duties. Operation, approval and administration roles are separated in the existing permission system. A single stolen credential is insufficient to complete the chain.
- Prompt injection defences. Tool whitelists, schema-constrained router output, external content tagged as data rather than instructions, and a red-team case library regressed in CI.
- Physical blast radius. The platform generates proposals only within contractually agreed adjustable ranges, and every command passes both permit verification and device-level protection at the edge. The physical blast radius is bounded by contracted interruptible capacity. This is the essential difference from a traditional dispatch system.
13. Evaluation
Because every decision is replayable from retained evidence, the audit architecture is also the evaluation infrastructure. Evaluation is layered so that a problem at the top can be localised to its cause.
| Layer | Object | Method |
|---|---|---|
| L4 | End-to-end decision quality | Shadow-run comparison, historical backtest revenue, counterfactual attribution |
| L3 | Safety chain correctness | Policy-pack tests, red-team cases, invariant tests |
| L2 | Professional-model skills | Forecast accuracy and interval calibration, solver quality, assessment accuracy (the core KPIs) |
| L1 | LLM tasks | Intent routing accuracy, structured-output compliance, retrieval faithfulness, numeric consistency between explanation text and lineage |
An LLM-as-judge is used only for readability regression and never to adjudicate numeric correctness. Interval calibration is assessed separately from point accuracy, because an accurate point forecast with an over-narrow interval makes risk estimation systematically optimistic, which is a more dangerous failure than point error.
Evaluation is a change gate. Prompt changes require the L1 task set and an L4 backtest without regression. Backend switches require the full L1 set and an L4 shadow comparison. Skill upgrades require the skill's L2 metrics. Policy-pack upgrades require all pack tests green plus impact analysis. Envelope widening is approved only on L4 shadow or online evidence. No data, no widening.
The programme's headline KPIs map to measurable definitions: forecast mean error within 8 percent, flexibility assessment accuracy of at least 90 percent, collaborative decision response within 3 minutes (P95 from event trigger to a proposal reaching pending or auto-approved, excluding human approval time), dispatch command execution success of at least 98 percent, overall market revenue improvement of 15 percent against a frozen human baseline, and cross-regional resource matching accuracy of at least 85 percent measured at the federation artefact level. The exact statistical definitions are to be confirmed with the operations team before acceptance.
14. Risk management and kill-switch hierarchy
The five risks most likely to cause major loss, in order:
- Automation complacency. After months of correct AI output, human approval degrades to a rubber stamp. Countermeasures: approval-duration analysis with spot checks on suspiciously fast approvals, periodic drill samples (deliberately bad proposals inserted into the approval flow), approver rotation and dual review, and, most effectively, envelope widening that removes routine items from the human queue so humans see only real exceptions.
- Correlated extreme days. Hot-day forecast error, customer non-response and price spikes occur together. Countermeasures: extreme-weather warning linkage, automatic envelope contraction to a conservative tier, D-2 manual rehearsal, and extreme-day scenarios in L4 regression.
- Stale resource profiles leading to systematic over-commitment. Countermeasures: a commitment buffer coefficient, fresh-state check interception of known failed resources, and monitoring of bid-to-verified capacity ratios.
- Rule drift between market rule revisions and the policy pack. Countermeasures: rule announcement monitoring, mandatory feedback of rejected bids, rejection-rate alerts.
- Dependency slippage on the GuangMing LLM integration and external interface coordination. Countermeasure: phase-1 acceptance explicitly does not depend on GuangMing integration.
The graduated kill-switch hierarchy, from precise to total:
| Level | Scope | Effect |
|---|---|---|
| L0 | Single action: permit revocation | Gateway refuses the command, edge falls back (seconds) |
| L1 | Single envelope: suspension | That action class returns to full manual approval (minutes) |
| L2 | Financial: loss circuit breaker | All auto-approval frozen, only position-reducing proposals allowed (minutes) |
| L3 | Channel: control or bid channel trip | Edge enters disconnection strategy, bids switch to the manual channel |
| L4 | Whole platform: AI proposals disabled | Pure manual operation. Scheduled workflows still run and produce data but generate no proposals. Business continuity at pre-launch level. |
Each level specifies who may trigger it, quantified trigger conditions and quantified recovery conditions, and at least one level is drilled each quarter. After any level trips, the system state is explainable from the ledger and the event log, so recovery never requires guesswork.
15. Inter-provincial federation (phase 2)
Cross-provincial collaboration is federation, not centralisation. Each provincial platform is an independent trust and responsibility domain. What crosses the boundary is signed, versioned business artefacts, never raw telemetry, customer-level data or control authority.
| Exchanged artefact | Content |
|---|---|
FlexibilityEnvelope |
Adjustable capability per interval (capacity, ramp, duration, price intent), without composition details |
Commitment |
A commitment to a capacity in a window, with constraints |
| Bids and awards | Inter-provincial market artefacts via the national and regional dispatch trading system |
| Delivery and settlement facts | Metered delivery confirmations and settlement statements, after the fact |
A federation gateway is the only inter-provincial channel: it signs and verifies artefacts, negotiates schema versions, deduplicates and records. Artefacts received from another province are treated as external evidence. They enter the event log and may trigger local workflows, but they never enter the local safety chain directly. A commitment request from another province goes through the full local check, approval and permit chain. The same artefact semantics used between the province and its aggregation units are reused between provinces with a peer trust relationship instead of a hierarchical one.
Phase 1 builds no federation but reserves three no-rework points: FlexibilityEnvelope is a first-class domain object from the start, the ledger supports an external-commitment entry type, and the policy engine supports multiple coexisting packs selected by scenario.
16. Delivery status and roadmap
Phase 1 (system development and in-province capability, December 2025 to May 2026) is organised as five milestones, sequenced so that the demonstrable bidding loop lands earliest and nothing in M1 to M3 depends on external-party scheduling.
| Milestone | Scope | Status |
|---|---|---|
| M1 | Domain schemas, cross-language contract pipeline, ingestion skeleton, time-series and relational stores, immutable snapshots, ledger with constraint cascade | Delivered |
| M2 | Load, PV and price forecasting with calibrated quantiles, bid optimisation MILP, report generation, L2 evaluation harness with committed baseline | Delivered (baseline on synthetic data) |
| M3 | Agent runtime, analysis and trading agents, the full safety chain with durable suspend/resume, policy engine with the first Hubei bidding pack, envelope and authority services, Case Desk with approval inbox, bid release as file export | Delivered; all eight invariant tests pass, including the LLM-down run and AI self-approval rejection |
| M4 | Resource dispatch agent, flexibility assessment, award decomposition, envelopes live end to end (bounded auto-approval, deviation-streak suspension, re-approval workflow), review workflow with ReviewFinding write-backs, AI insight cards |
Delivered |
| M5 | Shadow run: full loop on live data with all external effects simulated, daily shadow-vs-human-vs-hindsight report, KPI dashboard, loss circuit breaker and kill-switch levels L0 to L2 drilled | Next |
Phase-1 acceptance is the shadow run: 20 or more consecutive shadow days with complete lineage and an auto-generated KPI report. Storage in the current build uses file-backed reference semantics; PostgreSQL and TimescaleDB adapters and the programmatic trading-platform channel follow.
Phase 2 (resource onboarding and roll-out) adds the customer interaction and load control agents, live edge control links, programmatic bid submission, live dispatch and metering interfaces, the federation gateway, and the path from shadow run to controlled live bidding with per-item human approval and then gradual envelope widening.
Open items requiring business-side input
The architecture deliberately does not invent business parameters. The following are tracked as named configuration with placeholders until confirmed by the operations, trading and governance teams: Hubei day-ahead window and clearing times; whether proxy purchase and VPP bids share a channel; intraday market mechanics; deviation assessment rules; the form of long-term position constraints on day-ahead bids; envelope baselines, tolerances and suspension thresholds; approval level mapping to roles; the daily loss budget; the commitment buffer coefficient; extreme-day triggers; kill-switch authorities; KPI statistical definitions including the frozen baseline for the 15 percent revenue target; and the GuangMing LLM interface specification.
Appendix A. Canonical business objects
| Object | Meaning | Producer → consumer |
|---|---|---|
Proposal |
Any action the AI wants taken, with immutable digest and lineage | Agents → safety chain |
Envelope |
Pre-approved autonomy bounds (authority, downward) | Human approval → safety chain |
FlexibilityEnvelope |
Adjustable capability report (capability, upward; federation artefact) | Aggregation units / resource agent → trading agent, federation |
ExecutionPermit |
Short-lived, revocable execution credential | Authority service → gateways |
Approval |
Human decision bound to a proposal digest and validity window | Case Desk → safety chain |
PositionLedger / PositionUpdate |
Shared per-timescale position and constraint cascade | Trading agent ↔ runtime |
DecisionCase |
Operator's unit of accountability | Triggers → Case Desk |
PolicyPack |
Versioned, tested, executable rule bundle | Knowledge operations → policy engine |
SituationReport |
State, anomalies and risk assessment | Analysis agent → all |
ResourceProfile |
Capacity, constraints and reliability score | Resource agent → trading, load control |
ForecastBundle |
Curves with quantile intervals and model version | Forecast skills → agents |
Commitment |
Customer or inter-provincial delivery commitment | Interaction agent / federation |
ExecutionReceipt / ExecutionReport |
Idempotent effect receipt / execution outcome | Gateways, execution engine → review |
ReviewFinding |
Structured lesson from review, written back to memory | Review workflow → memory, profiles, envelopes |
EventEnvelope |
Typed event wrapper with schema version and causation chain | Runtime → event log |
Appendix B. Programme context
| Item | Detail |
|---|---|
| Lead organisation | State Grid Hubei Integrated Energy Service Co., Ltd. |
| Partners | State Grid Hubei Electric Power Research Institute; Wuhan University Institute of Data Intelligence; Wuhan Zhiwang Xingdian Technology Development Co., Ltd. |
| Positioning | Hubei demonstration, Central China roll-out, nationally replicable |
| Foundation model | GuangMing Power LLM (State Grid), accessed through the provider abstraction layer |
| Compliance regime | Multi-Level Protection Scheme (MLPS), production and management zone separation |
This document is a distillation of the platform's internal design documentation. Field-level definitions live in the platform's domain schema package and its exported JSON Schema contracts.