power-market-trading-docs/markets/china/TsingRoc_verification_recap.md

141 lines
16 KiB
Markdown
Raw Normal View History

# TsingRoc (清鹏智能) — Claim Verification & Research Notes
*Reference recap of a fact-checking session. Covers: verification of the original TsingRoc write-up, the "保险杯" competition, public-information availability, the technical-disclosure question, timeline/terminology, and the RL-ABS academic lineage.*
---
## 1. Scope
The session began as a fact-check of a Chinese-language paragraph describing **北京清鹏智能 (TsingRoc)**, a Tsinghua-background startup positioning itself as an "AI power-trading agent" challenger. It expanded into: (a) how verifiable the competition result is, (b) whether any real technical detail has been disclosed, (c) whether the 2024 timeline is plausible, and (d) where this sits in the academic literature.
---
## 2. The original claims vs. verification verdicts
| # | Claim (as written) | Verdict | Notes |
|---|---|---|---|
| 1 | Beijing TsingRoc, Tsinghua-background, in 2022 bet on power-trading AI agents | **Mostly true** | Company founded 2021 (incubated from a Tsinghua EE AI lab); "2022 focus on power trading" is accurate. "2022 founded" would be wrong. |
| 2 | In the 1st "保险杯" AI power-trading contest, ranked **15th of 124** retailer teams (beat ~90% of human traders) | **Core figures confirmed** | Caveats below. "首届/first" not explicitly documented in most sources; "124 retailer teams" is right for edition 1. |
| 3 | Deployed a **conservative** agent; aggressive back-test had lower procurement cost (**236 vs 307 元/MWh**) | **Confirmed** | Four risk-tuned agents: aggressive 236.00, weak-pref 265.38, deployed 307.04, most-conservative 316.47 元/MWh. Lower cost = better (they buy power). The 236 agent did **not** actually compete. |
| 4 | Won championships in **other AI *trading* contests** | **Not confirmed** | They won other AI/energy contests (e.g., offshore-wind output forecasting 1st place), but no confirmed win in a dedicated AI *trading* contest. Recommend softening to "other AI energy contests." |
| 5 | In Shanxi market, added ~**0.02 元** (wind) and ~**0.005 元/kWh** (solar) per-kWh revenue | **Confirmed** | Reported via 36Kr / 半熟财经 (2分 and 0.5分). This is a **real-deployment** claim, separate from the contest. |
| 6 | Running **trade-custody** cooperation with a top-tier retailer | **Confirmed** | Stated in the same 36Kr piece. |
| 7 | Strategy-subscription service covers **~30 energy firms** | **Unverified** | No public source found. |
| 8 | Tech route: **imperfect-information game theory + deep learning**; builds a "virtual market" simulator; **infers supply/demand & competitor bids from clearing prices** | **Partly supported / partly inferred** | Self-play + market simulation is supported (simulator "水镜", solver "卧龙", AlphaGo analogy). The exact phrasing "imperfect-information game theory" and "reverse-infer competitor bids from clearing prices" are **not the company's own words** — they read as a third-party technical inference. |
**Cross-cutting caveats surfaced during checking:**
- **Edition mix-up risk:** the 15th-place result was **Edition 1 (Dec 2024), based on Shandong (山东) rules***not* Shanxi. Shanxi was Edition 2's rule base. Don't let the (real-deployment) Shanxi revenue numbers bleed into the contest description.
- **Rank discrepancy:** most coverage says **15th**; a Renmin University (人大) commentary describes an all-engineer AI team at **23rd**. The "beat ~90% of humans" line is attached to both. Possibly two different AI teams, or divergent accounting.
---
## 3. The "保险杯" (Insurance Cup) AI Power-Trading Competition
**Nature & organizer.** A State-Grid-ecosystem, **industry-run *simulated* trading contest** — not a neutral academic benchmark or live-money event. Organized by **英大泰和财产保险 (Yingda P&C)**, State Grid's property-insurance arm, via the **英大售电业务共保体** (electricity-retail co-insurance consortium; members include 太保产险, 国寿财险, 中华联合, 永诚, 鼎和, 都邦). Framed as a value-added / risk-management service for retailer clients. Organizer therefore has a **promotional stake** in the "AI beat 90% of humans" narrative.
**Three editions:**
| Edition | Date / venue | Teams | Rule base | Top finishers |
|---|---|---|---|---|
| 1st | Nov→Dec 20 2024, Guangzhou (run week Dec 713) | **124** retailers | **Shandong** | 26 teams awarded; TsingRoc self-reported 15th |
| 2nd | Aug 7 2025, Guangzhou | **207** (retailers/universities/institutes) | **Shanxi** | 北京鑫泰能源, 电管家集团, 潮流能源科技 |
| 3rd | Reg. Mar → May 2026, Xi'an | **247** | **Shaanxi (陕西)** | 电管家集团 took 1st & 2nd |
**Data & rules (Edition 1).** Based on **real Shandong boundary-condition data**; teams got the most recent historical cycle's boundary data pre-contest, plus AI tools. Weather inputs: **ECMWF (EC) + open-source meteorological data**. Market structure: "中长期合约曲线三选一 + 日前全电量市场 + 实时全电量市场." Remote online; accounts issued per team. Prizes: 1st ¥50k, 2×2nd ¥30k, 3×3rd ¥10k, 20×merit ¥8k.
**Scoring.** Edition 1 ranked by summed spot + medium/long-term performance → total return; scores said to follow a normal distribution. Per-team metric visible via TsingRoc: **average procurement cost (元/MWh)**. TsingRoc's described mechanic: predict day-ahead vs. real-time price level, then flex day-ahead purchase volume **80120%**. Editions 23 add deviation (偏差) management to the assessment. Edition 3 fixed **identical load curves across teams** for fairness.
---
## 4. Is the result independently verifiable? (vs. Kaggle)
**Public information exists, but no rigorous primary source.**
**Available:** recruitment/registration notices (esp. the detailed Edition-3 rules notice via co-organizer 中能国宏 on Sina); award-ceremony coverage on industry outlets; **participant-side disclosures** — the closest to independent corroboration, e.g. a **Hohai University (河海大学)** notice that its PhD student placed *individual #2, overall #8 (top 5%)* in Edition 2; 益美国际 (HK:01870) subsidiary's Edition-1 二等奖; 电管家集团's Edition-3 sweep.
**Not available:** no official standalone website / public leaderboard; **no full ranked results** (only top-3/top-6 + self-reports); no published scoring formula, rules spec, or underlying datasets/platform (gated behind issued accounts); no third-party audit.
**Why it is *not* like Kaggle:**
| | Kaggle | 保险杯 |
|---|---|---|
| Access | Open to anyone | Invitational / eligibility-gated |
| Data | Published to all entrants | Proprietary, not released |
| Scoring | Automated vs. hidden test set | Undisclosed formula |
| Leaderboard | Public + private, full standings | Only top few announced |
| Reproducibility | Notebooks, often open-sourced | None |
| Host | Neutral platform | Stakeholder (insurer selling to participants) |
| Task type | Usually static prediction (but Kaggle also runs open agent/sim comps) | Interactive weekly trading in a simulator |
**Upshot:** you can establish the contest is real, who runs it, its structure, and who won. You **cannot** independently verify any specific placement or metric (incl. TsingRoc's 15th and the 236/307 figures) against an authoritative public record. Treat those as **reported by participant/organizer PR**, not audited results.
---
## 5. Technical disclosure — "is there any real substance, or is it all marketing?"
**Finding: no verifiable technical disclosure exists.** No paper, no findable patent, no architecture, no algorithm details, no ablations, no benchmark beyond the gated contest. Public "technology" = labels (deep learning, RL, self-play, game theory, "fused language + time-series + spatio-temporal big-models") + product code-names (**simulator 水镜 / Shuijing**, **solver 卧龙 / Wolong**) + an **AlphaGo analogy**. Labels + analogy ≠ disclosure.
**Important distinction — *undisclosed**nonexistent*.** "All marketing" can mean (a) the public technical story is marketing not evidence [well-supported], or (b) there's probably nothing real underneath [**stronger than the evidence**]. Points against (b): they fielded a **working agent that produced period-by-period results** (needs a running system); the founder gave some **falsifiable specifics** (four risk-tuned agents with cost numbers; the 80120% volume mechanic); the chief scientist is a real Tsinghua EE professor (李勇 / Li Yong — but his published record is urban computing / data science, **not** power-market trading agents).
**Defensible line for a write-up:** *"No technical substance about their trading agent has been publicly disclosed or independently verified; all capability claims trace to the company's own statements and a closed contest with unpublished results."* This is consistent both with real-but-trade-secret tech (common for AI startups) and with overselling — and the **inability to distinguish those two is itself the finding.** Do **not** overreach into a fraud claim.
---
## 6. Timeline & the "agent" terminology trap
A skepticism was raised: *"This was 2024; frontier models weren't smart enough to power agents, which didn't make real-world impact until late 2025."*
**Resolution — two different meanings of "agent":**
- **LLM agent** (reason in language, call tools, autonomous multi-step): the "real by late 2025" story is basically fair here (o1-style reasoning late 2024, computer-use previews ~Oct 2024, production usefulness across 2025).
- **RL / game-theory / control agent** (an autonomous decision *policy*, not a chatbot): **old and mature** — AlphaGo 2016, AlphaZero 2017, OpenAI Five / AlphaStar 201819, Pluribus 2019. None depend on frontier-LLM intelligence.
A power-trading "智能体" of this kind is the **second** sense. TsingRoc's own AlphaGo framing + a forecasting-plus-optimization mechanic confirms it. So a working trading agent in **Dec 2024 is not anachronistic**.
**Where the timeline skepticism *does* bite:** the **"powered by fused *large models*"** garnish (esp. "use a big model to *generate* the simulator"). If any real LLM-in-the-loop capability is implied for the Dec 2024 result, that's the doubtful, undisclosed part. Honest reading: the thing that actually traded was the **RL/forecasting core**; the "large model" language is likely dressing.
---
## 7. Academic lineage — RL-based electricity-market bidding is a decades-old, review-covered field
Hard evidence that "RL electricity-market bidding agents have been a standard academic topic for years":
- **1990s2000s roots:** agent-based England & Wales pool models where each generator is an autonomous adaptive agent developing bids via naive RL (Price 1997; Bunn & Oliveira ~2006). RothErev and Q-learning were standard bidding-agent rules.
- **2016 — game theory + RL fused:** Rashedi et al., *"Markov game approach for multi-agent competitive bidding strategies in electricity market,"* IET GTD (100+ citations).
- **2019 — DRL as state of the art:** *"Deep Reinforcement Learning for Strategic Bidding in Electricity Markets,"* IEEE Trans. — states bi-level optimization and RL are the state-of-the-art frameworks; proposes DDPG + prioritized experience replay.
- **2020 — flagship journal:** Liang, Guo, Ding & Hua, *"Agent-Based Modeling in Electricity Market Using DDPG,"* IEEE Trans. Power Systems 35(6):41804192.
- **2021 — learning equilibria:** Du, Li, Zandi & Xue, *"Approximating Nash Equilibrium in Day-ahead Electricity Market Bidding with Multi-agent Deep RL,"* J. Modern Power Systems & Clean Energy.
- **2022 — the field has systematic reviews:** *"Machine learning applications for electricity market agent-based models: A systematic literature review"* (arXiv 2206.02196). Review papers imply an established topic.
- **20232026 — still active:** incl. a **2025** paper, *"DRL-based strategic bidding … via variational-autoencoder-assisted competitor behavior learning"* — jointly models price-quantity offers **and competitors' dynamic behavior**; and a 2026 arXiv piece on validity-assessment for RL agent-based simulation of electricity markets.
**Two conclusions this settles:**
1. **"For years" is not loose** — 20+ years, top journals (IEEE TPWRS, Applied Energy, IET), review-covered; DRL versions predate the 2024 contest by 5+ years. A working RL trading agent in Dec 2024 is established practice, not a breakthrough.
2. **The "proprietary edge" is the field's standard playbook** — the *exact* technique the write-up sold as TsingRoc's differentiator (build a simulator; infer competitors' bids from clearing outcomes) is itself a **published academic method** (the 2025 competitor-behavior-learning paper is nearly one-to-one). Consistent with competent application of known methods, not a demonstrated moat.
---
## 8. What kind of research line is this — ABM / simulation-based?
Yes — properly **RL-ABS (reinforcement-learning agent-based simulation)**. The literature splits the space into **agent-based simulation** vs. **game-theoretic/equilibrium** approaches, with RL supplying the agents' learning rule. But note **two purposes** often conflated:
- **ABM-as-analysis (academic default):** many learning agents in a sim to study *emergent market behavior* (price formation, rule-change effects, tacit collusion — cf. the RUC regulator concern). Unit of interest = the *market*; output = insight/policy.
- **Simulation-as-training-gym (the product):** build a sim so *one* agent self-plays vs. simulated competitors, learns a policy, then is **deployed to trade real money** for a single participant. Unit of interest = *your* P&L. This is the AlphaGo pattern.
**TsingRoc sits in the second (deployment mode)** — confirmed by its own "energy-market chessboard + self-play" description. And the verifiable evidence (the 保险杯 result) is itself a **one-week replay on historical data** — i.e., the agent evaluated *inside a simulation*.
**Key nuances:**
- **"It's simulation-based" is not a knock.** Every superhuman decision agent (AlphaGo/AlphaStar/Pluribus) is trained by self-play in a simulator. The real questions are **simulator fidelity** (does 水镜 reproduce real clearing, congestion, competitor behavior, weather-driven supply?) and **sim-to-real transfer** (does the policy work in the live market?). That's a data/engineering problem — and it's the undisclosed, unaudited part where any real moat would live.
- **The one arguably-newer pitch element** — "use a **large model to *generate* the simulator**" (vs. classic hand-built RL-ABS environments) — is exactly the piece with **zero disclosed detail**, so its substance can't be assessed.
---
## 9. Bottom line
- The **hard, outcome-type data** (rank, 236/307, 0.02/0.005 元, custody cooperation) map to public reporting — but **rest on participant/organizer PR**, not an auditable record.
- **Wording to fix** in any downstream write-up: drop/verify "首届"; "124 *retailer* teams" is edition-specific; "won other AI *trading* contests" → "other AI energy contests"; **"~30-firm subscription" is unverified**; the "imperfect-information game theory / reverse-infer competitor bids" tech route should be flagged as **inference, not company disclosure**, and downgraded to *"per its own statements, a market-simulation + self-play RL approach (method not publicly disclosed)."*
- **Method placement:** standard **RL-ABS / self-play-in-a-market-simulator**, used in **deployment mode**. Well-trodden ground; the only potentially-differentiated parts — **simulator fidelity, forecasting quality, and the "LLM-built simulator" claim** — are undisclosed and unverified.
- **Overall:** a real, State-Grid-adjacent industry contest result and a plausibly-real system — but the technical claims are **asserted, not demonstrated, and not independently verifiable**. Weight it as a **controlled, self-reported demonstration**, not an open, audited benchmark.
---
*Note: dates/figures reflect public reporting available as of the session (July 2026). Competition editions and standings may update; participant self-disclosures are not independently audited.*