AI forecasting for energy

  • Event forecasting: on short-horizon questions, the best AI systems now score close to the best human forecasters, and the remaining differences are within statistical noise. ForecastBench (Jul 2026) cannot distinguish the top systems from superforecasters, and the top bot team in Metaculus's bot benchmark trails the Pros by a margin that is no longer significant. In September a bot won the Summer 2026 Metaculus Cup outright for the first time.
  • That evidence does not carry over to load, price and renewables forecasting. It comes from one-off questions ("will this regulation pass by June") where the information sits in documents and news. Time-series foundation models (TSFMs) have been tested on numeric series with long histories. Many utility problems sit in between, such as prices around policy changes, gas, or hydro in a drought year; there the two approaches combine.
  • TSFMs on energy data: zero-shot with weather covariates, Chronos-2 and TiRex-2 beat tuned per-task XGBoost on a 54-dataset energy benchmark. On electricity prices they do not consistently beat specialist models, and I found no study on utility hydro inflow, reservoir or gas-price forecasting.
  • Licences decide what you can use in production. Chronos-2, TiRex-2 and TimesFM 2.5 are Apache-2.0. TimesFM 3.0 and TabPFN-3.5 weights exclude production and commercial use.
  • First experiments: a TSFM run against your production model on one target, plus a small pilot of LLM forecasts on 20 to 30 of your own event questions, scored against your analysts.

Jump to: problem map · experiments · tools and data

Which approach for which problem

Forecasting problems range from data-driven to text-driven. For short-term load or solar, the series history plus weather forecasts carry most of the information. For a regulatory decision, almost none of it is in a series. Longer horizons move a problem toward the text end. The table lists typical problems at a utility with a large hydro share, with the AI approach worth testing, the baseline it has to beat, and how much published evidence exists for that specific problem.

ProblemTypeAI approach to testBaseline to beatEvidence
Day-ahead / intraday pricesSeriesTSFM with fundamentals and weather as covariates; ensemble with the specialist modelProduction EPF model; LEAR / DNN from epftoolbox; MSTLMixed three day-ahead studies; none on intraday found
LoadSeriesTSFM with weather covariatesProduction modelModerate benchmarks only (FETS, ERCOT); gains larger at high aggregation
Wind / solar outputSeriesTSFM with NWP forecasts as known-future covariatesProduction modelModerate benchmarks only; weather inputs clearly help
Hydro inflow, reservoir levelsSeriesTSFM with precipitation, temperature and snow covariates (inflow); reservoir levels also depend on dispatch decisionsProduction modelLimited a few streamflow and water-level studies, e.g. arXiv:2511.11849; none on utility inflow found
Gas, CO2 pricesSeries, event-drivenTSFM for the series; LLM forecaster for the events that move itForward curveNone found
Regulation, market rules, subsidy schemesEventLLM forecaster alongside analystsAnalyst forecastsIndirect general tournaments only
Project timing (grid, interconnectors, plants)DateLLM forecaster with date-type outputAnalyst or project-plan datesNone found
Supply shocks, geopoliticsEventLLM forecasterAnalysts; prediction markets where they existIndirect general tournaments only

Rules of thumb:

  • For a numeric series with long history, start with a TSFM-versus-production comparison.
  • A binary or date question with crisp resolution criteria is where LLM forecasters have a track record.
  • If a liquid market prices the quantity, the market price is the baseline. A fine-tuned forecasting model, Foresight-32B, trailed Polymarket in its own evaluation (Brier 0.199 vs. 0.170 on 251 questions).
  • Text that moves a number (outage notices, regulatory decisions, transit news) is best handled by an LLM that turns it into flags or scenarios, with a conventional model doing the arithmetic. ServiceNow's Context is Key benchmark (ICML 2025) shows LLMs can use such context in numeric forecasts, with notable failures.

Two experiments

1. TSFM comparison on one target

  • Target: pick one series and horizon you already forecast, e.g. day-ahead price for your bidding zone or regional load, next 24 to 48 hours. Day-ahead prices have been in 15-minute steps since October 2025; decide whether to forecast quarter-hours or hourly averages.
  • Arms: your production model; Chronos-2 with the same weather and fundamental inputs as known-future covariates; MSTL as a floor; an equal-weight average of Chronos-2 and the production model. For prices add TabPFN-TS in local mode as an evaluation-only arm; TiRex-2 if time allows.
  • Inputs: fix forecast issue times and use the weather forecasts that were available at those times, not reanalysis.
  • Test period: after the models' release dates (Chronos-2: Oct 2025; TiRex-2: Jul 2026), to avoid testing on series the models may have seen in pretraining.
  • Metrics: the ones you already use (CRPS or pinball, MAE overall and in spike hours), plus a decision metric if one exists (dispatch or trading value).
  • Effort: getting Chronos-2 forecasts from prepared data takes hours; a comparison you would trust takes the time your normal backtesting takes.

2. Pilot: LLM forecasts on your own event questions

  • Write 20 to 30 questions with resolution criteria and dates 1 to 6 months out.
  • Analysts forecast blind; an LLM system forecasts the same questions on the same day. FutureSearch's forecast operation is the quickest option: a table of questions in, probability and rationale out, about $0.25 to $5 per question.
  • Log probabilities, sources and model versions. Score analysts, AI and their average with the Brier score as questions resolve. Keep the questions binary (thresholds, deadlines) so one score covers all of them.
  • 30 questions will not prove superiority either way. The pilot tests the workflow and whether the AI rationales are useful, and starts a track record.
Starter questions (templates; fill in thresholds and dates)
  1. Will [national regulator] publish its final decision on [named tariff or market-rule consultation] by [date]?
  2. Will the European Commission publish a legislative proposal amending [named act] by [date]?
  3. Will [named interconnector or grid project] start commercial operation by [date]?
  4. Will [named support scheme] be open for applications by [date]?
  5. Will [named plant] be back in operation by [date], according to its REMIT urgent market messages?
  6. Will [country] adopt [named market-design measure] by [date]?
  7. Will the ICE Endex TTF front-month contract settle above €[X]/MWh on [date]? (Compare with the forward price and your own model.)
  8. Will EU aggregate gas storage be at least [X]% full on [date], per AGSI+?
  9. Will the [bidding zone] day-ahead market have more than [N] hours with negative prices in [month]? (State how 15-minute prices are aggregated to hours.)
  10. Will [country]'s weekly reservoir storage reported to ENTSO-E (MWh) be above its 2015 to 2024 same-week median in week [N] of [year]?

Time-series foundation models on energy data

TSFMs are pretrained on large collections of time series and produce forecasts without task-specific training ("zero-shot"), usually as quantiles. Covariate support matters for weather-sensitive targets: a model that cannot take weather forecasts is handicapped on load, wind and solar.

FETS benchmark: median NRMSE across 54 energy datasets (lower is better)
Chronos-2 0.472, TiRex-2 0.474, both zero-shot with covariates; XGBoost 0.611 and random forest 0.696, both trained per task. 0 0.25 0.5 0.75 Chronos-2 (zero-shot, covariates): 0.472 Chronos-2 0.472 TiRex-2 (zero-shot, covariates): 0.474 TiRex-2 0.474 XGBoost (trained per task): 0.611 XGBoost 0.611 Random forest (trained per task): 0.696 Random forest 0.696
Blue: zero-shot foundation models with covariates. Grey: per-task models (SHAP-based feature and lag selection, 250 Optuna trials per quantile). Obermeier et al., Energy and AI 2026. Errors were lower on aggregated targets such as national load and district heating. The tree baselines are tuned, but a production model built by people who know the series is a harder baseline.

Electricity prices: mixed

  • Hornek et al. 2025: 2024 day-ahead prices in DE, FR, NL, AT, BE. Chronos-Bolt and Time-MoE matched traditional methods; no TSFM statistically beat a biseasonal MSTL decomposition. No covariates were used, which handicaps the TSFMs.
  • Lipiecki & Weron, Aug 2026: DE, PL, ES, 2021 to 2025, nine zero-shot variants from five model families against two state-of-the-art price-forecasting benchmarks. Only the TabPFN models (TabPFN-2, TabPFN-3, TabPFN-TS-3) beat both consistently and significantly. In a battery-arbitrage backtest, the distributional neural-network benchmark earned more at lower risk tolerance: better error scores did not always mean better decisions. Their conclusion: foundation models "cannot universally replace market-specific models".
  • Pan & Ezzat, ICML 2026 workshop (revised 7 Oct 2026): TSFMs beat general-purpose baselines but not consistently the specialist methods, and depend heavily on covariate support. Simple ensembles of TSFMs with specialist models look promising. They also flag contamination: public price series may be in the pretraining data, which flatters backtests.

On day-ahead prices TSFMs are worth a comparison on your own data and look most promising as one member of an ensemble. None of the studies supports replacing a specialist model.

Load, wind, solar

Besides FETS: Za'ter & Hodge 2026 (ERCOT, 2018 data, older models such as Chronos-Bolt, Moirai and TimesFM) found zero-shot TSFMs often lost to models trained on the domain data; fine-tuning closed much of the gap. Weather inputs helped solar and wind clearly and load slightly (0.1 to 0.3 points nMAE). I found no published production use at a utility or TSO, though my search was not exhaustive.

Models worth testing

ModelWeights licenceCovariatesNotes
Chronos-2 (Amazon)Apache-2.0Past and known-future, numeric or categorical120M parameters, runs on modest hardware. Easiest start, also via AutoGluon 1.5.
TiRex-2 (NXAI)Apache-2.0Past and known-future82.5M parameters. Level with Chronos-2 on FETS. The original TiRex has a custom licence.
TimesFM 2.5 (Google)Apache-2.0Via the XReg add-on200M parameters, 16k context.
TimesFM 3.0 (Google)Non-commercialNative, past and futureReleased Aug 2026. Production use only through Google Cloud (BigQuery ML).
TabPFN-TS (Prior Labs)Evaluation onlyKnown-future only; past-only covariates droppedEarlier TabPFN versions were best in the 2026 price study. Current default weights (TabPFN-3.5) exclude "internal commercial decision-making" without a paid licence. Default client sends data to Prior Labs' cloud; local mode needs a GPU.

Left out: Moirai 2.0 (CC-BY-NC-4.0, research only). General leaderboards: GIFT-Eval, fev-bench; both are vendor-heavy and not energy-specific.

AI vs. top human forecasters on event questions

All of this comes from live tournaments, where questions resolve after the forecast is made. Backtests on past questions are unreliable because search results leak post-cutoff information.

Metaculus bot benchmark: top bot team minus Pro team, head-to-head score per question
Bot team minus Pro team score: Q3 2024 minus 11.3, Q4 2024 minus 8.9, Q1 2025 minus 17.7, Q2 2025 minus 20.0, Spring 2026 minus 1.25 with 95% interval from minus 4.87 to plus 2.37. 0 -5 -10 -15 -20 Q3 2024: -11.3 -11.3 Q3 2024 Q4 2024: -8.9 -8.9 Q4 2024 Q1 2025: -17.7 -17.7 Q1 2025 Q2 2025: -20.0 (p = 0.00001) -20.0 Q2 2025 Spring 2026: -1.25, 95% CI -4.87 to +2.37, p = 0.25 -1.25 Spring 2026
Below zero means the Pros scored better. The whisker on Spring 2026 is the 95% confidence interval, which includes zero. Head-to-head points are a relative log-score measure, not percentage points. Source: Metaculus FutureEval Spring 2026 results (Sep 2026).
  • Metaculus bot benchmark (300 to 500 questions per season, binary, numeric and multiple choice): ten of Metaculus's best human forecasters ("Pros") against the ten best bots, compared on the questions both forecast (99 in Spring 2026). The gap shrank from highly significant in 2025 to not significant in Spring 2026. Ranked individually, 9 of the top 10 were still Pros, and 7 of the top 10 bot makers were hobbyists.
  • ForecastBench (Forecasting Research Institute, binary questions only). Jan 2026: best LLM at difficulty-adjusted Brier 0.102, superforecasters about 0.017 better (FRI). Jul 2026: Cassi AI, xAI and Google DeepMind entries are statistically indistinguishable from superforecasters (one-sided p from 0.14 to 0.41). FRI calls this "likely parity". Caveats from FRI itself: superforecaster forecasts were last collected in 2024 and are extrapolated, and the intervals overlap heavily. On 8 Oct 2026 the live leaderboard has three Google DeepMind entries at or just above the superforecaster median, none significantly better. A new superforecaster round and quantile questions are planned.
  • Metaculus Cup (the main human tournament; bots compete but cannot win prizes). Summer 2025: Mantic 8th of 549, the first bot in the top 10 (TIME). Summer 2026 (885 forecasters, 58 questions): bots took 1st (laertes, by an individual builder), 2nd (Mantic) and 5th (FutureSearch). 58 questions is a small sample. The Summer 2026 bot-vs-Pro analysis is due after 5 November.

What this does not show

  • Long horizons. Almost all questions resolve within weeks to months. There is no good evidence for multi-year questions, where much energy strategy sits.
  • Equivalence or superiority. A non-significant difference leaves both open. The strongest individual humans still top the Metaculus rankings.
  • Domain questions. Tournament questions are general (politics, economics, science). None of the results is specific to energy.

Earlier "superhuman" claims did not survive replication. The Center for AI Safety's FiveThirtyNine bot (Sep 2024) was claimed superhuman on a 177-question backtest; Halawi re-ran it on 324 questions opened after November 2023 and got Brier 0.195 vs. the crowd's 0.141 (discussion). Polymarket "bot win-rate" stories are anecdotes; one study estimates about $40M in arbitrage profits on markets resolving Apr 2024 to Apr 2025 (Saguillo et al.), which is not forecasting skill.

What the strong bots do

From Metaculus's tournament analyses (Spring 2026, synthesis of 11 analyses) and FRI's description of top ForecastBench entries:

  • Current frontier reasoning model. In 8 of 8 paired comparisons the higher-reasoning variant of a bot beat its standard twin (pairs not fully independent).
  • Agentic search. Several rounds of search and reading, not a single retrieval step. Which provider (AskNews, Exa, Perplexity) made no clear difference.
  • Outside view first. Base rates and similar resolved questions, then case specifics.
  • Ensembling. Aggregate 3 to 7 runs, ideally from different models. Aggregating bots helped up to about 10 bots, then got worse as weaker ones were added.
  • Calibration. Platt scaling improved Brier by 0.016 on binary questions in Metaculus's Q1 to Q2 2025 data. Capping extreme probabilities was the strongest correlate of winning (r = 0.48, observational).
  • Use the market where there is one. Cassi's adjustment toward market prices was worth about 0.01 Brier on ForecastBench market questions.

Metaculus estimates a good harness is worth about 9 months of base-model progress (top scaffolded bots beat their plain baselines by 5 to 11 peer-score points, against roughly 0.9 points per month of model improvement; one tournament, rough figure).

Fine-tuning on resolved questions works for small open models: Turtel et al. trained a 14B model with outcome rewards to o1-level Brier with better calibration. The Metaculus synthesis puts fine-tuned open models at parity with the previous generation of closed models, not the current frontier. Start with a hosted frontier model and a forecasting harness.

Combining humans and AI: Halawi et al. 2024 (paper) found their system alone at Brier 0.179 vs. the crowd's 0.149, but averaging the two gave 0.146. Schoenegger et al. (Science Advances) found similar gains from mixing human and LLM forecasts on a small question set. Test human-AI averaging on your own questions.

Tools and data

Event forecasting

  • FutureSearch: API, Python SDK and MCP server. Binary, numeric (p10 to p90), date, categorical and conditional forecasts. Founded by former Metaculus staff. Default effort is HIGH; set LOW for large batches.
  • Metaculus forecasting-tools and bot template (MIT, labelled experimental): open-source harness pieces (search, multi-model calls, cost limits) for building your own bot. Built around Metaculus questions; for internal questions you reuse the parts.
  • Mantic: 2nd in the Summer 2026 Cup, no public API. Customers undisclosed; its lead investor says hedge funds and trading firms are most interested (Reuters, Sep 2026).
  • Radiant (Metaculus, Feb 2026): a visual map where forecasts are nodes and links show how they depend on each other, with AI to suggest questions and links. Prototype with a limited invited cohort (demo via hello@metaculus.com); no public API or MCP server mentioned. The launch post does not say it computes joint or conditional probabilities across the map.
  • Squiggle and Squiggle Hub (open source): a language for explicit probabilistic models with distributions, the option for branching scenario models today.
  • Prediction-market data: Polymarket, Kalshi and Manifold have public APIs; community MCP servers for them exist but are small hobby projects. Calling the APIs directly is simpler.

Time series and data

  • entsoe-py: Python client for the ENTSO-E Transparency Platform (prices, load, generation, forecasts, hydro reservoir filling). Needs a free API token; the README explains how to request one.
  • National TSO portals often go further than ENTSO-E, e.g. APG (Austria, API from 2023) or Energy-Charts (Fraunhofer ISE, Germany and Europe).
  • Open-Meteo: weather forecasts, archived forecasts and historical model runs. Free tier is non-commercial only; commercial use of historical forecasts needs the Professional plan.
  • epftoolbox: standard LEAR and DNN price-forecasting baselines and benchmark datasets.
  • chronos-forecasting and AutoGluon-TimeSeries: the quickest way to run Chronos-2 alongside classical and tree models in one backtest.
  • OpenSTEF (LF Energy, started by Alliander): open-source short-term forecasting pipeline used by grid operators; a reference for production setups.

Glossary

Tournament and LLM-forecasting terms
Superforecaster
A forecaster selected for a sustained record of accurate probability forecasts (Good Judgment Project lineage).
Metaculus Pros
Metaculus's paid panel of top forecasters, used as the human benchmark in its bot tournaments.
Brier score
Mean squared error between forecast probability and the 0/1 outcome. Lower is better; always saying 50% scores 0.25.
Difficulty-adjusted Brier
ForecastBench's adjustment for forecasters answering different question sets.
Peer score, head-to-head score
Metaculus log-score measures relative to other forecasters on the same questions. Higher is better; not percentage points.
Outside view / inside view
Start from base rates in comparable cases, then adjust for the specifics.
Platt scaling
Recalibrating probabilities with a logistic fit to past outcomes.
Extremizing
Pushing an aggregate probability away from 50%. Contested for bots.
Harness, scaffolding
The code around the model: search, prompting, multiple runs, aggregation, calibration.
Resolution criteria
The exact rule that decides how a question resolves. Vague criteria make scoring meaningless.
Zero-shot
Using a pretrained model on a new series without training on it.
Known-future vs. past-only covariates
Inputs known for the forecast horizon (weather forecasts, calendar) vs. inputs observed only up to now (realized generation).
MSTL
Seasonal-trend decomposition with multiple seasonal periods (daily, weekly).
EPF
Electricity price forecasting.
MCP
Model Context Protocol, a standard way for an AI assistant to call external tools and data sources.

Sources

  1. FRI, "AI models have likely reached parity with superforecasters", 16 Jul 2026. substack
  2. FRI, "LLMs are closing the gap on human forecasters", Jan 2026. substack
  3. ForecastBench live leaderboard, accessed 8 Oct 2026. forecastbench.org
  4. Metaculus, "FutureEval Spring results: pros beat bots, but the gap is nearly gone", Sep 2026. LessWrong
  5. "AI forecasting in 2026: what 11 analyses say", 2026. LessWrong
  6. Metaculus, "Announcing the Metaculus Cup Summer 2026 winners", 18 Sep 2026. metaculus.com
  7. TIME, Mantic in the Summer 2025 Metaculus Cup. time.com
  8. Reuters, "AI startup Mantic raises $25 million", 18 Sep 2026 (syndicated). thestar.com.my
  9. Halawi et al., "Approaching human-level forecasting with language models", 2024. arXiv:2402.18563
  10. Schoenegger et al., "Wisdom of the silicon crowd", Science Advances. doi:10.1126/sciadv.adp1528
  11. Turtel et al., "Outcome-based reinforcement learning to predict the future", TMLR 2025. arXiv:2505.17989
  12. Lightning Rod Labs, Foresight-32B model card (self-reported). Hugging Face
  13. "Is superhuman AI forecasting BS? Some experiments on the 539", 2024. LessWrong
  14. Discussion of FiveThirtyNine replication (Halawi's re-run). LessWrong
  15. Saguillo et al., arbitrage on Polymarket, AFT 2025. arXiv:2508.03474
  16. Obermeier et al., "FETS Benchmark", Energy and AI 2026. arXiv:2604.22328
  17. Hornek et al., "Benchmarking pre-trained time series models for electricity price forecasting", 2025. arXiv:2506.08113
  18. Lipiecki & Weron, "Foundation models for electricity price forecasting and battery arbitrage", Aug 2026. arXiv:2609.00089
  19. Pan & Ezzat, "Evaluating time series foundation models for electricity price forecasting", ICML 2026 FMSD workshop. arXiv:2607.02623
  20. Za'ter & Hodge, ERCOT power-system benchmark, 2026. arXiv:2604.22077
  21. Williams et al., "Context is Key", ICML 2025. arXiv:2410.18959
  22. Model cards and repos: Chronos-2, TiRex-2, TimesFM, TabPFN-TS, TabPFN-3.5 licence, Moirai 2.0.
  23. Metaculus, "Map the future before you build it" (Radiant), 26 Feb 2026. metaculus.com