Archive/Week-Ahead Electricity Price Forecasting for Battery Arbitrage: Benchmarking ML/DL Models and Interpreting Feature Importance Through Merit-Order Pricing in Spain
Week-Ahead Electricity Price Forecasting for Battery Arbitrage: Benchmarking ML/DL Models and Interpreting Feature Importance Through Merit-Order Pricing in Spain
Amgad Khamis, Francesco Crespi, David Sánchez
21 de julio de 2026
en

Abstract

Accurate electricity price forecasting is essential for market participants seeking to optimise bidding and arbitrage strategies. This paper presents a week-ahead (168 h) hourly electricity price forecasting study for the Spanish day-ahead market. Nine competing models—two naïve baselines (a Seasonal Naïve and a Day-of-Week persistence), a Lasso-estimated auto-regressive (LEAR) statistical benchmark, and six machine- and deep-learning models (CatBoost, Random Forest, LSTM, GRU, CNN, and a hybrid CNN–LSTM)—are benchmarked; the two leading models, CNN–LSTM and CatBoost, are then compared under exogenous-feature configurations. The analysis is complemented by an ex-post Add-One-In and Leave-One-Out feature-importance analysis, a controlled comparison of weather-input scenarios, and a rolling battery-arbitrage backtest that translates forecast quality into economic value. Under an endogenous benchmark of weekly rolling origins across 2024 (with a rotating start weekday) and Diebold–Mariano testing, a recursive CatBoost and the hybrid CNN–LSTM are statistically indistinguishable and both significantly outperform a direct multi-horizon CatBoost; once an operational (forecasted) weather input is added, recursive CatBoost becomes significantly the most accurate while remaining simpler and more stable to train, a ranking confirmed on a fully out-of-sample 2025 year. Operational weather forecasts are found to be the best weather input, recovering about 84% of the perfect-foresight weather improvement over a no-weather baseline, with the advantage concentrated at longer lead times. Natural-gas-fired generation emerged as the dominant explanatory feature, consistent with the marginal-pricing mechanism governing the Spanish market. In a rolling battery-arbitrage backtest on the out-of-sample 2025 year, a deployable forecast-driven 4-h grid-scale unit (200 MW/800 MWh) captured about 89% of perfect-foresight value at a 168 h optimisation horizon and about 87% at 24 h; extending the horizon from 24 h to 168 h added about 2.4% of profit, an optimisation-horizon (look-ahead) effect bounded at +4.5% under perfect foresight.

IPC Classification

H01

Keywords

week-aheadelectricitypriceforecastingbatteryarbitragebenchmarkingmodelsinterpretingfeatureimportancethroughmerit-orderpricingspainaccurateessentialmarketparticipantsseekingoptimisebiddingstrategiespaper
Citar esta publicación

€ 4.00