Sheng Yin, Vivek Teja Tanjavooru
This paper introduces Representation-to-Decision (R2D), an end-to-end imitation learning framework that maps heterogeneous multi-horizon time-series inputs directly to battery control decisions through modular Temporal Feature Extractors (TFEs) and a shared latent representation, without an explicit load and PV forecasting step. While accurate forecasting improves prediction quality, optimal control performance remains unguaranteed in prediction-then-optimization approaches. The proposed method effectively encodes load, photovoltaic (PV) generation, and time-of-use pricing signals into an implicit temporal representation, enabling the framework to clone the dispatch behavior of an aging-aware Mixed Integer Linear Programming (MILP) expert. The results indicate that R2D achieves 62–77% of the global optimum, outperforming tested baselines in Model Predictive Control (MPC) and Reinforcement Learning (RL). The ablation studies conducted cover essential aspects such as architecture, imitation learning expert, time horizon considerations, and cross-site generalization, contributing to the robustness of the proposed approach. Overall, the findings underline the effectiveness of R2D in enhancing battery energy management systems by incorporating temporal features and control policies.
@article{418c3739-9801-4237-a963-6068be7bfb00,
title={Learning Control Policies from Heterogeneous Multi-Horizon Time Series in Battery Energy Management Systems},
author={Sheng Yin and Vivek Teja Tanjavooru},
year={2023},
language={English}
}TY - JOUR TI - Learning Control Policies from Heterogeneous Multi-Horizon Time Series in Battery Energy Management Systems AU - Sheng Yin AU - Vivek Teja Tanjavooru PY - 2023 LA - English ER -
Yaya Dagal D
This document presents an educational module on electrochemical energy storage systems—primary cells, accumulators, and batteries—with specific emphas