# Day-Ahead Electricity Price Forecasting Logic using LightGBM

This document outlines the price data description, engineering logic, feature pipeline, and model architecture for forecasting day-ahead electricity prices ($1$ to $24$ hours ahead) for the DK1 pricing zone, referencing the data paradigms of PEPF_part1 and PEPF_part2, but replacing slow algorithms (Quantile Regression and Quantile Random Forests) with a high-performance **LightGBM** implementation combined with **Binned Residual Bootstrapping**.

---

## 1. Electricity Price Data & Target

The data is sourced from the Danish **Energi Data Service** and covers the period from **2024-01-01 to 2025-09-30 UTC** (639 days, or exactly **15,336 hourly observations**).

### A. Target Variable
* **`DK1_EUR/MWh`** (hourly electricity spot price in the DK1 pricing zone, expressed in EUR per MWh).

### B. Input Features (Data.csv Schema)
The system incorporates 15 historical and forecast exogenous columns from the Danish power grid:

| Feature Name | Category | Description |
| :--- | :--- | :--- |
| **`DK1_EUR/MWh`** | Target | Hourly elspot price in the DK1 zone (€/MWh). |
| **`LocalPowerMWhDK1`** | Production | Electricity generated by local power plants (MWh). |
| **`LocalPowerSelfConMWhDK1`** | Consumption | Locally produced electricity consumed without entering the grid (MWh). |
| **`CentralPowerMWhDK1`** | Production | Electricity generated by large centralized power plants (MWh). |
| **`CommercialPowerMWhDK1`** | Production | Electricity produced for commercial purposes by private operators (MWh). |
| **`HydroPowerMWhDK1`** | Production | Electricity generated from hydroelectric plants in DK1 (MWh). |
| **`OffshoreWindGe100MW_MWhDK1`** | Production | Electricity generated by offshore wind farms $\ge 100\text{ MW}$ (MWh). |
| **`OffshoreWindLt100MW_MWhDK1`** | Production | Electricity generated by offshore wind farms $< 100\text{ MW}$ (MWh). |
| **`OnshoreWindGe50kW_MWhDK1`** | Production | Electricity generated by onshore wind farms $\ge 50\text{ kW}$ (MWh). |
| **`OnshoreWindLt50kW_MWhDK1`** | Production | Electricity generated by small onshore wind farms $< 50\text{ kW}$ (MWh). |
| **`SolarPowerGe10Lt40kW_MWhDK1`**| Production | Solar electricity generated from systems between $10–40\text{ kW}$ (MWh). |
| **`SolarPowerGe40kW_MWhDK1`** | Production | Solar electricity generated from systems $> 40\text{ kW}$ (MWh). |
| **`SolarPowerLt10kW_MWhDK1`** | Production | Solar electricity from small residential systems $< 10\text{ kW}$ (MWh). |
| **`SolarPowerSelfConMWhDK1`** | Consumption | Solar power generated and consumed on-site without exporting (MWh). |
| **`PowerToHeatMWhDK1`** | Conversion | Electricity converted to heat for district heating systems (MWh). |
| **`GrossConsumptionMWhDK1`** | Consumption | Total electricity consumed, including losses and self-consumption (MWh). |

---

## 2. Why LightGBM?

While Quantile Regression Forest (QRF) is statistically robust, it is computationally expensive to train and predict over large multi-year datasets, especially when scaling across multiple lead times. **LightGBM** (Light Gradient Boosting Machine) is chosen for the operational pipeline due to:
* **High Efficiency**: 10x to 100x faster training and inference compared to random forests.
* **Frugal Memory Footprint**: Uses histogram-based algorithms to bucket continuous values, making it highly memory-efficient.

---

## 3. Forecasting Architecture: Direct Multi-Step Forecasting

To predict a full $24$-hour ahead horizon without accumulation errors, we use the **Direct Multi-step Forecasting** method.
* Instead of training a single recursive model (which propagates errors step-by-step), we train **24 independent LightGBM models**, one for each lead time $k \in \{1, 2, \dots, 24\}$:
  $$\hat{P}(t + k) = f_k(X_t)$$
* At forecast origin $t$, model $k$ predicts the price for $t+k$ directly using the feature vector $X_t$ containing variables known at or before time $t$.

---

## 4. Feature Engineering Pipeline

For each lead time model $k$, the input feature vector $X_t$ is dynamically constructed:

### A. Calendar & Time Features
* **Hour of Day** ($0–23$): Captures diurnal pricing peaks.
* **Day of Week** ($0–6$): Captures weekday industrial loads vs. weekend demand drops.
* **Month** ($1–12$): Captures seasonal temperature and lighting differences.
* **Is Weekend** ($0/1$): Boolean modifier for weekend profiles.
* **Is Peak Hours** ($0/1$): Boolean for peak business hours (typically 07:00 to 20:00).

### B. Historical Price Lags (Autoregressive Features)
* **Lead-Time Lag**: $P(t - k)$ — the most recent actual price available.
* **Day-Ahead Lag**: $P(t - k - 24)$ — the price at the same hour yesterday.
* **Weekly Lag**: $P(t - 168)$ — the price at the same hour, same day last week (vital for capturing weekly demand cycles).
* **Rolling Trend**: $24$-hour rolling average of the price lags, representing short-term price level shifts.

### C. Exogenous Forecast Features (System Variables)
These variables are sourced from system-operator forecasts or simulated forecast profiles:
* **Total System Load Forecast** (kW)
* **Local PV Generation Forecast** (kW)
* **Wind Generation Forecast** (kW)
* **Grid Import/Export Congestion Signals**

---

## 5. Probabilistic Forecasting via Binned Residual Bootstrapping

To support uncertainty-aware decision-making in the BESS controller (e.g., Chance-Constrained Optimization), the agent requires quantile forecasts: **P10** (pessimistic low), **P50** (median/base), and **P90** (optimistic high).

Rather than training three independent LightGBM models per lead time, we use a hybrid approach:
1. Train **one single LightGBM model** per lead time $k$ to predict the mean/median point forecast:
   $$\hat{P}_{\text{point}}(t + k) = f_k(X_t)$$
2. Extract out-of-fold validation residuals:
   $$\text{Residuals}_{\text{val}} = Y_{\text{val}} - \hat{P}_{\text{val}}$$
3. Construct prediction intervals using a **Binned Residual Bootstrap** on these validation residuals.

### How the Binned Residual Bootstrap Works:
* **Value-Based Binning**: Validation rows are grouped into $N$ bins (e.g., 15 bins) based on the model's own *predicted value* ($\hat{P}_{\text{val}}$).
* **Heteroscedasticity Capture**: This ensures that a test prediction $\hat{P}_{\text{test}}$ draws its bootstrap sample only from the validation bin that had similar predicted values. During calm price regimes, the residuals will be narrow; during high/volatile price regimes, the residuals will be wide.
* **Quantile Construction**:
  * For a test point prediction $\hat{P}_{\text{test}}$, locate its matching prediction bin.
  * Extract the residual pool for that bin and draw $N_{\text{bootstrap}}$ (e.g., 5000) samples with replacement.
  * Calculate the required quantiles of the drawn residuals: $q_{10}$ and $q_{90}$.
  * Compute final prediction intervals:
    $$\hat{P}_{\text{P10}}(t+k) = \hat{P}_{\text{test}}(t+k) + q_{10}$$
    $$\hat{P}_{\text{P50}}(t+k) = \hat{P}_{\text{test}}(t+k) \quad (\text{Point Forecast})$$
    $$\hat{P}_{\text{P90}}(t+k) = \hat{P}_{\text{test}}(t+k) + q_{90}$$

---

## 6. Post-Processing: Enforcing Quantile Monotonicity

Because quantiles are constructed statistically from empirical residual distributions, we apply a sorting/monotonicity operator at test runtime to guarantee physical consistency:
$$\hat{P}_{\text{P10}}^{\text{corrected}}(t) = \hat{P}_{\text{P10}}(t)$$
$$\hat{P}_{\text{P50}}^{\text{corrected}}(t) = \max\left(\hat{P}_{\text{P50}}(t),\ \hat{P}_{\text{P10}}^{\text{corrected}}(t)\right)$$
$$\hat{P}_{\text{P90}}^{\text{corrected}}(t) = \max\left(\hat{P}_{\text{P90}}(t),\ \hat{P}_{\text{P50}}^{\text{corrected}}(t)\right)$$

This ensures that the predicted intervals are logically consistent before they are sent to the BESS logic engine.
