Skip to content

Wind farm generation forecast and losses, tested on open data

Without using any customer data, we measured our method on the public SCADA data of La Haute Borne, a four-turbine wind farm in France. Below we report what worked and where it fell short, as it is.

ENGIE La Haute Borne4 × Senvion MM82 (2.05 MW)2014–201510-min SCADA
9.3%
Lower error than the persistence baseline at the 24-hour horizon
83–87%
Measured prediction-interval coverage (target 90%)
10.6 MWh
Below-expected output flagged in the test period (0.23% of production)
17,287
Complete hourly plant records used

01

The question

An operator wants to know two things: how much will the farm produce tomorrow, and where is production being lost today? This case tests the method side of both: time-ordered forecasting, an uncertainty interval, and loss detection that needs no fault labels.

02

The data

10-minute SCADA records from ENGIE's La Haute Borne wind farm (France) for 2014–2015. Records were averaged to hourly values and summed to plant level; 17,287 complete hours remained (all four turbines reporting). The data was split in time order: 70% training, 15% calibration, 15% test (September–December 2015). The model never saw the test period during training.

Plant
La Haute Borne, France
Turbines
4 × Senvion MM82, 2.05 MW
Period
2014–2015, 10-min SCADA
Split
70 / 15 / 15, time-ordered
Test period
Sep–Dec 2015
Licence
Etalab Open Licence 2.0

03

The method

  1. A

    Time-ordered forecast

    Plant output 1, 6 and 24 hours ahead is forecast from past SCADA values only (power, wind speed, temperature, lags, time of day), averaging three tree-based models (HistGB, Random Forest, Extra Trees). The baseline is persistence: “output stays as it is now”. In wind, it is the simple baseline that is hardest to beat at short horizons.

  2. B

    Conformal interval

    From the errors in the calibration period, an interval targeting 90% coverage is drawn around every forecast (split conformal). What matters is whether it actually holds 90% in the test period.

  3. C

    Unlabelled anomaly and loss

    A power-curve model learns expected output from wind speed and temperature; Isolation Forest flags turbine-hours that deviate. Hours more than 10% of rated power below expectation with wind above 4 m/s are summed as lost MWh. No fault labels are used.

04

Results

Error is reported as nMAE: mean absolute error as a share of installed capacity. Lower is better.

Forecast error by horizon

ModelPersistence

nMAE, % of capacity

On par with persistence at 1 hour; the gap widens as the horizon grows.

Prediction-interval coverage

Target 90%. The test period fell in a different season from calibration, so coverage stayed below target.

Measured values (test period)
HorizonnMAE modelnMAE persistenceSkillInterval coverageInterval half-width
1 h4.76%4.73%−0.6%87.0%±10.3%
6 h11.13%11.56%3.7%85.9%±21.8%
24 h15.52%17.10%9.3%83.0%±24.8%

Loss estimate

Of 10,447 turbine-hours in the test period, 175 were flagged as anomalous; 15 of them were clearly below expected output. That adds up to 10.6 MWh, 0.23% of the period's 4,651.1 MWh.

10.6 / 4,651.1 MWh0.23%

05

Limits, stated openly

  • 01No real weather forecast (NWP) was used; inputs are past SCADA only. That is why persistence is not beaten at the 1-hour horizon.
  • 02Interval coverage stayed below the 90% target because of distribution shift between calibration and test periods. In the field, rolling recalibration is needed.
  • 03The anomaly step uses no fault labels. The result is an estimate of below-expected output, not a claim of detecting faults.
  • 04One wind farm and a single test period. Generalising to another site means measuring again with that site's data.

06

What changes on your site

We repeat the same measurement on your plant's data, with the same openness. Where improvement is expected:

  • Real weather forecasts (Open-Meteo NWP extrapolated to hub height): where the main gain at longer horizons is expected.
  • Site-specific calibration and a rolling interval: to bring coverage closer to target.
  • Matching with maintenance logs: to tell whether flagged hours are faults, curtailment or sensor errors.

07

Reproducibility

The measurement is a single Python script with a fixed seed (42); the same data and library versions give the same result. The script and result file are available on request.

Data: ENGIE, La Haute Borne open data ↗ · Etalab Open Licence 2.0

Let’s run the same test on your plant’s data.