Skip to content

Gaps and bad records in SCADA data: what to do before modelling

Timestamps, gaps, frozen sensors, physically impossible values, curtailment and maintenance periods, sensor replacement: what to check in SCADA data before building a predictive maintenance or power forecasting model.

2 min readSCADA · Data quality · Predictive maintenance · Wind

In predictive maintenance and power forecasting projects, most of the time goes into the data, not the model. SCADA systems are built for operations, not for analysis: records are usually stored as 10-minute averages, communication outages leave gaps, and sensor replacements and recalibrations leave no trace in the data. A model trained before these problems are dealt with learns data errors, not faults.

The checks below are the ones we run at the start of every project on wind farm and rotating equipment data.

1. Timestamps

The first question is which time zone a timestamp belongs to, and whether it marks the start or the end of the averaging interval. In data recorded in local time, one hour appears twice or not at all when daylight saving time changes. Data from different systems (SCADA, alarm logs, maintenance work orders) should all be converted to UTC before they are joined.

2. Gaps

A communication outage deletes data; some systems fill the gap with zeros or with the last value. A power record filled with zeros tells the model "the turbine produced nothing". A gap should stay a gap, and coverage should be reported per channel.

3. Frozen values

A faulty or disconnected sensor can repeat the same value for hours. Values that do not change at all over a given period should be flagged. This matters most for temperature and vibration channels: a frozen temperature hides real heating.

4. Physically impossible values

A negative wind speed, a direction above 360 degrees, output far above rated power, or a bearing temperature below ambient are data errors. Define physical lower and upper limits for every channel and flag the records that exceed them.

5. Curtailment, maintenance and fault periods

A normal behaviour model should learn from periods when the machine runs healthy and unconstrained. If grid curtailment, planned maintenance, icing and fault stops stay in the training data, the model treats them as "normal" and misses a real deviation. These periods should be derived from status codes and alarm logs and flagged.

6. Sensor replacement and level shifts

When a sensor is replaced or recalibrated, the measured level can shift suddenly, and the model mistakes it for degradation. Sudden level changes that maintenance records do not explain deserve a separate look.

Flag instead of delete

Rather than deleting bad records, adding a flag that says why each record is questionable has two benefits: clean windows can be selected for training, and loss calculations can state plainly which part of the data could not be assessed. We describe our power forecasting and loss study on an open wind farm dataset in our case study, and how to proceed without labelled failure data in this note. You can run most of these checks on your own CSV file, without uploading it, with our SCADA data quality check. For a first assessment with your own data, get in touch.