Skip to content

Starting predictive maintenance without labeled failure data

Predictive maintenance does not need a dataset of labeled past failures. Learning normal behavior, linking alarms to a physical fault class and collecting labels along the way is a sound start.

2 min readPredictive maintenance · Anomaly detection · SCADA

Most predictive maintenance projects stall on the same sentence: "We have no labeled failure data." Machine learning brings to mind a large dataset in which every past failure is marked. Yet in most plants failures are rare, their records are scattered and every machine runs a little differently. The good news: you do not need failure labels to get started.

Learn what normal looks like first

The first step is to model the machine's normal behavior, not its failures. Four to eight weeks of vibration, temperature, current and SCADA/PLC records show how a machine behaves under different loads and conditions. Unsupervised methods such as Isolation Forest or autoencoders learn this "normal envelope"; as new measurements move away from it, the health score drops.

The most time-consuming part is not the model but the data: separating downtime, repairing sensor dropouts and telling apart the operating modes of the same machine (idle, part load, full load).

Make alarms meaningful

A health score on its own tells the maintenance team little. The signal and frequency in which a deviation starts point to a physical fault class: bearing, gear, alignment, imbalance or electrical. This classification turns "machine 7 is abnormal" into "machine 7, possible bearing damage" and sends the right team with the right part.

Collect labels along the way

Every alarm should be closed with a technician's finding: was there really a problem, and what was it? Once these findings are written back to the model, the system produces its own labeled data within a few months. Labels become an output of the process, not a precondition.

Show the uncertainty

A remaining-life or output forecast given as a single number creates either too much confidence or too little. Methods such as conformal prediction intervals put next to each forecast the range in which the actual value falls with a stated probability. The maintenance planner then sees the "earliest" and "latest" scenarios together.

Start small, measure

A single line or a group of similar machines is enough for a first pilot. Two things need measuring: how well alarms match technician findings, and how unplanned downtime changes. We have shared how we tested our approach on an open wind farm dataset, including forecast error and interval coverage, on our case study page.