Spotting machine failures before they happen
Ten thousand milling cycles, five failure modes and one very clear warning sign: tool wear past 200 minutes.
Which sensor readings signal that a machine is about to fail, and what should maintenance do about it?
In my Predictive Maintenance project I built and tuned Random Forest and Neural Network models on a 5,000-record dataset (80% accuracy) and turned the results into business-facing reports. This page explores the public AI4I 2020 dataset that is widely used for the same problem.
What breaks: failures by mode
Failed cycles only; one failure can have more than one mode
- Heat dissipation115
- Overstrain98
- Power95
- Tool wear46
- Random1
Failure rate climbs with tool wear
Share of cycles that failed, by minutes of tool wear
Failure rates spike at the torque extremes
All 339 failures vs a 1,200-cycle sample of healthy runs
- Healthy cycle
- Failure
Failure rate by product grade and torque
Darker = more failures per cycle
Average sensor readings: failed vs healthy cycles
| Failed | 339 | 300.9 | 310.3 | 1,496 | 50.2 | 144 |
| Healthy | 9,661 | 300 | 310 | 1,540 | 39.6 | 107 |
- Only 339 of 10,000 cycles failed — a 3.4% failure rate. On data this imbalanced, a model that always predicts "no failure" scores 96.6% accuracy, so recall and precision on the failure class are the numbers to watch.
- Heat dissipation is the most common failure mode (115 failures); 9 failures carry no mode label at all.
- Failure risk stays low until tool wear reaches 200 minutes, then jumps from 2.2% to 15.4% — about 7× higher.
- Failed cycles run at 50.2 Nm of torque on average vs 39.6 Nm for healthy ones. Most failures (60.5%) still happen at normal torque, but in the 60+ Nm band up to 46.3% of cycles fail.
- Schedule tool changes at about 200 minutes of wear instead of running to failure.
- Alert operators in the 60+ Nm torque band, especially on low grade (l) products.
- Judge any predictive model on recall and precision for the failure class, not overall accuracy.
- Profiled 10,000 cycles: temperatures, speed, torque, tool wear and five labelled failure modes.
- Counted failures per mode (a single failure can trigger several modes).
- Banded tool wear and torque to see where failure risk jumps.
- Compared average sensor readings for failed vs healthy cycles.
- The data is synthetic, so real machines will be noisier.
- Findings are descriptive; a production model would need time-based validation.
- Class-imbalance analysis
- Failure-mode breakdown
- Feature banding
- Group comparison