← All work
Manufacturing analytics

Spotting machine failures before they happen

Ten thousand milling cycles, five failure modes and one very clear warning sign: tool wear past 200 minutes.

Original project · 2025PythonScikit-learnRandom ForestNeural Networks
The question

Which sensor readings signal that a machine is about to fail, and what should maintenance do about it?

The original project

In my Predictive Maintenance project I built and tuned Random Forest and Neural Network models on a 5,000-record dataset (80% accuracy) and turned the results into business-facing reports. This page explores the public AI4I 2020 dataset that is widely used for the same problem.

At a glance
Machine cycles10,000
Failures339
Failure rate3.4%
Top failure mode115Heat dissipation
The analysis

What breaks: failures by mode

Failed cycles only; one failure can have more than one mode

  • Heat dissipation115
  • Overstrain98
  • Power95
  • Tool wear46
  • Random1
Heat dissipation failures lead, followed closely by overstrain and power failures.

Failure rate climbs with tool wear

Share of cycles that failed, by minutes of tool wear

0%5%10%15%20%2.2%0–492.3%50–992.3%100–1492.7%150–19915.4%200+
From 200 minutes of wear the failure rate is 7× the baseline.

Failure rates spike at the torque extremes

All 339 failures vs a 1,200-cycle sample of healthy runs

  • Healthy cycle
  • Failure
0204060801,0001,5002,0002,5003,000Rotational speed (rpm)Torque (Nm)
Most failures (60.5%) happen at normal torque, but the failure rate spikes at the edges — high torque at low speed, low torque at high speed.

Failure rate by product grade and torque

Darker = more failures per cycle

< 20 Nm20–39 Nm40–59 Nm60+ Nm
Low grade (L)16.7%0.6%4.5%46.3%
Medium grade (M)16.2%0.6%2.5%35.1%
High grade (H)3.6%0.7%2.4%33.3%
High torque is dangerous for every grade — up to 46.3% of those cycles fail.

Average sensor readings: failed vs healthy cycles

Average sensor readings: failed vs healthy cycles
Failed339300.9310.31,49650.2144
Healthy9,6613003101,54039.6107
What the data shows
  1. Only 339 of 10,000 cycles failed — a 3.4% failure rate. On data this imbalanced, a model that always predicts "no failure" scores 96.6% accuracy, so recall and precision on the failure class are the numbers to watch.
  2. Heat dissipation is the most common failure mode (115 failures); 9 failures carry no mode label at all.
  3. Failure risk stays low until tool wear reaches 200 minutes, then jumps from 2.2% to 15.4% — about 7× higher.
  4. Failed cycles run at 50.2 Nm of torque on average vs 39.6 Nm for healthy ones. Most failures (60.5%) still happen at normal torque, but in the 60+ Nm band up to 46.3% of cycles fail.
Recommendations
  • Schedule tool changes at about 200 minutes of wear instead of running to failure.
  • Alert operators in the 60+ Nm torque band, especially on low grade (l) products.
  • Judge any predictive model on recall and precision for the failure class, not overall accuracy.
How the analysis works
  1. Profiled 10,000 cycles: temperatures, speed, torque, tool wear and five labelled failure modes.
  2. Counted failures per mode (a single failure can trigger several modes).
  3. Banded tool wear and torque to see where failure risk jumps.
  4. Compared average sensor readings for failed vs healthy cycles.
Caveats
  • The data is synthetic, so real machines will be noisier.
  • Findings are descriptive; a production model would need time-based validation.
Methods
  • Class-imbalance analysis
  • Failure-mode breakdown
  • Feature banding
  • Group comparison

Want to work together?

I'm open to full-time roles and internships across the U.S. The best way to reach me is email.

Say hello ✉📍 Dallas, TX

Designed with care · 2026