AI says itβs 92% confident. But 92% of what? In Episode 61 of Chaos & Caffeine, we strip away the marketing language around artificial intelligence and machine learning in reliability and maintenance. We start with something many people may not realize: AI is broader than machine learning. Rule-based systems, expert systems, machine learning, and hybrid combinations of these can all fall under the AI umbrella. From there, we break down what accuracy, precision, recall, specificity, F1 and F2 scores, ROC/AUC, anomaly scores, health scores, risk, criticality, thresholds, model drift, sensor drift, and ground truth actually mean.
If you are evaluating predictive maintenance, wireless monitoring, CMMS/EAM platforms, condition monitoring, or any product advertised as “AI-powered,β this episode is about knowing what questions to ask before believing the score on the screen. A system can claim 95% accuracy and still generate mostly false alarms, and a β92% confidenceβ score may not mean there is a 92% probability that anything is actually wrong. We also dig into why synthetic data, missing ground truth, poorly defined health scores, and claims that AI can simply βfill inβ missing maintenance history deserve serious scrutiny. Donβt buy AI because the dashboard looks intelligent. Buy evidence that it actually works in the field.















