Evaluates diagnostic accuracy achievable by LLM agents for bearing fault diagnosis when restricted to free, offline commodity hardware.
The study benchmarks eight small, publicly accessible LLMs (1B–9B parameters) running locally via Ollama on three standard datasets (Paderborn, CWRU, HUST) to assess their diagnostic accuracy on vibration-derived features for three-class bearing fault diagnosis.
Key Findings: Top Performer: The model gemma2:9b (5.4 GB) emerged as the most accurate and reliable, achieving approximately 0.70–0.87 accuracy with quantified reliability using a training-free, zero-shot or few-shot approach. Comparison to Frontiers: On identical prompts and data, the best free local models matched or exceeded a paid frontier model (Claude Sonnet) in accuracy, with no advantage gained from cloud-hosted services. Baseline Limitations: While the local LLMs are useful for resource-constrained SMEs due to their offline, free, and training-free nature, they are less accurate than simple supervised baselines (e.g., logistic regression, random forest) which achieved 0.97–1.00 accuracy on the same features when labeled data was available. Reliability Mechanism: The study proposes ensemble agreement gating, a training-free rule that withholds predictions when multiple local models disagree, raising accuracy on the retained subset (e.g., from 0.77 to 0.87 on Paderborn).
The study concludes that while local LLMs do not outperform supervised classifiers in pure accuracy, they offer a practical, expertise-free solution for SMEs lacking labeled data or ML staff, delivering useful diagnostic capabilities at zero cost and fully offline.
The study examines whether locally runnable large language models can perform useful bearing fault diagnosis without dependence on cloud services, paid APIs, or high-end accelerators. It focuses on a practical deployment scenario in which an LLM agent must reason about bearing condition using free, offline, commodity hardware, making it a test of the gap between frontier LLM capabilities and what is achievable in resource-constrained industrial or edge environments. Rather than treating LLMs as general-purpose diagnostic oracles, the work frames them as components in a diagnostic workflow, where model choice, inference constraints, and agent design all affect the reliability of the final fault assessment.
Its main contribution is an evaluation of the diagnostic accuracy that can realistically be obtained under these constraints, rather than an optimistic demonstration using unconstrained model access. The material is useful for understanding which aspects of bearing fault diagnosis are tractable with local LLMs—such as high-level triage, interpretation of preprocessed condition indicators, generation of maintenance guidance, or explanation of likely failure modes—and which remain difficult, including fine-grained defect classification, quantitative severity estimation, and consistent decision-making across closely related fault types. It also highlights the trade-offs among model size, memory footprint, inference speed, and diagnostic performance when the system must run entirely on locally available hardware.
This matters because bearing faults are a common and consequential failure mode in rotating machinery, and many operational settings prioritize data locality, low cost, offline operation, and integration with existing maintenance workflows. By grounding the analysis in locally runnable models and commodity hardware, the study provides a realistic benchmark for teams considering LLM-assisted condition monitoring. It also supports a broader design implication: in industrial fault diagnosis, LLMs are likely to be most valuable not as standalone sensors or classifiers, but as reasoning and communication layers atop more deterministic signal-processing or specialized machine-learning components.