PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools
Abstract
LLM agents are beginning to invoke industrial asset-management tools through the Model Context Protocol (MCP), yet whether they can act reliably on this substrate for safety-critical Prognostics and Health Management (PHM) is unanswered. Prior benchmarks conflate protocol fluency with reasoning, instrumentation failures with agent failures, and tool use with tool retrieval. We introduce PHMForge, an evaluation environment that closes each conflation. PHMForge ships 99 SME-authored scenarios across eight industrial asset classes spanning rotating equipment, aero-engines, and lithium-ion battery cells, on real public datasets including NASA PCoE. The benchmark is served through 39 MCP-native tools that wrap published PHM algorithms (e.g., C-MAPSS, ISO 10816, Arrhenius capacity-fade models, and time-series foundation models for sequence forecasting). Krippendorff's α ∈ [0.74, 0.82] on a 30-scenario stratified rotating-equipment/aero-engine sample; the lithium-ion battery extension is single-rater. Across three agentic frameworks and six LLM backbones, the strongest configuration reaches 80.8% pass@1 on the full 99-scenario set, with the residual gap concentrated in orchestration and tool-sequencing errors. Crucially, an architectural ablation reveals that replacing MCP tool execution with text-based Retrieval-Augmented Generation (RAG) over telemetry-equivalent evidence collapses Remaining Useful Life (RUL) prediction pass-all-3 from 100% to 20% (5/5 vs. 1/5 scenarios) on the lithium-ion battery class, exposing the structural limits of static retrieval for prognostic computation. Trajectory-level decomposition shows orchestration errors dominate failures across backbones, while schema-invalid tool calls are concentrated in smaller open-weight models and rare in frontier configurations. Frontier LLMs are stronger at calling tools than at planning when to call them. PHMForge is open-sourced with deterministic evaluators, a public leaderboard, and a datasheet. A full-suite evaluation costs approximately 20–50 USD per backbone in API spend.