Reliability & Fallback Design Patterns for Multi-Agent LLM Systems
Abstract
When a multi-agent LLM pipeline fails mid-run, most systems stop or blindly retry the same stage. Neither is a reliability strategy. We evaluate seven conditions on 150 HotpotQA multi-hop questions across three trials (n=450 per condition) on a fixed Planner-Retriever-Synthesizer-Formatter pipeline: no fallback, blind retry, four targeted recovery patterns (tool-grounded retry, deterministic checkpointing, graceful degradation, cross-agent verification), and a composed checkpointing-plus-retry pattern. Four targeted conditions significantly beat the no-fallback baseline on pass rate after Bonferroni correction (p<0.001). Blind retry does not. A cross-agent verifier gate costs 2.5x more for strictly inferior reliability. Composing two targeted patterns does not outperform the best single pattern, suggesting a reliability ceiling. A P1 check ablation shows that schema/JSON validity carries most of the single-check lift; retrieval-only and citation-only alone do not beat B0. We release code, data, and logs