An Agentic Approach with Evidence-Grounded Verification for Multimodal Emotion Reasoning in Conversation
Abstract
\begin{abstract} LLM-based multimodal emotion reasoning systems perform well on standard benchmarks. Instruction-tuning the base model can improve performance, but it remains computationally expensive. Agentic pipelines, usually built from frozen models, avoid this cost, but they introduce a new challenge, since a wrong prediction must be routed back to the specific agent responsible for it. We address this with a three-part agentic framework. Modality-specialist perception agents post evidence-cited observations to a shared board, which an integration agent turns into a description, label set, and citation list. An MLLM critic reviews this draft inline and can trigger one revision before the draft is finalized. A separate evidence-grounded verification layer, made of deterministic re-derivation checks and an entailment check, then audits the finalized draft and attributes any failure to a specific agent. A verifier-guided repair loop then clusters attributed failures by defect type and turns them into targeted prompt patches, validated on validation set before admission, with no weight updates anywhere in the loop. On MER2025 dataset, targeted repair beats the frozen static pipeline and two undifferentiated reflection baselines across four MLLMs spanning 2B to 38B parameters, VideoLLaMA3-2B, phi-3.5-vision, Qwen2.5-VL-7B, and InternVL2.5-38B, raising F-score by 0.6 to 2.2 points depending on scale (0.523 to 0.545 on Qwen2.5-VL-7B). A component ablation study further assesses the contribution of each part of the pipeline. Code is available at \url{https://anonymous.4open.science/r/MultiAgentEmotion_Reasoning-CF5A/README.md}. \end{abstract}