SuperSycophantic: Stress-Testing Frontier LLMs from Single- to Multi-Turn Sycophancy
Abstract
AI sycophancy is gaining increasing prevalence as models optimized to maximize user satisfaction may tend to unconditionally agree with users even at the expense of factuality, which poses great risk as AI are increasingly used for decision support in high-stake scenarios. We present \ourWork, a systematic stress-testing framework encompassing objective (OBJ) questions with verifiable answers and subjective (SUB) scenarios without ground truth on sycophancy induced from first-turn context framing to multi-turn user pressure. Evaluation of 9 frontier models revealed high sycophancy in even the best-performing models such as GPT-5.4 changes answers from right to wrong to please users in 23.8\% of OBJ scenarios and blindly follows the pressured user view in more than half of the SUB scenarios. We found that strength of user tone is one of the most impactful yet previously overlooked factor for AI sycophancy and discover an interesting pattern where Claude models are uniquely more sycophantic under moderate pressure because strong triggers often lead them to rethink questions from scratch to arrive at the correct impartial answer. These findings highlight the need for anti-sycophancy training in future model development, where training models to rethink from scratch when facing pressure may serve as a promising paradigm towards more truthful AI systems. We provide our code and data in the supplementary material.