Towards Unsupervised Martingale Training Against Belief Entrenchment
Abstract
Rational agents should change their beliefs in response to new information, but the martingale property of Bayesian beliefs says that the direction of these changes should not be predictable from their initial beliefs alone. We use the Martingale Score, which measures this predictability, to introduce Martingale Training (MT), an unsupervised training signal, and find that it shows promise. On a forecasting model deliberately trained to exhibit confirmation bias, MT improves held-out Brier score from 0.291 to 0.234 on Qwen3-32B, roughly matching training with outcome labels. In simulated human–AI conversations, MS detects when a sycophantic chatbot entrenches its users, an effect an order of magnitude larger than any entrenchment in the chatbot’s own beliefs. Training the chatbot on its user’s MS improves the user’s forecasts and reduces entrenchment, but not better than a simple prompt baseline. As language models increasingly shape what people believe, methods that detect and reduce entrenchment could help build assistants that improve reasoning and human rationality, rather than reinforce existing views and beliefs.