Learning to Think from Multiple Thinkers: How Heterogeneity in Chain of Thoughts Affects Efficient Learnability
Abstract
Chain-of-Thought (CoT) supervision can make complex reasoning efficiently learnable, yet existing theory largely assumes that all traces come from a single homogeneous source. In practice, CoT data is heterogeneous, coming from different sources, and requires substantial efforts (such as filtering) to produce homogeneous supervision. Motivated by this gap, we introduce learning-theoretic models of CoT collection from multiple sources, spanning passive collection to active and adaptive curation. Under passive collection, we show that input-dependent contribution biases can make learning cryptographically hard, even though learning is provably easy under homogeneous CoT supervision from any single source. Good coverage, however, can still be achieved efficiently. Even without such biases, filtering can be computationally necessary: learning is easy if traces from any one source can be isolated, whereas their unfiltered mixture can remain cryptographically hard to learn from. Finally, inspired by iterative and adaptive fine-tuning, we give a generic efficient algorithm that collects homogeneous batches across rounds and combines information from different, adaptively chosen sources to learn to arbitrary accuracy, while requiring from each source a number of CoT examples independent of the desired accuracy.