AstroThink: Can Large Language Models Reason Their Way Through Astrophysics Olympiad Problems?
Abstract
State-of-the-art (SOTA) benchmarks and models exist for solving olympiad problems from many domains such as mathematics, informatics and physics. However, there are only some benchmarks available for the astrophysics domain, which only partially support astronomy. These efforts incorporate multiple-choice questions and single-word answers, which often only test memory instead of conceptual understanding and breadth of knowledge. The International Olympiad of Astronomy and Astrophysics (IOAA) 2024 question paper includes problems that require in-depth thinking and robust conceptual foundations. We present a comprehensive evaluation of Large Language Models’ (LLMs) performance on the IOAA questions, using a two-fold validation scheme, encompassing both human annotators and an LLM judge, along with a qualitative grading criteria to identify the most common points of failure in numerical calculations and reasoning steps. None of the chosen LLMs could achieve medal-level performance (the top 50th percentile or above, calculated from a reference score for the theory examination), and struggled significantly with the reasoning part of each problem. The best performance was achieved by DeepSeek-V3.1 with 44.83%. Hence, these LLMs are incapable of reasoning correctly through astrophysics olympiad problems, since they suffer from logical errors, utilisation of irrelevant approaches and contextual misunderstanding.