AI’s Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review
Anamaria Hell ⋅ Kateryna Vovk ⋅ Veena Krishnaraj ⋅ Jia Liu ⋅ Kosuke Aizawa ⋅ Adrian Bayer ⋅ Linda Blot ⋅ Jessica A Cowell ⋅ SUYOG GARG ⋅ Jonathan Grée ⋅ Benjamin A Horowitz ⋅ Masaya Ichikawa ⋅ Kanyuni Iemoto ⋅ Keigo Kondo ⋅ Zacharie Lorsin ⋅ Kevin S McCarthy ⋅ James Robinson ⋅ Miguel Ruiz-Granda ⋅ Leander Thiele ⋅ Ievgen Vovk ⋅ Mingshen Zhou
Abstract
We investigate how well large language models (LLMs) can assist with literature reviews for scientific research. We perform a controlled study of eight expert-conceived research projects across the areas of physics, astrophysics, and cosmology. Each project has a defined background and goal, and human experts and AI prompters are asked to perform identical literature review tasks in parallel. We compare the relevant literature selected by humans with that selected by mid-2025 LLMs. We find the overlap between human- and AI-selected references to be small ($<$6\%), indicating that AI models do not yet reproduce a competent expert search on their own, though they have the potential to complement literature searches by humans. We then assess the reliability and completeness of AI-generated candidate references. We find that while fabricated references make up 3\% of the AI-generated references, 64\% are real papers with at least one metadata mismatch, indicating that the mid-2025 models require systematic verification. However, the performance is significantly improved for the 2026 model ChatGPT Pro 5.5, with a single-project test showing zero fabrication or metadata mismatches.
Chat is not available.
Successful Page Load