AI’s Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation
Abstract
We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three mid-2025 LLMs (ChatGPT, Claude, and DeepSeek). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude Opus 4.8 and ChatGPT Pro 5.5) using a four-aspect evaluation rubric, and were asked to judge whether each proposal was written by a human or an AI. Human reviewers rated human- and AI-written proposals similarly, whereas both AI reviewers scored AI-written proposals about one point higher (on a five-point scale) than human-written proposals. Human reviewers correctly identified human- and AI-written proposals 72\% and 79\% of the time, respectively, while both AI reviewers correctly classified all 32 proposals (100\%). This suggests that LLMs can produce project plans comparable to human-written ones in the eyes of human reviewers, but AI reviewers show a systematic pro-AI preference---a concrete risk for AI-assisted proposal evaluation.