ProSearch: Benchmarking Multi-Constraint Protocol Retrieval in Experimental Science
Abstract
In real-world AI for Science scenarios, retrieval systems are increasingly expected to support experimental workflows such as experimental design and method reproduction. These workflows require retrieving protocols that satisfy a set of interdependent constraints, such as specific materials and operating conditions. However, systematic evaluation of retrievers in such multi-constraint scenarios remains lacking. To address this gap, we introduce ProSearch, a scientific retrieval benchmark for multi-constraint experimental protocol retrieval. ProSearch has three key features: (1) realism, as queries are instantiated from real method paragraphs using expert-defined experimental fields, query templates, and filtering rules; (2) complexity, involving 60 distinct fields, such as experimental subjects, materials, and operating conditions, together with dependencies among these fields; and (3) diagnosticity, as each query contains 3 to 5 explicit hard constraints and covers challenging cases such as numeric constraints and negation constraints. We construct ProSearch through LLM-based extraction, data-driven constraint selection, controlled query generation, and evidence-based relevance validation. The final benchmark contains 1,551 high-quality queries across six experimental science domains. We use ProSearch to systematically evaluate 19 representative retrievers, spanning lexical matching, general dense retrieval, scientific retrieval, and reasoning-oriented retrieval. The results reveal that current retrievers generally struggle with multi-constraint protocol retrieval, with the best-performing model achieving an average nDCG@10 of 0.290. Further constraint-wise and error analyses reveal key failure modes of existing retrievers.