KeplerBench: A Benchmark for Scientific Discovery from Observations
Abstract
Artificial intelligence is now expected to make scientific discoveries, and general relativity has been proposed as the test of whether it can take a conceptual leap of the highest order. We argue for an earlier and simpler test that current systems have not passed. Kepler's laws were among the first non-obvious mathematical statements about the physical world, the ellipse with the Sun at one focus overturned two thousand years of circular astronomy, and the laws led directly to Newton's. Kepler worked from Tycho Brahe's sparse and noisy records of where the planets appeared in the sky, and Astronomia Nova documents his reasoning in unusual detail. We introduce KeplerBench, which gives a learner historically faithful observations of this kind and asks whether it can propose a law as general as Kepler's, in whatever form, without being told the variables in which the law is simple. Simulated star systems with different masses, orbits, and force laws then test whether the discovery procedure generalizes beyond the Solar System. A pilot on Tycho's Mars record shows that a flexible interpolator fits the observations and fails on later years, while a Keplerian orbit degrades gracefully. The benchmark scores three levels of capability, prediction, global law, and local law, and we offer it as a test for systems that claim to be scientists.