Autonomous Research Project Management as an Agent Skill
Alexander Chen ⋅ Jeffrey Meng ⋅ Bram Hoex ⋅ Tong Xie
Abstract
This work presents an end-to-end demonstration of autonomous machine learning research conducted by an agent skill on consumer hardware. The paper's primary qualifying result – autonomously formulated, implemented, and benchmarked by the agent – establishes that exact Kernel Ridge Regression (KRR) on regular spatial grids can be solved in closed form in $O(N \log N)$ time and $O(N)$ memory. Using Block Circulant with Circulant Blocks (BCCB) embeddings and 2D-FFT spectral shrinkage, this approach eliminates the prohibitive $O(N^3)$ computational bottleneck of dense kernel methods without relying on basis approximations. Evaluated on 2,005 monthly fields of NOAA Kaplan SST v2 climate anomalies ($36 \times 72$ grid), the spectral solver matches a dense floored-torus reference within $2.62 \times 10^{-12}$ relative infinity-norm precision, achieving masked reconstruction RMSEs of $0.082\text{–}0.088$ ($\sim 14\%$ of field standard deviation). The autonomous workflow systematically characterised the non-periodic free-boundary gap (0.95 for Matérn-3/2, 0.12 for RBF), diagnosed localised split-conformal coverage breakdowns under spatial autocorrelation ($0.739\text{–}0.836$ at halo boundaries against a 0.90 target), and pre-registered a transfer-forecasting hypothesis that was rigorously refuted across 0/4 horizons. The research was autonomously executed by DeepSeek V4 Flash, orchestrated by our agent skill suite within DeepSeek Harness (DSH). Experiments were executed on CPU-only hardware (Apple M2 Pro; 78.7 s solver time, 1.57 GB peak RSS). Long-horizon state was decoupled into a file-based epic- and issue-tracking substrate. Across 74 sub-agent sessions, the agent demonstrated closed-loop scientific resilience: routing two failed hypothesis review gates back to literature retrieval, patching bootstrap indexing bugs, and executing with only four discrete human steering events. Finally, we reflect on autonomous research governance, arguing that scientific credibility requires inspectable state, falsifiable review gates, and transparent reporting of negative results, urging the machine learning community to favour agent-accessible structured formats over static PDF manuscripts.
Chat is not available.
Successful Page Load