AUTOEMBED: Can Coding Agents Train Text Embedding Models?
Abstract
Can coding agents autonomously train useful text embedding models? We introduce AutoEmbed, a configurable framework and benchmark for developing and evaluating embedding models under fixed budgets. Each run provides a base model, one H100, and ten hours. We split every evaluation task in half: the agent develops on one half and is scored on the other. Across 58 runs by four agents, six settings test building from a raw encoder, improving a strong embedder, and adapting an embedder to legal, financial, medical, and code retrieval. The base model's headroom dominates performance. All four agents raise the raw encoder's score from 38.0 to 58.2--62.8, and the best comes within 1.5 points of a same-size published embedder, although retrieval remains fourteen points behind it. None meaningfully improves the strong embedder; domain gains range from 0.2 to 8.5 points. The method an agent chooses is associated with its outcome: full fine-tuning gains 8.9 points on average against 2.2 for low-rank adaptation, though the comparison is observational. Every submission declares its training data, and the audit finds no run trained on a hidden query with its relevant document. Reviewing traces and checkpoints reveals what the audit cannot: six submissions use prompts keyed to evaluation task names, and one submitted encoder is byte-identical to its base, with prompts supplying two thirds of its gain. Whether such conditioning is adaptation or reward hacking depends on what a benchmark sets out to measure, a line contamination auditing cannot draw. Code and artifacts will be released upon publication.