Source-Grounded Executable Evaluation of Software-Migration Agents
Abstract
Software-migration agents are often evaluated through build outcomes, static constraints, or rubric-based inspection, none of which directly establishes preservation of interactive behavior. We introduce \method, a source-grounded and applicability-aware evaluation harness. The current implementation combines deterministic artifact checks with source-conditioned GUI execution over framework-agnostic user goals, while supporting unit-test evidence when source and target expose compatible test interfaces. A fixed rule-based aggregator maps applicable evidence to \Pre, \Deg, \Reg, or \Inc{} without introducing a final LLM judge. We evaluate the framework in two settings. On a 13-case controlled probe suite containing 10 injected faults and three semantics-preserving controls, the combined static and GUI evidence detects all faults without false positives, whereas an output-only LLM rubric detects seven faults and flags two controls. In an initial AIRUN jQuery-to-React pilot~\citep{airun} using outputs from Kimi K2.5, DeepSeek-V3.2, and Qwen3-Coder, all three targets satisfy two shallow output constraints, but only one builds and completes the evaluated workflow. These results constitute a limited pilot rather than a broad model comparison; cases without sufficient executable evidence are reported as inconclusive.