LLM-in-the-Loop Feedback-Driven ETL: Toward Trustworthy Data Pipelines
Abstract
We present a modular architecture for Extract-Transform-Load (ETL) pipelines augmented by large language models (LLMs): each pipeline stage is executed by an LLM "Worker" agent and validated by a paired "LLM-as-a-Checker," with a feedback loop that consolidates human corrections and model-flagged discrepancies. An early pilot suggested a circular validation failure mode, a checker from the same model family as the worker appearing blind to that family's errors, echoing known limits of LLM self-correction and self-preference in LLM evaluators. We then subjected this hypothesis to a controlled ablation: six checker configurations (same-family, same-family with Reflexion-style self-critique, deterministic rules only, cross-architecture, and hybrids of deterministic rules with each LLM channel) under identical seeds, a shared worker transform, prompt-controlled conditions, and explicit ground truth, on both UCI Adult and real data-lake files from the KramaBench benchmark. The ablation does not reproduce family blindness; instead it reveals a precision/recall structure the pilot's single configuration could not see: the same-family checker in our pairing over-flags (recall 0.84-0.93 at precision 0.49-0.64), cross-architecture checkers are conservative but precise (recall 0.22-0.28 at precision 0.94-1.00), self-critique shows no detectable benefit over single-pass same-family checking at this scale, and deterministic rules are precision-guaranteed on rule-covered classes by construction. The rules+cross hybrid detects more than either of its channels alone (recall 0.45-0.51 at precision ≥0.96, with one false positive across 210 checked rows), showing that rule-based and cross-model validation catch complementary errors; a rules+same-family variant attains the highest observed recall (0.96 on UCI Adult; between-run) but inherits the same-family false-positive flood, confirming that hybridization complements precision rather than repairing it. This is the decisive property for human-in-the-loop pipelines, where every false flag costs reviewer time. We trace the pilot's apparent blindness to an under-specified checker prompt, identifying prompt specification as a first-order variable that LLM-validation studies must control, and release a fully seeded harness with per-verdict path attribution from which every reported ablation number regenerates.