Your Ground Truth Is Running the Same Biased Estimator as Your Method: The Missing Convergence Criterion
Abstract
We built a pipeline turning segmentation masks of a banana leaf into agronomic parameters, validated it against a synthetic leaf with analytically exact ground truth, and reported a perimeter error of −0.01 %. The number was an artefact: the generator computed its reference with the same binary marching-squares call the pipeline used, so two instances of one bias were measuring each other. The true error was +4.55 %. That such coupling is a known failure mode -- the inverse crime -- we concede. We report that its standard antidote fails here, and what that cost. Refining the reference grid cannot separate reference from method when the shared estimator has a bias floor: binary marching squares self-converges at observed order 2.79 while its true error sits flat at +8.20 %, and the refinement study ranks the estimators backwards. The defective reference then inverted the ranking of two contour extractors, favouring the more elaborate route while it was 2.9× less accurate. Auditing the stack a practitioner inherits by default -- scikit-image, OpenCV, PlantCV, HistomicsTK, PyRadiomics, IBSI -- we find the bias stated almost nowhere at the point of use, the one standard that does state it prescribing an estimator of the same class, and the recommended remedy not independent at all: four-direction discrete Crofton is exactly marching squares divided by Kulpa's constant on saddle-free masks. More pixels do not help, so none of this can be bought off with a better sensor. We give the three-grid acceptance test that catches it, which needs no reference instrument and no annotated data, and a convergent estimator giving −0.26 % on our leaf mask where marching squares gives +4.08 %.