Silent Metrics: Compiler-Dependent Resource Accounting and Its Consequences for Q-HPC Benchmarking and Scheduling
Abstract
Benchmarks for AI-generated quantum code grade circuit depth and two-qubit gate count alongside scientific correctness, and the same counts feed cost models and placement decisions in hybrid Quantum-HPC (Q-HPC) runtimes. These quantities depend on the tuple (construction path, target basis, coupling map, compiler version), not the algorithm alone, a dependence long familiar from quantum compilation. What we report is how quietly it corrupts the systems that consume them: treated as invariants, these counts fail silently in graders and schedulers alike. In a 96-task benchmark, four of seventeen QAOA Max-Cut tasks reported zero two-qubit gates: all 17 build the same cost layer from RZZ gates, which Qiskit Aer keeps native rather than decomposing to CX, and only those four extracted the metric by gate name. Over four construction paths and four targets, replicated on five graph instances, a two-name cx, cz extractor returns zero in 5 of 16 configurations every time and the benchmark's shipped cx, cz, ecr extractor in 1 of 16, while structural counting spans a factor of 3.5 to 6.1 for identical scientific output, widening with instance size. Feeding these counts to a standard fidelity-based placement rule, we show that the compilation target alone decides hardware-versus-simulation over a wide band of thresholds, and that the extraction defect reports the maximum possible fidelity, routing to hardware any workload whose true count implies otherwise. We derive four requirements for resource-metric assertions and audit our benchmark's graded fields against them.