Goal Retrieval versus Action-Value Ordering: Contrastive Critics under Best-of-K Selection
Abstract
Goal-conditioned action generators can sample many plausible actions, but best-of-K control requires a critic to order them for a fixed state and goal. Standard contrastive critics are trained and evaluated on a different ranking: distinguishing future goals from unrelated negatives for a fixed state-action pair. Although the population-optimal unnormalized contrastive score is monotone in behavior-policy value, this guarantee does not extend to cosine-normalized scores and need not be realized by a finite trained network. We audit the resulting gap using shared state-goal queries and frozen candidate pools, so critics differ only in their action ordering. Across eight OGBench tasks, raw and cosine critics achieve standard retrieval AUC of 0.96-1.00; on navigation, however, their graded progress ordering is weak or inverted. In an exact-value controlled problem, all three contrastive selectors have higher mean regret than uniform random selection and can worsen as K grows, while TD-Q remains near the pool oracle. On PointMaze, the only task with a reliable learned value reference, TD-Q also provides substantially better fixed-query ordering, although executed-return differences are smaller and task-dependent. Matched controls point to graded value supervision, rather than score bounding or capacity alone, as the main source of improvement. Standard retrieval performance therefore does not certify the action ordering consumed by best-of-K selection; that ordering should be evaluated directly before deployment.