Reproducibility study of FACTER: Fairness-Aware Conformal Thresholding and Prompt Engineering
Abstract
We present a reproducibility study of FACTER, a post-hoc framework that combines confor- mal thresholding with iterative prompt engineering to mitigate demographic bias in black- box LLM-based recommender systems. Using the released codebase and experimental set- ting from the original paper, we evaluate FACTER on MovieLens-1M and Amazon Movies & TV with various LLM backbones. We assess fairness using the reported violation-based criterion group and counterfactual metrics (SNSR, CFR), and measure recommendation quality via catalog-mapped ranking metrics (NDCG@10, Recall@10) alongside the validity rate of generated items (Valid@10). Across datasets and supported backbones, we repro- duce FACTER's key qualitative behavior, with fairness violations decreasing sharply and converging within a small number of calibration rounds. However, unlike the original study, we observe substantially lower recommendation quality for both Zero-Shot and FACTER, with low validity under open-vocabulary generation, item mapping, and evaluation-protocol differences providing plausible explanations for this discrepancy. We further identify and resolve multiple implementation and reproducibility issues in the released code, providing a cleaner and easier-to-run codebase to support future replication. Overall, our findings sup- port FACTER's effectiveness in reducing measured fairness violations, but the near-floor utility leaves utility preservation unverifiable in our reproduction.