Static Recovery Is Not Dynamic Stability: Dynamics-Aware Benchmarking of Protein Motif Scaffolding
Abstract
Protein motif scaffolding benchmarks aim to measure whether generative models can build scaffolds around functionally important residue sets. Their conclusions depend on two coupled choices: which motifs are tested and which criteria define success. Existing benchmarks typically use small hand-curated motif sets and static refolding-based criteria such as motif recovery, self-consistency, and structural novelty. However, protein function often depends on conformational ensembles and transitions, raising the question of whether static refolding success translates to dynamic fidelity of motifs and scaffolds. Here, we address both benchmark scope and success criteria. We construct the largest systematically derived benchmark of structurally conserved functional motifs from PROSITE-linked experimental structures, yielding 220 cases from 174 PROSITE motif-pattern entries of varying conformations. We evaluate six recent scaffold-generation models and find that performance separates into target coverage and unique-solution yield: RFdiffusion and Proteina solve more targets, whereas ESM3 produces more unique successful scaffolds on the targets they solve. We then perform all-atom molecular dynamics simulations for a set of statically solved designs, comparing motif drift, scaffold drift, motif-contact preservation, and motif-dihedral free-energy landscapes against reference simulations. Of 55 statically solved model–problem pairs, 31 satisfy our MD-supported criterion and only 8 pass all checks. Refolding success is therefore useful evidence for local motif recovery, but not sufficient evidence for dynamical fidelity of the full scaffold, exposing misalignment between static benchmark success and downstream dynamical behavior. The mismatch is motif-dependent: compact, self-contained motifs are often MD-supported, whereas interface-like or context-dependent motifs fail more often. We show that motif-scaffolding models have varying performance on three complementary axes: target coverage, unique-solution yield, and MD-supported dynamical fidelity.