Deciding What a Polymer Representation Should Encode: Establishing What the Benchmark Can Test
Abstract
Polymers are extended molecules with a structural hierarchy that affects structure--property relationships. While repeat-unit (monomer) chemistry can be represented as a finite molecular structure, polymer identity additionally depends on higher-level information such as composition (the proportion of monomer types), molecular size, and chain architecture (monomer arrangement and chain structure). Polymer benchmark data sets are smaller and less consistent than established small molecule benchmarks. To drive model development that successfully encodes the relevant hierarchical structural information, we ask whether predictive performance on available benchmarks can support claims about the polymer information a representation captures. We analyse public polymer data sets independently of trained models in two stages to establish what polymer-specific information a benchmark data set can meaningfully test. First, we test the distinguishability of the association of monomer chemistry, chain architecture, and chain composition with the target. Second, we estimate how much target property variation is associated with each factor, accounting for apparent explanatory power that arises from dividing observations into smaller groups. (i) Architecture cannot be independently evaluated. In a widely used copolymer benchmark of 42,966 polymers with computed EA and IP, architecture and composition are incompletely crossed, and architecture has no replication once monomer identity and composition are fixed. Its association with the target properties therefore cannot be cleanly separated from composition. Predictive performance cannot establish that an architecture-aware representation succeeds because it captures architecture. (ii) Chemistry carries most of the readily predictable signal. The predictors below are group means, not trained models: each returns the training mean over polymers sharing the relevant factor values, so its error measures what that factor alone makes recoverable. On EA, a monomer-pair lookup using no polymer-level structural information reaches MAE 0.171 eV, compared with 0.475 eV for a constant predictor. Conversely, composition and architecture without monomer identity yield 0.472 eV. (iii) The limitation is widespread. Of 11 public polymer datasets, identified from a survey of 19 polymer data resources, only two permit chemistry to be compared with independently varying polymer-level structure. Eight are homopolymer datasets with no composition or architecture variation, and seven contain only one observation per distinct monomer type, preventing chemistry from being separated from polymer identity for this analysis. (iv) Preliminary - what matters appears to depend on the property. On the same 42,966 polymers, composition has a measurably different association with EA and IP, while chemistry dominates both. With the molecular sample fixed, this preliminary result is observed for only one pair of targets. This work relates to construct validity in ML benchmarking: whether a score provides evidence for the claim drawn from it. While existing work considers task content, shortcuts, and controlled evaluation, we address a complementary question for polymer representation learning - does the recorded structural variation permit the claimed polymer-specific factor to be independently evaluated? Our results motivate asking what polymer information is relevant to the target and whether the benchmark can evaluate it before deciding what complexity a representation should encode.