SolvMix: Learning Formulation-State Landscapes for Liquid Electrolyte Conductivity Prediction
Abstract
Electrolyte conductivity is not a property of isolated molecules, but a response over a formulation landscape defined by solvents, salts, composition, salt mole ratio, and temperature. The same molecular component can play different effective roles as its amount and operating conditions change, while the underlying formulation state is observed only through macroscopic conductivity measurements. This makes conductivity prediction a distinctive mixture learning problem: a model must learn how molecular identities become formulation-specific states, rather than merely how molecules are represented. We introduce SolvMix, a framework for learning formulation-state landscapes for electrolyte conductivity prediction. SolvMix constructs component tokens for solvent and salt molecules, calibrates them into amount-aware component states before mixture interaction, and predicts conductivity through a temperature-formatted readout. We evaluate SolvMix on three curated electrolyte conductivity benchmarks, CALiSol-23, DiffMix, and Bamboo-mixer, covering diverse solvents, salts, compositions, salt mole ratios, and temperatures. Across standard splits on all three benchmarks and binned out-of-distribution (OOD) splits on Bamboo-mixer, SolvMix outperforms strong molecular-representation, mixture-aggregation, and geometric-interaction baselines under the reported protocols. Ablations show that the best-performing formulation-state construction combines amount-and-type-aware component tokens, multiplicative amount calibration, token-level interaction followed by mean pooling, and late-concat temperature formatting. OOD results further show clear gains under formulation-level shifts, where static component descriptors and local interaction patterns are least aligned with the required generalization. These results suggest that electrolyte conductivity prediction benefits from treating formulation-state construction as a first-class modeling problem, beyond molecular representation learning or explicit local interaction modeling.