Chonghyo Joo, Ye Seol Lee
Solubility prediction across temperature is crucial for solvent selection and process design, yet remains challenging for data-driven models because temperature is typically treated as a simple scalar input. Such treatment limits the ability of data-driven methods to represent temperature-dependent solute–solvent interactions within their latent feature spaces. Here, we introduce a temperature-conditioned (T-conditioned) molecular representation that embeds temperature directly into learned features through feature-wise linear modulation (FiLM) after dimension expansion. We evaluate this approach using two representative solubility prediction architectures: Chemprop, based on learned graph-based representations, and Fastprop, based on fixed descriptors. Trained on BigSolDB and assessed under both interpolation and extrapolation to unseen solutes using the Leeds and SolProp datasets, T-conditioning consistently reduces prediction error by up to 38% for Chemprop, with particularly strong improvements at elevated temperatures (up to 51% R M S E reduction above 330 K), while gains for Fastprop remain modest. Analysis of the T-conditioning parameters reveals that temperature-dependent modulation increases with temperature, indicating that the T-conditioning layer captures temperature-sensitive latent features relevant to solubility. An application study on pyrazinamide solubility prediction further demonstrates improved agreement with experimental trends and highlights the utility of representation-level conditioning when mechanistic models are constrained by parameters. These results demonstrate that T-conditioning enables accurate and physically meaningful prediction of temperature-dependent properties by embedding temperature directly into molecular representations.