materials-chemistry/ai produced the result/arXiv 2025 · v2
Machine learning fitted to flexible molecules predicts oxidation potentials and hydration energies
Researchers rebuilt a standard statistical fitting method so it could handle molecules that bend into many shapes, then trained it on measured oxidation potentials and hydration energies. The model did all the property prediction.
spectrum · one line per step, placed by what the step does · bright lines used AI
Kernel Ridge Regression for conformer ensembles made easy with Structured Orthogonal Random Features
arXiv, 2025
doi:10.48550/arxiv.2505.21247 · record aix-00205 v2 · checked 2026-10-09
- AI was for
- Property prediction
- Model family
- Multilayer perceptron
- Checked by
- Held-out
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Molecules are not rigid. A single compound can twist and rotate about its bonds into many different shapes, called conformers, and at room temperature it flickers between them. Each shape is present in a certain proportion, set by its energy. That makes life awkward for anyone trying to predict a molecule's behaviour from its structure, because there is no single structure to point at. Two properties of practical interest are the oxidation potential, the voltage at which a molecule gives up an electron in a solvent, and the hydration energy, how readily it dissolves in water. Both are measured in the laboratory, one compound at a time.
The researchers set out to predict these two properties from structure alone, while treating each molecule as a weighted collection of its conformers rather than a single frozen shape. They also wanted the method to stay cheap to run, so they rewrote an established fitting technique, kernel ridge regression, in a form built from layers of random features. They then tested it on published experimental measurements of oxidation potentials in the solvent acetonitrile and on hydration energies from a public collection.
Where AI came in
The learned model is the whole result here. Conformers for each molecule were generated with conventional chemistry software, and each one was turned into a list of numbers describing its structure. Those numbers were fed through the researchers' layered model, which was then fitted to the experimental values of the training molecules. Settings such as the widths of the model's internal functions were tuned automatically, partly with a Bayesian optimisation package, by checking how well the model did when each training molecule was held back in turn.
The fitted model was then asked to predict the properties of molecules it had never seen, standing in for measuring those compounds in the laboratory. Prediction error fell as the training set grew for every structural description tried. For oxidation potentials, several of those descriptions gave errors close, within the statistical scatter, to a figure of 0.2 eV reported previously. Using only each molecule's lowest-energy conformer instead of the full ensemble shifted the errors by about as much as that scatter.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors rewrite Kernel Ridge Regression over Boltzmann ensembles of molecular conformers in terms of Structured Orthogonal Random Features, giving a layered model (Multilevel SORF) with trigonometric activations and layers for ensembles, conformers, atoms and localized orbitals. Conformers were generated with the MMFF94 forcefield, encoded with several existing molecular representations, and models were fitted to two experimental datasets: oxidation potentials in acetonitrile (largest training set 474 molecules) and FreeSolv hydration energies (largest training set 514 molecules). Mean absolute error decreased with training set size for every representation tested, and for oxidation potentials the combinations with aSLATM, SLATM, FCHL19 and cMBDF gave errors similar, within statistical error, to the 0.2 eV reported earlier for SLATM with KRR. Replacing each ensemble with only the lowest-MMFF94-energy conformer changed the errors by an amount comparable to the statistical error. The work also introduces self-consistent versions of the Huber and LogCosh loss functions used during hyperparameter optimization.
How AI was used
Supervised regression models were fitted in this study to predict two experimental molecular properties. For each molecule, conformers were generated with the Morfeus implementation of a conformer-generation algorithm using the MMFF94 forcefield in RDKit at 298.15 K, with low-weight conformers cut off at rho_cut = 0.05; Boltzmann weights were retained as conformer importance weights. Each conformer was encoded with one of several molecular representations (Coulomb Matrix, SLATM, aSLATM, FCHL19, cMBDF, SOAP, and FJK, the latter from Huckel calculations in PySCF with a 6-311G basis, Boys localized orbitals and SMD solvation). These representations were fed through stacks of SORF-based layers (feature, sum, switch, weighted-sum, normalization and mixed-extensive layers) whose dot products reproduce the corresponding kernel functions, with Nf = 32768 features at each feature layer and three Hadamard transforms. Regression coefficients were obtained from a singular value decomposition of the feature matrix. Hyperparameters (kernel widths, the regularizer lambda, and optional quantity shifts) were tuned by minimizing a self-consistent squared LogCosh loss on leave-one-out cross-validation errors, using L-BFGS and SLSQP from SciPy and the BOSS Bayesian optimization package for lambda. Trained models were then run over the held-out test molecules, and the whole workflow was rerun using only the lowest-energy conformer of each molecule as a comparison.
The shape of the work
Structural · the record, drawn
no AI
Obtain experimental property datasets
Obtaining raw data, whether by measurement, download or retrieval.
this work used the collection of experimental oxidation potentials of molecules in acetonitrile Eox compiled in Ref.where the paper describes this · verbatim
no AI
Generate Boltzmann conformer ensembles
Numerical or physics simulation, including where a learned surrogate replaces it.
We generated molecular conformers with the algorithm from Ref., as implemented in the Morfeus packagewhere the paper describes this · verbatim
no AI
Compute molecular representations
Encoding data into features, descriptors, embeddings or graphs.
Almost all representation functions used in this work were calculated using the QML2 codewhere the paper describes this · verbatim
AI
Fit MSORF regression coefficients
Fitting model parameters, including fine-tuning an existing model.
trained the model using these hyperparameterswhere the paper describes this · verbatim
AI
Optimize hyperparameters on leave-one-out errors
Iterative search over a space. Its result feeds back into an earlier step.
Our hyperparameter optimization procedure is based on optimizing leave-one-out cross-validation errorswhere the paper describes this · verbatim
AI
Predict properties of test molecules
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
model’s predictions for the test set moleculeswhere the paper describes this · verbatim
no AI
Evaluate errors and compare with single-conformer and literature results
Testing outputs against ground truth.
For each dataset and method the procedure was repeated 4 times, providing both the average MAEs and their dispersionswhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is the predictive performance of the proposed machine learning method itself; every reported quantity is a model prediction error
we divided a dataset into a training set (approximately 80% of the total number of points, making it 474 for LESwhere the paper describes this · verbatim
The version used in this work together with the ML scripts has been uploaded as the 0.1.6 releasewhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- DataWhether the data are available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of MSORF (Multilevel Structured Orthogonal Random Features)Which version of the model was used is not stated.
- Version of BOSS (Bayesian Optimization Structure Search)Which version of the model was used is not stated.
- What step 4 replacedThe paper gives no basis for what the AI stood in for.
- What step 5 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00205, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error