~/aixsci
200 records · all checked

materials-chemistry/ai produced the result/arXiv 2025 · v2

Machine learning fitted to flexible molecules predicts oxidation potentials and hydration energies

Researchers rebuilt a standard statistical fitting method so it could handle molecules that bend into many shapes, then trained it on measured oxidation potentials and hydration energies. The model did all the property prediction.

1. Obtain experimental property datasets2. Generate Boltzmann conformer ensembles3. Compute molecular representations4. Fit MSORF regression coefficients5. Optimize hyperparameters on leave-one-out errors6. Predict properties of test molecules7. Evaluate errors and compare with single-conformer and literature results

spectrum · one line per step, placed by what the step does · bright lines used AI

Kernel Ridge Regression for conformer ensembles made easy with Structured Orthogonal Random Features
arXiv, 2025

doi:10.48550/arxiv.2505.21247 · record aix-00205 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Property prediction
Model family
Multilayer perceptron
Checked by
Held-out
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Molecules are not rigid. A single compound can twist and rotate about its bonds into many different shapes, called conformers, and at room temperature it flickers between them. Each shape is present in a certain proportion, set by its energy. That makes life awkward for anyone trying to predict a molecule's behaviour from its structure, because there is no single structure to point at. Two properties of practical interest are the oxidation potential, the voltage at which a molecule gives up an electron in a solvent, and the hydration energy, how readily it dissolves in water. Both are measured in the laboratory, one compound at a time.

The researchers set out to predict these two properties from structure alone, while treating each molecule as a weighted collection of its conformers rather than a single frozen shape. They also wanted the method to stay cheap to run, so they rewrote an established fitting technique, kernel ridge regression, in a form built from layers of random features. They then tested it on published experimental measurements of oxidation potentials in the solvent acetonitrile and on hydration energies from a public collection.

Where AI came in

The learned model is the whole result here. Conformers for each molecule were generated with conventional chemistry software, and each one was turned into a list of numbers describing its structure. Those numbers were fed through the researchers' layered model, which was then fitted to the experimental values of the training molecules. Settings such as the widths of the model's internal functions were tuned automatically, partly with a Bayesian optimisation package, by checking how well the model did when each training molecule was held back in turn.

The fitted model was then asked to predict the properties of molecules it had never seen, standing in for measuring those compounds in the laboratory. Prediction error fell as the training set grew for every structural description tried. For oxidation potentials, several of those descriptions gave errors close, within the statistical scatter, to a figure of 0.2 eV reported previously. Using only each molecule's lowest-energy conformer instead of the full ensemble shifted the errors by about as much as that scatter.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors rewrite Kernel Ridge Regression over Boltzmann ensembles of molecular conformers in terms of Structured Orthogonal Random Features, giving a layered model (Multilevel SORF) with trigonometric activations and layers for ensembles, conformers, atoms and localized orbitals. Conformers were generated with the MMFF94 forcefield, encoded with several existing molecular representations, and models were fitted to two experimental datasets: oxidation potentials in acetonitrile (largest training set 474 molecules) and FreeSolv hydration energies (largest training set 514 molecules). Mean absolute error decreased with training set size for every representation tested, and for oxidation potentials the combinations with aSLATM, SLATM, FCHL19 and cMBDF gave errors similar, within statistical error, to the 0.2 eV reported earlier for SLATM with KRR. Replacing each ensemble with only the lowest-MMFF94-energy conformer changed the errors by an amount comparable to the statistical error. The work also introduces self-consistent versions of the Huber and LogCosh loss functions used during hyperparameter optimization.

How AI was used

Supervised regression models were fitted in this study to predict two experimental molecular properties. For each molecule, conformers were generated with the Morfeus implementation of a conformer-generation algorithm using the MMFF94 forcefield in RDKit at 298.15 K, with low-weight conformers cut off at rho_cut = 0.05; Boltzmann weights were retained as conformer importance weights. Each conformer was encoded with one of several molecular representations (Coulomb Matrix, SLATM, aSLATM, FCHL19, cMBDF, SOAP, and FJK, the latter from Huckel calculations in PySCF with a 6-311G basis, Boys localized orbitals and SMD solvation). These representations were fed through stacks of SORF-based layers (feature, sum, switch, weighted-sum, normalization and mixed-extensive layers) whose dot products reproduce the corresponding kernel functions, with Nf = 32768 features at each feature layer and three Hadamard transforms. Regression coefficients were obtained from a singular value decomposition of the feature matrix. Hyperparameters (kernel widths, the regularizer lambda, and optional quantity shifts) were tuned by minimizing a self-consistent squared LogCosh loss on leave-one-out cross-validation errors, using L-BFGS and SLSQP from SciPy and the BOSS Bayesian optimization package for lambda. Trained models were then run over the held-out test molecules, and the whole workflow was rerun using only the lowest-energy conformer of each molecule as a comparison.

The shape of the work

Structural · the record, drawn

ACQUISITIONSIMULATIONREPRESENTATIONTRAININGOPTIMISATIONINFERENCEVALIDATION1234567AIAIAIObtainexperimentalproperty datasetsGenerateBoltzmannconformer ensemb…Compute molecularrepresentationsFit MSORFregressioncoefficientsOptimizehyperparameterson leave-one-out…Predictproperties oftest moleculesEvaluate errorsand compare withsingle-conformer…↤ physical experimentloops back
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Obtain experimental property datasets

Obtaining raw data, whether by measurement, download or retrieval.

this work used the collection of experimental oxidation potentials of molecules in acetonitrile Eox compiled in Ref.where the paper describes this · verbatim
in the paper
2Simulation
no AI

Generate Boltzmann conformer ensembles

Numerical or physics simulation, including where a learned surrogate replaces it.

We generated molecular conformers with the algorithm from Ref., as implemented in the Morfeus packagewhere the paper describes this · verbatim
in the paper
3Representation
no AI

Compute molecular representations

Encoding data into features, descriptors, embeddings or graphs.

Almost all representation functions used in this work were calculated using the QML2 codewhere the paper describes this · verbatim
in the paper
4Training
AI

Fit MSORF regression coefficients

Fitting model parameters, including fine-tuning an existing model.

trained the model using these hyperparameterswhere the paper describes this · verbatim
in the paper
5Optimisation
AI

Optimize hyperparameters on leave-one-out errors

Iterative search over a space. Its result feeds back into an earlier step.

Our hyperparameter optimization procedure is based on optimizing leave-one-out cross-validation errorswhere the paper describes this · verbatim
in the paper
6Inference
AI

Predict properties of test molecules

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

model’s predictions for the test set moleculeswhere the paper describes this · verbatim
in the paper
7Validation
no AI

Evaluate errors and compare with single-conformer and literature results

Testing outputs against ground truth.

For each dataset and method the procedure was repeated 4 times, providing both the average MAEs and their dispersionswhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's result is the predictive performance of the proposed machine learning method itself; every reported quantity is a model prediction error

~What the AI was for
~Model families
~How it was taught
Supervisedour reading
~Models named
MSORF (Multilevel Structured Orthogonal Random Features) · Trained from scratchBOSS (Bayesian Optimization Structure Search) · Trained from scratchour reading
+How results were checked
Held-outin the paper
we divided a dataset into a training set (approximately 80% of the total number of points, making it 474 for LESwhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata not reportedin the paper
The version used in this work together with the ML scripts has been uploaded as the 0.1.6 releasewhere the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 8 items
  • Trained model weightsWhether the trained model is available is not stated.
  • DataWhether the data are available is not stated.
  • ComputeThe hardware or time used is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of MSORF (Multilevel Structured Orthogonal Random Features)Which version of the model was used is not stated.
  • Version of BOSS (Bayesian Optimization Structure Search)Which version of the model was used is not stated.
  • What step 4 replacedThe paper gives no basis for what the AI stood in for.
  • What step 5 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00205, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error