~/aixsci
200 records · all checked

materials-chemistry/ai produced the result/arXiv 2024 · v2

A compact way to describe molecules speeds up machine learning of their properties

Researchers built cMBDF, a fixed-size numerical description of each atom's surroundings, and used it to train kernel ridge regression models that predict quantum properties of small molecules from their geometry alone.

1. Assemble quantum chemistry datasets2. Encode atoms as cMBDF feature vectors3. Fit kernel ridge regression models4. Predict quantum properties5. Evaluate learning curves against baseline representations6. Correlate representation distances with energetic differences

spectrum · one line per step, placed by what the step does · bright lines used AI

Generalized convolutional many body distribution functional representations
arXiv, 2024

doi:10.48550/arxiv.2409.20471 · record aix-00114 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Property prediction
Model family
Gaussian process
Checked by
Benchmark1000 tested
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Chemists often want to know a molecule's properties — how much energy is released when its atoms come together, how it responds to an electric field, how it absorbs light. These can be worked out by solving the equations of quantum mechanics on a computer, but that is slow, and slower still for every molecule in a large library. One alternative is to learn the mapping from structure to property from examples. The obstacle is how to hand a molecule to a learning algorithm. A list of atomic positions is not enough on its own: rotate or shift the molecule, or relabel identical atoms, and the numbers change while the chemistry does not.

The usual answer is a 'representation': a recipe that turns each atom's local surroundings into a fixed list of numbers that ignores such irrelevant changes. Many such recipes exist, and they differ in how long they take to compute, how large they grow and how much training data a model needs on top of them. The researchers set out to define a new one, cMBDF, and to measure how models built on it compared with models built on established alternatives.

Where AI came in

The learning itself was done by kernel ridge regression, a standard statistical method that predicts a new molecule's property by comparing it with training molecules and weighting their known values by similarity. Models were fitted to molecules drawn from the public QM7b, QM9 and VQM24 collections, with the similarity scale and a smoothing term chosen by grid search. The fitted models were then run on molecules held out of training to predict atomization energies, dipole moments, HOMO-LUMO gaps, heat capacity, polarizability and two settings used in quantum chemistry calculations.

In that step the models stood in for the quantum chemistry simulation that would otherwise have produced each number. The representation at the heart of the paper is not itself learned; it is a hand-written formula, giving each atom a vector of 40 numbers, or 60 when four-body terms are included, whatever the molecule's size. Everything reported about it — the prediction errors at each training set size, and the timings — comes from training and running these models, and from comparing them with models using the ACSF, LMBTR, FCHL19, SOAP, Coulomb matrix, SLATM and MORDRED representations.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The paper introduces cMBDF, a compact hand-crafted atomic representation that encodes a local chemical environment as invariant functionals of a Gaussian electron-density proxy, evaluated as convolutions on pre-computed grids. Kernel ridge regression models built on these vectors were trained to predict atomization energies and other quantum properties for molecules from the QM7b, QM9 and VQM24 datasets, and compared with models using ACSF, LMBTR, FCHL19, SOAP, Coulomb matrix, SLATM and MORDRED representations. The atomic feature vector stays fixed at 40 dimensions (60 with four-body terms) regardless of system size or composition, and the reported models showed lower prediction error than the compared representations in the low-training-data regime. On QM9 the model reached 1 kcal/mol mean absolute error after training on 4,000 molecules; generating the VQM24 learning curves took 23 hours with FCHL19 against 8 minutes with cMBDF.

How AI was used

Molecular geometries from public quantum chemistry datasets were converted into atomic feature vectors by the cMBDF scheme, in which translationally and rotationally invariant two-, three- and pseudo four-body functionals of a smooth atom-centred density are weighted by interaction potentials and evaluated by indexing pre-computed convolution grids obtained via fast Fourier transforms; comparison representations were generated with the QMLcode and Dscribe libraries. Kernel ridge regression models were then fitted to the training labels using a screened atomic Gaussian kernel, or a global kernel over bagged molecular vectors for intensive properties, with the kernel length-scale and the regulariser chosen by grid search and the weighting-function hyper-parameter set by a grid search on 1k random QM7 molecules. The fitted models were run over out-of-sample molecules to predict atomization energies, dipole moments, HOMO-LUMO gaps, heat capacity, polarizability, optimal exact-exchange mixing fractions and basis-set scaling factors, and errors were recorded as a function of training set size alongside training and prediction timings. Representation matrix distances were also computed and correlated with energetic differences.

The shape of the work

Structural · the record, drawn

ACQUISITIONREPRESENTATIONTRAININGINFERENCEVALIDATIONINTERPRETATION123456AIAIAssemble quantumchemistrydatasetsEncode atoms ascMBDF featurevectorsFit kernel ridgeregression modelsPredict quantumpropertiesEvaluate learningcurves againstbaseline represe…Correlaterepresentationdistances with e…↤ simulation
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Assemble quantum chemistry datasets

Obtaining raw data, whether by measurement, download or retrieval.

Applicability for organic and inorganic chemistry is tested as represented by the QM7b, QM9 and VQM24 data sets.where the paper describes this · verbatim
in the paper
2Representation
no AI

Encode atoms as cMBDF feature vectors

Encoding data into features, descriptors, embeddings or graphs.

For generating the representation vector of an atom within a molecule, only the molecular internal coordinates are required to be calculated.where the paper describes this · verbatim
in the paper
3Training
AI

Fit kernel ridge regression models

Fitting model parameters, including fine-tuning an existing model.

The regression weights α are obtained from the set of training labels 𝐲train via the following equationwhere the paper describes this · verbatim
in the paper
4Inference
AI

Predict quantum properties

Running a trained model over new data to predict, classify or score. The AI stood in for simulation.

Mean absolute errors (MAE) for prediction of atomization energies using a kernel ridge regression (KRR) model are shown from the QM7b dataset.where the paper describes this · verbatim
in the paper
5Validation
no AI

Evaluate learning curves against baseline representations

Testing outputs against ground truth.

Figure 3 shows learning curves (ML model prediction error as a function of training set size) of atomization energieswhere the paper describes this · verbatim
in the paper
6Interpretation
no AI

Correlate representation distances with energetic differences

Extracting understanding from model behaviour.

Figure 6A shows a correlation plot between atomization energy and representation matrix distances between all pairs of molecules from the QM7b dataset.where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's results are the predictive accuracy, data efficiency and timing of kernel machine-learning models built on the introduced representation; the reported findings exist only because the models were trained and run.

~What the AI was for
the ML model used throughout this work with all representations including cMBDF is the established Kernel Ridge Regressionwhere the paper describes this · verbatim
~Model families
Gaussian processour reading
~How it was taught
Supervisedour reading
~Models named
Kernel ridge regression with cMBDF representation (3-body, m=4, n=2) · Trained from scratchKernel ridge regression with cMBDF (4-body) representation · Trained from scratchKernel ridge regression with FCHL19 representation · Trained from scratchKernel ridge regression with SOAP representation · Trained from scratchKernel ridge regression with ACSF representation · Trained from scratchKernel ridge regression with LMBTR representation · Trained from scratchKernel ridge regression with Coulomb matrix representation · Trained from scratchKernel ridge regression with SLATM representation · Trained from scratchKernel ridge regression with MORDRED representation · Trained from scratchour reading
+How results were checked
Benchmark1000 testedin the paper
cMBDF reaches chemical accuracy on the entire QM9 dataset after training on only 4,000 moleculeswhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata not reportedin the paper
Python implementation for generating cMBDF representations along with gradients is openly available at https://github.com/dkhan42/cMBDFwhere the paper describes this · verbatim
+Compute
36 core 4.8GHz Intel Xeon W9-3475X compute node with 1 TB DDR5 ECC RAM; generation of the VQM24 learning curves took 23 hours with FCHL19 versus 8 minutes with cMBDF; 1.3 minutes of compute for training and prediction on the QM9 taskin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 12 items
  • Trained model weightsWhether the trained model is available is not stated.
  • DataWhether the data are available is not stated.
  • Version of Kernel ridge regression with cMBDF representation (3-body, m=4, n=2)Which version of the model was used is not stated.
  • Version of Kernel ridge regression with cMBDF (4-body) representationWhich version of the model was used is not stated.
  • Version of Kernel ridge regression with FCHL19 representationWhich version of the model was used is not stated.
  • Version of Kernel ridge regression with SOAP representationWhich version of the model was used is not stated.
  • Version of Kernel ridge regression with ACSF representationWhich version of the model was used is not stated.
  • Version of Kernel ridge regression with LMBTR representationWhich version of the model was used is not stated.
  • Version of Kernel ridge regression with Coulomb matrix representationWhich version of the model was used is not stated.
  • Version of Kernel ridge regression with SLATM representationWhich version of the model was used is not stated.
  • Version of Kernel ridge regression with MORDRED representationWhich version of the model was used is not stated.
  • What step 3 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00114, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error