materials-chemistry/ai produced the result/Journal of Chemical Theory and Computation 2023 · v2
Machine-learned potential makes long simulations of CO hydrogenation on rhodium affordable
Researchers wanted the free energy barrier for a single step in CO hydrogenation on a rhodium surface. A machine-learned model of the atomic forces stood in for quantum chemistry, making the 16 nanoseconds of biased molecular dynamics tractable.
spectrum · one line per step, placed by what the step does · bright lines used AI
Estimating Free Energy Barriers for Heterogeneous Catalytic Reactions with Machine Learning Potentials and Umbrella Integration
Journal of Chemical Theory and Computation, 2023
doi:10.1021/acs.jctc.3c00541 · record aix-00130 v2 · checked 2026-10-09
- AI was for
- Simulation surrogate
- Model family
- Gaussian process
- Checked by
- Held-out
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about

Many industrial reactions happen on the surface of a solid metal, which grips the reacting molecules and helps them rearrange. To predict how fast such a step goes, chemists need the free energy barrier: the height of the hill the atoms must climb, averaged over all the jiggling that real atoms do at working temperatures. The usual way to get the forces between atoms is density functional theory, or DFT, a quantum calculation. It is accurate but slow. Running it for the millions of steps that proper statistical sampling needs is, as the paper puts it, prohibitive at the first-principles level.
The system here is carbon monoxide gaining a hydrogen atom on a rhodium surface, forming a fragment called CHO, and the reverse step in which CHO falls apart again. The team set out to measure the free energy barrier for that decomposition by umbrella integration, a method that nudges the system through a series of overlapping windows along the reaction path and stitches the pieces together, and then to turn the barrier into a reaction rate.
Where AI came in
The AI was a Gaussian Approximation Potential: a model fitted to DFT energies and forces that then predicts them itself, far faster, for configurations it has not seen. It is a stand-in for the quantum calculation rather than for the chemist. Training was a loop. The current model ran the biased dynamics, a selection step picked out the most distinct new structures, those were labelled with single DFT calculations and fed back in. Energies and forces were reported converged after 17 rounds. Two models were fitted, one per choice of DFT settings.
The converged models then drove the production runs, 100 picoseconds per window at 523 kelvin, and unbiased runs used to check how often trajectories turned back at the barrier. A second, smaller piece of machine learning, Gaussian process regression, was fitted to the free energy slopes from each window and integrated to give the free energy curve along the reaction path together with an uncertainty band. The rate expressions built on that curve were ordinary analytic formulae. The models were checked against DFT on fresh configurations each round, and the resulting barriers compared with simpler harmonic estimates.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record

Gaussian Approximation Potentials were trained on DFT energies and forces for the hydrogenation of CO to CHO on Rh(111) and used in place of DFT to run the biased molecular dynamics required for umbrella integration. Training was iterative: biased dynamics with the current potential produced new configurations, which were selected by farthest point sampling, labelled with single-point DFT and added to the training set, with energies and forces reported converged after 17 iterations. Production umbrella integration used 100 ps trajectories per window at 523 K, amounting to 16 ns of dynamics overall, with separate potentials trained on revPBE+vdWsurf and BEEF-vdW reference data. The free energy barrier for CHO decomposition nearly vanishes at the revPBE+vdWsurf level and a barrier of 0.13 eV remains at the BEEF-vdW level, lower than the harmonic approximation estimates, while unbiased trajectories seeded near the barrier gave transmission coefficients of 0.87 and 0.81.
How AI was used
A Gaussian Approximation Potential, combining a two-body term with a SOAP many-body representation, was fitted to DFT energies and forces for the Rh(111) surface with CHO and CO+H adsorbates, serving as a surrogate for the DFT potential energy surface so that extensive biased sampling became tractable. The initial training set comprised CI-NEB path images, light-element dimer curves and optimised and rattled pristine surface configurations. Training then proceeded in an automated loop: the current potential ran biased molecular dynamics in randomly chosen umbrella windows, diverse new structures were extracted by farthest point sampling on kernel distances, evaluated with single-point DFT, used first as a validation set and then appended to the training set for the next fit; from iteration twelve onward the potential also ran NEB calculations so its minimum energy path could be compared with the DFT one. A one-dimensional collective variable was constructed as a linear combination of the C-H distance and the CO tilt angle along the DFT NEB path, and harmonic biasing potentials were placed along it. The converged potentials drove the production biased dynamics with a Langevin thermostat; free energy gradients from the window averages were integrated with Gaussian process regression to give the free energy surface and its uncertainty, from which transition state theory rate constants were computed alongside harmonic and Pitzer-Gwinn estimates built on DFT vibrational frequencies. The same potentials ran unbiased trajectories seeded in the initial state and near the barrier to estimate half-lives and transmission coefficients. A second potential was fitted to the same final configurations recalculated with BEEF-vdW, without further training iterations.
The shape of the work
Structural · the record, drawn
no AI
Generate initial DFT reference data
Obtaining raw data, whether by measurement, download or retrieval.
we define an initial training set of 50 configurations, consisting of the images from the DFT based NEB calculation shown in Figure 2where the paper describes this · verbatim
no AI
Define one-dimensional collective variable
Encoding data into features, descriptors, embeddings or graphs.
we define the CV ξ used in the following as a linear combination of d and θwhere the paper describes this · verbatim
AI
Train GAP interatomic potentials
Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.
two separate GAP models were trained with different reference datawhere the paper describes this · verbatim
AI
Iteratively generate and label new configurations
Numerical or physics simulation, including where a learned surrogate replaces it. The AI stood in for simulation. Its result feeds back into an earlier step.
We find that energies and forces are well converged in 17 iterationswhere the paper describes this · verbatim
no AI
Validate potential against DFT on unseen configurations
Testing outputs against ground truth.
the validation errors are for unseen configurations from the exact kind of simulation that we intend to run with this modelwhere the paper describes this · verbatim
AI
Run production biased molecular dynamics
Numerical or physics simulation, including where a learned surrogate replaces it. The AI stood in for simulation.
The converged potentials were used to run extensive US calculations with 100 ps trajectories per window, at 523 Kwhere the paper describes this · verbatim
AI
Integrate gradients to free energy surface and rate constants
Extracting understanding from model behaviour. The AI stood in for conventional algorithm.
Stecher et al. showed that this can be done in an uncertainty aware fashion using Gaussian Process Regression (GPR)where the paper describes this · verbatim
AI
Test TST assumptions with unbiased trajectories
Testing outputs against ground truth. The AI stood in for simulation.
240 MD trajectories were started by using 120 random velocities drawn from the Maxwell–Boltzmann distributionwhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The reported free energy barriers and rate constants come from umbrella integration run on machine-learned interatomic potentials; the paper states such sampling is computationally prohibitive at the first-principles level, so the central result depends on the learned surrogate.
Machine-learning potentials can provide fast and accurate surrogate models of the DFT PESwhere the paper describes this · verbatim
we can compare the rate constants obtained from the HA with those obtained via the UI free energy barrierswhere the paper describes this · verbatim
All hyperparameters as well as the potential itself are provided as Supporting Information in this article.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- DataWhether the data are available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of Gaussian Approximation Potential (revPBE+vdWsurf reference data)Which version of the model was used is not stated.
- Version of Gaussian Approximation Potential (BEEF-vdW reference data)Which version of the model was used is not stated.
- Version of Gaussian process regression model for uncertainty-aware umbrella integrationWhich version of the model was used is not stated.
About this article
Record aix-00130, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error