materials-chemistry/ai in a supporting role/The Journal of Physical Chemistry C 2023 · v2
Simulations map how surface atoms rearrange in alloys when molecules attach
Researchers used quantum-chemistry calculations to work out whether lone dopant atoms in metal surfaces stay put or sink inwards, bare and with three small molecules attached. A neural network was then trained on those results to predict the same energies quickly.
spectrum · one line per step, placed by what the step does · bright lines used AI
Single Atom Alloys Segregation in the Presence of Ligands
The Journal of Physical Chemistry C, 2023
doi:10.1021/acs.jpcc.3c05827 · record aix-00071 v2 · checked 2026-10-08
- AI was for
- Property prediction
- Model family
- Multilayer perceptron, Random forest, Support vector machine, Gradient-boosted trees, Linear model
- Checked by
- Held-out27 tested
- Code
- available
AI processed or interpreted data, but the main finding does not rest on it.
What this research was about

A single atom alloy is a metal surface in which lone atoms of one metal sit scattered within a host made of another. Those isolated atoms are what makes such materials interesting as catalysts, because a lone atom behaves differently from a patch of the same metal. The difficulty is that the lone atoms do not necessarily stay on the surface. Depending on the pair of metals involved, a dopant atom may prefer to sink into the bulk below, or the host may prefer to cover it over. The quantity that captures this preference is the segregation energy: the energy change when the dopant moves between the surface and the interior.
Real catalysts rarely sit in a vacuum. Their surfaces are usually coated with molecules, called ligands, that bind to the metal. The researchers calculated segregation energies for nickel, palladium or platinum dopants in silver, gold or copper hosts, on two different crystal faces, both bare and with methylamine, methylamide or methylthiolate attached, giving 240 systems. They then set out to build a faster way of estimating the same numbers.
Where AI came in
The calculations themselves used density functional theory, a standard quantum-mechanical method, and no machine learning. The learning came afterwards. The 180 systems with ligands supplied the answers, and the inputs were ordinary tabulated properties of the two metals, such as how tightly each binds in bulk and how large its atoms are, together with calculated ligand binding strengths. A random forest, an ensemble of decision trees, ranked which inputs mattered and cut them to four.
Five kinds of regression model were then trained and tuned on those four inputs; a small neural network was chosen. On the held-out 27 systems it predicted segregation energies with a mean absolute error of 0.107 electronvolts, and it matched 8 of 10 segregation behaviours reported in the experimental literature. The model stands in for the calculations, offering a quicker route to the same energies rather than supporting the paper's physical conclusions, which rest on the calculations.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record

Density functional theory was used to compute segregation energies for single atom alloys combining Ni, Pd or Pt dopants with Ag, Au or Cu hosts on (111) and (100) surfaces, both bare and with methylamine, methylamide or methylthiolate adsorbed, giving 240 systems in total. The 180 ligated systems were used to fit a four-feature neural network multilayer perceptron regressor, whose features were the difference in bulk cohesive energy divided by dopant coordination number, the difference in ligand binding energy divided by the adsorbate coordination number, and the differences in Wigner-Seitz radius and electron affinity. On the held-out test set the model gave a mean absolute error of 0.107 eV and a root mean square error of 0.137 eV, and it matched 8 of 10 experimental observations compiled from the literature. The calculations indicate that the presence of ligands narrows the range of segregation energies relative to bare surfaces, and that the ligand adsorption configuration and binding strength shift the trends.
How AI was used
Segregation energies calculated with DFT for ligated single atom alloy slabs supplied the regression targets, and features were drawn from tabulated elemental properties obtained via the Mendeleev package together with DFT binding energies of each ligand on a single metal atom; all features were standardised to zero mean and unit variance. A random forest regression variable importance analysis, supported by a variance inflation factor check for multicollinearity, reduced the candidate features to four. An 85/15 train/test split with 5-fold cross-validation on the training portion was used, and hyperparameters for a neural network multilayer perceptron, kernel ridge regression, support vector regressor, random forest regressor and extreme gradient boosting regressor were tuned with GridSearchCV by minimising validation mean absolute error. The multilayer perceptron was selected and evaluated on the held-out split, with the procedure repeated over 100 random train/test splits to report mean and standard deviation of the errors; model predictions were then compared with segregation behaviour reported for ten experimental systems. Implementation used Scikit-Learn, and the DFT calculations used CP2K.
The shape of the work
Structural · the record, drawn
no AI
Compute segregation energies with DFT
Numerical or physics simulation, including where a learned surrogate replaces it.
The four different cases (nonligated and 3 ligated systems) resulted in a total of 240 different systems studied in this work.where the paper describes this · verbatim
no AI
Assemble and standardise candidate features
Encoding data into features, descriptors, embeddings or graphs.
tabulated elemental properties of the host and dopant such as the covalent radius, electronegativity, electron affinity, and first ionization potentialwhere the paper describes this · verbatim
AI
Select features by random forest importance
Reducing a candidate set by filtering or ranking, in a single pass.
For feature selection, a variable importance plot based on the random forest regression was employedwhere the paper describes this · verbatim
AI
Train and tune regression models
Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.
A 85/15% train/test split was chosen, and a 5-fold cross-validation was implemented using the training datawhere the paper describes this · verbatim
AI
Evaluate predictions against held-out DFT data
Testing outputs against ground truth.
The 15% test data was used in the final step to evaluate the accuracy of the model in predicting Esegwhere the paper describes this · verbatim
AI
Compare predictions with reported experimental systems
Testing outputs against ground truth.
we compare it against experimental observations (10 different experimental systems reported in Table S8)where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's physical conclusions about segregation trends come from the DFT calculations; the trained regression model is built on those data and offered as a faster route to predicting segregation energy, so the findings do not rest on the model.
We applied a supervised machine learning approach to develop an accurate Eseg regression model.where the paper describes this · verbatim
27 data points are the test set, and 153 points are the training setwhere the paper describes this · verbatim
the Python code utilized to develop the model are available free of charge on our GitHub repositorywhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- Version of Neural network multilayer perceptron regressor (NN MLP)Which version of the model was used is not stated.
- Version of Kernel ridge regression (second-order polynomial kernel)Which version of the model was used is not stated.
- Version of Support vector regressor (SVR)Which version of the model was used is not stated.
- Version of Random forest regressorWhich version of the model was used is not stated.
- Version of Extreme gradient boosting regressor (XGB)Which version of the model was used is not stated.
- What step 3 replacedThe paper gives no basis for what the AI stood in for.
- What step 5 replacedThe paper gives no basis for what the AI stood in for.
- What step 6 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00071, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error