~/aixsci
200 records · all checked

materials-chemistry/ai in a supporting role/The Journal of Physical Chemistry C 2023 · v2

Simulations map how surface atoms rearrange in alloys when molecules attach

Researchers used quantum-chemistry calculations to work out whether lone dopant atoms in metal surfaces stay put or sink inwards, bare and with three small molecules attached. A neural network was then trained on those results to predict the same energies quickly.

1. Compute segregation energies with DFT2. Assemble and standardise candidate features3. Select features by random forest importance4. Train and tune regression models5. Evaluate predictions against held-out DFT data6. Compare predictions with reported experimental systems

spectrum · one line per step, placed by what the step does · bright lines used AI

Single Atom Alloys Segregation in the Presence of Ligands
The Journal of Physical Chemistry C, 2023

doi:10.1021/acs.jpcc.3c05827 · record aix-00071 v2 · checked 2026-10-08

ai-supportingrole of AI
AI was for
Property prediction
Model family
Multilayer perceptron, Random forest, Support vector machine, Gradient-boosted trees, Linear model
Checked by
Held-out27 tested
Code
available

AI processed or interpreted data, but the main finding does not rest on it.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Structural diagrams showing bulk structure, dopant on a metal surface, and ligand adsorption configurations used for segregation calculations.
Structures used to calculate segregation energies, including dopant placement and ligand adsorption on metal surfaces.Figure 1 from Salem et al., The Journal of Physical Chemistry C 2023 · source · CC BY · resized

A single atom alloy is a metal surface in which lone atoms of one metal sit scattered within a host made of another. Those isolated atoms are what makes such materials interesting as catalysts, because a lone atom behaves differently from a patch of the same metal. The difficulty is that the lone atoms do not necessarily stay on the surface. Depending on the pair of metals involved, a dopant atom may prefer to sink into the bulk below, or the host may prefer to cover it over. The quantity that captures this preference is the segregation energy: the energy change when the dopant moves between the surface and the interior.

Real catalysts rarely sit in a vacuum. Their surfaces are usually coated with molecules, called ligands, that bind to the metal. The researchers calculated segregation energies for nickel, palladium or platinum dopants in silver, gold or copper hosts, on two different crystal faces, both bare and with methylamine, methylamide or methylthiolate attached, giving 240 systems. They then set out to build a faster way of estimating the same numbers.

Where AI came in

The calculations themselves used density functional theory, a standard quantum-mechanical method, and no machine learning. The learning came afterwards. The 180 systems with ligands supplied the answers, and the inputs were ordinary tabulated properties of the two metals, such as how tightly each binds in bulk and how large its atoms are, together with calculated ligand binding strengths. A random forest, an ensemble of decision trees, ranked which inputs mattered and cut them to four.

Five kinds of regression model were then trained and tuned on those four inputs; a small neural network was chosen. On the held-out 27 systems it predicted segregation energies with a mean absolute error of 0.107 electronvolts, and it matched 8 of 10 segregation behaviours reported in the experimental literature. The model stands in for the calculations, offering a quicker route to the same energies rather than supporting the paper's physical conclusions, which rest on the calculations.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

Structural diagrams showing bulk structure, dopant on a metal surface, and ligand adsorption configurations used for segregation calculations.
Structures used to calculate segregation energies, including dopant placement and ligand adsorption on metal surfaces.Figure 1 from Salem et al., The Journal of Physical Chemistry C 2023 · source · CC BY · resized

Density functional theory was used to compute segregation energies for single atom alloys combining Ni, Pd or Pt dopants with Ag, Au or Cu hosts on (111) and (100) surfaces, both bare and with methylamine, methylamide or methylthiolate adsorbed, giving 240 systems in total. The 180 ligated systems were used to fit a four-feature neural network multilayer perceptron regressor, whose features were the difference in bulk cohesive energy divided by dopant coordination number, the difference in ligand binding energy divided by the adsorbate coordination number, and the differences in Wigner-Seitz radius and electron affinity. On the held-out test set the model gave a mean absolute error of 0.107 eV and a root mean square error of 0.137 eV, and it matched 8 of 10 experimental observations compiled from the literature. The calculations indicate that the presence of ligands narrows the range of segregation energies relative to bare surfaces, and that the ligand adsorption configuration and binding strength shift the trends.

How AI was used

Segregation energies calculated with DFT for ligated single atom alloy slabs supplied the regression targets, and features were drawn from tabulated elemental properties obtained via the Mendeleev package together with DFT binding energies of each ligand on a single metal atom; all features were standardised to zero mean and unit variance. A random forest regression variable importance analysis, supported by a variance inflation factor check for multicollinearity, reduced the candidate features to four. An 85/15 train/test split with 5-fold cross-validation on the training portion was used, and hyperparameters for a neural network multilayer perceptron, kernel ridge regression, support vector regressor, random forest regressor and extreme gradient boosting regressor were tuned with GridSearchCV by minimising validation mean absolute error. The multilayer perceptron was selected and evaluated on the held-out split, with the procedure repeated over 100 random train/test splits to report mean and standard deviation of the errors; model predictions were then compared with segregation behaviour reported for ten experimental systems. Implementation used Scikit-Learn, and the DFT calculations used CP2K.

The shape of the work

Structural · the record, drawn

SIMULATIONREPRESENTATIONSCREENINGTRAININGVALIDATIONVALIDATION123456AIAIAIAIComputesegregationenergies with DFTAssemble andstandardisecandidate featur…Select featuresby random forestimportanceTrain and tuneregression modelsEvaluatepredictionsagainst held-out…Comparepredictions withreported experim…↤ simulation
AI stepNo AI↤ what the AI stood in for
1Simulation
no AI

Compute segregation energies with DFT

Numerical or physics simulation, including where a learned surrogate replaces it.

The four different cases (nonligated and 3 ligated systems) resulted in a total of 240 different systems studied in this work.where the paper describes this · verbatim
in the paper
2Representation
no AI

Assemble and standardise candidate features

Encoding data into features, descriptors, embeddings or graphs.

tabulated elemental properties of the host and dopant such as the covalent radius, electronegativity, electron affinity, and first ionization potentialwhere the paper describes this · verbatim
in the paper
3Screening
AI

Select features by random forest importance

Reducing a candidate set by filtering or ranking, in a single pass.

For feature selection, a variable importance plot based on the random forest regression was employedwhere the paper describes this · verbatim
in the paper
4Training
AI

Train and tune regression models

Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.

A 85/15% train/test split was chosen, and a 5-fold cross-validation was implemented using the training datawhere the paper describes this · verbatim
in the paper
5Validation
AI

Evaluate predictions against held-out DFT data

Testing outputs against ground truth.

The 15% test data was used in the final step to evaluate the accuracy of the model in predicting Esegwhere the paper describes this · verbatim
in the paper
6Validation
AI

Compare predictions with reported experimental systems

Testing outputs against ground truth.

we compare it against experimental observations (10 different experimental systems reported in Table S8)where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI in a supporting roleour reading

The paper's physical conclusions about segregation trends come from the DFT calculations; the trained regression model is built on those data and offered as a faster route to predicting segregation energy, so the findings do not rest on the model.

+What the AI was for
We applied a supervised machine learning approach to develop an accurate Eseg regression model.where the paper describes this · verbatim
+How it was taught
Supervisedin the paper
+Models named
Neural network multilayer perceptron regressor (NN MLP) · Trained from scratchKernel ridge regression (second-order polynomial kernel) · Trained from scratchSupport vector regressor (SVR) · Trained from scratchRandom forest regressor · Trained from scratchExtreme gradient boosting regressor (XGB) · Trained from scratchin the paper
+How results were checked
Held-out27 testedin the paper
27 data points are the test set, and 153 points are the training setwhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
the Python code utilized to develop the model are available free of charge on our GitHub repositorywhere the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 10 items
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • Version of Neural network multilayer perceptron regressor (NN MLP)Which version of the model was used is not stated.
  • Version of Kernel ridge regression (second-order polynomial kernel)Which version of the model was used is not stated.
  • Version of Support vector regressor (SVR)Which version of the model was used is not stated.
  • Version of Random forest regressorWhich version of the model was used is not stated.
  • Version of Extreme gradient boosting regressor (XGB)Which version of the model was used is not stated.
  • What step 3 replacedThe paper gives no basis for what the AI stood in for.
  • What step 5 replacedThe paper gives no basis for what the AI stood in for.
  • What step 6 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00071, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error