~/aixsci
200 records · all checked

materials-chemistry/ai produced the result/Nature Communications 2024 · v2

Machine-learned force model simulates silicon-oxygen structures from glass surfaces to monoxide grains

Researchers fitted a machine-learning model of the forces between silicon and oxygen atoms, trained on quantum-mechanical calculations chosen by the model's own uncertainty, then used it to simulate silica under pressure, porous structures and amorphous silicon monoxide.

1. Seed database from existing silicon and silica datasets2. Fit moment tensor potentials to current database3. Sample new environments by uncertainty-driven MD and amorphous matrix embedding4. Label selected structures with SCAN DFT and extend database5. Fit final non-linear ACE potential6. Run large-scale ACE molecular dynamics for silica, surfaces, aerogels and SiO7. Evaluate potential against held-out DFT data, DFT single points and literature experiments

spectrum · one line per step, placed by what the step does · bright lines used AI

Modelling atomic and nanoscale structure in the silicon–oxygen system through active machine learning
Nature Communications, 2024

doi:10.1038/s41467-024-45840-9 · record aix-00012 v2 · checked 2026-10-07

ai-resultrole of AI
AI was for
Simulation surrogate, Experimental design
Model family
Linear model, Gaussian process
Checked by
Held-out
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Silicon and oxygen combine into a wide family of materials: quartz crystals, window glass, the thin oxide layers in microchips, and odd in-between compounds such as silicon monoxide, where there is one oxygen for every silicon. Understanding any of them means knowing how the atoms sit relative to one another, and that in turn means knowing the forces between them. Those forces can be computed from quantum mechanics, by a method called density-functional theory, but the cost grows steeply with the number of atoms. In practice that limits such calculations to a few hundred atoms over a few picoseconds, while the features researchers care about, such as pores or separate grains, are far larger.

The alternative is a cheap approximate formula for the forces, known as an interatomic potential, which can push millions of atoms around for far longer. The difficulty is making one accurate across the whole silicon-oxygen range at once, rather than tuned to quartz alone or pure silicon alone. The authors set out to build a single model covering the full binary system, and then to use it on high-pressure forms of silica, on quartz and glass surfaces, on porous aerogel-like structures, and on models of amorphous silicon monoxide.

Where AI came in

The force model itself is the machine learning. It was trained, in the ordinary supervised sense, to reproduce energies and forces that density-functional theory had computed for a library of atomic arrangements, standing in for those quantum calculations wherever a simulation needed them. The final model, in a framework called atomic cluster expansion, was fitted to a database of 11,428 structures and reached errors of 16.7 millielectronvolts per atom on energies for arrangements held back from training.

Learning also chose what to compute. Small committees of simpler models were run against one another: where they disagreed about the force on an atom, that atom's surroundings were flagged as unfamiliar, cut out, embedded in a melted matrix and sent for quantum labelling, and the cycle repeated. That replaced a person deciding by hand which structures the model still needed to see. Every structure, phase diagram and grain-size figure reported came from simulations driven by the fitted model.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors built a machine-learning interatomic potential for the full binary silicon–oxygen system, using an active-learning loop in which high-uncertainty atomic environments found in large-scale simulations were cut out, embedded in an amorphous matrix, and labelled with SCAN density-functional theory. The final database contains 11,428 structures with a total of about 1.3 million atoms, and the fitted atomic cluster expansion model reaches test-set errors of 16.7 meV per atom for energies and 306 meV per ångström for forces. Simulations with the potential covered high-pressure silica polymorphs and compressed amorphous silica, quartz and amorphous silica surfaces, porous structures, and melt-quenched models of amorphous silicon monoxide, which segregated into silicon-rich and silica-like regions with estimated average grain diameters between 24 and 54 ångström, compared with 30–40 ångström from published transmission electron microscopy. Heating the silicon monoxide models to 1400 K and quenching to 1200 K produced crystallisation in the silicon-rich regions.

How AI was used

Two families of machine-learned interatomic potentials were fitted to density-functional-theory energies and forces. During database construction, moment tensor potentials implemented in the MLIP package were fitted to the current database and then used to drive molecular dynamics in three separate tracks (high-pressure silica, silica surfaces, non-stoichiometric SiOx); structures exceeding an extrapolation threshold were selected, labelled with SCAN DFT in VASP, and added to the database, and the cycle was repeated until the threshold was no longer exceeded. In a further stage, committees of two to four moment tensor potentials trained on the same database gave per-atom committee errors in large-scale simulations; environments of high-uncertainty atoms were cut into DFT-sized cubes, the atoms inside the potential cut-off were held fixed, and the surrounding region was melted in a machine-learning molecular-dynamics run to form an amorphous matrix before DFT labelling. The merged database was then used to fit the final potential in the nonlinear atomic cluster expansion framework with PACEMAKER, using 600 basis functions, 5700 parameters, Bessel radial functions, a force-to-energy weight ratio of 0.01 and 2000 BFGS steps; linear and Finnis–Sinclair-like variants were fitted to the same data for comparison. The final potential then drove large-scale molecular dynamics and static calculations in LAMMPS, with ASE and the OVITO Python interface used for protocols and analysis, for melt-quenching, compression, aerogel densification, surface relaxation, thermodynamic integration in calphy and thermal treatment of the silicon monoxide models.

The shape of the work

Structural · the record, drawn

PREPARATIONTRAININGSIMULATIONACQUISITIONTRAININGSIMULATIONVALIDATION1234567AIAIAIAISeed databasefrom existingsilicon and sili…Fit moment tensorpotentials tocurrent databaseSample newenvironments byuncertainty-driv…Label selectedstructures withSCAN DFT and ext…Fit finalnon-linear ACEpotentialRun large-scaleACE moleculardynamics for sil…Evaluatepotential againstheld-out DFT dat…↤ simulation↤ manual curation↤ simulation↤ simulationloops back
AI stepNo AI↤ what the AI stood in for
1Preparation
no AI

Seed database from existing silicon and silica datasets

Cleaning, filtering, normalising or labelling data already obtained.

We initialised the protocol with two existing datasets for siliconwhere the paper describes this · verbatim
in the paper
2Training
AI

Fit moment tensor potentials to current database

Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.

we fitted moment tensor potential (MTP) models to the databasewhere the paper describes this · verbatim
in the paper
3Simulation
AI

Sample new environments by uncertainty-driven MD and amorphous matrix embedding

Numerical or physics simulation, including where a learned surrogate replaces it. The AI stood in for manual curation.

We used 2–4 MTPs trained on the same database to estimate a per-atom committee errorwhere the paper describes this · verbatim
in the paper
4Acquisition
no AI

Label selected structures with SCAN DFT and extend database

Obtaining raw data, whether by measurement, download or retrieval. Its result feeds back into an earlier step.

Energies and forces for new structures were computed with DFT and added to the database.where the paper describes this · verbatim
in the paper
5Training
AI

Fit final non-linear ACE potential

Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.

The final potential is a complex non-linear ACE model, obtained by summation of one linear and seven non-linear ACE termswhere the paper describes this · verbatim
in the paper
6Simulation
AI

Run large-scale ACE molecular dynamics for silica, surfaces, aerogels and SiO

Numerical or physics simulation, including where a learned surrogate replaces it. The AI stood in for simulation.

In contrast, we created our models by melt–quench simulations.where the paper describes this · verbatim
in the paper
7Validation
no AI

Evaluate potential against held-out DFT data, DFT single points and literature experiments

Testing outputs against ground truth.

For validation, we held out 5% of these structures from training, selected at random.where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

All reported structures, phase diagrams, coordination numbers and SiO nanostructure models are produced by molecular dynamics driven by the fitted machine-learning potential; the findings do not exist without it.

~What the AI was for
For the final potential fit, we used the nonlinear Atomic Cluster Expansion (ACE) as implemented in PACEMAKER.where the paper describes this · verbatim
~Model families
~How it was taught
SupervisedActive learningour reading
~Models named
Si–O ACE, complex non-linear embedding (8 terms) · Trained from scratchSi–O ACE, linear embedding · Trained from scratchSi–O ACE, Finnis–Sinclair-like embedding · Trained from scratchMoment tensor potentials (MLIP) · Trained from scratchSiO2-GAP-22 · Off the shelfour reading
+How results were checked
Held-outin the paper
The resulting potential has a test-set root mean square error (RMSE) of 16.7 meV atom−1 for energieswhere the paper describes this · verbatim
+Code · weights · data
code not reportedweights not reporteddata availablein the paper
Source data are provided as a Source Data file.where the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 9 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of Si–O ACE, complex non-linear embedding (8 terms)Which version of the model was used is not stated.
  • Version of Si–O ACE, linear embeddingWhich version of the model was used is not stated.
  • Version of Si–O ACE, Finnis–Sinclair-like embeddingWhich version of the model was used is not stated.
  • Version of Moment tensor potentials (MLIP)Which version of the model was used is not stated.
  • Version of SiO2-GAP-22Which version of the model was used is not stated.

About this article

Record aix-00012, version 2, checked by a person on 2026-10-07. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY-4.0; quotations are at most 25 words. How we work · Report an error