materials-chemistry/ai produced the result/ACS Applied Materials & Interfaces 2024 · v2
Neural networks predict how zeolites take up carbon dioxide from structure alone
Researchers trained graph neural networks to predict two carbon dioxide uptake properties of aluminium-substituted zeolites straight from their atomic structure, replacing slow Monte Carlo simulations and then steering a search for structures hitting chosen uptake targets.
spectrum · one line per step, placed by what the step does · bright lines used AI
Graph Neural Networks for Carbon Dioxide Adsorption Prediction in Aluminum-Substituted Zeolites
ACS Applied Materials & Interfaces, 2024
doi:10.1021/acsami.4c12198 · record aix-00160 v2 · checked 2026-10-09
- AI was for
- Property prediction, Simulation surrogate, Candidate generation
- Model family
- Graph neural network, Multilayer perceptron
- Checked by
- Held-out
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Zeolites are crystalline minerals, usually built from silicon and oxygen, riddled with pores and channels of a fixed size and shape. Swapping some silicon atoms for aluminium changes the chemistry of those pores, and extra positively charged ions such as sodium are added to balance the charge. The result is a material that can hold gas molecules, which is why zeolites are studied for capturing carbon dioxide. Two numbers describe how strongly a gas sticks: the heat of adsorption, the energy released when a molecule settles into a pore, and the Henry coefficient, which measures uptake when the gas is very dilute. Both normally come from simulations that track molecules inserted into the pores at random, and those take hours per structure.
The difficulty is that the aluminium atoms can sit in a vast number of different arrangements within the same framework, and each arrangement behaves slightly differently. The authors set out to learn that relationship instead of simulating it each time. They built a dataset of aluminium-substituted structures for four zeolite frameworks, known as MOR, MFI, RHO and ITW, labelled each with simulated heat of adsorption and Henry coefficient values, and then trained models to predict those labels from structure.
Where AI came in
Each structure was turned into a graph: the silicon and aluminium sites became points, marked one for aluminium and zero for silicon, linked where they share an oxygen, with extra points standing for the pores themselves and carrying the pore's area and ring size. A graph neural network passes information along those links so each site learns about its surroundings. One model was trained per framework, on ninety per cent of the data, and tested on the remaining tenth. Predictions took milliseconds against hours for a Monte Carlo simulation of a single zeolite.
The authors also altered the architecture so that every pore outputs its own number, and these are added up to give the overall prediction, letting the pore-by-pore contributions be read off directly. For four MOR structures those per-pore values tracked how often a carbon dioxide molecule actually sat in each pore in simulation. Finally the trained model served as the scoring function inside a genetic algorithm, an evolutionary search that breeds and mutates candidate aluminium arrangements. It produced 260 MOR and 310 MFI structures aimed at chosen heat-of-adsorption targets, which were then checked by simulation.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors built graph neural networks that predict the CO2 heat of adsorption and Henry coefficient of aluminum-substituted zeolites directly from their structure, trained on Monte Carlo simulation data they generated for the MOR, MFI, RHO and ITW topologies. They modified the Equivariant Porous Crystal Networks architecture so that each pore produces a scalar that is summed into the prediction, and report that the two architectures' performance metrics have overlapping confidence intervals on both properties. The per-pore values tracked how often a CO2 molecule occupied each pore in Monte Carlo simulations of four MOR structures. Using the trained model as the fitness function of a genetic algorithm, they generated 260 MOR and 310 MFI configurations at target heats of adsorption and then simulated them, obtaining a mean absolute error of 1.20 for MOR and 2.69 for MFI between the model and Monte Carlo values.
How AI was used
Zeolite configurations were generated with the rule-based ZEORAN program and labelled with CO2 heat of adsorption and Henry coefficients from Widom-insertion Monte Carlo simulations in RASPA. Each configuration was encoded as a periodic graph in which T-atoms are nodes marked 1 for aluminum and 0 for silicon, pore nodes carry pore area and ring size, edges join T-atoms sharing an oxygen and T-atoms to their pores, and edge distances are expanded in radial basis functions under the minimum image convention. The EPCN message-passing architecture, which shares parameters between symmetry-equivalent nodes and edges, was trained separately for each topology on a 90/10 random split, for 200 epochs with AdamW, a learning rate of 0.001 and a batch size of 32, with ten runs per model under random weight initialisation to obtain confidence intervals. The authors' variant drops the pooling of pore hidden states and instead passes each pore's hidden state through a shared MLP, summing the resulting per-pore scalars, so the pore-level terms can be read off as contributions. For inverse design, the trained model's prediction enters the fitness function of a PyGAD genetic algorithm over binary aluminum/silicon gene vectors, run for 50 generations with 2 parents mating, single-point crossover at probability 0.2, 5 elite solutions kept and an initial population of 32 seeded from the training-set aluminum distribution; the resulting structures were then simulated by Monte Carlo.
The shape of the work
Structural · the record, drawn
no AI
Generate Al/Si configurations of four topologies
Producing candidate objects that did not previously exist.
The configurations were generated using the ZEORAN program.where the paper describes this · verbatim
no AI
Compute reference properties by Monte Carlo
Numerical or physics simulation, including where a learned surrogate replaces it.
Monte Carlo (MC) simulations using the Widom particle insertion method in the canonical ensemble (NVT) were performed.where the paper describes this · verbatim
no AI
Encode zeolites as graphs with pore nodes
Encoding data into features, descriptors, embeddings or graphs.
we represent the T-atoms (Al/Si) as nodes in the graphwhere the paper describes this · verbatim
AI
Train one model per topology on both properties
Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.
All models were trained for 200 epochs, using the AdamW optimizer with a learning rate of 0.001 and a batch size of 32.where the paper describes this · verbatim
AI
Predict adsorption properties on held-out structures
Running a trained model over new data to predict, classify or score. The AI stood in for simulation.
In Figure 4, we compare the heat of adsorption distributions on the test set between our ML algorithm and MC simulations.where the paper describes this · verbatim
AI
Attribute adsorption to individual pores
Extracting understanding from model behaviour. The AI stood in for simulation.
the model now predicts a scalar value per feature for each porewhere the paper describes this · verbatim
AI
Inverse design with genetic algorithm
Iterative search over a space. The AI stood in for exhaustive search. Its result feeds back into an earlier step.
This is achieved using a genetic algorithm which can generate zeolites satisfying a target heat of adsorption.where the paper describes this · verbatim
no AI
Simulate generated structures to check targets
Testing outputs against ground truth.
Following this, we simulated the generated structures using MC.where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's object is the model itself: the reported adsorption-property values, the pore-level adsorption-site attribution and the inverse-designed configurations are all produced by the trained network.
To model the heat of adsorption of the different zeolites, we make use of Graph Neural Networks (GNNs).where the paper describes this · verbatim
each testing set consists of 10% of the data points from a zeolitewhere the paper describes this · verbatim
The generated structures and their simulated properties are available on GitHub.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of EPCN (Equivariant Porous Crystal Networks)Which version of the model was used is not stated.
- Version of EPCN extension with per-pore scalar outputs (this work)Which version of the model was used is not stated.
About this article
Record aix-00160, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error