materials-chemistry/ai produced the result/Nature Communications 2024 · v2
Transformer model predicts how gas-storing crystals take up gases, replacing slow simulations
Researchers built Uni-MOF, a transformer trained first on hundreds of thousands of porous crystal structures and then on around 3,000,000 adsorption measurements, so that a structure file plus a gas, temperature and pressure yields a predicted uptake.
spectrum · one line per step, placed by what the step does · bright lines used AI
A comprehensive transformer-based approach for high-accuracy gas adsorption predictions in metal-organic frameworks
Nature Communications, 2024
doi:10.1038/s41467-024-46276-x · record aix-00015 v2 · checked 2026-10-07
- AI was for
- Property prediction, Simulation surrogate
- Model family
- Transformer
- Checked by
- Held-out
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Metal-organic frameworks, or MOFs, are crystalline solids built from metal hubs joined by organic struts. The result is a sponge-like material riddled with pores, and gases such as methane, argon or krypton can collect on the inner surfaces, a process called adsorption. How much gas a given framework holds depends on its particular architecture and on the conditions, since temperature and pressure both change the answer. Millions of frameworks can be imagined on paper, so the practical question is which ones to make. Measuring each in the laboratory is slow, and the usual computational stand-in, a Grand Canonical Monte Carlo simulation that samples gas molecules moving in and out of the pores, costs considerable processor time for every material, gas and condition.
The researchers set out to replace that per-case calculation with a single model covering many frameworks, several gases and a range of operating conditions. They wanted predictions that need nothing more than the crystallographic file describing where the atoms sit, together with the gas and the temperature and pressure of interest.
Where AI came in
The AI is a transformer, a network that weighs how parts of an input relate to one another, here applied to atoms in a crystal rather than words in a sentence. It was trained in two stages. First came self-supervised pre-training on more than 631,000 collected and computer-generated framework structures with no adsorption labels attached: the model had to guess masked atom types, restore coordinates that had been jittered with noise, and reproduce the lattice matrix that describes the repeating unit cell. Learning from the structures themselves in this way needs no human annotation. The model was then fine-tuned on around 3,000,000 labelled adsorption data points, with extra components encoding which gas was involved and the temperature and pressure.
In use, the fine-tuned model stands in for the Monte Carlo simulation. Given held-out structures it had not seen in training, it output predicted uptakes and also geometric descriptors such as pore diameters and void fraction. On test sets split so that no structure was shared with training, the coefficient of determination, a measure of how closely predictions track the reference values, was 0.98, 0.92 and 0.83 across three databases, and 0.85 for krypton when the split was made by gas instead. A variant trained without the pre-training stage scored 0.70 rather than 0.83. Fitting the predicted low-pressure methane uptakes placed three laboratory materials in the same order as published high-pressure experiments, while absolute low-pressure predictions departed from experiment for two others.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
Uni-MOF is a transformer model pre-trained with self-supervised masked-atom, coordinate-denoising and lattice-prediction tasks on over 631,000 collected and computationally generated MOF and COF structures, then fine-tuned on around 3,000,000 labelled adsorption data points so that a crystallographic file plus gas, temperature and pressure yields a predicted uptake. On test sets split by material so no structure was shared with training, the coefficient of determination was 0.98 for hMOF_MOFX_DB, 0.92 for CoRE_MOFX_DB and 0.83 for CoRE_MAP_DB, and 0.85 for krypton when the data were instead split by adsorbate gas. Removing pre-training lowered the CoRE_MAP_DB value from 0.83 to 0.70, and predicted structural features reached above 0.99 for hMOFs. Langmuir fits to predicted low-pressure methane uptakes ordered Zn2(bdc)2(dabco), MIL-101 and MOF-177 the same way as published high-pressure experimental values, while absolute low-pressure predictions deviated from experiment for Mg-dobdc and MOF-5.
How AI was used
MOF and COF structures were collected from existing databases and generated with the ToBaCCo.3.0 constructor, parsed with pymatgen into atom types, coordinates and lattice information, and labelled with adsorption uptakes taken from MOFXDB or produced by Grand Canonical Monte Carlo simulation in RASPA. A transformer encoder adapted from Uni-Mol was pre-trained on the unlabelled structures with three self-supervised objectives: predicting masked atom types, recovering coordinates corrupted by uniform noise applied to 15% of atoms, and regressing the 3x3 lattice matrix, using an edge gated distance kernel for spatial positional encoding. The pre-trained weights were then fine-tuned for supervised prediction of adsorption uptake, with added blocks that embed gas identity together with gas descriptors, temperature via equal-distance discretisation and pressure via logarithmic discretisation, so one model covers multiple gases and operating conditions. Data were split 8:1:1 by MOF structure (and in a separate experiment 5:1:1 by adsorbate gas) for training, validation and testing, with a separately trained variant omitting pre-training and per-sub-dataset single-system models as comparisons. The fine-tuned models were run over held-out structures to predict uptakes and geometric structural features; predictions were Langmuir-fitted and ranked for adsorbent screening, and embeddings and attention weights were visualised with t-SNE and heat maps.
The shape of the work
Structural · the record, drawn
no AI
Assemble MOF and COF structure library
Obtaining raw data, whether by measurement, download or retrieval.
we employed the ToBaCCo.3.0 program to generate over 306,773 MOF structureswhere the paper describes this · verbatim
no AI
Generate adsorption labels by GCMC simulation
Numerical or physics simulation, including where a learned surrogate replaces it.
we conducted Grand Canonical Monte Carlo (GCMC) simulations using the RASPA software to produce another 99,000+ gas adsorption uptake datasetwhere the paper describes this · verbatim
no AI
Parse structures and split datasets
Cleaning, filtering, normalising or labelling data already obtained.
Materials properties, including lattice vectors, lattice angles, unit cell volume, atoms, and coordinates, are extracted using the pymatgen.where the paper describes this · verbatim
AI
Self-supervised pre-training on 3D structures
Fitting model parameters, including fine-tuning an existing model. The AI stood in for expert judgement.
In the pre-training stage, we devised two tasks. 1) reconstructing the pristine three-dimensional positions from the noisy datawhere the paper describes this · verbatim
AI
Fine-tune with gas and operating-condition blocks
Fitting model parameters, including fine-tuning an existing model.
we trained the model using around 3,000,000 labeled data points across various adsorption conditions of MOFs and COFswhere the paper describes this · verbatim
AI
Predict uptake and structural features for unseen materials
Running a trained model over new data to predict, classify or score. The AI stood in for simulation.
the properties under varied working conditions can be predicted based solely on the crystallographic information file (CIF) of MOF materialswhere the paper describes this · verbatim
no AI
Rank adsorbents and screen high-performance materials
Reducing a candidate set by filtering or ranking, in a single pass.
the Uni-MOF framework is capable of accurately screening high performance adsorbents under high pressure based solely on prediction adsorption capacity under low pressurewhere the paper describes this · verbatim
no AI
Inspect learned representations and attention
Extracting understanding from model behaviour.
we visualize the structural features, which are 512-dimensional vectors, using the t-SNE methodwhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The reported results are the model's own predictions of gas adsorption uptake and structural features; there is no separate non-AI finding the paper rests on.
Compared with other Transformer-based models such as MOFormer and MOFTransformer, our Uni-MOF, as a Transformer-based frameworkwhere the paper describes this · verbatim
the prediction results in the never-before-seen test set represent the final performance (R shown here) of the modelwhere the paper describes this · verbatim
Code to run the Uni-MOF model is available in GitHubwhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of Uni-MOFWhich version of the model was used is not stated.
- Version of Uni-MOF w/o pre-trainingWhich version of the model was used is not stated.
- What step 5 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00015, version 2, checked by a person on 2026-10-07. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY-4.0; quotations are at most 25 words. How we work · Report an error