materials-chemistry/ai produced the result/arXiv 2024 · v2
Four million simulated X-ray patterns used to test 21 crystal-symmetry classifiers
Researchers simulated over four million powder X-ray diffraction patterns from 119,569 crystal structures, then trained 21 neural networks from scratch to read each pattern's crystal system and space group.
spectrum · one line per step, placed by what the step does · bright lines used AI
SimXRD-4M: Big Simulated X-ray Diffraction Data Accelerate the Crystal Symmetry Classification
arXiv, 2024
doi:10.48550/arxiv.2406.15469 · record aix-00036 v2 · checked 2026-10-08
- AI was for
- Classification
- Model family
- Convolutional neural network, Recurrent neural network, Transformer, Multilayer perceptron
- Checked by
- Held-out
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Shine X-rays at a powder of ground-up crystal and the beam scatters into a pattern of peaks. The positions and heights of those peaks depend on how the atoms inside the crystal are arranged and repeat. From that pattern, researchers try to work out the crystal's symmetry: which of the seven broad crystal systems it belongs to, and which of the 230 space groups, the full catalogue of ways a repeating pattern of atoms can be arranged in three dimensions. This is hard because real measurements are messy. Grain size, orientation, stress, heat-driven atomic vibration, instrument misalignment, background signal and noise all smear, shift and reshape the peaks.
The team wanted a large, labelled collection of patterns on which methods for this task could be compared. They took crystal structures from the Materials Project database, filtered them for symmetry problems, duplication and size, and ran each surviving structure through their own simulation code, Pysimxrd, 33 times with randomly combined physical conditions. The result, SimXRD-4M, holds 4,065,346 simulated patterns, each a sequence of 3,501 numbers labelled with its true crystal system and space group.
Where AI came in
The patterns themselves were not produced by a learned model; they come from physics-based simulation of diffraction. The machine learning sits downstream, as the thing being measured. The authors trained 21 neural networks from scratch on the simulated patterns, each learning to map a pattern to one of 7 crystal systems or one of 230 space groups. The line-up included eleven convolutional networks drawn from earlier symmetry-identification work, recurrent networks of several kinds, transformers and a simple multilayer perceptron. In each case the network stands in for the conventional step-by-step analysis a crystallographer or a rule-based algorithm would apply to assign symmetry from peak positions.
The trained networks were then run in several settings: on held-out patterns of the same crystals under different simulated conditions, on patterns of crystals the networks had never seen, and on real measured patterns of well-characterised minerals from the RRUFF collection, where performance matched what the networks achieved on simulated test data. Accuracy fell for symmetry classes that appear rarely in the data, and in the unseen-crystals setting several models, including most of the convolutional networks, the perceptron and the plain transformer, landed at around 24% accuracy. The authors also masked peaks band by band to see which parts of a pattern the networks leaned on.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors built SimXRD-4M, a dataset of 4,065,346 simulated powder X-ray diffraction patterns covering 119,569 crystal structures, each simulated 33 times under randomly combined physical conditions using their own simulation code, Pysimxrd. They then trained 21 sequence models — CNN architectures from prior symmetry-identification work, recurrent networks, transformers and an MLP — to classify each pattern into one of 7 crystal systems or 230 space groups, under an in-library split (same crystals, different simulated conditions) and an out-of-library split (disjoint crystals). Accuracy was lower for low-frequency symmetry classes, and several models (CNN1-4, CNN6, CNN8, MLP and the raw transformer) reached accuracy around 24% in the out-of-library setting; across the retraining experiments, label smoothing and focal loss gave better results than weighted classification. Models trained on SimXRD and tested on the experimental RRUFF patterns gave performance consistent with their performance on the simulated test data.
How AI was used
Learned models appear only as the subject of the benchmark, not in dataset construction: the patterns themselves come from a physics-based simulation of diffraction intensity, peak profile, thermal and instrumental factors. Crystal structures were pulled from the Materials Project (MP-2024.1), filtered with Spglib for broken symmetry, duplication, space group discrepancies and a 500-atom cell limit, then passed to Pysimxrd with randomly sampled condition parameters to produce 3501-dimensional d-I patterns labelled by crystal system and space group. The labelled patterns were split two ways — by simulated environment for in-library evaluation and by crystal type for out-of-library evaluation — and used to train 11 CNN architectures taken from prior symmetry-identification work, RNN, LSTM and GRU in unidirectional and bidirectional form, a raw transformer, iTransformer, PatchTST, and an MLP baseline, each from scratch in PyTorch with a batch size of 128, learning rate of 2.5e-4, 50 epochs and early stopping patience of 3, using cross-entropy and, in separate runs, weighted classification, label smoothing and focal loss. For recurrent models and transformers a sequence representation is learned first and an MLP head predicts the symmetry. The trained models were then run over held-out simulated splits, over parameter-interval subsets of the test data, over the experimental RRUFF patterns, and over patterns with peak features masked by intensity band for a feature-importance analysis.
The shape of the work
Structural · the record, drawn
no AI
Retrieve crystal structures from Materials Project
Obtaining raw data, whether by measurement, download or retrieval.
We utilize the latest dataset, denoted as MP-2024.1, which includes a total of 154,718 crystallographic structures as of January 2024.where the paper describes this · verbatim
no AI
Filter structures for symmetry consistency and size
Cleaning, filtering, normalising or labelling data already obtained.
A total of 119,569 crystal structures were screened.where the paper describes this · verbatim
no AI
Simulate powder XRD patterns under varied physical conditions
Numerical or physics simulation, including where a learned surrogate replaces it.
For each input crystal, we randomly combine these physical conditions within reasonable ranges, repeating the process 33 times to simulate XRD patternswhere the paper describes this · verbatim
no AI
Construct in-library and out-of-library splits
Cleaning, filtering, normalising or labelling data already obtained.
the dataset is randomly split according to the types of simulated environmentswhere the paper describes this · verbatim
AI
Train sequence classifiers on simulated patterns
Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.
All models are trained for 50 epochs with an early stopping patience of 3.where the paper describes this · verbatim
AI
Classify symmetry on simulated test splits
Running a trained model over new data to predict, classify or score. The AI stood in for conventional algorithm.
Table 2 displays the baseline performance and inference time for the classification of crystal systems and space groups.where the paper describes this · verbatim
AI
Test trained models on experimental RRUFF patterns
Testing outputs against ground truth. The AI stood in for conventional algorithm.
We train and validate all baseline models on SimXRD and then test their performance on RRUFF.where the paper describes this · verbatim
AI
Peak-masking feature importance analysis
Extracting understanding from model behaviour.
By masking features in powder XRD patterns, we observe the impact of the masked features on the model’s inferencewhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's reported findings are the classification behaviour of 21 neural sequence models on the released dataset, so the results consist of trained-model outputs; the dataset itself is produced by physics simulation rather than by a learned model
We benchmark 21 sequence models in both in-library and out-of-library scenarioswhere the paper describes this · verbatim
119,569 × 30 training instances, 119,569 × 1 validation instances, and 119,569 × 2 testing instanceswhere the paper describes this · verbatim
we have made the SimXRD database, simulation code, benchmark models, evaluation process and tutorial notebooks into a repositorywhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of CNN 1-11 (existing CNN architectures for symmetry identification, classified by layer counts, pooling, ensembling and dropout)Which version of the model was used is not stated.
- Version of RNNWhich version of the model was used is not stated.
- Version of LSTMWhich version of the model was used is not stated.
- Version of GRUWhich version of the model was used is not stated.
- Version of Bidirectional RNNWhich version of the model was used is not stated.
- Version of Bidirectional LSTMWhich version of the model was used is not stated.
- Version of Bidirectional GRUWhich version of the model was used is not stated.
- Version of Transformer (raw)Which version of the model was used is not stated.
- Version of iTransformerWhich version of the model was used is not stated.
- Version of PatchTSTWhich version of the model was used is not stated.
- Version of MLPWhich version of the model was used is not stated.
- What step 8 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00036, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error