materials-chemistry/ai produced the result/npj Computational Materials 2022 · v2
Machine learning sifts 25,000 known materials for quantum-technology host candidates
Researchers assembled a database of over 25,000 computed materials, labelled some as promising or unpromising hosts for quantum devices, then trained four classifiers to sort the rest. All three labelling schemes and all four methods agreed on 47 candidates.
spectrum · one line per step, placed by what the step does · bright lines used AI
Predicting solid state material platforms for quantum technologies
npj Computational Materials, 2022
doi:10.1038/s41524-022-00888-3 · record aix-00066 v2 · checked 2026-10-08
- AI was for
- Classification
- Model family
- Linear model, Random forest, Gradient-boosted trees
- Checked by
- Held-out
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Many quantum technologies rest on a single flaw inside a crystal. A missing atom, or a foreign one, can trap an electron whose spin behaves as a quantum bit, holding information and emitting light that can be read out. The best known example sits in diamond, but diamond is not the only option, and the quality of such a defect depends heavily on the crystal around it. A good host is usually a semiconductor or insulator, meaning it has a band gap: an energy range electrons cannot occupy, which keeps the defect's own energy levels cleanly separated from the bulk material. Which of the many thousands of known crystals qualify is not obvious.
Searching by hand is slow, because each candidate ordinarily needs careful quantum-mechanical calculation. The authors instead set out to screen broadly. They pulled material entries from several public databases built on density functional theory, a standard method for computing electronic structure, and described each one with nearly 5,000 numerical features drawn from composition, crystal structure, atomic sites and electronic properties. They then marked subsets as suitable or unsuitable hosts using three different rule- and literature-based schemes, including one that counted only materials with reported quantum-compatible behaviour.
Where AI came in
The sorting itself was done by software that learns from examples. For each labelling scheme, the authors trained four standard classifiers — logistic regression, decision trees, random forests and gradient boosting — on the labelled materials, after compressing the thousands of features down with principal component analysis, a technique that finds the combinations of measurements that vary most. Settings were chosen by repeatedly holding back part of the data and testing on it. The trained classifiers were then run across the unlabelled materials, giving each a verdict and a confidence number. Taking only the materials where all schemes and all four methods agreed left 47: eight single elements, 29 two-element and 10 three-element compounds.
In place of a specialist weighing each crystal against known requirements, or running fresh calculations for every entry, the trained models supplied that judgement at the scale of the whole database. An off-the-shelf learned model, AFLOW-ML, also estimated band gaps straight from crystal structure files, standing in for simulation at that step. The authors additionally inspected which feature combinations their models leaned on: the literature-based scheme weighted symmetry and structural measures such as bond length and orientation, while the other two leaned on band gap and ionic character. No experimental testing of the predictions is reported.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors assembled a dataset of over 25,000 materials from DFT databases with nearly 5,000 physics-informed features, and labelled subsets of it as suitable or unsuitable semiconductor hosts for quantum technologies using three different labelling schemes (Ferrenti, extended Ferrenti and empirical). Logistic regression, decision trees, random forests and gradient boosting were trained on each labelled set and then used to classify the remaining materials. The empirical scheme, which labelled only materials with reported quantum-compatible behaviour, yielded fewer predicted candidates, and its models weighted features related to symmetry and crystal structure such as bond length, orientation and radial distribution, whereas the two Ferrenti schemes weighted band gap and ionic character. All three approaches and all four methods agreed on 47 candidates at 0.5 confidence, comprising 8 elemental, 29 binary and 10 tertiary compounds; no experimental testing of the predictions is reported.
How AI was used
Entries were extracted from Materials Project and matched against OQMD, JARVIS-DFT, AFLOW, AFLOW-ML and Citrination, with AFLOW-ML run on crystallographic information files to return a predicted metal/insulator verdict and band gap. Matminer and Pymatgen were used to featurize composition, structure, atomic sites, density of states and band structure. Three rule- and literature-based labelling procedures assigned suitable and unsuitable labels to subsets of the data, including a manual crystallographic screening step in the empirical procedure. For each labelled set, principal component analysis was applied to reduce dimensionality and four supervised classifiers — logistic regression, decision trees, random forests and gradient boosting via XGBoost — were fitted, with hyperparameters selected by 5x5 stratified cross-validation. The trained classifiers were then run over the unlabelled materials and test entries to assign suitability with a confidence value; agreement between methods and approaches was taken at confidence cut-offs of 0.5, 0.75 and 0.85, and the fitted models' principal components were inspected to identify which descriptors drove classification.
The shape of the work
Structural · the record, drawn
no AI
Query materials databases
Obtaining raw data, whether by measurement, download or retrieval.
We start by extracting all entries in the Materials Project (MP) database that match a specific query.where the paper describes this · verbatim
AI
Predict band gaps from structure with AFLOW-ML
Running a trained model over new data to predict, classify or score. The AI stood in for simulation.
use the CIF files as input to AFLOW-ML, which then returns an anticipated band gapwhere the paper describes this · verbatim
no AI
Featurize materials with Matminer
Encoding data into features, descriptors, embeddings or graphs.
After material extraction we apply tools from the open-source library Matminer to generate thousands of features from the data.where the paper describes this · verbatim
no AI
Label suitable and unsuitable candidates by three approaches
Cleaning, filtering, normalising or labelling data already obtained.
Below we follow three separate procedures for labeling materials as suitable or unsuitable candidates for QT.where the paper describes this · verbatim
AI
Reduce dimensionality and train four classifiers per approach
Fitting model parameters, including fine-tuning an existing model.
we apply a 5×5 stratified cross-validation when searching for the optimal hyperparameter combinationswhere the paper describes this · verbatim
AI
Classify unlabelled materials as suitable or unsuitable
Running a trained model over new data to predict, classify or score. The AI stood in for expert judgement.
the ML methods were employed on the test setswhere the paper describes this · verbatim
no AI
Intersect method agreement at confidence thresholds
Reducing a candidate set by filtering or ranking, in a single pass.
All approaches and their corresponding ML methods agree on a total of 47 potential candidates to 0.5 confidencewhere the paper describes this · verbatim
no AI
Analyse which features the classifiers weighted
Extracting understanding from model behaviour.
Importantly, our findings also reveal which material properties are weighted by the ML methods upon predicting a material as suitable for QT applicationswhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is the set of predicted candidate host materials, which is produced by the four trained classifiers; no non-AI route to the same output is reported.
In this work our focus is on supervised learning only with labeled data for classification problems.where the paper describes this · verbatim
We note that all four ML methods have high evaluation metrics for the optimal hyperparameters.where the paper describes this · verbatim
The codes employed to develop the machine learning results are available online at.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of logistic regressionWhich version of the model was used is not stated.
- Version of decision treesWhich version of the model was used is not stated.
- Version of random forestsWhich version of the model was used is not stated.
- Version of gradient boosting (XGBoost)Which version of the model was used is not stated.
- Version of AFLOW-ML (Property Labeled Material Fragments)Which version of the model was used is not stated.
- What step 5 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00066, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error