materials-chemistry/ai produced the result/arXiv 2024 · v2
Machine learning ranks stable materials by predicted superconducting temperature
Researchers trained a simple statistical model on measured superconductors, then used it to estimate the critical temperature of about 153,000 known compounds from chemical composition alone. Sixty-four stable candidates were predicted above 250 K.
spectrum · one line per step, placed by what the step does · bright lines used AI
High-Tc superconductor candidates proposed by machine learning
arXiv, 2024
doi:10.48550/arxiv.2406.14524 · record aix-00146 v2 · checked 2026-10-09
- AI was for
- Property prediction
- Model family
- Linear model
- Checked by
- Held-out13661 tested
- Code
- available
The finding the paper is about came from the AI.
What this research was about
A superconductor carries electricity with no resistance at all, but only below a certain temperature, known as the critical temperature, or Tc. For most materials that temperature is very low, so the search is on for ones that work closer to room temperature. The difficulty is that Tc cannot be read off a chemical formula. It emerges from how electrons and the vibrations of the atomic lattice interact, and calculating it from first principles is costly. Measuring it means making the material and cooling it, one compound at a time, which limits how much of the space of possible compounds anyone can check.
The researchers set out to estimate Tc from chemical composition alone, with no information about how the atoms are arranged, and then to apply that estimate across a large public database of known and computed materials. Compounds predicted to have a high Tc, and calculated to be thermodynamically stable, were kept and ranked.
Where AI came in
The learning part was a ridge regression, a standard method for fitting a straight-line relationship between numbers while discouraging extreme fitted values. Rather than one model for everything, a fresh model was fitted for each material being asked about, trained only on its ten closest matches in a cleaned set of experimentally measured superconductors drawn from the SuperCon database. Closeness was judged using 147 numerical descriptors generated from each formula, built from statistics of the properties of the elements present and from the proportions in which they appear.
Accuracy was checked by predicting Tc for materials held back from training and comparing with the measured values, including against simpler baselines. The models then supplied the predicted Tc for every candidate in the Materials Project database. Those predictions stand in for measurements that were not made: no superconductivity experiments were carried out on the proposed candidates, so their reported temperatures, including the highest at 316 K for LiCuF4, exist only as model outputs.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors predicted superconducting critical temperature from chemical composition alone by training a ridge regression model separately for each query material on its ten nearest neighbours in the curated SuperCon data set. On out-of-sample SuperCon test materials the similarity-based models gave mean absolute errors of about 5 K across the 0–250 K range, and about 3 K for leave-one-out predictions after removing samples with large spreads in their feature-weight products. Applying the approach to about 153k Materials Project materials and keeping those within 0.030 eV/atom of the convex hull, the ambient-pressure model placed sixty-four materials above 250 K, thirty-four of which have DFT-computed band gaps below 1 eV; the highest prediction was 316 K for LiCuF4. No superconductivity measurements were performed on the proposed candidates.
How AI was used
SuperCon entries were cleaned by averaging repeated measurements and removing contentious, single-element, ten-element and arbitrarily doped stoichiometries, and a separate ambient-pressure set was formed by removing samples that SuperCon2 indicated were measured under applied pressure. Each composition was encoded as 147 Matminer features built from statistics of elemental properties and stoichiometric norms, with no structural information. For a given query material, the Euclidean nearest neighbours in the training set were retrieved and a ridge regression model was fitted on those neighbours alone, with the regularisation strength chosen by cross-validation on that training subset and the matrix inversion done by Cholesky decomposition in scikit-learn; absolute values of predictions were taken. Learning curves over similarity-selected versus random training samples, and against k-nearest-neighbour regression baselines, were used to set n=10, and leave-one-out predictions were made for every SuperCon sample. The same procedure was then run over the Materials Project set, after which predictions with large feature-weight product spreads, materials above the convex hull threshold, and in a second pass materials with larger band gaps, were discarded before ranking by predicted Tc.
The shape of the work
Structural · the record, drawn
no AI
Clean SuperCon data set
Cleaning, filtering, normalising or labelling data already obtained.
We cleaned the data set by assigning to stoichiometries with multiple Tc measurements their mean values.where the paper describes this · verbatim
no AI
Separate ambient-pressure subset
Cleaning, filtering, normalising or labelling data already obtained.
These were removed from SuperCon to create a separate data setwhere the paper describes this · verbatim
no AI
Query Materials Project candidates
Obtaining raw data, whether by measurement, download or retrieval.
We apply our similarity-based ML method to ∼ 153k samples listed in the Materials Project databasewhere the paper describes this · verbatim
no AI
Generate composition-based features
Encoding data into features, descriptors, embeddings or graphs.
147 features were generated for each sample from its composition using the materials informatics Python library Matminerwhere the paper describes this · verbatim
AI
Train query-aware ridge models and predict SuperCon Tc
Fitting model parameters, including fine-tuning an existing model. The AI stood in for physical experiment.
are then used to train a ridge regression model, from which the test sample’s Tc is predictedwhere the paper describes this · verbatim
no AI
Evaluate prediction error against measured Tc
Testing outputs against ground truth.
the learning curves show the prediction error (mean absolute error, MAE) on the test set after training on n training sampleswhere the paper describes this · verbatim
AI
Predict Tc across Materials Project
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
Each sample’s predictions under implicit pressure (n=10) and ambient pressure (n=10) were made by training on its nearest neighborswhere the paper describes this · verbatim
no AI
Filter and rank high-Tc candidates
Reducing a candidate set by filtering or ranking, in a single pass.
Those with computed energies above their convex hulls of greater than 0.030 eV/atom are also disregarded as being thermodynamically unstable.where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The reported candidates and their Tc rankings exist only as outputs of the trained ridge regression models; no measurement supports them.
Training of query-aware similarity-based ridge regression models on experimental SuperCon datawhere the paper describes this · verbatim
predictions under unknown/ambient pressure are made for each of the 13,661/13,624 materialswhere the paper describes this · verbatim
Refer to https://zenodo.org/records/14052692 for: Python code to generate ML features and to implement our similarity-based ML modelswhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- Version of Similarity-based ridge regression, ambient-pressure modelWhich version of the model was used is not stated.
- Version of Similarity-based ridge regression, implicit-pressure modelWhich version of the model was used is not stated.
- Version of k-nearest neighbors regression (baseline)Which version of the model was used is not stated.
About this article
Record aix-00146, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error