~/aixsci
200 records · all checked

materials-chemistry/ai produced the result/ACS Applied Materials & Interfaces 2025 · v2

Language models suggest ingredients and firing temperatures for inorganic materials recipes

Researchers tested seven off-the-shelf language models on predicting the starting chemicals and heating temperatures for making solid materials, then used the models to write 28,548 synthetic recipes and trained a smaller model, SyntMTE, on them.

1. Curate and split literature synthesis dataset2. Prompt LMs for ranked precursor sets3. Prompt LMs for calcination and sintering temperatures4. Aggregate predictions across model ensembles5. Generate synthetic solid-state recipe dataset6. Train SyntMTE and baselines in two stages7. Extract LLZO reference conditions from literature corpus8. Predict sintering temperatures for doped LLZO

spectrum · one line per step, placed by what the step does · bright lines used AI

Language Models Enable Data-Augmented Synthesis Planning for Inorganic Materials
ACS Applied Materials & Interfaces, 2025

doi:10.1021/acsami.5c11229 · record aix-00210 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Candidate generation, Property prediction
Model family
Large language model, Transformer, Gradient-boosted trees, Multilayer perceptron
Checked by
Held-out1000 tested
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Making a new inorganic material is not like mixing a drink. A chemist must choose which powders to start from, then heat them in stages: calcination, a first bake that drives off gases and forms the compound, and sintering, a hotter bake that fuses the grains into a solid. The right powders and the right temperatures are rarely obvious. They depend on which elements are involved, what intermediate compounds form along the way, and a great deal of accumulated laboratory habit. That knowledge sits scattered across decades of published papers rather than in any neat table, which makes planning a route for an untried compound largely a matter of experience and trial and error.

The researchers asked whether language models, trained on text at large, carry enough of that scattered know-how to propose a recipe. They also asked whether such models could be used to manufacture training data for a smaller, purpose-built model, so that the know-how might be transferred into a system that predicts processing temperatures directly.

Where AI came in

Seven commercial language models were each shown forty worked examples and asked to name the starting chemicals for a thousand held-out target materials, and to give calcination and sintering temperatures for another thousand reactions. GPT-4.1 matched the published starting chemicals exactly 53.8% of the time on its first guess. Reported average errors on temperature were 126 °C for calcination and 101 °C for sintering. Combining three models by averaging their temperatures, or by merging their ranked ingredient lists, improved on the single best model and cut the cost per prediction by around 70% compared with GPT-4.1 alone.

The models were then put to a second use: generating 28,548 complete recipes for compounds known to have been made in a laboratory. A transformer called SyntMTE, built from an encoder already trained on calculated materials properties, was trained first on these machine-written recipes and then on recipes mined from the literature. That two-stage training improved its average error by over 4%. Tested on forty reported routes for doped lithium lanthanum zirconium oxide, a solid battery electrolyte, with all such compounds kept out of training, the model reproduced the reported ordering of sintering temperatures across dopants, including the drop for bismuth.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

Seven off-the-shelf language models were prompted with 40 in-context examples to recommend precursors and predict calcination and sintering temperatures for inorganic solid-state synthesis, using held-out 1,000-entry subsets of a text-mined literature dataset. GPT-4.1 reached 53.8% Top-1 exact-match accuracy on precursor prediction, and reported mean absolute errors were 126 °C for calcination and 101 °C for sintering temperatures; averaging or rank-fusing three models improved results beyond Top-1 and cut inference cost by around 70% relative to GPT-4.1. The authors then used the models to produce 28,548 complete synthetic recipes and trained SyntMTE, a transformer derived from the DFT-pretrained MTEncoder, first on the synthetic data and then on literature data; a model trained only on synthetic data was 6% worse than the literature-trained model, while two-stage training improved mean absolute error by over 4% for SyntMTE. Applied to 40 reported routes for doped Li7La3Zr2O12 with all LLZO compounds withheld from training, the model reproduced the reported ordering of dopant-dependent sintering temperatures, including the drop for Bi substitution.

How AI was used

A text-mined inorganic synthesis dataset was filtered and split, with two 1,000-entry subsets held out and 40 in-context examples drawn from the validation fraction. Seven language models were queried through the OpenRouter API at temperature 0.1 with a parser-aware retry policy: one task asked each model to generate ranked precursor sets without being told how many precursors to use, with outputs canonicalised via pymatgen and scored by Top-k exact match; the other asked for calcination and sintering temperatures as structured JSON, scored by MAE, RMSE and R². Predictions from three-model ensembles were then combined without further learning, by min-, average- and max-rank fusion for precursor lists and by averaging for temperatures. Separately, lab-synthesized compounds from the Materials Project were reduced by maximum-entropy sampling over MTEncoder representations, non-solid-state entries were flagged out by GPT-4.1, and precursor sets and processing conditions were generated for the remainder, with incomplete routes and sub-threshold temperatures removed. SyntMTE encodes the target and all precursors with shared MTEncoder weights pretrained on the Alexandria DFT database, mean-pools and concatenates the embeddings, and regresses both temperatures through a two-layer MLP head; it and three baselines were trained under synthetic-only, literature-only and sequential synthetic-then-literature regimes on a chronological split. For the case study, an o3 model extracted temperatures from an LLZO literature corpus, LLZO and commonly doped variants were excluded from training and validation, and the fine-tuned model was applied to the extracted compositions.

The shape of the work

Structural · the record, drawn

PREPARATIONGENERATIONINFERENCESCREENINGGENERATIONTRAININGPREPARATIONINFERENCE12345678AIAIAIAIAIAICurate and splitliteraturesynthesis datasetPrompt LMs forranked precursorsetsPrompt LMs forcalcination andsintering temper…Aggregatepredictionsacross model ens…Generatesyntheticsolid-state reci…Train SyntMTE andbaselines in twostagesExtract LLZOreferenceconditions from …Predict sinteringtemperatures fordoped LLZO↤ statistical model↤ statistical model↤ manual curation↤ manual curation↤ physical experiment
AI stepNo AI↤ what the AI stood in for
1Preparation
no AI

Curate and split literature synthesis dataset

Cleaning, filtering, normalising or labelling data already obtained.

After filtering for elemental consistency and removing ambiguous entries, 18,804 reactions remainwhere the paper describes this · verbatim
in the paper
2Generation
AI

Prompt LMs for ranked precursor sets

Producing candidate objects that did not previously exist. The AI stood in for statistical model.

We submitted prompts to an LM provider (via OpenRouter) without specifying the number of precursorswhere the paper describes this · verbatim
in the paper
3Inference
AI

Prompt LMs for calcination and sintering temperatures

Running a trained model over new data to predict, classify or score. The AI stood in for statistical model.

we prompt models to output calcination and sintering temperatures in °C with 40 in-context exampleswhere the paper describes this · verbatim
in the paper
4Screening
no AI

Aggregate predictions across model ensembles

Reducing a candidate set by filtering or ranking, in a single pass.

For the precursor recommendation, we aggregated ranked lists using three rank-fusion strategies.where the paper describes this · verbatim
in the paper
5Generation
AI

Generate synthetic solid-state recipe dataset

Producing candidate objects that did not previously exist. The AI stood in for manual curation.

Incorporating minimum temperature thresholds of 300 °C for calcination and 500 °C for sintering yields 28,548 plausible solid-state recipes.where the paper describes this · verbatim
in the paper
6Training
AI

Train SyntMTE and baselines in two stages

Fitting model parameters, including fine-tuning an existing model.

We first adapt the MTEncoder on a large LM-generated data set to bias it toward solid-state reaction conditions, then fine-tune on experimental literature recipes.where the paper describes this · verbatim
in the paper
7Preparation
AI

Extract LLZO reference conditions from literature corpus

Cleaning, filtering, normalising or labelling data already obtained. The AI stood in for manual curation.

Calcination and sintering temperatures were extracted using OpenAI’s o3 model, followed by manual spot checks.where the paper describes this · verbatim
in the paper
8Inference
AI

Predict sintering temperatures for doped LLZO

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

we withheld any LLZO-based compounds from the training sets and predicted the sintering temperatureswhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's results are the language models' own predictions, the LM-generated recipe dataset, and a transformer trained on it; without the models there is no finding.

+What the AI was for
we evaluate whether they enable knowledge transfer by training a transformer, SyntMTE, on 28,548 LM-generated reaction recipeswhere the paper describes this · verbatim
+How it was taught
Zero-shotSupervisedTransfer / fine-tuningSemi-supervisedin the paper
+Models named
OpenAI GPT-4.1 · Off the shelfGoogle Gemini 2.0 Flash · Off the shelfMeta Llama 4 Maverick · Off the shelfxAI Grok 3 mini Beta · Off the shelfDeepSeek Chat v3 · Off the shelfAlibaba Qwen 2.5 VL · Off the shelfMistral Small 3.1 · Off the shelfOpenAI o3 · Off the shelfSyntMTE (derived from MTEncoder, pretrained on Alexandria DFT data) · Fine-tunedCrabNet · Trained from scratchXGBoost regressor with mean-pooled reaction features · Trained from scratchCompositional feedforward network · Trained from scratchin the paper
+How results were checked
Held-out1000 testedin the paper
we randomly sample 1,000 entries for our LM evaluationwhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata not reportedin the paper
Prompts, which are available in the repository and Supporting Informationwhere the paper describes this · verbatim
+Compute
Experiments were run five times each on two NVIDIA RTX A6000 GPUs; LM inference via the OpenRouter API, with per-prediction cost estimated from token rates.in the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 15 items
  • Trained model weightsWhether the trained model is available is not stated.
  • DataWhether the data are available is not stated.
  • Version of OpenAI GPT-4.1Which version of the model was used is not stated.
  • Version of Google Gemini 2.0 FlashWhich version of the model was used is not stated.
  • Version of Meta Llama 4 MaverickWhich version of the model was used is not stated.
  • Version of xAI Grok 3 mini BetaWhich version of the model was used is not stated.
  • Version of DeepSeek Chat v3Which version of the model was used is not stated.
  • Version of Alibaba Qwen 2.5 VLWhich version of the model was used is not stated.
  • Version of Mistral Small 3.1Which version of the model was used is not stated.
  • Version of OpenAI o3Which version of the model was used is not stated.
  • Version of SyntMTE (derived from MTEncoder, pretrained on Alexandria DFT data)Which version of the model was used is not stated.
  • Version of CrabNetWhich version of the model was used is not stated.
  • Version of XGBoost regressor with mean-pooled reaction featuresWhich version of the model was used is not stated.
  • Version of Compositional feedforward networkWhich version of the model was used is not stated.
  • What step 6 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00210, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error