~/aixsci
200 records · all checked

materials-chemistry/ai produced the result/Communications Chemistry 2022 · v2

Deep-learning model designs petrol blends from octane and soot targets

Researchers trained a neural network to predict three combustion properties of fuels and their mixtures, then searched the model's internal representation for blends matching chosen targets. Eighty-six candidate mixtures came out; one blend of 22 components was put forward.

1. Curate RON, MON and YSI database2. Stratified train/validation/test split3. Encode molecules as SMILES and Mordred descriptors4. Train joint-property model with linear mixing operator5. Predict properties for held-out species and blends6. Compare predictions with measurements and baselines7. Search latent space for mixtures matching targets8. Post-screen candidates on physical properties

spectrum · one line per step, placed by what the step does · bright lines used AI

Artificial intelligence-driven design of fuel mixtures
Communications Chemistry, 2022

doi:10.1038/s42004-022-00722-3 · record aix-00058 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Property prediction, Candidate generation
Model family
Recurrent neural network, Multilayer perceptron, Graph neural network
Checked by
Held-out
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Petrol is not one substance but a mixture of many hydrocarbons, and how it behaves in an engine depends on the whole blend. Two standard measures, the research octane number and the motor octane number, describe how well a fuel resists knocking, the premature ignition that damages engines. A third, the yield sooting index, describes how much soot a fuel tends to produce when it burns. All three are measured in the laboratory, which is slow and needs sizeable samples, so only a small fraction of possible blends has ever been tested.

Mixtures are harder still than single compounds. A simple assumption, that a blend's octane number is just the average of its components weighted by how much of each is present, does not always hold, because molecules interact as they burn. The researchers set out to build a computer model that predicts all three properties for pure compounds and for mixtures at once, and then to run that model backwards: instead of asking what a given blend would do, asking which blends would hit a chosen combination of octane numbers and sooting behaviour.

Where AI came in

A neural network was trained from scratch on a database of published laboratory measurements for pure hydrocarbons, laboratory blends and real fuels. Each molecule entered the model twice over, once as a text string spelling out its structure and once as a list of calculated numerical descriptors; the network compressed both into a single string of numbers, a kind of coded fingerprint. A mixture's fingerprint was formed by combining its components' fingerprints in proportion to how much of each was present, and a final layer of the network mapped fingerprints to the three properties. Tested on data held back from training, it tracked the measurements with a correlation above 0.92 for all three.

The design step depended on the model entirely. Because the space of fingerprints is continuous, the researchers could treat the trained network as a mathematical function and use gradient-based optimisation to hunt for blends whose predicted properties sat near targets of 95 research octane, 85 motor octane and 60 sooting index, while obeying European petrol specifications. That search stood in for trying candidates one by one, which the number of possible blends makes impractical. It produced 86 mixtures. Conventional thermodynamic calculations, not machine learning, then screened these for vapour pressure, leaving five. The candidates and their properties exist as model output; experimental confirmation was left to future work.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors curated a literature database of research octane number, motor octane number and yield sooting index measurements for pure hydrocarbons, surrogate blends and real fuels, and trained a deep-learning model that encodes each molecule from its SMILES string and Mordred descriptors and represents a mixture as a composition-weighted linear combination of its components' latent vectors. On an independent test set the model reached R above 0.92 for all three properties, and its mixture octane-number mean absolute errors were lower than a linear-by-mole mixing rule across the blend sizes reported. Searching the model's latent space under gasoline specification constraints for targets of RON 95, MON 85 and YSI 60 produced 86 candidate mixtures, of which five had estimated Reid vapour pressure in the stated acceptable range; one blend of 22 components was put forward as the most promising candidate, with experimental confirmation left to future work.

How AI was used

A deep-learning model was trained from scratch to predict three combustion properties jointly. SMILES strings were one-hot encoded and passed through three stacked LSTM layers, while Mordred molecular descriptors were min-max normalised and passed through fully connected layers; the two resulting fingerprints were concatenated into a per-component latent vector. A mixing operator inside the training loop formed a mixture's latent vector as a matrix-vector product of component latent vectors with their compositions, and a fully connected predictor mapped latent vectors to the target properties, with the sooting index predicted via the measured soot-volume-fraction quantities and converted using the scale endpoints. The data were split by hierarchical stratified sampling, training used a weighted mean-squared-error loss, and hyperparameters were tuned with a Bayesian optimisation platform on a validation split. Because the latent space is continuous and differentiable, the trained predictor was used as the objective of a constrained optimisation solved with SciPy using PyTorch automatic differentiation gradients, started from database points closest to the target vector; a greedy depth-first search then reduced those solutions to smaller blends, with constraints enforced by Dykstra's method. A published graph-neural-network model was also run on test-set components for comparison. Non-learned thermodynamic correlations were applied afterwards to screen the resulting candidates.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONREPRESENTATIONTRAININGINFERENCEVALIDATIONOPTIMISATIONSCREENING12345678AIAIAICurate RON, MONand YSI databaseStratifiedtrain/validation/testsplitEncode moleculesas SMILES andMordred descript…Trainjoint-propertymodel with linea…Predictproperties forheld-out species…Comparepredictions withmeasurements and…Search latentspace formixtures matchin…Post-screencandidates onphysical propert…↤ conventional algorithm↤ physical experiment↤ exhaustive search
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Curate RON, MON and YSI database

Obtaining raw data, whether by measurement, download or retrieval.

The database of experimentally obtained measurements for the three combustion-related properties (RON, MON, and YSI) for single hydrocarbons and mixtures was curatedwhere the paper describes this · verbatim
in the paper
2Preparation
no AI

Stratified train/validation/test split

Cleaning, filtering, normalising or labelling data already obtained.

each subset was randomly split into 85% train/validation and 15% test set using stratified sampling in the scikit-learn librarywhere the paper describes this · verbatim
in the paper
3Representation
no AI

Encode molecules as SMILES and Mordred descriptors

Encoding data into features, descriptors, embeddings or graphs.

Generated SMILES strings were converted to a binary matrix using one-hot encoding.where the paper describes this · verbatim
in the paper
4Training
AI

Train joint-property model with linear mixing operator

Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.

the weighted loss function is used to train the modelwhere the paper describes this · verbatim
in the paper
5Inference
AI

Predict properties for held-out species and blends

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

Figure 2 shows the parity plots for the model’s independent test setwhere the paper describes this · verbatim
in the paper
6Validation
no AI

Compare predictions with measurements and baselines

Testing outputs against ground truth.

We compared the predictive model’s performance with (1) three data-driven models developed for predicting RON, MON, and YSI of pure componentswhere the paper describes this · verbatim
in the paper
7Optimisation
AI

Search latent space for mixtures matching targets

Iterative search over a space. The AI stood in for exhaustive search.

From the results, 20 mixtures with 5-26 components were reported using a full-scope search, whereas the greedy search generated 66 mixtureswhere the paper describes this · verbatim
in the paper
8Screening
no AI

Post-screen candidates on physical properties

Reducing a candidate set by filtering or ranking, in a single pass.

Five of 86 mixtures exhibited RVP in an acceptable range (50 kPa ≤ RVP ≤ 100 kPa).where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The designed fuel mixtures that the paper reports are produced by searching the deep-learning model's latent space; the candidate blends and their predicted properties exist only as model output

+What the AI was for
the AI fuel design tool was built on top of an end-to-end DL model based on recurrent and fully connected (FC) layerswhere the paper describes this · verbatim
+How it was taught
Supervisedin the paper
+Models named
Joint-properties predictive DL model (Extractor 1 LSTM encoder, Extractor 2 fully connected encoder, mixing operator, predictor) · Trained from scratchGNN model by Schweidtmann et al. (RON/MON/DCN baseline) · Off the shelfin the paper
+How results were checked
Held-outin the paper
MAEs were calculated for the ON predictions of 69 mixtures of varying sizes in the independent test set.where the paper describes this · verbatim
+Code · weights · data
code not reportedweights not reporteddata availablein the paper
Training and test datasets for pure components and mixtures are provided in Supplementary Data 2 and Supplementary Data 3.where the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 6 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of Joint-properties predictive DL model (Extractor 1 LSTM encoder, Extractor 2 fully connected encoder, mixing operator, predictor)Which version of the model was used is not stated.
  • Version of GNN model by Schweidtmann et al. (RON/MON/DCN baseline)Which version of the model was used is not stated.

About this article

Record aix-00058, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error