materials-chemistry/ai produced the result/Journal of Cheminformatics 2023 · v2
Neural network predicts crystal density of candidate explosives from molecular structure
Researchers trained a graph-based neural network, FFiTrNet, on 12,072 compounds to predict how densely a molecule packs in its crystal. The model produced every density figure reported for the test compounds, in place of measurement or quantum-chemical calculation.
spectrum · one line per step, placed by what the step does · bright lines used AI
Force field-inspired transformer network assisted crystal density prediction for energetic materials
Journal of Cheminformatics, 2023
doi:10.1186/s13321-023-00736-6 · record aix-00155 v2 · checked 2026-10-09
- AI was for
- Property prediction
- Model family
- Transformer, Graph neural network, Random forest
- Checked by
- Held-out109 tested
- Code
- available
The finding the paper is about came from the AI.
What this research was about

Energetic materials are the chemicals used in explosives and propellants. How well one performs depends heavily on its crystal density: how much mass is squeezed into a given volume once the molecules settle into a regular, repeating solid. A denser crystal packs more stored energy into the same space. The trouble is that density is a property of the packed solid, not of the single molecule drawn on paper. Finding it normally means growing a crystal and measuring it, or running demanding quantum-chemical calculations. Neither is quick, and for a compound nobody has yet made, neither is straightforward.
The researchers set out to predict crystal density from molecular structure alone. They assembled a set of 12,072 compounds made only of carbon, hydrogen, oxygen and nitrogen, drawn from the Cambridge Structural Database, each with a text description of its structure and a known crystal density. They then built and trained a model on this set and tested it on a separate collection of 109 proposed energetic materials compiled by Huang and Massa.
Where AI came in
The AI did the predicting. Each compound's structure was turned into a three-dimensional arrangement of atoms using the chemistry toolkit RDKit, and the distances, angles and twists between nearby atoms were converted into quantities resembling the energy terms used in classical models of molecular forces. Those quantities told the network which atoms to pay attention to when judging a molecule. A transformer layer, the component that weighs each part of an input against the others, combined these views into a single prediction of density. The model, called FFiTrNet, was trained from scratch, alongside four comparison models including a simpler neural network and a random forest.
Every density figure the paper reports for the held-out test compounds and for the Huang and Massa set came from one of these models rather than from an experiment or a simulation. On the 109 proposed energetic materials, FFiTrNet gave a mean absolute error of 0.0489 g/cm, against 0.0446 g/cm for the densest region of its own test split; the authors link part of the gap to fluorine-containing molecules, which were not in the training data. Errors were also broken down by compound family to see which kinds of structure the model handled less accurately.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record

The authors assembled a dataset of 12,072 CHON compounds with SMILES strings and crystal densities from the Cambridge Structural Database, generated 3D conformers with RDKit, and trained a 3D-aware graph network, FFiTrNet, that replaces the axial attention of the earlier FFiNet with a Transformer encoder over force-field-inspired energy terms. On a 0.8:0.1:0.1 split, FFiTrNet reported the lowest MAE and RMSE of the five models tested in the density regions above 1.8 g/cm and between 1.6 and 1.8 g/cm, while D-MPNN had lower errors in the two lower-density regions. Applied to the 109-compound Huang & Massa set of putative energetic materials, FFiTrNet gave an MAE of 0.0489 g/cm, compared with 0.0446 g/cm in the above-1.8 g/cm region of the CSD test split; the authors attribute part of the difference to fluorine-containing molecules absent from training.
How AI was used
A supervised regression model predicted crystal density from molecular structure alone. SMILES strings curated from the CSD were converted to 3D conformers with RDKit's ETKDG method, per-atom features were extracted, and interatomic distances, angles and dihedrals over 1-, 2- and 3-hop neighbourhoods were converted into force-field-style energy terms used as attention scores. The 1-, 2- and 3-hop attention outputs, stacked with a learnable output token, were passed through a single Transformer encoder layer in place of the original model's axial attention, and the token's representation was used for the density output. FFiTrNet and FFiNet were trained from scratch on the curated dataset alongside GATv2, D-MPNN and a random forest on 208 RDKit descriptors, with the same training strategy, hyperparameter optimisation and three independent runs per model. The trained models were then run over the held-out test split and, as a separate out-of-distribution test, over the Huang & Massa dataset after removing overlapping training points, with errors broken down by density region and by compound family.
The shape of the work
Structural · the record, drawn
no AI
Curate CHON crystal dataset from CSD
Obtaining raw data, whether by measurement, download or retrieval.
we established a dataset with 12,072 compounds containing CHON elements with their Simplified Molecular-Input Line-Entry System (SMILES) strings and crystal densitywhere the paper describes this · verbatim
no AI
Generate 3D conformers and atom features
Encoding data into features, descriptors, embeddings or graphs.
The fast ETKDG method from RDkit is applied to generate atom positions.where the paper describes this · verbatim
AI
Train FFiTrNet and baseline models
Fitting model parameters, including fine-tuning an existing model.
New 3D-aware GNNs models FFiNet and its upgraded version FFiTrNet are then trained and tested in this CSD curated dataset.where the paper describes this · verbatim
AI
Predict densities for held-out test split
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
randomly splitting the data into training, validation, and test dataset with a ratio of 0.8:0.1:0.1where the paper describes this · verbatim
AI
Predict densities for Huang & Massa compounds
Running a trained model over new data to predict, classify or score. The AI stood in for simulation.
we use this pretrained model to predict the potential energetic materials dataset: Huang & Massa datasetwhere the paper describes this · verbatim
no AI
Evaluate errors overall and per density region
Testing outputs against ground truth.
We adopt three different metrics to evaluate the regression model, like mean absolute error (MAE), root mean square error (RMSE)where the paper describes this · verbatim
no AI
Analyse error by molecular family
Extracting understanding from model behaviour.
By listing out the MAE of each group, we can further investigate the relationship between molecular structure and model accuracy.where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is the predictive performance of a trained neural network; every reported density value for the test and Huang & Massa sets comes from a model.
the self-attention mechanism from Transformer is used to replace the axial attention in original modelwhere the paper describes this · verbatim
we use another small dataset from Huang & Massa, who obtain explosive properties against 109 putative energetic materialswhere the paper describes this · verbatim
Source code and dataset is available at GitHub page: https://github.com/jjx-2000/FFiTrNet.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- Version of FFiTrNet (force field-inspired Transformer network)Which version of the model was used is not stated.
- Version of FFiNet (force field-inspired neural network)Which version of the model was used is not stated.
- Version of GATv2Which version of the model was used is not stated.
- Version of D-MPNN (Directed Message Passing Neural Network)Which version of the model was used is not stated.
- Version of Random forest on RDKit descriptorsWhich version of the model was used is not stated.
- What step 3 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00155, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error