astronomy/ai produced the result/arXiv 2022 · v2
Neural network predicts gas pressure in galaxy clusters from dark matter alone
Researchers trained a neural network on the IllustrisTNG-300 simulation to predict the electron pressure field inside galaxy clusters directly from dark matter particles, standing in for a much costlier simulation of the gas physics.
spectrum · one line per step, placed by what the step does · bright lines used AI
Predicting the Thermal Sunyaev-Zel'dovich Field using Modular and Equivariant Set-Based Neural Networks
arXiv, 2022
doi:10.48550/arxiv.2203.00026 · record aix-00093 v2 · checked 2026-10-08
- AI was for
- Simulation surrogate
- Model family
- Multilayer perceptron, Autoencoder
- Checked by
- Held-out
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Galaxy clusters are the largest objects held together by gravity: hundreds or thousands of galaxies, plus a vast cloud of hot, thin gas. That gas leaves a mark on the afterglow of the Big Bang, because light passing through it picks up a small, characteristic change in energy. To predict that mark, astronomers need to know the pressure of the electrons in the gas throughout the cluster. The trouble is that dark matter, which supplies most of a cluster's mass, is far cheaper to simulate than gas. Gravity-only simulations are relatively quick; adding the gas, with its cooling, heating and outflows, is much more demanding.
The researchers set out to bridge that gap. Using two versions of the same simulation, one with gravity alone and one with the full gas physics, they asked whether the electron pressure from the expensive run could be predicted from the dark matter particles of the cheap one. They compared their predictions against a standard analytic formula, known as a GNFW profile, which describes a cluster's pressure as a smooth function of distance from its centre and was fitted to the same simulation data.
Where AI came in
The AI was the method that produced the result: a neural network trained from scratch on clusters from the simulation, with the electron pressure values from the full-physics run as its targets. It was built to read the dark matter particles as an unordered set, so the answer does not depend on the order the particles are listed in, and to respect rotation, so turning a cluster around turns the prediction with it. The network was split into parts with distinct jobs: one shifted the centre of the analytic profile, one read the particles near the point being predicted, one combined these with cluster-wide quantities, and one modelled the leftover scatter.
On held-out clusters never seen in training, the paper reports its loss measure improving by 70 per cent over the analytic profiles fitted to the same data, with a further 7 per cent from the part that models scatter; the authors describe that part as limited by their small training set. By adding and removing modules the researchers examined what the network relied on, concluding that a cluster's elongated shape mattered little and that local regions already carried enough information to infer cluster-wide properties. Training and prediction ran in PyTorch on a GPU, with hyperparameters searched using Optuna.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors trained a rotationally and permutation equivariant set-based neural network (a DeepSets architecture) on the IllustrisTNG-300 simulation to predict the continuous electron pressure field inside galaxy clusters directly from the dark matter particles of a gravity-only run, as a cheaper stand-in for a full hydrodynamic simulation. The architecture is split into separate modules: one that corrects mis-centering of an analytic GNFW profile, one that reads the local particle environment, an aggregator, and a conditional-VAE module that models stochasticity in the mapping. On the held-out clusters the paper reports its loss metric improving by 70% over analytic profiles fitted to the same simulation data, with a further 7% from the conditional-VAE extension, which the authors describe as limited by their small training set. Module-level tests indicated that cluster triaxiality had negligible impact and that local regions already carried enough information to infer global cluster properties.
How AI was used
Clusters above a mass cut were identified in the present-day gravity-only IllustrisTNG 300-1 snapshot with the Rockstar halo finder, and target electron pressure fields were voxelised from the matching full-physics run with Voxelize at a voxel size fixed in units of R200. Cluster-scale scalars (mass, offsets, angular momentum, inertia-tensor eigenvalues) and unit vectors (angular momentum, offset, inertia-tensor eigenvectors) were constructed as SO(3) scalars and vectors, self-similarly normalised, and perturbed with noise during training. A composed network was then fitted to the voxelised targets with a loss normalised by the characteristic pressure scale: a vector DeepSet (Origin) shifted the centre of an analytic GNFW profile, a scalar DeepSet (Local) produced features from particles within a radius of the target position, an Aggregator MLP with dropout combined these with target-position and cluster-scale information, and a conditional-VAE encoder (Stochastic) mapped shell-averaged residuals to a one-dimensional latent code. Particle sets were sparsely sampled during training as regularisation and sampled more densely at test time. Training used Adam with a one-cycle schedule, per-module learning rates, weight decay and gradient clipping, with annealed KL weighting for the stochastic variants; hyperparameters were searched with Optuna, using multi-objective optimisation and Pareto-frontier selection where the Stochastic module was present. The trained network was then run over all voxels of the held-out clusters, and individual modules were added or removed to study their contributions.
The shape of the work
Structural · the record, drawn
no AI
Identify clusters in the gravity-only snapshot
Obtaining raw data, whether by measurement, download or retrieval.
We use the state-of-the-art cluster finder code Rockstar to identify clusters with masses M200>5×1013M⊙/h in the gravity-only snapshot.where the paper describes this · verbatim
no AI
Voxelise electron pressure targets from the full-physics run
Cleaning, filtering, normalising or labelling data already obtained.
We produce electron pressure fields from the full-physics simulation using Voxelize, with a voxel sidelength of 5R200/64.where the paper describes this · verbatim
no AI
Build and normalise cluster-scale scalar and vector features
Encoding data into features, descriptors, embeddings or graphs.
Since our training set is relatively small, we find it crucial to add noise to the cluster-scale properties.where the paper describes this · verbatim
AI
Train the modular set-based network and its conditional-VAE module
Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.
Training is performed using the Adam optimizer and a one-cycle learning rate schedule.where the paper describes this · verbatim
AI
Search hyperparameters and select a model from the Pareto frontier
Iterative search over a space. Its result feeds back into an earlier step.
For hyperparameter searches we use the Optuna packagewhere the paper describes this · verbatim
AI
Predict the electron pressure field for held-out clusters
Running a trained model over new data to predict, classify or score. The AI stood in for simulation.
testing is of course performed on all available voxelswhere the paper describes this · verbatim
no AI
Compare predictions against targets and the GNFW benchmark
Testing outputs against ground truth.
Network losses evaluated on testing set and compared against the GNFW benchmark model.where the paper describes this · verbatim
AI
Ablate and add modules to interpret what the network uses
Extracting understanding from model behaviour.
we can separately study the influence of local and cluster-scale environment, determine that cluster triaxiality has negligible impactwhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is the predicted electron pressure field itself, produced by the trained network; without the network there is no result.
we choose to employ a rotationally equivariant DeepSets architecture to operate directly on the set of dark matter particleswhere the paper describes this · verbatim
The resulting 463 clusters are randomly assigned to training (70 %), validation (20 %), and testing (10 %) sets.where the paper describes this · verbatim
We make our code publicly available at this URL.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- DataWhether the data are available is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of Modular equivariant set-based network (Origin, Local and Aggregator modules)Which version of the model was used is not stated.
- Version of Conditional-VAE extension (Stochastic module)Which version of the model was used is not stated.
- Version of GNFW analytic electron pressure profile (benchmark, fitted to the same simulation data)Which version of the model was used is not stated.
- What step 5 replacedThe paper gives no basis for what the AI stood in for.
- What step 8 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00093, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error