~/aixsci
200 records · all checked

astronomy/ai produced the result/Monthly Notices of the Royal Astronomical Society 2023 · v2

Neural networks weigh cluster X-ray maps to estimate galaxy cluster masses

Researchers trained convolutional neural networks on simulated eROSITA X-ray images of 3285 galaxy clusters to predict cluster mass, then used saliency maps to see which pixels the networks relied on.

1. Generate mock eROSITA observations2. Build redshift-normalised photon maps3. Encode spectroscopic galaxies as phase-space cubes4. Fit scalar proxy baselines and search combinations5. Train CNN mass estimators6. Predict masses on held-out folds7. Saliency-based interpretability study

spectrum · one line per step, placed by what the step does · bright lines used AI

Benchmarks and explanations for deep learning estimates of X-ray galaxy cluster masses
Monthly Notices of the Royal Astronomical Society, 2023

doi:10.1093/mnras/stad2005 · record aix-00223 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Property prediction
Model family
Convolutional neural network
Checked by
Held-out3285 tested
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Galaxy clusters are the largest gravitationally bound objects in the universe, holding hundreds of galaxies in a vast cloud of hot gas that glows in X-rays. Most of a cluster's mass is dark matter, which cannot be seen directly, so its mass has to be inferred from what can be seen. The usual approach reduces an X-ray image to a single number, such as total brightness or the count of X-ray photons, and feeds that into a scaling relation fitted beforehand. Clusters are messy and individual, though, so such single numbers predict mass only roughly, leaving a spread between the estimate and the truth.

The researchers built mock observations of 3285 clusters drawn from the Magneticum cosmological simulation, where the true mass of every cluster is known. The mocks were made to resemble what the eROSITA X-ray telescope would actually record, including background emission, the instrument's response, its blurring of fine detail, and contaminating light from active black holes. They also produced matching mock galaxy catalogues describing how cluster galaxies move. The aim was to compare mass estimates made from whole images against the conventional single-number proxies, and then to ask what the image-based estimators were paying attention to.

Where AI came in

Convolutional neural networks, which learn directly from images rather than from summary numbers, were trained from scratch on the mock X-ray maps with the simulation's true masses as the answer key. Each network was tested on clusters it had not been trained on. From single-band maps the spread in mass estimates was 17.8 per cent, against about 26 per cent for an idealised brightness proxy and about 34 per cent for the best realistic single-number proxy. Splitting the images into soft, medium and hard X-ray bands gave 16.2 per cent, and adding the galaxy motion data gave 15.9 per cent. The networks stood in for the fitted scaling relations.

The researchers then examined the trained networks with saliency maps, which measure how much each pixel of an image nudges the predicted mass. The networks gave much less weight to the crowded cluster centre than the photon-count scaling relation does; near the centre, the proxy's importance was about ten times the networks'. Pixels dominated by light from active black holes, which is not part of the cluster gas, carried lower importance than pixels dominated by the hot gas itself.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

Convolutional neural networks were trained to predict galaxy cluster masses (M500c) from mock eROSITA X-ray photon maps built on the Magneticum hydrodynamical simulation, which include background emission, instrument response, point spread function and AGN contamination. Single-band bolometric photon maps gave a predictive mass scatter of 17.8%, compared with about 26% for the idealised bolometric X-ray luminosity proxy and about 34% for the best realistic scalar proxy, the photon count within R500c; splitting the maps into soft, medium and hard energy bands reduced the scatter to 16.2%, and adding mock spectroscopic galaxy dynamics gave 15.9%. A gradient-based saliency analysis found that the networks assign much lower per-pixel importance to cluster centres than the photon-count scaling relation does, with the median proxy saliency about ten times the CNN saliency at 0.07 R500c, and assign AGN-dominated pixels lower median saliency than ICM-dominated pixels.

How AI was used

Mock eROSITA observations of 3285 Magneticum clusters were produced with PHOX for intra-cluster medium emission, a projection of simulated AGN sources, and SIXTE for background, instrument response and point spread function. Photon lists were binned into 128x128 maps over a 2.3 h-1 Mpc aperture, normalised for redshift and luminosity distance and log-scaled, in both a single 0.5-10 keV band and three separate soft, medium and hard bands; matching mock spectroscopic galaxy catalogues were smoothed with Gaussian kernel density estimators and sampled onto 64x64x64 phase-space grids. A feed-forward convolutional network of four convolutional layers, two max-pooling layers and four dense layers was trained from scratch with a mean squared error loss on logarithmic M500c, using the Adam optimiser at learning rate 1e-3, ReLU activations, L2 regularisation, a 90/10 training-validation split and early stopping after 20 epochs without improvement. The same architecture was applied to three-channel multi-band inputs, and a joint model combined a 2D convolutional X-ray extractor with a 3D convolutional dynamical extractor feeding a shared dense network trained end to end. Cross-validation folds were assigned by splitting the simulation volume into eight equal quadrants to avoid shared formation environments. Scalar proxy baselines were fitted with a power-law multi-property scaling and covariance model, with an exhaustive search over observable combinations. Trained networks were then interrogated with absolute-gradient saliency maps, compared against a closed-form saliency derived for the photon-count proxy, binned radially, and averaged over binary masks separating ICM-dominated from AGN-dominated pixels.

The shape of the work

Structural · the record, drawn

SIMULATIONPREPARATIONREPRESENTATIONTRAININGTRAININGINFERENCEINTERPRETATION1234567AIAIAIGenerate mockeROSITAobservationsBuildredshift-normalisedphoton mapsEncodespectroscopicgalaxies as phas…Fit scalar proxybaselines andsearch combinati…Train CNN massestimatorsPredict masses onheld-out foldsSaliency-basedinterpretabilitystudy↤ statistical model↤ statistical model
AI stepNo AI↤ what the AI stood in for
1Simulation
no AI

Generate mock eROSITA observations

Numerical or physics simulation, including where a learned surrogate replaces it.

we have generated realisations of background emission, instrument response, and point spread function that are consistent with the eROSITA telescope designwhere the paper describes this · verbatim
in the paper
2Preparation
no AI

Build redshift-normalised photon maps

Cleaning, filtering, normalising or labelling data already obtained.

we apply a logarithmic scaling to the images before they are passed as input to the neural networkwhere the paper describes this · verbatim
in the paper
3Representation
no AI

Encode spectroscopic galaxies as phase-space cubes

Encoding data into features, descriptors, embeddings or graphs.

we use kernel density estimators to estimate the distribution of galaxies in the 3D dynamical phase spacewhere the paper describes this · verbatim
in the paper
4Training
no AI

Fit scalar proxy baselines and search combinations

Fitting model parameters, including fine-tuning an existing model.

we measure the predictive variance of each individual observable in our baseline suite and list them in Table 3where the paper describes this · verbatim
in the paper
5Training
AI

Train CNN mass estimators

Fitting model parameters, including fine-tuning an existing model. The AI stood in for statistical model.

Eight independent models are then each trained on seven folds (∼85% of the data) and tested on the last eighth fold.where the paper describes this · verbatim
in the paper
6Inference
AI

Predict masses on held-out folds

Running a trained model over new data to predict, classify or score. The AI stood in for statistical model.

Figure 4 shows the distribution of mass predictions made by our multi-band model on the independent test set.where the paper describes this · verbatim
in the paper
7Interpretation
AI

Saliency-based interpretability study

Extracting understanding from model behaviour.

The specific method we use is a gradient-based saliency map, which measures the pixel-wise sensitivity to model outputswhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's reported result is the mass scatter achieved by the CNN estimators themselves, together with an interpretability study of those same models; the finding does not exist without the networks.

+What the AI was for
The specific neural network model used to process X-ray images for this task is a convolutional neural network.where the paper describes this · verbatim
+Model families
+How it was taught
Supervisedin the paper
+Models named
single-band X-ray CNN · Trained from scratchmulti-band X-ray CNN · Trained from scratchjoint X-ray and dynamical CNN (2D plus 3D convolutional extractors) · Trained from scratchpower-law multi-property scalar proxy model · Trained from scratchin the paper
+How results were checked
Held-out3285 testedin the paper
We use an eight-fold cross-validation procedure to evaluate the predictive performance of these models on our mock catalogue.where the paper describes this · verbatim
+Code · weights · data
code availableweights availabledata availablein the paper
All code for the data analysis and machine learning models investigated in this analysis has been made available on Github.where the paper describes this · verbatim
+Compute
Training performed on a single Nvidia Ampere A100 GPU, with the 3D convolutional model taking about 15 times longer to train than the 2D convolutional models; computing resources provided by the Pittsburgh Supercomputing Center, and the X-ray mock data set generated at MARCC.in the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 5 items
  • Version of single-band X-ray CNNWhich version of the model was used is not stated.
  • Version of multi-band X-ray CNNWhich version of the model was used is not stated.
  • Version of joint X-ray and dynamical CNN (2D plus 3D convolutional extractors)Which version of the model was used is not stated.
  • Version of power-law multi-property scalar proxy modelWhich version of the model was used is not stated.
  • What step 7 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00223, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error