~/aixsci
200 records · all checked

astronomy/ai in a supporting role/Astronomy and Astrophysics 2022 · v2

Five codes measured galaxy shapes in simulated Euclid images, one using neural networks

The Euclid Morphology Challenge tested five pieces of software on about 1.5 million simulated galaxies. One of the five measured galaxy shapes with neural networks, and a generative model produced one set of the test images.

1. Generate input galaxy parameter catalogues2. Render analytic Sérsic galaxy fields3. Train generative galaxy image model4. Generate realistic galaxy images5. Train CNN morphology estimator6. Fit structural parameters with five codes7. Score codes against true parameters

spectrum · one line per step, placed by what the step does · bright lines used AI

Euclidpreparation
Astronomy and Astrophysics, 2022

doi:10.1051/0004-6361/202245042 · record aix-00132 v2 · checked 2026-10-09

ai-supportingrole of AI
AI was for
Property prediction, Simulation surrogate
Model family
Autoencoder, Normalising flow, Convolutional neural network
Checked by
Benchmark212000 tested
Code
available

AI processed or interpreted data, but the main finding does not rest on it.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Galaxies are not simply blobs of light. Each one has a shape: a bright central bulge, perhaps, surrounded by a flatter disc, and a certain degree of flattening depending on how it is tilted towards us. Astronomers summarise those shapes with a handful of numbers, such as the radius inside which half the light falls, the ratio of the short axis to the long one, and an index describing how steeply the brightness fades from the centre outwards. Measuring those numbers is hard because a telescope blurs every galaxy, the sky adds noise, and faint galaxies are only a smudge of a few pixels.

The Euclid space telescope is expected to image enormous numbers of galaxies, so the shape-measuring software has to work automatically and be trusted at scale. The Euclid Morphology Challenge set out to check that. Teams ran five codes on a common set of simulated galaxies whose true shapes were known in advance, and the predictions were scored against those known values using completeness, bias, scatter and the fraction of badly wrong answers, split up by how bright each galaxy was.

Where AI came in

Machine learning entered in two places. One of the five codes, DeepLeGATo, estimates a galaxy's shape numbers with convolutional neural networks, a kind of model that learns patterns directly from image pixels, rather than by fitting a mathematical light profile to the image as the other four codes do. It was trained on simulated galaxy images whose true shapes were known, with separate versions trained for brighter and fainter objects, and then applied to the challenge images.

The other use was in making the test images themselves. Most of the simulated galaxies were drawn from tidy mathematical formulas, which real galaxies do not obey. So a generative model, combining a variational auto-encoder with a normalising flow, was trained on Hubble Space Telescope pictures of real galaxies to learn what their light profiles actually look like. It was then asked to produce galaxy images matching a list of chosen shape numbers, standing in for a formula-based simulator and giving the codes a messier set of images to be tested on.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The Euclid Morphology Challenge compared five surface-brightness-fitting codes (DeepLeGATo, Galapagos-2, Morfometryka, ProFit and SourceXtractor++) on about 1.5 million simulated Euclid VIS galaxies, of which about 350 000 lie above 5σ. The simulated fields comprised analytic single-Sérsic and double-Sérsic profiles rendered with Galsim and a third set of non-analytic galaxies produced by a variational auto-encoder combined with a normalising flow trained on HST images. Measured against the known input parameters, all codes recovered single-Sérsic radius, axis ratio and Sérsic index with bias and dispersion below 10% down to about IE=23, bulge-disc decompositions were less stable, and results on the non-analytic galaxies were typically degraded by a factor of about 3. The authors report that all codes underestimate their own measurement uncertainties by a factor of around 2, and attribute much of the faint-end difference between codes to the priors or initial guesses each adopts.

How AI was used

Two learned components were used. First, a deep generative model built from a variational auto-encoder and a normalising flow (the FVAE) was trained on HST galaxy images to learn a latent representation of real galaxy light profiles, with the latent distribution conditioned on structural parameters and magnitude; it was then run over a parameter catalogue to render a field of non-analytic galaxy images, which were given the same PSF convolution and noise post-processing as the analytic Galsim fields. Second, one of the five participating morphology codes, DeepLeGATo, estimates structural parameters with convolutional neural networks rather than by fitting an analytic model; it was trained on simulated stamps with known true parameters, with separate weights fitted for objects brighter and fainter than magnitude 24.5, and then run over the challenge fields using fixed 64x64 pixel stamps. The remaining four codes are parametric or non-learned fitters. Predictions from all codes were pooled into a common catalogue and compared with the simulation inputs using completeness, bias, dispersion, outlier fraction and a weighted global score computed in bins of apparent magnitude; uncertainty calibration was assessed by checking how often the true value fell inside each code's reported 1σ interval.

The shape of the work

Structural · the record, drawn

ACQUISITIONSIMULATIONTRAININGGENERATIONTRAININGINFERENCEVALIDATION1234567AIAIAIAIGenerate inputgalaxy parametercataloguesRender analyticSérsic galaxyfieldsTrain generativegalaxy imagemodelGeneraterealistic galaxyimagesTrain CNNmorphologyestimatorFit structuralparameters withfive codesScore codesagainst trueparameters↤ simulation↤ simulation↤ conventional algorithm↤ conventional algorithm
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Generate input galaxy parameter catalogues

Obtaining raw data, whether by measurement, download or retrieval.

The input catalogues were created using the EGG simulator, which outputs a double-Sérsic components catalogue.where the paper describes this · verbatim
in the paper
2Simulation
no AI

Render analytic Sérsic galaxy fields

Numerical or physics simulation, including where a learned surrogate replaces it.

The galaxy images were then created using the Galsim software.where the paper describes this · verbatim
in the paper
3Training
AI

Train generative galaxy image model

Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.

Using HST images the model learns how to simulate real 2D noiseless galaxy profiles at a VIS-like resolution.where the paper describes this · verbatim
in the paper
4Generation
AI

Generate realistic galaxy images

Producing candidate objects that did not previously exist. The AI stood in for simulation.

The resulting architecture, called the Flow-Variational AutoEncoder (FVAE), can therefore simulate galaxies directly from a catalogue of parameterswhere the paper describes this · verbatim
in the paper
5Training
AI

Train CNN morphology estimator

Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.

the DeepLeGATo algorithm was trained separately for two sets of objects, objects fainter and brighter than magnitude 24.5where the paper describes this · verbatim
in the paper
6Inference
AI

Fit structural parameters with five codes

Running a trained model over new data to predict, classify or score. The AI stood in for conventional algorithm.

All of them fitted at least a single profile to each galaxy in the IE bandwhere the paper describes this · verbatim
in the paper
7Validation
no AI

Score codes against true parameters

Testing outputs against ground truth.

we use four main indicators to evaluate and compare the different codes: completeness (𝒞); bias (ℬ); dispersion (𝒟); and outlier fraction (𝒪)where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI in a supporting roleour reading

Four of the five evaluated fitting codes are non-learned parametric fitters, and the headline conclusion about single-Sérsic accuracy rests on analytic Galsim simulations; learned models enter as one of the evaluated codes (DeepLeGATo) and as the generator of the 'realistic' simulation set. A reviewer may prefer 'instrument', since the conclusions about non-analytic profiles depend on the neural-network-generated images.

+What the AI was for
DeepLeGATo bases its photometric galaxy profile modelling on convolutional neural networks.where the paper describes this · verbatim
+How it was taught
Self-supervisedSupervisedin the paper
+Models named
Flow-Variational AutoEncoder (FVAE) · Trained from scratchDeepLeGATo · Trained from scratchin the paper
+How results were checked
Benchmark212000 testedin the paper
Each team tested the performance of their codes on a common set of simulated Euclid galaxies that was provided to them.where the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata not reportedin the paper
We provide information about the precise version of the software packages used for the EMC, along with the configuration files and READMEswhere the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 5 items
  • Trained model weightsWhether the trained model is available is not stated.
  • DataWhether the data are available is not stated.
  • ComputeThe hardware or time used is not stated.
  • Version of Flow-Variational AutoEncoder (FVAE)Which version of the model was used is not stated.
  • Version of DeepLeGAToWhich version of the model was used is not stated.

About this article

Record aix-00132, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error