~/aixsci
200 records · all checked

astronomy/ai produced the result/arXiv 2024 · v2

Neural networks sort supernovae and estimate their distances from brightness alone

A thesis built machine learning tools for supernova cosmology using only brightness measurements. Convolutional networks classified supernova types and predicted redshift distributions, standing in for the spectra and expert judgement normally needed.

1. Simulate supernova lightcurves and host catalogue2. Apply quality and lightcurve-fit selection cuts3. Encode lightcurves as wavelength-time heatmaps4. Train classification and redshift CNNs5. Predict SN types and redshift PDFs6. Match supernovae to host galaxies and estimate host photo-z7. Fit cosmological parameters and quantify mismatch bias8. Pretrain with generic augmentations and fine-tune with targeted augmentations

spectrum · one line per step, placed by what the step does · bright lines used AI

Towards Precision Photometric Type Ia Supernova Cosmology with Machine Learning
arXiv, 2024

doi:10.48550/arxiv.2406.04529 · record aix-00129 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Classification, Property prediction, Denoising
Model family
Convolutional neural network, Gaussian process, Transformer, Recurrent neural network, Multilayer perceptron, Clustering, Random forest
Checked by
Benchmark4057 tested
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Exploding stars called type Ia supernovae are used to measure how the universe expands. They brighten and fade in a way astronomers understand well enough to work out how far away they are, so a large collection of them can reveal how dark energy behaves. The trouble is telling them apart from other kinds of exploding star. The reliable way is to take a spectrum, spreading the light into its colours, but spectra are slow and expensive and new surveys will find far more supernovae than can be followed up this way. The same problem applies to redshift, the stretching of light that tells you distance.

The work set out to do both jobs from photometry alone, meaning repeated brightness measurements through a handful of colour filters. That gives a lightcurve: a sparse, noisy record of how bright the object was, in each filter, on each night it was observed. The aim was to classify supernovae and estimate their redshifts from such lightcurves, and then to check what errors in those steps do to the cosmology drawn from them.

Where AI came in

Lightcurves were first turned into something a network could read. A Gaussian process, a statistical way of filling gaps smoothly while tracking its own uncertainty, was fitted across wavelength and time, then sampled on a grid to make a pair of images: one of brightness, one of uncertainty. Convolutional neural networks, the image-recognition architecture, were trained from scratch on these images. SCONE separated type Ia from other supernovae and also attempted a six-way typing; Photo-zSNthesis produced a full probability distribution for redshift. A self-organizing map assigned photometric redshifts to catalogue galaxies when matching supernovae to their host galaxies.

In place of a spectrum and an expert eye, the networks returned a probability for each object. Those probabilities then fed the statistical fit for the cosmological parameters. Training used simulated supernovae with known answers, and the models were also run on real observations and on samples deliberately unlike the training data. A separate part of the thesis, Connect Later, pretrained models on unlabelled data and then fine-tuned them with augmentations chosen for the expected mismatch, tested on astronomical, wildlife and medical-image benchmarks.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

This thesis develops machine learning tools for supernova cosmology from photometry alone. SCONE, a convolutional neural network applied to Gaussian-process-interpolated wavelength-time heatmaps of supernova lightcurves, reached 99.73±0.26% test accuracy separating simulated type Ia from non-Ia supernovae and 75% accuracy on 6-way typing at the date of trigger with redshift information. Photo-zSNthesis, a related CNN, predicts full redshift probability distributions from lightcurves and was compared against LCFIT+Z on simulated LSST and SDSS data and on 489 observed SDSS supernovae. A study of directional-light-radius host galaxy matching on DES-SN5YR-like simulations found 1.7% of simulated supernovae matched to the wrong host and a shift in the dark energy equation of state parameter of Δw=0.0013±0.0026 with a CMB prior, and the Connect Later framework fine-tunes pretrained models with targeted augmentations on astronomical, wildlife and histopathology benchmarks.

How AI was used

Simulated supernova lightcurves were generated with SNANA following PLAsTiCC, SDSS and DES survey models, passed through quality and SALT lightcurve-fit selection cuts, and class-balanced into training, validation and test splits. Each lightcurve was encoded by fitting a two-dimensional Gaussian process in wavelength and time with a Matérn 3/2 kernel at a fixed 6000 Å wavelength length scale, fitting the time length scale, and sampling the fitted model on a 32×180 grid to produce stacked flux and uncertainty heatmaps normalised to [0,1]. Convolutional networks with full-height kernels were trained from scratch on these heatmaps: SCONE with binary cross-entropy for Ia versus non-Ia and sparse categorical cross-entropy for 6-way typing, with a variant concatenating redshift and redshift error into the fully connected classifier; Photo-zSNthesis with residual blocks and a softmax layer over discretised redshift bins, calibrated afterwards by temperature scaling. Trained models were run over held-out splits, truncated early-time lightcurves, bright subsets, out-of-distribution spectroscopic-like samples and observed SDSS and DES lightcurves. In the host-mismatch study, a self-organizing map trained on griz fluxes assigned photometric redshifts to catalogue galaxies, the directional light radius method matched supernovae to hosts in both data and simulations, and SuperNNova and SCONE supplied Ia probabilities that weighted the BEAMS likelihood before cosmology fitting with wfit. Connect Later pretrained an Informer encoder by masked autoencoding of lightcurve observations and SwAV-pretrained ResNet-50 and DenseNet121 image models, then fine-tuned with linear probing followed by fine-tuning using targeted augmentations that resample each object to a new redshift or randomise image background and stain colour.

The shape of the work

Structural · the record, drawn

SIMULATIONPREPARATIONREPRESENTATIONTRAININGINFERENCEPREPARATIONINTERPRETATIONTRAINING12345678AIAIAIAIAISimulatesupernovalightcurves and …Apply quality andlightcurve-fitselection cutsEncodelightcurves aswavelength-time …Trainclassificationand redshift CNNsPredict SN typesand redshift PDFsMatch supernovaeto host galaxiesand estimate hos…Fit cosmologicalparameters andquantify mismatc…Pretrain withgenericaugmentations an…↤ conventional algorithm↤ expert judgement↤ expert judgement↤ statistical model↤ conventional algorithm
AI stepNo AI↤ what the AI stood in for
1Simulation
no AI

Simulate supernova lightcurves and host catalogue

Numerical or physics simulation, including where a learned surrogate replaces it.

All simulations for this work are produced with the SuperNova ANAlysis (SNANA) software.where the paper describes this · verbatim
in the paper
2Preparation
no AI

Apply quality and lightcurve-fit selection cuts

Cleaning, filtering, normalising or labelling data already obtained.

In order to ensure that the model is learning only from high-quality information, we have instituted some additional quality-based cutswhere the paper describes this · verbatim
in the paper
3Representation
AI

Encode lightcurves as wavelength-time heatmaps

Encoding data into features, descriptors, embeddings or graphs. The AI stood in for conventional algorithm.

we use the approach described by to apply 2-dimensional Gaussian process regression to the raw lightcurve datawhere the paper describes this · verbatim
in the paper
4Training
AI

Train classification and redshift CNNs

Fitting model parameters, including fine-tuning an existing model. The AI stood in for expert judgement.

Both classification modes use the Adam optimizer at a constant 1e-3 learning rate for 400 epochs.where the paper describes this · verbatim
in the paper
5Inference
AI

Predict SN types and redshift PDFs

Running a trained model over new data to predict, classify or score. The AI stood in for expert judgement.

Each classifier outputs PIa values, the predicted probability of each SN to be a type Ia.where the paper describes this · verbatim
in the paper
6Preparation
AI

Match supernovae to host galaxies and estimate host photo-z

Cleaning, filtering, normalising or labelling data already obtained. The AI stood in for statistical model.

we train a Self-Organizing Map (SOM) to characterize and discretize the photometric space of host galaxieswhere the paper describes this · verbatim
in the paper
7Interpretation
no AI

Fit cosmological parameters and quantify mismatch bias

Extracting understanding from model behaviour.

We fit for w and Ωm using wfit, a fast cosmology grid-search program in SNANAwhere the paper describes this · verbatim
in the paper
8Training
AI

Pretrain with generic augmentations and fine-tune with targeted augmentations

Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.

after pretraining with generic augmentations, fine-tune with targeted augmentations designed with knowledge of the distribution shiftwhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

Chapters 2, 3, 5 and 6 report classification, redshift estimation and robustness results that are themselves the findings; the models produce the results the thesis is about

+What the AI was for
we present a machine learning method for photometric classification of SNe, Supernova Classification with a COnvolutional Neural Network (SCONE)where the paper describes this · verbatim
+How it was taught
SupervisedSelf-supervisedUnsupervisedSemi-supervisedTransfer / fine-tuningin the paper
+Models named
SCONE · Trained from scratchSCONE with redshift · Trained from scratchPhoto-zSNthesis · Trained from scratch2D Gaussian process regression (Matern 3/2 kernel) · Trained from scratchSelf-Organizing Map (host galaxy photo-z) · Trained from scratchSuperNNova (SNN+Z and SNN-NoZ) · Trained from scratchMulti-layer perceptron baseline · Trained from scratchInformer encoder (masked autoencoding pretraining, Connect Later) · Fine-tunedResNet-50 pretrained with SwAV on ImageNet · Fine-tunedDenseNet121 pretrained with SwAV on unlabeled Camelyon17-WILDS · Fine-tunedCIGALE (host galaxy stellar mass and SFR fits) · Off the shelfin the paper
+How results were checked
Benchmark4057 testedin the paper
We evaluate our model and our baseline for comparison, LCFIT+Z, on a test set of 4,057 simulated PLAsTiCC-like lightcurveswhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
The documented source code has been released on Github (github.com/helenqu/scone) to ensure reproducibilitywhere the paper describes this · verbatim
+Compute
First training epoch on one NVIDIA V100 Volta GPU takes approximately 2 seconds; subsequent epochs approximately 1 second; Haswell node epochs approximately 26 secondsin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 12 items
  • Trained model weightsWhether the trained model is available is not stated.
  • Version of SCONEWhich version of the model was used is not stated.
  • Version of SCONE with redshiftWhich version of the model was used is not stated.
  • Version of Photo-zSNthesisWhich version of the model was used is not stated.
  • Version of 2D Gaussian process regression (Matern 3/2 kernel)Which version of the model was used is not stated.
  • Version of Self-Organizing Map (host galaxy photo-z)Which version of the model was used is not stated.
  • Version of SuperNNova (SNN+Z and SNN-NoZ)Which version of the model was used is not stated.
  • Version of Multi-layer perceptron baselineWhich version of the model was used is not stated.
  • Version of Informer encoder (masked autoencoding pretraining, Connect Later)Which version of the model was used is not stated.
  • Version of ResNet-50 pretrained with SwAV on ImageNetWhich version of the model was used is not stated.
  • Version of DenseNet121 pretrained with SwAV on unlabeled Camelyon17-WILDSWhich version of the model was used is not stated.
  • Version of CIGALE (host galaxy stellar mass and SFR fits)Which version of the model was used is not stated.

About this article

Record aix-00129, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error