astronomy/ai produced the result/arXiv 2025 · v2
Astronomers test pre-trained image models for sifting telescope alerts
Researchers compared off-the-shelf image-recognition networks, pre-trained on everyday photographs or on galaxy images, against a purpose-built network for deciding which nightly survey alerts are real transients. The models did the sifting and the outlier hunting themselves.
spectrum · one line per step, placed by what the step does · bright lines used AI
Pre-training vision models for the classification of alerts from wide-field time-domain surveys
arXiv, 2025
doi:10.48550/arxiv.2512.11957 · record aix-00100 v2 · checked 2026-10-08
- AI was for
- Classification, Anomaly detection
- Model family
- Convolutional neural network, Transformer, Multilayer perceptron
- Checked by
- Held-out
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Telescopes that scan the whole sky night after night look for things that change: stars that explode, objects that flare, anything that was not there before. These are called transients. Each night a survey such as the Zwicky Transient Facility sends out hundreds of thousands of alerts, small image cutouts flagging a spot of sky that has altered. Most are not real astronomy. They are artefacts of the camera, satellite trails, bad subtractions between a new image and an older reference one. Deciding which handful deserve a follow-up telescope has traditionally meant people looking at pictures, which does not scale with the alert rate.
The researchers set out to test whether general-purpose image-recognition networks, taken off the shelf and adapted, can do this vetting as well as a network built specially for the job. They rebuilt the training set used by the existing tool, BTSbot, from the survey's alert broker and its human inspectors' accept-and-reject labels: 769,056 alerts across 25,609 sources. They then asked how performance and running cost changed with how much training data was available.
Where AI came in
Two standard vision networks, ConvNeXt-pico and MaxViT-tiny, were each started three ways: from weights learned on ImageNet, a large collection of ordinary labelled photographs; from Zoobot, whose weights come from about 842,000 galaxy images annotated by volunteers; and from scratch, with random starting values. Pre-training means a network first learns general visual structure on one large picture collection, so that less data is needed for the real task. Each was then fine-tuned to answer a yes-or-no question about an alert, and compared with the retrained custom network. Some versions also took in numerical metadata alongside the images.
The trained models then scored held-out alerts, standing in for the nightly human inspection, and their speed and memory use were measured. Separately, the internal representations the networks had formed were squeezed down to two dimensions with UMAP, and an Isolation Forest, which looks for points sitting apart from the crowd, picked out unusual alerts. Astronomers then looked at those by eye and found rare sources plus two sources that had been labelled wrongly in the training set.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors compared vision models for vetting bright transient candidates in the Zwicky Transient Facility alert stream, using an updated BTSbot training set of 769,056 alerts across 25,609 sources. ConvNeXt-pico and MaxViT-tiny were fine-tuned from ImageNet-1k pre-training, from Zoobot weights pre-trained on about 842,000 annotated Galaxy Zoo images, and from random initialisation, and compared against a retrained custom CNN baseline across four training-set sizes. For image-only models, pre-training gave higher ROC AUC than training from scratch, and Galaxy Zoo pre-training generally gave higher ROC AUC than ImageNet pre-training, with the Galaxy Zoo MaxViT highest; for multi-modal models the baseline scored higher on trigger F1 at training sets below roughly 100,000 alerts. In the reported inference tests ConvNeXt was 6.4 times faster on CPU throughput and used 5.5 times less memory than the custom CNN, while MaxViT was 3.2 times slower and used 3.0 times more memory than it. UMAP embeddings combined with an Isolation Forest surfaced rare sources and two mislabelled sources in the training set.
How AI was used
The BTSbot training set was recompiled from the ZTF alert broker and marshal, and a parallel fine-tuning set of DESI Legacy Survey DR10 colour images was assembled at matching field of view and pixel scale, with reduced splits created by randomly keeping 5, 10 and 50 per cent of alerts. Two off-the-shelf architectures, ConvNeXt-pico and MaxViT-tiny, were initialised either from ImageNet-1k weights obtained from timm, from Zoobot weights pre-trained on Galaxy Zoo volunteer annotations and obtained from Hugging Face, or from PyTorch default initialisation; each had its head replaced with a randomly initialised three-layer MLP, and all backbone and head parameters were updated during fine-tuning on the binary transient-candidate vetting task. Multi-modal variants concatenated a metadata-branch embedding with the image embedding, and the custom CNN used by BTSbot was retrained as a baseline. A separate Bayesian hyperparameter sweep was run for each configuration on the Weights and Biases platform, after which five seeded trials were run with the best hyperparameters and scored on the test split. Trained models were then run over the test split to produce scores and to measure CPU and GPU throughput and peak memory. Penultimate-layer representations were extracted and reduced with UMAP at default settings, and an Isolation Forest was run on each labelled subset of each embedding to select candidate outliers for visual inspection.
The shape of the work
Structural · the record, drawn
no AI
Recompile BTSbot alert training set
Obtaining raw data, whether by measurement, download or retrieval.
The updated training set contains 769,056 alerts across 25,609 sourceswhere the paper describes this · verbatim
no AI
Build fine-tuning datasets and reduced-size splits
Cleaning, filtering, normalising or labelling data already obtained.
We create smaller versions of the train split by randomly selecting 5, 10, and 50% of alerts to preservewhere the paper describes this · verbatim
AI
Fine-tune off-the-shelf architectures and retrain baseline
Fitting model parameters, including fine-tuning an existing model.
During training, all weights and biases in the vision backbone and the MLP head are unfrozen and allowed to be updatedwhere the paper describes this · verbatim
AI
Score test-split alerts with trained models
Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.
the time it takes to compute predictions for each of them using a batch size of 32 and 32-bit floating point precisionwhere the paper describes this · verbatim
no AI
Evaluate against labels and baseline
Testing outputs against ground truth.
Model performance is evaluated using the area under the receiver operating characteristic curve (ROC AUC)where the paper describes this · verbatim
AI
Embed latent representations with UMAP
Encoding data into features, descriptors, embeddings or graphs.
We extract high-dimensional representations from the penultimate layer of the models and fit a UMAP transformation to reduce their dimensionalitywhere the paper describes this · verbatim
AI
Flag outliers with Isolation Forest
Reducing a candidate set by filtering or ranking, in a single pass. The AI stood in for manual curation.
We run an Isolation Forest independently on each subsetwhere the paper describes this · verbatim
no AI
Inspect flagged outliers
Extracting understanding from model behaviour.
We visually inspect these sources and identify a number of interesting outlierswhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The study's results are the classification performance and computational cost of the models themselves; the alert-vetting output is produced entirely by the trained models
We select two pre-training regimens to compare against training from scratch: pre-training on ImageNet and pre-training on Galaxy Zoowhere the paper describes this · verbatim
The performance metrics we report are the median and standard deviation for these final trials computed on the test splitwhere the paper describes this · verbatim
we release BTSbot models presented here publicly on the Hugging Face platform, along with open-source training code on GitHubwhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- DataWhether the data are available is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of BTSbot custom CNN baseline (uni-modal and multi-modal)Which version of the model was used is not stated.
- Version of UMAPWhich version of the model was used is not stated.
- Version of Isolation ForestWhich version of the model was used is not stated.
- What step 3 replacedThe paper gives no basis for what the AI stood in for.
- What step 6 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00100, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error