astronomy/ai in a supporting role/The Astronomical Journal 2025 · v2
Astronomers hunt dust-buried young star clusters in eleven nearby galaxies
A team combed JWST and Hubble images of 11 nearby star-forming galaxies by eye, finding 292 candidate star clusters still wrapped in their birth dust. That human catalogue was then used to train image-recognition networks to pick out similar objects.
spectrum · one line per step, placed by what the step does · bright lines used AI
PAH Marks the Spot: Digging for Buried Clusters in Nearby Star-forming Galaxies
The Astronomical Journal, 2025
doi:10.3847/1538-3881/ae10ad · record aix-00041 v2 · checked 2026-10-08
- AI was for
- Classification
- Model family
- Convolutional neural network
- Checked by
- Held-out862 tested
- Code
- not reported
AI processed or interpreted data, but the main finding does not rest on it.
What this research was about
Stars are born in dense clouds of gas and dust, and for their first few million years the youngest clusters stay hidden inside that cloud. Dust blocks visible light, so a cluster can be bright and yet leave almost no trace in an ordinary optical image. Infrared light passes through dust more easily, which is why telescopes such as JWST can see into these cocoons. Useful signposts include emission at a wavelength of 3.3 micrometres from large carbon-rich molecules known as polycyclic aromatic hydrocarbons, or PAHs, which glow in warm, irradiated gas, and the relative strength of two hydrogen emission lines, which indicates how much dust lies in front.
The researchers set out to assemble a catalogue of these deeply embedded clusters across 11 nearby star-forming galaxies, using JWST infrared imaging alongside Hubble ultraviolet and optical maps. Sixteen experts inspected image cutouts on the Zooniverse platform, looking for compact PAH emission, a high hydrogen line ratio and no detectable optical counterpart. That produced 292 candidates. Comparing their brightnesses with stellar population models gave a median age of 4.5 million years, an average dust extinction of 6.0 magnitudes along the line of sight, and a median lower limit on cluster mass of about a thousand times the mass of the Sun.
Where AI came in
The AI came in after the catalogue was complete. The 292 by-eye candidates were matched against a general list of PAH-bright peaks, keeping 233 as positive examples, and 4,779 other PAH-bright sources were drawn as negatives. Small image cutouts of each object, stacked across filters in four configurations, were fed to two convolutional neural networks — ResNet18 and VGG19-bn, both pattern-recognition models first trained on everyday photographs and then adjusted on this astronomical data. Eighty per cent of the examples were used for training and the rest held back for testing, with ten models trained per configuration and their majority vote taken as the answer.
In effect the networks stood in for the human eye, attempting the sorting job the 16 experts had done by hand. On the 862 held-back objects, the models agreed with the human labels on 805 negatives and 24 positives across all four configurations, with 27 false positives and 6 false negatives. Accuracy reached 95 per cent for the plentiful PAH-bright class, but the score for the rarer embedded clusters averaged 0.36 in the configurations lacking certain filters; the authors put this down to the lopsided numbers and to the two classes looking alike. The study's reported findings rest on the human catalogue, not the networks.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors searched JWST and HST imaging of 11 nearby star-forming galaxies for deeply embedded young star clusters, using a Zooniverse by-eye inspection by a team of 16 experts to identify 292 candidates on the basis of compact 3.3 micron PAH emission, a high Pa alpha/H alpha ratio and no detectable broadband optical source. Comparing photometry to stellar population models gave a median age of 4.5 Myr, an average line-of-sight extinction of AV = 6.0 mag and a median stellar mass lower limit of 10^3 solar masses. That by-eye catalogue was then used to train ImageNet-pretrained ResNet18 and VGG19-bn convolutional neural networks to separate embedded cluster candidates from other PAH-bright sources, across four filter configurations. On the held-out test split the models reached up to 95% accuracy for the majority PAH-bright class, while F1 scores for the embedded cluster class averaged 0.36 for the configurations without Pa alpha and F150W data, which the authors attribute to class imbalance and feature overlap.
How AI was used
The 292 human-identified candidates were cross-matched against a general catalogue of F335M PAH peaks, retaining 233 as the positive class (Label 1), and other PAH-bright peaks were randomly sampled at a 20:1 ratio to form 4779 negatives (Label 0), with source extraction performed separately for galaxy centres and outer regions to limit spatial bias. For every object, 299x299 pixel cutouts in each filter were stacked into multiextension FITS files in four filter configurations reflecting PHANGS Cycle 1 and Cycle 2 data availability and the presence or absence of F770W; 80% of each label was used for training and 20% reserved for testing. Two deep CNNs, VGG19-bn and ResNet18, were used with ImageNet 'IMAGENET1K_V1' pretrained weights, with one three-channel model trained per three image layers and the models concatenated into a single classifier; unfilled layers in the last model were set to zero arrays. Inputs were cropped to the central 50x50 pixels and resized back to 299x299, with random rotation between 0 and 360 degrees and alternating vertical flips as augmentation. Training used cross-entropy loss, an Adam optimiser and a fixed learning rate of 10^-4, with batch size 32 for 4000 batches for ResNet18 and batch size 14 for 3000 batches for VGG19-bn, on GPUs at the University of Wyoming. No class weights were applied. Ten models were trained per configuration and the mode of their predictions taken as each test object's predicted label, with performance reported through confusion matrices, F1 scores and model agreement fractions.
The shape of the work
Structural · the record, drawn
no AI
Assemble and align multiband imaging
Cleaning, filtering, normalising or labelling data already obtained.
We first resample all data to the NIRCam F335M grid using the reproject_exact function from Astropy.where the paper describes this · verbatim
no AI
By-eye candidate identification on Zooniverse
Obtaining raw data, whether by measurement, download or retrieval.
we generated 3402 tags across our 2674 boxed regions for the 11 Cycle 2 galaxies in our samplewhere the paper describes this · verbatim
no AI
Group tags and clean catalogue
Cleaning, filtering, normalising or labelling data already obtained.
we identify 292 embedded cluster candidates across the 11 available PHANGS Cycle 2 galaxieswhere the paper describes this · verbatim
no AI
Photometry and derivation of ages, masses, extinctions
Extracting understanding from model behaviour.
we determine the ages of the embedded clusters via their Pa α equivalent widthswhere the paper describes this · verbatim
no AI
Build labelled machine learning sample
Cleaning, filtering, normalising or labelling data already obtained.
we produce a sample of 4779 other PAH-bright objects and 233 human-identified embedded clusterswhere the paper describes this · verbatim
AI
Train CNN classifiers by transfer learning
Fitting model parameters, including fine-tuning an existing model. The AI stood in for manual curation.
In total, we train 10 models for all eight model configurations (four data variations for each of the two CNNs).where the paper describes this · verbatim
AI
Classify held-out objects and compare models
Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.
the mode identification of the 10 models becomes an object's 'predicted label'where the paper describes this · verbatim
no AI
Compare model-selected objects to known populations
Testing outputs against ground truth.
we compare the results of testing to expectations of other stellar cluster populations in the Cycle 1 galaxy, NGC 628where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The catalogue of 292 embedded cluster candidates and all derived ages, masses and extinctions come from by-eye identification and photometry; the CNNs are trained on that catalogue and assessed as a possible route for future cataloguing, so the paper's reported findings do not rest on them
we utilize two CNNs: VGG19-bn and ResNet18where the paper describes this · verbatim
there are 805 true negatives, 27 false positives, 6 false negatives, and 24 true positives consistent across all four model configurationswhere the paper describes this · verbatim
Data from the High-Level Science Project PHANGS-HST were obtained from the Mikulski Archive for Space Telescopeswhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
About this article
Record aix-00041, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error