~/aixsci
200 records · all checked

astronomy/ai produced the result/arXiv 2025 · v2

Random forest sorts 130 million Magellanic Cloud sources into ten classes

Astronomers trained a probabilistic random forest on spectroscopically classified stars and galaxies, then used it to assign a class and a probability to each of the roughly 130 million sources in a near-infrared survey of the Magellanic Clouds.

1. Obtain new optical spectra of candidate sources2. Build multi-wavelength feature table3. Assemble, label and balance training sets4. Train and test separate SMC and LMC classifiers5. Classify all VMC sources6. Compute and inspect feature importances7. Split output into confidence tiers8. Check classifications against independent catalogues

spectrum · one line per step, placed by what the step does · bright lines used AI

The VMC Survey : LI. Classifying extragalactic sources using a probabilistic random forest supervised machine learning algorithm
arXiv, 2025

doi:10.48550/arxiv.2501.08196 · record aix-00235 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Classification
Model family
Random forest
Checked by
Held-out
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

The Magellanic Clouds are two small galaxies near our own, and surveys of them collect light from enormous numbers of points on the sky. Each point, or source, might be a star inside the Clouds, a star in the foreground of our own Galaxy, or something far beyond: a distant galaxy, or an active galactic nucleus, meaning a galaxy whose central black hole is drawing in gas and shining brightly. Telling these apart reliably usually needs a spectrum, which spreads a source's light out into its component colours. Spectra take a lot of telescope time, so only a small fraction of sources ever get one.

What is available for almost every source is photometry: brightness measurements in a handful of filters, from optical light through to the far infrared. Differences between those brightnesses, known as colours, carry clues about what a source is. The researchers set out to turn the small set of sources with known spectroscopic identities into a way of labelling every source in the VISTA Survey of the Magellanic Clouds, and to say how confident each label was.

Where AI came in

The AI here is a probabilistic random forest, a method that grows many decision trees, each asking a series of questions about a source's brightnesses and colours, and combines their votes into a class plus a probability for each class. The probabilistic version also takes the measurement errors into account. The team built a table of 237 features per source by cross-matching the near-infrared catalogue with optical, infrared and far-infrared data, labelled a training set using new and published spectra, and trained separate classifiers for each Cloud with 100 trees.

The trained classifiers then ran over both full catalogues, standing in for the spectroscopy and hand-sorting that would otherwise be needed to identify each source. The paper's catalogues are the model's output: the classifications and their probabilities. Those probabilities were used to divide the results into low-, mid- and high-confidence tiers. On held-out test data the classifiers reached average accuracies of about 79 per cent for the Small Magellanic Cloud and 87 per cent for the Large, rising to about 90 and 98 per cent for the 56,696,719 sources with class probabilities above 80 per cent.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

A probabilistic random forest was trained on optical-to-far-infrared photometry of spectroscopically classified sources and applied to the roughly 130 million sources of the VISTA Survey of the Magellanic Clouds, assigning each source one of ten classes or an 'Unknown' label together with class probabilities. The classifiers reached average accuracies of about 79 per cent (SMC) and 87 per cent (LMC) on held-out data, and about 90 per cent (SMC) and 98 per cent (LMC) for the 56,696,719 sources with class probabilities above 80 per cent. After removing sources classed as Unknown, 707,939 sources in the SMC field and 397,899 in the LMC field were classified, including more than 77,600 extragalactic sources behind the Clouds. Cross-matching with independent X-ray and radio catalogues found 554 of 883 X-ray sources classed as AGN, and 1756 of 2694 radio sources classed as AGN with 659 classed as galaxies.

How AI was used

The authors cross-matched the VMC near-IR PSF catalogue with SMASH optical, Gaia DR3, Spitzer SAGE, AllWISE and unWISE photometry within a 1 arcsec radius, added a smoothed Herschel SPIRE 250 μm background flux measurement at each position, and formed colours between all photometric bands with propagated errors, giving 237 features with measurement errors and homogenised NaN values for missing data. Labels came from spectroscopically classified sources, including 26 sources from new SAAO 1.9m observations and 22 from new SALT observations, literature catalogues such as Milliquas, 6dFGS, SDSS in the GAMA09 field and SAGE-spec, plus a Simbad search for proper-motion stars; ten classes were used, together with an Unknown class built from randomly selected VMC sources. Training sets were split 75/25 into training and test sets with the split randomised across runs, and the training portion was upsampled by random duplication to the size of the largest class. Separate classifiers were trained for the SMC and LMC, sharing extragalactic and foreground training sources while keeping Magellanic stellar classes Cloud-specific, using the probabilistic random forest implementation with a probability threshold of 0.05 and 100 trees chosen from a scan over tree numbers; runs were repeated across ten random seed states and confusion matrices averaged. The trained classifiers were then run over the full SMC and LMC PSF catalogues, mean-decrease-impurity feature importances were averaged over ten repeats, and outputs were divided into low-, mid- and high-confidence catalogues by class probability, with sources whose combined AGN and galaxy probabilities crossed a threshold promoted between tiers. X-ray, radio, Quaia and published YSO catalogues were held out of the feature set and used only for cross-matched comparison.

The shape of the work

Structural · the record, drawn

ACQUISITIONREPRESENTATIONPREPARATIONTRAININGINFERENCEINTERPRETATIONSCREENINGVALIDATION12345678AIAIAIObtain newoptical spectraof candidate sou…Buildmulti-wavelengthfeature tableAssemble, labeland balancetraining setsTrain and testseparate SMC andLMC classifiersClassify all VMCsourcesCompute andinspect featureimportancesSplit output intoconfidence tiersCheckclassificationsagainst independ…↤ manual curation↤ manual curation
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Obtain new optical spectra of candidate sources

Obtaining raw data, whether by measurement, download or retrieval.

We observed 174 new optical spectrawhere the paper describes this · verbatim
in the paper
2Representation
no AI

Build multi-wavelength feature table

Encoding data into features, descriptors, embeddings or graphs.

New features were created by subtracting each feature from all the other features to create colourswhere the paper describes this · verbatim
in the paper
3Preparation
no AI

Assemble, label and balance training sets

Cleaning, filtering, normalising or labelling data already obtained.

the ‘resample’ function of Python’s Scikit-learn module was used to upsample all the class samples to the same size as the majority classwhere the paper describes this · verbatim
in the paper
4Training
AI

Train and test separate SMC and LMC classifiers

Fitting model parameters, including fine-tuning an existing model. The AI stood in for manual curation.

Individual classifiers are trained for the SMC and LMC due the different stellar populationswhere the paper describes this · verbatim
in the paper
5Inference
AI

Classify all VMC sources

Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.

The classifiers were used on the entirety of the LMC and SMC PSF catalogueswhere the paper describes this · verbatim
in the paper
6Interpretation
AI

Compute and inspect feature importances

Extracting understanding from model behaviour.

The classifiers are trained on the full datasets for SMC and LMC, from which the feature importances were calculated.where the paper describes this · verbatim
in the paper
7Screening
no AI

Split output into confidence tiers

Reducing a candidate set by filtering or ranking, in a single pass.

The catalogues of sources are separated into high-confidence sources (Pclass > 80%), mid-confidence sources (60% < Pclass < 80%)where the paper describes this · verbatim
in the paper
8Validation
no AI

Check classifications against independent catalogues

Testing outputs against ground truth.

we use the radio and X-ray detected sources as an independent check to test the classificationswhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's product is the set of source classifications and class probabilities produced by the trained classifiers; the reported catalogues exist only as model output.

+What the AI was for
Classificationin the paper
We used a supervised machine learning algorithm (probabilistic random forest) to classify ∼ 130 million sourceswhere the paper describes this · verbatim
+Model families
Random forestin the paper
+How it was taught
Supervisedin the paper
+Models named
Probabilistic Random Forest (PRF) · Trained from scratchin the paper
+How results were checked
Held-outin the paper
each dataset was split into training and testing sets, where 75 % of the data were trained onwhere the paper describes this · verbatim
+Code · weights · data
code not reportedweights not reporteddata availablein the paper
The training sets for the SMC and LMC are made available alongside this paper as online supplementary material.where the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 6 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of Probabilistic Random Forest (PRF)Which version of the model was used is not stated.
  • What step 6 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00235, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error