~/aixsci
200 records · all checked

astronomy/ai produced the result/Astronomy and Astrophysics 2025 · v2

Machine learning sifts a million candidate moving objects to find 258 new cool stars

Astronomers trawled repeated infrared sky surveys for objects that shift position, using a random forest classifier to rank more than a million candidate tracks so that only the most promising ones reached a human eye.

1. Obtain epochal W2 detection catalogues2. Apply quality cuts and astrometric error model3. Link detections into candidate motion tracks4. Train artefact-rejection classifier iteratively5. Score all candidates and queue the highest-ranked6. Visually vet queued candidates7. Cross-match against existing catalogues to isolate new objects8. Estimate temperatures, spectral types and distances

spectrum · one line per step, placed by what the step does · bright lines used AI

New ultracool dwarf candidates from multi-epoch WISE data
Astronomy and Astrophysics, 2025

doi:10.1051/0004-6361/202452011 · record aix-00227 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Classification
Model family
Random forest
Checked by
Held-out
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Most stars appear fixed, but a few drift measurably across the sky from year to year. That drift, called proper motion, is a clue that an object lies close to us. Some of the nearest neighbours of the Sun are ultracool dwarfs: objects too small and faint to shine like proper stars, glowing mostly in infrared light. Finding them means comparing images of the same patch of sky taken at different times and spotting the dots that move. The difficulty is that sky surveys contain vast numbers of spurious detections — noise, image defects, the glare of bright stars — and these fake dots can mimic a moving object convincingly.

The researchers worked through the unTimely catalogue, a uniform reprocessing of repeated infrared images from the WISE space telescope, looking in one infrared band for objects moving faster than roughly 0.3 arcseconds a year. A rule-based step linked detections from different dates into candidate tracks, producing 1,271,921 of them. The task was then to separate genuine movers from artefacts, and to characterise anything that had not been catalogued before.

Where AI came in

The random forest classifier — a method that combines many simple decision trees into one vote — was the filter between the automated track-finding and the human inspectors. It judged each candidate track as a real moving object or an artefact, using statistical descriptions of the track and its surroundings as input. It began with a small set of examples checked by eye, then scored every candidate; the hundred most promising went to a human for inspection through an image-viewing dashboard, and those verdicts were fed back as fresh training data for the next round. The loop continued until the classifier no longer put forward anything it thought real.

The alternative would have been to eyeball all 1,271,921 candidates. In the end 59,322 were inspected, of which 21,885 were confirmed as real movers and 37,437 as artefacts. The classifier's performance was checked by leaving out single examples one at a time, and the wider search was tested by injecting simulated objects and by seeing how many already-known ultracool dwarfs came back. Everything after the vetting — matching against existing catalogues, fitting model atmospheres to the colours, estimating temperatures and distances — used no trained model.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors searched the unTimely catalogue of epochal unWISE detections in the W2 band for objects with proper motions above roughly 0.3 arcseconds per year. A rule-based track-linking algorithm produced 1,271,921 initial candidates, and a Random Forest classifier, retrained over successive rounds on visually checked examples, ranked them so that only part of the pool reached human inspection; 21,885 candidates were visually vetted as real moving objects and 37,437 as artefacts, out of 59,322 inspected in total. Of the confirmed objects, 258 had no SIMBAD or Gaia counterpart and had not been published before; fitting their multi-wavelength photometry with a BT-Settl model grid and comparing colours with known ultracool dwarfs indicated that all except 6 are compatible with being ultracool dwarfs, including at least 33 T dwarf candidates with estimated distances closer than about 40 parsecs and effective temperatures below 1300 K.

How AI was used

Machine learning was used for artefact rejection between a non-learned motion-detection step and human visual vetting. Candidate tracks were first built by pairing W2-band detections from different epochs in the unTimely catalogue, fitting proper motions with RANSAC robust linear regression, and computing statistical properties of the track points and their neighbours to serve as features. A Random Forest binary classifier from scikit-learn, with 100 estimators and automatically balanced class weights, was bootstrapped on candidates with the 60 largest and 60 smallest values of the coordinate-time correlation feature, all of which were visually checked. The classifier was then applied to all candidates, the hundred with the largest predicted probability of being real were inspected through a Jupyter dashboard built on the WISEView cutout API, and the resulting labels were added to the training set for the next iteration; the loop ran until the classifier stopped returning candidates with non-zero probability. Classifier performance was characterised with Leave One Out cross-validation, detection sensitivity with injection of simulated objects into the epoch catalogues, and combined detection plus filtering performance by recovery of known ultracool dwarfs from UltracoolSheet. Downstream characterisation of the surviving objects used catalogue cross-matching, BT-Settl model spectral fitting in VOSA, and nearest-neighbour and colour-based interpolation against UltracoolSheet spectral types, none of which involved a trained model.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONGENERATIONTRAININGINFERENCEPREPARATIONSCREENINGINTERPRETATION12345678AIAIObtain epochal W2detectioncataloguesApply qualitycuts andastrometric erro…Link detectionsinto candidatemotion tracksTrainartefact-rejectionclassifier itera…Score allcandidates andqueue the highes…Visually vetqueued candidatesCross-matchagainst existingcatalogues to is…Estimatetemperatures,spectral types a…↤ manual curation↤ manual curationloops back
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Obtain epochal W2 detection catalogues

Obtaining raw data, whether by measurement, download or retrieval.

we used recently released unTimely catalogue which is the result of uniform analysis of individual epochal unWISE coaddswhere the paper describes this · verbatim
in the paper
2Preparation
no AI

Apply quality cuts and astrometric error model

Cleaning, filtering, normalising or labelling data already obtained.

We applied basic quality cuts to the objects from epochal catalogues by selecting only the detections from primary parts of unWISE coaddswhere the paper describes this · verbatim
in the paper
3Generation
no AI

Link detections into candidate motion tracks

Producing candidate objects that did not previously exist.

The algorithm returned 1,271,921 initial candidates from all tiles of unTimely catalogue.where the paper describes this · verbatim
in the paper
4Training
AI

Train artefact-rejection classifier iteratively

Fitting model parameters, including fine-tuning an existing model. The AI stood in for manual curation.

then using derived labels to improve the classifier and repeat the whole processwhere the paper describes this · verbatim
in the paper
5Inference
AI

Score all candidates and queue the highest-ranked

Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.

We then applied the classifier to all candidates, selected a hundred ones with largest predicted probability of being truewhere the paper describes this · verbatim
in the paper
6Preparation
no AI

Visually vet queued candidates

Cleaning, filtering, normalising or labelling data already obtained. Its result feeds back into an earlier step.

we got 21,885 candidates visually vetted to be real moving objects, and 37,437 – as various artefactswhere the paper describes this · verbatim
in the paper
7Screening
no AI

Cross-match against existing catalogues to isolate new objects

Reducing a candidate set by filtering or ranking, in a single pass.

We checked for associated SIMBAD database entries for all 21,885 objects with high proper motions that we detectedwhere the paper describes this · verbatim
in the paper
8Interpretation
no AI

Estimate temperatures, spectral types and distances

Extracting understanding from model behaviour.

we fit the measurements with the grid of spectra for BT-Settl model atmospheres using VO SED Analyzer (VOSA)where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The reported sample of moving objects is the set that the Random Forest classifier ranked highly enough to be placed in front of a human inspector; of 1,271,921 algorithm candidates only 59,322 were ever inspected, so the published catalogue is defined jointly by the classifier and the visual vetting. Could also be read as 'analysis', since the final real/artefact decision was made by eye and the track-finding itself is rule-based.

+What the AI was for
Classificationin the paper
for extracting ‘‘true’’ moving objects from them based on Random Forest binary classifierwhere the paper describes this · verbatim
+Model families
Random forestin the paper
+How it was taught
SupervisedActive learningin the paper
+Models named
Random Forest binary classifier (scikit-learn, 100 estimators, balanced class weights) · Trained from scratchin the paper
+How results were checked
Held-outin the paper
We characterized the performance of final classifier using Leave One Out cross-validationwhere the paper describes this · verbatim
+Code · weights · data
code not reportedweights not reporteddata availablein the paper
The table is published at https://doi.org/10.5281/zenodo.14651236.where the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 5 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of Random Forest binary classifier (scikit-learn, 100 estimators, balanced class weights)Which version of the model was used is not stated.

About this article

Record aix-00227, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error