astronomy/ai produced the result/arXiv 2026 · v2
Classifiers label individual TESS brightness readings as planet transits or not
Researchers trained six kinds of machine learning classifier to judge, reading by reading, whether a star's brightness dip marks a planet crossing its face. The models did the judging; the experiment varied where their training labels came from.
spectrum · one line per step, placed by what the step does · bright lines used AI
Astro-Hunters: Machine Learning for Exoplanet Transit Detection in TESS Photometry
arXiv, 2026
doi:10.48550/arxiv.2608.18172 · record aix-00238 v2 · checked 2026-10-09
- AI was for
- Classification, Detection, Anomaly detection
- Model family
- Gradient-boosted trees, Random forest, Multilayer perceptron, Linear model
- Checked by
- Held-out
- Code
- available
The finding the paper is about came from the AI.
What this research was about
When a planet passes in front of its star, it blocks a sliver of light, and the star dims very slightly for a few hours. NASA's TESS satellite watches stars and records their brightness every two minutes, producing a long string of measurements called a light curve. Finding planets means spotting those shallow dips. The difficulty is that a single two-minute reading is mostly noise. The dip is far smaller than the random wobble in any one measurement, so the evidence for a planet only really accumulates when many readings are considered together.
The researchers built a pipeline that fetches TESS light curves for twelve stars already known to host planets, cleans and flattens them, and then reduces each individual reading to seven simple summary numbers describing it and its immediate neighbours. They then asked whether machine learning could tag each reading as in-transit or not, and how much the answer depended on where the training labels came from rather than on the choice of model.
Where AI came in
Learned models appear twice. Six families of classifier — logistic regression, a small two-layer neural network, a random forest, extremely randomised trees, histogram-based gradient boosting and XGBoost — were trained from scratch on the seven numbers per reading, and then run to give each reading a probability of being in transit. A cut-off was chosen on one group of stars and applied unchanged to stars the models had never seen. Separately, an unsupervised method called an isolation forest, which flags unusual points without being told what to look for, was used to generate one of the label sets being compared.
The classifiers stood in for the judgement of deciding, point by point, whether a reading belongs to a transit — work otherwise done by fixed, hand-written search algorithms. One such algorithm, a Box Least Squares period search, was run on the same twelve light curves by conventional code for comparison. The training labels themselves were not learned: they came from published planet orbital timings held in the NASA Exoplanet Archive. The reported numbers are measurements of how well the fitted models performed, so they exist only because the models were trained and run.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The study built a pipeline that retrieves TESS two-minute light curves, detrends them, reduces each cadence to seven sliding-window statistics, and trains tree- and network-based classifiers to label individual cadences as in-transit or not, using labels derived from published ephemerides for a corpus of 189,279 cadences from twelve confirmed planet hosts. Holding features, model and the star-disjoint protocol fixed, the authors report that area under the precision-recall curve moves from 0.0298 to 0.8584 across label sources, a factor of 29, while six architectures span a factor of 1.8; labels from an isolation forest fitted to the classifier's own features gave AUC 0.9915, an unconverted BJD transit epoch gave chance performance, and correctly referenced ephemerides gave AUC 0.788 at 5.3 times the prevalence baseline. A measured median single-cadence signal-to-noise ratio of 2.10 is reported to cap per-cadence AUC at 0.932, with the best of the six architectures at 0.814. A Box Least Squares search on the same twelve light curves recovered eight of twelve published orbital periods from a single sector.
How AI was used
Learned models appear at two points. Six supervised classifier families — L2-regularised logistic regression, a two-layer perceptron, a random forest, extremely randomised trees, histogram-based gradient boosting and XGBoost — were fitted from scratch to a seven-dimensional per-cadence feature vector (detrended flux, window mean, deviation, minimum, flux/mean ratio, standard deviation and skewness over a window of plus or minus 32 cadences) and run to assign each cadence a transit probability, with a threshold chosen on development hosts and applied unchanged to held-out hosts. Training labels were derived outside the model, from NASA Exoplanet Archive ephemerides converted from BJD to BTJD and applied over every transiting planet; label provenance was varied as an experimental condition, and one of the compared label sources was produced by an unsupervised isolation forest fitted to the same feature matrix the classifier consumed. Evaluation used star-disjoint GroupKFold with four folds and an explicit three-way host partition, with a Box Least Squares period search and standard false-positive diagnostics computed by conventional, non-learned code on the same light curves. Fitted models and the serving code are released behind a hosted inference endpoint that reuses the training preprocessing path.
The shape of the work
Structural · the record, drawn
no AI
Retrieve and stitch TESS light curves
Obtaining raw data, whether by measurement, download or retrieval.
Light curves are retrieved from MAST by target name, restricted to SPOC-processed two-minute products, and stitched across sectors.where the paper describes this · verbatim
no AI
Clean, normalise and detrend photometry
Cleaning, filtering, normalising or labelling data already obtained.
Invalid cadences are removed; flux is divided by its median to normalisewhere the paper describes this · verbatim
no AI
Compute seven sliding-window statistics per cadence
Encoding data into features, descriptors, embeddings or graphs.
The seven statistics of Section III-B are computed over a sliding window, yielding one row per cadence.where the paper describes this · verbatim
no AI
Derive per-cadence labels from published ephemerides
Cleaning, filtering, normalising or labelling data already obtained.
Ephemerides are queried from the NASA Exoplanet Archive, converted from BJD to BTJDwhere the paper describes this · verbatim
AI
Train six classifier families under star-disjoint splits
Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.
A gradient-boosted tree ensemble is trained under the star-disjoint partition of Section III-E.where the paper describes this · verbatim
AI
Score held-out cadences and apply decision threshold
Running a trained model over new data to predict, classify or score. The AI stood in for conventional algorithm.
The decision threshold is selected to maximise F1 on the development hosts and then applied unchanged to the held-out hosts.where the paper describes this · verbatim
no AI
Evaluate against gold labels and the single-cadence information bound
Testing outputs against ground truth.
Reported metrics are the area under the ROC curve; the area under the precision–recall curvewhere the paper describes this · verbatim
no AI
Classical Box Least Squares period-search baseline
Testing outputs against ground truth.
Table IX reports a Box Least Squares search over the same twelve light curves.where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's results are measurements of trained classifiers' per-cadence detection performance under different label sources; the reported findings exist only because the models were fitted and run.
a gradient-boosted tree classifier assigning each cadence a transit probabilitywhere the paper describes this · verbatim
six hosts for training, three for development, and three held out entirelywhere the paper describes this · verbatim
Code reproducing every table and figure: https://github.com/astral-fate/Astro-Hunters-Model-traningwhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of XGBoost (200 estimators, depth 5, learning rate 0.1)Which version of the model was used is not stated.
- Version of Histogram-based gradient boostingWhich version of the model was used is not stated.
- Version of Random forest (200 estimators, depth 12)Which version of the model was used is not stated.
- Version of Extremely randomised trees (200 estimators, depth 12)Which version of the model was used is not stated.
- Version of Multilayer perceptron (64, 32 hidden units)Which version of the model was used is not stated.
- Version of L2-regularised logistic regressionWhich version of the model was used is not stated.
- Version of Isolation forest (used to synthesise the circular label set in the ablation)Which version of the model was used is not stated.
About this article
Record aix-00238, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error