~/aixsci
200 records · all checked

astronomy/ai produced the result/Astronomy and Astrophysics 2023 · v2

Random Forest sorts 9,446 white dwarfs by type from Gaia spectra

Astronomers trained a Random Forest classifier on low-resolution Gaia spectra of nearby white dwarfs, using existing catalogue labels, then applied it to assign spectral types to 9,446 stars that had none.

1. Assemble labelled training sample and Gaia spectra2. Use Hermite spectral coefficients as features3. Train Random Forest classifiers4. Cross-validate and score classifications5. Classify unclassified white dwarfs into DA and non-DA6. Classify non-DAs into primary and secondary subtypes7. Rank spectral coefficients by feature importance8. Check classifications against HR diagram and an external catalogue

spectrum · one line per step, placed by what the step does · bright lines used AI

White dwarf Random Forest classification through Gaia spectral coefficients
Astronomy and Astrophysics, 2023

doi:10.1051/0004-6361/202347601 · record aix-00122 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Classification
Model family
Random forest
Checked by
Held-out2905 tested
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

A white dwarf is what remains when a star like the Sun runs out of fuel: a dense, slowly cooling cinder about the size of the Earth. Astronomers sort white dwarfs into spectral types according to which chemical elements show up in their light. Hydrogen-dominated ones are called DA; the rest, the non-DAs, are split further into classes such as DB, DC, DQ and DZ depending on whether helium, carbon or metals leave their mark, or whether no lines appear at all. Making that judgement normally needs a detailed spectrum of each star, taken one at a time, which is slow work.

The Gaia space telescope has recorded low-resolution spectra for a very large number of stars, including white dwarfs within 100 parsecs of the Sun. These spectra are coarse, and Gaia stores them not as plain curves but as sets of 110 mathematical coefficients. The researchers set out to see whether those coefficients alone carried enough information to assign a spectral type, without first converting them back into calibrated light measurements.

Where AI came in

The classifier was a Random Forest, built with the scikit-learn library and trained from scratch. A Random Forest grows many simple decision trees, each asking a series of yes-or-no questions about the input numbers, and lets them vote on the answer. Here the inputs were the 110 coefficients of each star's Gaia spectrum, and the answers it was taught to give came from spectral types already recorded in the Montreal White Dwarf Database for white dwarfs within 100 parsecs. The models were arranged in stages: first DA against non-DA, then non-DAs into DB, DC, DQ and DZ, then attempts at finer sub-labels.

Tested by holding back portions of the 2,905 labelled stars in turn, it reached an accuracy of 0.90 on the DA versus non-DA split and 0.81 on the non-DA classes, while the finer secondary types were mostly not recovered. The trained models were then run across the rest of the 100-parsec sample, producing spectral types for 9,446 white dwarfs that had not been classified before. In effect the classifier took over the case-by-case inspection a specialist would otherwise do by hand. Its labels were afterwards checked for consistency on a standard brightness-versus-colour diagram and compared with a published catalogue made by a neural network.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

A Random Forest classifier was trained on the 110 Hermite coefficients of Gaia DR3 low-resolution BP/RP spectra, using spectral types from the Montreal White Dwarf Database for white dwarfs within 100 pc as labels. In stratified ten-fold cross-validation over the 2 905 labelled objects it reached an accuracy of 0.90 for the DA versus non-DA split and 0.81 when non-DAs were separated into DB, DC, DQ and DZ, with recall above 80% for DA and DB and precision of 92% for DB, DQ and DZ; secondary spectral types were mostly not recovered. Applied to the remaining 100-pc sample, the classifier assigned spectral types to 9 446 previously unclassified white dwarfs, and a ranking of coefficient importance by mean decrease in impurity placed most of the weight on the first blue coefficients and on a red coefficient identified with H-alpha.

How AI was used

The 110 Hermite basis coefficients of each object's internally calibrated Gaia BP and RP spectrum were used directly as input features, with no external flux calibration applied. Spectral types from the Montreal White Dwarf Database for white dwarfs within 100 pc supplied the training labels. Random Forest models were built with scikit-learn and evaluated by stratified ten-fold cross-validation, in a hierarchical arrangement: a first model separating hydrogen-rich DA from non-DA, a second splitting non-DAs into DB, DC, DQ and DZ, and further models attempting secondary subtypes within each primary type, with training subsets chosen to stay as close to balanced as possible and, for the cool sample, restricted to objects matching the colour range being classified. The trained models were then run over the unclassified 100-pc population, first assigning DA or non-DA labels to the cool objects not covered by an earlier spectral-energy-distribution fitting method, then assigning subtypes. Coefficient importance was computed from the forests by mean decrease in impurity, and the resulting labels were checked for positional consistency in the Gaia Hertzsprung-Russell diagram and cross-matched against a published neural-network classification.

The shape of the work

Structural · the record, drawn

ACQUISITIONREPRESENTATIONTRAININGVALIDATIONINFERENCEINFERENCEINTERPRETATIONVALIDATION12345678AIAIAIAssemble labelledtraining sampleand Gaia spectraUse Hermitespectralcoefficients as …Train RandomForestclassifiersCross-validateand scoreclassificationsClassifyunclassifiedwhite dwarfs int…Classify non-DAsinto primary andsecondary subtyp…Rank spectralcoefficients byfeature importan…Checkclassificationsagainst HR diagr…↤ manual curation↤ manual curation↤ manual curation
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Assemble labelled training sample and Gaia spectra

Obtaining raw data, whether by measurement, download or retrieval.

the MWDD spectroscopic white dwarf classification within 100 pc is used as the input labelled sample for the training and validation testswhere the paper describes this · verbatim
in the paper
2Representation
no AI

Use Hermite spectral coefficients as features

Encoding data into features, descriptors, embeddings or graphs.

As input data for the Random Forest algorithm, we use the 110 Hermite coefficients.where the paper describes this · verbatim
in the paper
3Training
AI

Train Random Forest classifiers

Fitting model parameters, including fine-tuning an existing model. The AI stood in for manual curation.

For each fold, a Random Forest is trained with all nine remainder folds, and tested on it.where the paper describes this · verbatim
in the paper
4Validation
no AI

Cross-validate and score classifications

Testing outputs against ground truth.

For the cross-validation, we adopt the stratified k -fold method.where the paper describes this · verbatim
in the paper
5Inference
AI

Classify unclassified white dwarfs into DA and non-DA

Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.

To that sample we apply our Random Forest algorithm, once trained with those objects labelled in the MWDD (2 905 objects).where the paper describes this · verbatim
in the paper
6Inference
AI

Classify non-DAs into primary and secondary subtypes

Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.

the 3 041 found non-DAs were classified into DC, DQ and DZ categorieswhere the paper describes this · verbatim
in the paper
7Interpretation
no AI

Rank spectral coefficients by feature importance

Extracting understanding from model behaviour.

The method used to compute the feature importance was the mean decrease in impurity (MDI)where the paper describes this · verbatim
in the paper
8Validation
no AI

Check classifications against HR diagram and an external catalogue

Testing outputs against ground truth.

Both catalogues were cross-matched and 1 103 objects were found in both tables.where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The spectral types reported for 9 446 previously unclassified white dwarfs are the direct output of the Random Forest classifier; the paper's result is that classification

+What the AI was for
Classificationin the paper
we explore the possibility of utilising a machine learning approach, based on a Random Forest algorithmwhere the paper describes this · verbatim
+Model families
Random forestin the paper
+How it was taught
Supervisedin the paper
+Models named
Random Forest classifier (scikit-learn) · Trained from scratchin the paper
+How results were checked
Held-out2905 testedin the paper
white dwarfs with labels in the MWDD within 100 pc into DA and non-DA types (2,905 objects; 1993 as DAs and 912 as non-DAs)where the paper describes this · verbatim
+Code · weights · data
code not reportedweights not reporteddata availablein the paper
The whole catalogue can be found in the electronic version of the paper.where the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 4 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • Version of Random Forest classifier (scikit-learn)Which version of the model was used is not stated.

About this article

Record aix-00122, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error