~/aixsci
200 records · all checked

astronomy/ai produced the result/ · v2

Machine learning reads planet-forming disk masses from archival ALMA observations

Astronomers built a grid of physical models of the dusty discs around young stars, then trained decision-tree regressors to run the models backwards, reading gas mass, dust mass and disc size from telescope measurements of 34 discs.

1. Build self-consistent thermochemical disk model grid2. Compute synthetic ALMA observables for each model3. Train regressors mapping observables to disk parameters4. Infer disk parameters for archival ALMA sample5. Flag out-of-domain, gravitationally unstable and optically thick cases6. Compare inferred masses with independent literature estimates

spectrum · one line per step, placed by what the step does · bright lines used AI

DiskMINT-GARDEN: Self-consistent Models to Estimate Disk Masses

doi:not-stated · record aix-00090 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Property prediction, Simulation surrogate
Model family
Gradient-boosted trees
Checked by
Held-out
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Planets form in discs of gas and dust around young stars. How much gas a disc holds sets what kind of planets it can make, and how quickly. But the gas is mostly hydrogen, which barely shows up in these cold, faint discs. Astronomers instead measure things they can see: the glow of millimetre-wavelength light from dust grains, the light from rarer molecules such as a heavy form of carbon monoxide, and how far across the sky the dust emission extends. Turning those few measurements back into a mass is the hard part. It needs a model of the disc's density, its temperature and its chemistry, all of which depend on each other.

The researchers already had such a model, called DiskMINT, which works out a disc's vertical structure, how starlight travels through it and how carbon monoxide is created and destroyed. Running it forwards, from an assumed disc to predicted observations, is straightforward but slow. The authors wanted the reverse: given a real disc's observations, which disc properties produced them. They set out to build that inverse route and apply it to archival data from the ALMA radio telescope array.

Where AI came in

The learning stepped in only at that reversal. The team ran DiskMINT across 480 model discs, varying the star's mass, the disc's gas mass, the ratio of dust to gas and a characteristic radius, then used radiative transfer code to turn each one into a set of synthetic observations. On those pairs they trained gradient-boosted decision trees, a standard method that stacks many simple yes-or-no rules into one predictor. One regressor was fitted for each quantity: gas mass, dust mass and characteristic radius. The star's mass was supplied as known rather than predicted.

Run on real ALMA measurements for 34 discs, the regressors produced the masses and radii reported, with uncertainties estimated by feeding randomly jittered versions of each measurement through them. In effect the trained model stood in for searching the slow simulations for a match. Steps outside the learning checked whether each disc's observations fell inside the range the grid covered, and flagged cases where the physics made the estimates unreliable. The paper's comparisons against gas masses measured by other means rest on these predictions.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors built a grid of self-consistent protoplanetary disk models with DiskMINT, coupling hydrostatic vertical structure, continuum and line radiative transfer, and a reduced CO chemical network, and computed synthetic ALMA observables for each model with RADMC-3D. Gradient-boosted decision tree regressors were trained on this grid to map millimetre continuum luminosity, C18O line luminosity and the 90% dust radius to gas mass, dust mass and characteristic radius at fixed stellar mass. Applied to archival ALMA data for 34 disks, the inferred gas masses agree with published dynamical, HD-based and DALI-based estimates to within about a factor of two for the targets the authors judge reliable. The authors report that the CO depletion factors used in DALI modelling are not required once grain-surface conversion of CO to CO2 is included.

How AI was used

Machine learning was used only for the inverse mapping from observables to disk physical parameters. The training data came entirely from the authors' own thermochemical model grid: 480 DiskMINT models spanning six stellar masses, five gas disk masses, four dust-to-gas ratios and four characteristic radii, each post-processed with RADMC-3D into an observable vector of Band 6 continuum luminosity, C18O (2-1) and (3-2) line luminosities and the radius enclosing 90% of the continuum emission. Separate XGBoost gradient-boosted decision tree regressors, one per target parameter (gas mass, dust mass, characteristic radius), were fitted in log10 space on both observables and targets, with a reproducible 90/10 train-validation split, L1 and L2 regularisation of 1.0 each, learning rate 0.02, maximum tree depth 5, up to 2000 estimators and early stopping after 50 rounds without validation improvement. Stellar mass was treated as a known conditioning parameter rather than inferred. The trained regressors were then run on archival ALMA measurements for the 34-disk sample, with Monte Carlo realisations of each observable vector drawn from the reported Gaussian measurement uncertainties passed through the regressors to obtain confidence intervals. Non-learned steps handled domain control, clipping predictions for targets outside the grid envelope, and flagging models that were gravitationally unstable or optically thick.

The shape of the work

Structural · the record, drawn

SIMULATIONSIMULATIONTRAININGINFERENCESCREENINGVALIDATION123456AIAIBuildself-consistentthermochemical d…Compute syntheticALMA observablesfor each modelTrain regressorsmappingobservables to d…Infer diskparameters forarchival ALMA sa…Flagout-of-domain,gravitationally …Compare inferredmasses withindependent lite…↤ simulation↤ simulation
AI stepNo AI↤ what the AI stood in for
1Simulation
no AI

Build self-consistent thermochemical disk model grid

Numerical or physics simulation, including where a learned surrogate replaces it.

We generate a grid of disk models using DiskMINT following similar setups in our previous works.where the paper describes this · verbatim
in the paper
2Simulation
no AI

Compute synthetic ALMA observables for each model

Numerical or physics simulation, including where a learned surrogate replaces it.

For each model in the grid, we compute synthetic observables using RADMC-3D.where the paper describes this · verbatim
in the paper
3Training
AI

Train regressors mapping observables to disk parameters

Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.

The regression model is trained on the DiskMINT-GARDEN grid in log space.where the paper describes this · verbatim
in the paper
4Inference
AI

Infer disk parameters for archival ALMA sample

Running a trained model over new data to predict, classify or score. The AI stood in for simulation.

we infer the corresponding disk physical parameters (Mgas, Mdust, and Rc) for each target using the trained regression modelwhere the paper describes this · verbatim
in the paper
5Screening
no AI

Flag out-of-domain, gravitationally unstable and optically thick cases

Reducing a candidate set by filtering or ranking, in a single pass.

we perform domain checks by comparing each target’s observables to the range spanned by the training gridwhere the paper describes this · verbatim
in the paper
6Validation
no AI

Compare inferred masses with independent literature estimates

Testing outputs against ground truth.

We compare DiskMINT-GARDEN inferred Mgas to dynamical gas masses in Figure 2.where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The disk gas masses, dust masses and characteristic radii reported for the 34-disk sample are produced by the trained regression model mapping observables to physical parameters; the paper's comparisons with dynamical, HD-based and DALI masses rest on those regression outputs.

+What the AI was for
we implement this mapping using a supervised machine-learning regression model based on gradient-boosted decision treeswhere the paper describes this · verbatim
+Model families
+How it was taught
Supervisedin the paper
+Models named
XGBoost gradient-boosted decision tree regressor (one per predicted parameter) · Trained from scratchin the paper
+How results were checked
Held-outin the paper
We split the DiskMINT-GARDEN grid into a reproducible 90/10 train–validation split, using the 90% subset for trainingwhere the paper describes this · verbatim
+Code · weights · data
code availableweights availabledata availablein the paper
The trained regression model and inference tools are released as part of the public DiskMINT v1.7.0 on GitHub.where the paper describes this · verbatim
+Compute
University of Arizona High Performance Computing resources; no accelerator time or run time statedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 2 items
  • How many were testedThe paper gives no count of what was tested.
  • Version of XGBoost gradient-boosted decision tree regressor (one per predicted parameter)Which version of the model was used is not stated.

About this article

Record aix-00090, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error