astronomy/ai produced the result/arXiv 2024 · v2
Weighing the Milky Way's disk reveals a mass excess along the Local Arm
Astronomers used the motions of millions of Gaia stars to weigh a patch of our Galaxy's disk. Two machine-learning products supplied the inputs: a smoothed map of how densely stars sit, and predicted line-of-sight velocities for stars missing them.
spectrum · one line per step, placed by what the step does · bright lines used AI
First spiral arm detection using dynamical mass measurements of the Milky Way disk
arXiv, 2024
doi:10.48550/arxiv.2401.04571 · record aix-00099 v2 · checked 2026-10-08
- AI was for
- Property prediction
- Model family
- Gaussian process, Multilayer perceptron
- Checked by
- Held-out
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
Our Galaxy is a flat disk of stars, gas and dust, wound through with spiral arms. The arms are easy to see as places where bright stars crowd together, but counting bright stars is not the same as weighing matter. Much of the disk's mass is in faint stars, cold gas and dark matter, none of which shows up in a photograph. One way round this is to let gravity do the measuring: stars bob up and down through the disk plane, and how far they stray from the mid-plane depends on how much mass is pulling them back. Turning that bobbing into a mass requires knowing both where stars are and how fast they move vertically, everywhere in the patch being studied.
The researchers set out to map the gravitational pull of the disk around the Sun using this dynamical approach. They divided the disk plane into small square cells out to a few kiloparsecs, built four samples of Gaia stars selected by intrinsic brightness, and for each cell measured the difference in gravitational potential between the mid-plane and a height of 400 parsecs, which serves as a stand-in for the mass per unit area of the thin disk. They then asked whether the resulting map was a smooth outward decline or carried extra structure.
Where AI came in
Two learned models, both taken from previously published work rather than trained here, supplied the two ingredients the gravity calculation needs. The first was Gaussian process regression, a statistical method that fits a smooth, flexible surface through scattered data points, used to turn patchy star counts into a continuous and differentiable map of stellar number density. The gravity equation requires gradients of that density, so a smooth model stands in for raw, noisy counts.
The second was a Bayesian neural network that predicts a star's velocity along the line of sight. Gaia measures that velocity directly for only part of its catalogue, with around 60 per cent of the selected stars having it; for the rest only motion across the sky is known. The network's predictions, which come as a range of possible values rather than a single number, filled that gap so vertical velocities could be computed for every star. The authors drew 20 sets of values per star to carry the uncertainty through, substituting model output for a measurement that was never made.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors applied the vertical Jeans equation to four absolute-magnitude samples of Gaia DR3 stars, on a 100 pc grid of the Galactic disk plane out to 3 kpc, to measure the gravitational potential difference between the mid-plane and 400 pc as a proxy for thin disk surface density. The two inputs to the Jeans equation came from machine learning: a Gaussian Process regression model of the stellar number density field, and Bayesian neural network predictions of line-of-sight velocity for the stars lacking Gaia radial velocity measurements. The residual of the measured potential relative to a fitted exponential shows an over-density in all four samples whose position matches the Local Arm as traced by stellar age and chemistry; a fitted arm model gives a relative over-density of roughly 20 % and a width of 0.4 kpc. The inferred thin disk scale length is 3.3–4.2 kpc when the analysis is restricted to stars within 2 kpc.
How AI was used
Two learned data products, both taken from published work rather than trained here, supplied the ingredients of the vertical Jeans analysis. For the stellar number density field, the study used Gaussian Process regression maps of the same four magnitude samples, built from StarHorse spectro-astrometric distances with Gaia DR3, Pan-STARRS1, SkyMapper, 2MASS and AllWISE photometry, with correlation lengths of 300, 300 and 100 pc; the authors then fitted a mirror-symmetric three-component analytic function to the GP output in each area cell. For the vertical velocity field, around 60 % of the 23 551 383 selected stars have Gaia DR3 radial velocities; for the remainder the authors used a public catalogue of Bayesian neural network predictive posteriors for line-of-sight velocity, trained on the subset of DR3 stars that do have radial velocities. These posteriors were handled by multiple imputation: 20 realisations were drawn per star, giving 20 complete datasets, and a parametric vertical velocity variance model plus mean vertical velocity was fitted by maximum likelihood to individual stellar velocities in each cell and its adjacent cells, for each imputation separately. The resulting density and dispersion fits were combined through the vertical Jeans equation to map the potential difference, to which exponential and logarithmic-spiral models were then fitted using a 1-norm objective.
The shape of the work
Structural · the record, drawn
no AI
Construct magnitude-selected Gaia samples on a disk-plane grid
Obtaining raw data, whether by measurement, download or retrieval.
We construct four separate stellar samples by making cuts in Gaia G -band absolute magnitudewhere the paper describes this · verbatim
no AI
Clean samples and mask area cells
Cleaning, filtering, normalising or labelling data already obtained.
We remove possible (probability >0.1) members of the open clusters tabulated by and stars with StarHorse distance precisions worse than 10%.where the paper describes this · verbatim
AI
Obtain stellar number density field from Gaussian Process regression
Running a trained model over new data to predict, classify or score.
we use the results from, where the stellar number density fields of the four data samples were modelled with Gaussian Process (GP) regressionwhere the paper describes this · verbatim
AI
Supply missing radial velocities with Bayesian neural network predictions
Running a trained model over new data to predict, classify or score. The AI stood in for unresolved measurement.
For the remainder of the stars without radial velocity measurements, we use the Bayesian radial velocity predictions of.where the paper describes this · verbatim
no AI
Build multiple imputations and vertical velocity fields
Cleaning, filtering, normalising or labelling data already obtained.
In the present work, we construct 20 such imputationswhere the paper describes this · verbatim
no AI
Fit analytic vertical density and velocity dispersion profiles
Cleaning, filtering, normalising or labelling data already obtained.
we fit an analytic function to the stellar number density that was obtained via GP regression, separately for each data sample and area cellwhere the paper describes this · verbatim
no AI
Apply vertical Jeans equation to infer potential difference per cell
Extracting understanding from model behaviour.
we measure the gravitational potential difference between the disk mid-plane and a height of 400 pcwhere the paper describes this · verbatim
no AI
Fit exponential and spiral arm models to residuals and compare with tracer maps
Testing outputs against ground truth.
we fit a simple analytic model to the residual seen in Fig. 5where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The dynamically measured potential is computed from two machine-learning data products: a Gaussian-Process stellar number density field and Bayesian-neural-network radial velocity predictions that supply the missing velocity component for the proper-motion-only stars. The reported surface density map therefore rests on model output, although the Jeans analysis itself is analytic. 'analysis' was considered but rejected because the finding does depend on these products.
for the former, we use maps produced via Gaussian Process regression; for the latter, we use Bayesian Neural Network radial velocity predictionswhere the paper describes this · verbatim
This newer model was shown to be highly successful in validation tests with DR3 radial velocities excluded from the training datawhere the paper describes this · verbatim
The catalogue of radial velocity predictions is available via the Gaia mirror archivewhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of Gaussian Process regression stellar number density field (results taken from a cited study)Which version of the model was used is not stated.
- Version of Bayesian neural network radial velocity predictor (published Gaia DR3 prediction catalogue)Which version of the model was used is not stated.
- What step 3 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00099, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error