astronomy/ai produced the result/arXiv 2025 · v2
A search engine for radio galaxy shapes, built on a fine-tuned image-text model
Astronomers fine-tuned the open-source OpenCLIP model on labelled radio galaxy images and their descriptions, then used it to index about 170,000 extended radio sources so that a typed description or a sample picture returns similar objects.
spectrum · one line per step, placed by what the step does · bright lines used AI
EMUSE: Evolutionary Map of the Universe Search Engine
arXiv, 2025
doi:10.48550/arxiv.2506.15090 · record aix-00061 v2 · checked 2026-10-08
- AI was for
- Classification, Detection
- Model family
- Transformer
- Checked by
- Held-out
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Radio telescopes pick up the glow of charged particles spiralling in magnetic fields, and many galaxies appear in these maps not as neat blobs but as jets, lobes and plumes stretching far beyond the starlight. The shape matters, because it carries clues about what is happening around the galaxy's central black hole and in the gas surrounding it. Astronomers have long sorted these objects into families, such as the two Fanaroff-Riley classes, which differ in where along the jets the radio emission is brightest. The difficulty is scale. Modern surveys record far more sources than anyone can page through by eye, and the odd-looking ones are exactly the ones worth finding.
The work here draws on the Evolutionary Map of the Universe, a radio survey of the southern sky, using 160 tiles from its first year of observations covering 4,500 square degrees. The researchers set out to build a way of asking that archive a question in ordinary terms, by typing a description of a shape or by supplying a picture of one source and asking what else looks like it.
Where AI came in
The central component is OpenCLIP, an openly available model trained on large numbers of internet images paired with captions, which learns to place a picture and its description close together in the same mathematical space. The researchers fine-tuned it on 2,900 radio galaxies from the RadioGalaxyNET dataset, plus an added category for peculiar and rare shapes, feeding it three-channel images that combine radio and infrared views alongside text describing the morphology. On held-out test data, sampled ten times, it reached 84±3 per cent accuracy across the main shape categories, with FR-I and FR-x sources most often mistaken for each other. Fine-tuning took about 1.5 hours on a single graphics processor.
A second model, Gal-DINO, was used off the shelf to find radio sources in the survey tiles and pick out candidate host galaxies in the infrared. The fine-tuned image encoder then converted cutouts of roughly 170,000 extended sources into numerical fingerprints stored in a database. A query, whether typed or supplied as an image, is turned into a fingerprint the same way, and the closest matches are returned with their positions, brightness and candidate hosts. In effect the model stands in for a person scanning images and judging resemblance. The authors note that results shift with how a text query is phrased, and that source types missing from the fine-tuning data, such as supernova remnants, may not be found.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors fine-tuned the open-source OpenCLIP image-text model on 2,900 radio galaxies from the RadioGalaxyNET dataset, covering FR-I, FR-II, FR-x, R-type and peculiar morphologies, using three-channel radio-radio-infrared cutouts paired with expanded text descriptions. Evaluated on held-out splits drawn ten times, the model reached 84±3 % accuracy over the main morphological categories after 100 epochs, with the greatest confusion between FR-I and FR-x. The fine-tuned encoders were then used to embed about 170,000 extended radio sources drawn from 160 first-year EMU survey tiles, and these embeddings back a search engine, EMUSE, that returns similar sources for a text prompt or a query image. The authors report that text queries are sensitive to phrasing and that source types absent from the fine-tuning set, such as cluster relics and supernova remnants, may not be retrieved.
How AI was used
Radio image tiles from the EMU first-year observations and matching AllWISE W1 infrared mosaics were cut around Selavy-based source positions, and each cutout was passed through the Gal-DINO object detection model inside the RG-CAT pipeline to produce per-tile catalogues of bounding boxes, categories, confidence scores and candidate infrared hosts. A fine-tuning set was built from RadioGalaxyNET radio galaxies plus an added category of peculiar and rare morphologies, rendered as three-channel PNGs with two clipped radio channels and one AllWISE W1 channel, and paired with morphological text descriptions expanded from the class and subcategory labels. OpenCLIP, pre-trained on LAION image-text pairs, was adapter-fine-tuned on these pairs using only the contrastive objective, on a single NVIDIA H100 GPU for 100 epochs; accuracy and confusion were measured on 80:20 splits sampled ten times, while the deployed model was fine-tuned on the full dataset. The fine-tuned image encoder was then run over cutouts for the extended sources filtered from the catalogues to build an embedding database with catalogue metadata. At query time, a text prompt is tokenised and embedded with the text encoder, or an image is preprocessed and embedded with the image encoder, and cosine similarity against the stored embeddings returns the top-k sources above a probability threshold through a Streamlit application.
The shape of the work
Structural · the record, drawn
no AI
Assemble radio and infrared survey tiles
Obtaining raw data, whether by measurement, download or retrieval.
This study uses data from EMU’s first-year observations, covering 160 tiles (4,500 square degrees).where the paper describes this · verbatim
AI
Detect radio sources and infrared hosts with Gal-DINO
Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.
Each cutout is analysed with Gal-DINO to extract bounding boxes, categories, and confidence scoreswhere the paper describes this · verbatim
no AI
Build image-text fine-tuning set
Cleaning, filtering, normalising or labelling data already obtained.
we generate 4′×4′ image cutouts from the EMU-PS1 survey and corresponding cutouts from the AllWISE surveywhere the paper describes this · verbatim
AI
Adapter fine-tune OpenCLIP
Fitting model parameters, including fine-tuning an existing model.
we fine-tune the pre-trained OpenCLIP model on a single NVIDIA H100 GPU for 100 epochswhere the paper describes this · verbatim
AI
Embed extended EMU sources into searchable database
Encoding data into features, descriptors, embeddings or graphs.
The fine-tuned model is then used to generate image embeddings for each PNG.where the paper describes this · verbatim
AI
Encode user text or image query
Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.
the input is first tokenised using the OpenCLIP tokeniser, and its embedding is obtained through the fine-tuned model’s text encoderwhere the paper describes this · verbatim
no AI
Rank database by cosine similarity and return top-k
Reducing a candidate set by filtering or ranking, in a single pass.
we compute the similarity between the query embedding (either derived from a text or an image query) and the precomputed embeddingswhere the paper describes this · verbatim
no AI
Evaluate accuracy and inspect retrieved sources
Testing outputs against ground truth.
accuracy exceeds 50% after a single epoch and gradually increases to 84±3 % after 100 epochswhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's object is a fine-tuned multimodal model and the retrieval tool built on it; the reported results are the model's classifications and retrievals, so the contribution does not exist without the model.
we fine-tune OpenCLIP, a multimodal foundation model, using radio source images and their corresponding textual descriptionswhere the paper describes this · verbatim
we split the radio source dataset into an 80:20 ratio for training and testingwhere the paper describes this · verbatim
The OpenCLIP model with fine-tuning settings is available at https://github.com/Nikhel1/Finetune_OpenCLIPwhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- How many were testedThe paper gives no count of what was tested.
- Version of OpenCLIPWhich version of the model was used is not stated.
- Version of Gal-DINOWhich version of the model was used is not stated.
- What step 4 replacedThe paper gives no basis for what the AI stood in for.
- What step 5 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00061, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error