structural-biology/ai produced the result/Bioinformatics 2023 · v2
Neural network assigns protein domains to structural families from sequence alone
Researchers built CATHe, a small neural network that reads numerical summaries of protein sequences produced by a language model and sorts the sequences into CATH structural superfamilies. It annotated 4.62 million previously unassigned domains.
spectrum · one line per step, placed by what the step does · bright lines used AI
CATHe: detection of remote homologues for CATH superfamilies using embeddings from protein language models
Bioinformatics, 2023
doi:10.1093/bioinformatics/btad029 · record aix-00166 v2 · checked 2026-10-09
- AI was for
- Classification
- Model family
- Protein language model, Multilayer perceptron, Linear model
- Checked by
- Held-out197 tested, 179 worked
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Proteins are chains of amino acids that fold into shapes, and the shape largely decides what a protein does. Biologists group folded units, called domains, into families of relatives that share an ancestor and a common architecture. The CATH database is one such classification, sorting domains into superfamilies. The difficulty is that evolution erases similarity in the sequence long before it changes the shape. Two domains can fold the same way while sharing almost none of their letters. Standard tools compare sequences letter by letter, or against statistical profiles built from known families, and those tools lose the trail once the resemblance grows faint. Such distant relatives are known as remote homologues.
The researchers set out to assign domain sequences to CATH superfamilies without relying on visible sequence similarity, and then to use that method on domains in the Pfam database that CATH's existing profile models could not place.
Where AI came in
A protein language model, ProtT5, was used as released to turn each domain sequence into a fixed-length list of numbers, an embedding. Such models are trained on large collections of sequences without labels, and the embedding is a compressed summary of what the model has learned about a sequence. The researchers then trained a shallow artificial neural network, CATHe, on those embeddings to predict the superfamily. Training, validation and test sets were kept below 20% sequence identity, so the network could not succeed by spotting near-identical sequences. Accuracy was 85.6 ± 0.4% across 1773 superfamilies and 98.2 ± 0.3% across the 50 largest.
The embedding and the network stood in for sequence comparison: CATHe was measured against homology-based inference with BLAST, logistic regression, a network trained on ProtBERT embeddings, a network given only sequence lengths, and a random baseline. The paper's main output exists only because the network produced it, namely superfamily labels for 4.62 million Pfam domains kept above a 0.9 probability threshold. Those labels were then checked by non-AI means. For 197 human domains with good AlphaFold2 models, structure comparison with SSAP plus expert inspection confirmed 179.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
CATHe assigns protein domain sequences to CATH superfamilies using an artificial neural network trained on embeddings from the ProtT5 protein language model. Training, validation and test sets were built so that no sequence shared more than 20% identity with another, and the models reached an accuracy of 85.6 ± 0.4% on a set of 1773 superfamilies and 98.2 ± 0.3% on the 50 largest superfamilies. Applied to Pfam domains in UniProt that CATH HMMs could not assign, CATHe produced CATH superfamily annotations for 4.62 million domains at an expected error rate below 0.5%. For 197 human Pfam domains with good-quality AlphaFold2 models, 179 (90.86%) of the CATHe assignments were confirmed by SSAP structure comparison together with manual curation.
How AI was used
Domain sequences from CATH v4.3 were split into a CATH-Gene3D training set and PDB-derived validation and test sets, filtered to below 20% sequence identity within and between splits, yielding the TOP 1773 SUPERFAMILIES and TOP 50 SUPERFAMILIES datasets. Each domain sequence was converted into a fixed-length embedding using the pre-trained ProtT5 and ProtBERT protein language models, which were used as released. A shallow artificial neural network (CATHe) was then trained on the ProtT5 embeddings to classify domains into superfamilies, alongside six comparison models the authors ran themselves: an ANN on ProtBERT embeddings, an ANN on sequence lengths alone, logistic regression on each embedding type, homology-based inference via BLAST, and a random baseline. Performance was measured on the held-out test splits with accuracy, F1-score, MCC and balanced accuracy, with confidence intervals from bootstrapping. For application, UniProt regions that did not match CATH-HMMs but did match Pfam-HMMs were clustered at 60% identity, embedded with ProtT5, and scored by the trained CATHe model, keeping only predictions above a 0.9 probability threshold. Assignments for a human subset were then compared by SSAP against structural relatives in the predicted superfamily, using AlphaFold2 models from the EBI database, with non-matching cases inspected manually. Embeddings from the trained network's final layer were projected with t-SNE to examine what the classifier had learned.
The shape of the work
Structural · the record, drawn
no AI
Build non-redundant superfamily datasets
Cleaning, filtering, normalising or labelling data already obtained.
we ensured there was less than 20% sequence identity between and within the three setswhere the paper describes this · verbatim
AI
Encode domain sequences as language-model embeddings
Encoding data into features, descriptors, embeddings or graphs. The AI stood in for conventional algorithm.
we generated sequence embeddings for all of themwhere the paper describes this · verbatim
AI
Train CATHe classifier and baseline models
Fitting model parameters, including fine-tuning an existing model. The AI stood in for statistical model.
CATHe model (which is an ANN model trained on ProtT5 embeddings)where the paper describes this · verbatim
AI
Evaluate on held-out test splits
Testing outputs against ground truth.
The performance of CATHe and the six other baseline models was measured on the testing set of the TOP 1773 SUPERFAMILIES dataset.where the paper describes this · verbatim
no AI
Assemble Pfam domains missed by CATH-HMMs
Obtaining raw data, whether by measurement, download or retrieval.
Any regions not matching the CATH-HMMs were scanned against the library of Pfam-HMMswhere the paper describes this · verbatim
AI
Predict CATH superfamilies for Pfam domains
Running a trained model over new data to predict, classify or score. The AI stood in for statistical model.
Applying the threshold probability of 0.9 to CATHe predictions for the PFAM-S60 setwhere the paper describes this · verbatim
no AI
Structurally check assignments with SSAP
Testing outputs against ground truth.
we used our in-house protein structure comparison method, SSAP to compare them against structural relativeswhere the paper describes this · verbatim
no AI
Manually curate domains below SSAP thresholds
Testing outputs against ground truth.
manual curation on the 55 domains that did not cross the SSAP thresholds confirmed that 37 more domains were valid superfamily matcheswhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is the set of CATH superfamily assignments produced by the trained neural network; the superfamily predictions for 4.62 million Pfam domains exist only because the model made them.
uses embeddings from ProtT5 as input to train machine learning models to classify protein sequences into CATH superfamilieswhere the paper describes this · verbatim
a total of 179 domains out of 197 (90.86%) from Pfam-human that matched CATH superfamilies using CATHewhere the paper describes this · verbatim
The code for the developed models is available on https://github.com/vam-sin/CATHewhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- Version of ProtT5Which version of the model was used is not stated.
- Version of ProtBERTWhich version of the model was used is not stated.
- Version of CATHe (ANN on ProtT5 embeddings)Which version of the model was used is not stated.
- Version of ANN on ProtBERT embeddingsWhich version of the model was used is not stated.
- Version of ANN on sequence lengthsWhich version of the model was used is not stated.
- Version of Logistic regression on ProtT5 embeddingsWhich version of the model was used is not stated.
- Version of Logistic regression on ProtBERT embeddingsWhich version of the model was used is not stated.
About this article
Record aix-00166, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error