~/aixsci
200 records · all checked

structural-biology/ai produced the result/Bioinformatics 2023 · v2

Neural network assigns protein domains to structural families from sequence alone

Researchers built CATHe, a small neural network that reads numerical summaries of protein sequences produced by a language model and sorts the sequences into CATH structural superfamilies. It annotated 4.62 million previously unassigned domains.

1. Build non-redundant superfamily datasets2. Encode domain sequences as language-model embeddings3. Train CATHe classifier and baseline models4. Evaluate on held-out test splits5. Assemble Pfam domains missed by CATH-HMMs6. Predict CATH superfamilies for Pfam domains7. Structurally check assignments with SSAP8. Manually curate domains below SSAP thresholds

spectrum · one line per step, placed by what the step does · bright lines used AI

CATHe: detection of remote homologues for CATH superfamilies using embeddings from protein language models
Bioinformatics, 2023

doi:10.1093/bioinformatics/btad029 · record aix-00166 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Classification
Model family
Protein language model, Multilayer perceptron, Linear model
Checked by
Held-out197 tested, 179 worked
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Proteins are chains of amino acids that fold into shapes, and the shape largely decides what a protein does. Biologists group folded units, called domains, into families of relatives that share an ancestor and a common architecture. The CATH database is one such classification, sorting domains into superfamilies. The difficulty is that evolution erases similarity in the sequence long before it changes the shape. Two domains can fold the same way while sharing almost none of their letters. Standard tools compare sequences letter by letter, or against statistical profiles built from known families, and those tools lose the trail once the resemblance grows faint. Such distant relatives are known as remote homologues.

The researchers set out to assign domain sequences to CATH superfamilies without relying on visible sequence similarity, and then to use that method on domains in the Pfam database that CATH's existing profile models could not place.

Where AI came in

A protein language model, ProtT5, was used as released to turn each domain sequence into a fixed-length list of numbers, an embedding. Such models are trained on large collections of sequences without labels, and the embedding is a compressed summary of what the model has learned about a sequence. The researchers then trained a shallow artificial neural network, CATHe, on those embeddings to predict the superfamily. Training, validation and test sets were kept below 20% sequence identity, so the network could not succeed by spotting near-identical sequences. Accuracy was 85.6 ± 0.4% across 1773 superfamilies and 98.2 ± 0.3% across the 50 largest.

The embedding and the network stood in for sequence comparison: CATHe was measured against homology-based inference with BLAST, logistic regression, a network trained on ProtBERT embeddings, a network given only sequence lengths, and a random baseline. The paper's main output exists only because the network produced it, namely superfamily labels for 4.62 million Pfam domains kept above a 0.9 probability threshold. Those labels were then checked by non-AI means. For 197 human domains with good AlphaFold2 models, structure comparison with SSAP plus expert inspection confirmed 179.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

CATHe assigns protein domain sequences to CATH superfamilies using an artificial neural network trained on embeddings from the ProtT5 protein language model. Training, validation and test sets were built so that no sequence shared more than 20% identity with another, and the models reached an accuracy of 85.6 ± 0.4% on a set of 1773 superfamilies and 98.2 ± 0.3% on the 50 largest superfamilies. Applied to Pfam domains in UniProt that CATH HMMs could not assign, CATHe produced CATH superfamily annotations for 4.62 million domains at an expected error rate below 0.5%. For 197 human Pfam domains with good-quality AlphaFold2 models, 179 (90.86%) of the CATHe assignments were confirmed by SSAP structure comparison together with manual curation.

How AI was used

Domain sequences from CATH v4.3 were split into a CATH-Gene3D training set and PDB-derived validation and test sets, filtered to below 20% sequence identity within and between splits, yielding the TOP 1773 SUPERFAMILIES and TOP 50 SUPERFAMILIES datasets. Each domain sequence was converted into a fixed-length embedding using the pre-trained ProtT5 and ProtBERT protein language models, which were used as released. A shallow artificial neural network (CATHe) was then trained on the ProtT5 embeddings to classify domains into superfamilies, alongside six comparison models the authors ran themselves: an ANN on ProtBERT embeddings, an ANN on sequence lengths alone, logistic regression on each embedding type, homology-based inference via BLAST, and a random baseline. Performance was measured on the held-out test splits with accuracy, F1-score, MCC and balanced accuracy, with confidence intervals from bootstrapping. For application, UniProt regions that did not match CATH-HMMs but did match Pfam-HMMs were clustered at 60% identity, embedded with ProtT5, and scored by the trained CATHe model, keeping only predictions above a 0.9 probability threshold. Assignments for a human subset were then compared by SSAP against structural relatives in the predicted superfamily, using AlphaFold2 models from the EBI database, with non-matching cases inspected manually. Embeddings from the trained network's final layer were projected with t-SNE to examine what the classifier had learned.

The shape of the work

Structural · the record, drawn

PREPARATIONREPRESENTATIONTRAININGVALIDATIONACQUISITIONINFERENCEVALIDATIONVALIDATION12345678AIAIAIAIBuildnon-redundantsuperfamily data…Encode domainsequences aslanguage-model e…Train CATHeclassifier andbaseline modelsEvaluate onheld-out testsplitsAssemble Pfamdomains missed byCATH-HMMsPredict CATHsuperfamilies forPfam domainsStructurallycheck assignmentswith SSAPManually curatedomains belowSSAP thresholds↤ conventional algorithm↤ statistical model↤ statistical model
AI stepNo AI↤ what the AI stood in for
1Preparation
no AI

Build non-redundant superfamily datasets

Cleaning, filtering, normalising or labelling data already obtained.

we ensured there was less than 20% sequence identity between and within the three setswhere the paper describes this · verbatim
in the paper
2Representation
AI

Encode domain sequences as language-model embeddings

Encoding data into features, descriptors, embeddings or graphs. The AI stood in for conventional algorithm.

we generated sequence embeddings for all of themwhere the paper describes this · verbatim
in the paper
3Training
AI

Train CATHe classifier and baseline models

Fitting model parameters, including fine-tuning an existing model. The AI stood in for statistical model.

CATHe model (which is an ANN model trained on ProtT5 embeddings)where the paper describes this · verbatim
in the paper
4Validation
AI

Evaluate on held-out test splits

Testing outputs against ground truth.

The performance of CATHe and the six other baseline models was measured on the testing set of the TOP 1773 SUPERFAMILIES dataset.where the paper describes this · verbatim
in the paper
5Acquisition
no AI

Assemble Pfam domains missed by CATH-HMMs

Obtaining raw data, whether by measurement, download or retrieval.

Any regions not matching the CATH-HMMs were scanned against the library of Pfam-HMMswhere the paper describes this · verbatim
in the paper
6Inference
AI

Predict CATH superfamilies for Pfam domains

Running a trained model over new data to predict, classify or score. The AI stood in for statistical model.

Applying the threshold probability of 0.9 to CATHe predictions for the PFAM-S60 setwhere the paper describes this · verbatim
in the paper
7Validation
no AI

Structurally check assignments with SSAP

Testing outputs against ground truth.

we used our in-house protein structure comparison method, SSAP to compare them against structural relativeswhere the paper describes this · verbatim
in the paper
8Validation
no AI

Manually curate domains below SSAP thresholds

Testing outputs against ground truth.

manual curation on the 55 domains that did not cross the SSAP thresholds confirmed that 37 more domains were valid superfamily matcheswhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's result is the set of CATH superfamily assignments produced by the trained neural network; the superfamily predictions for 4.62 million Pfam domains exist only because the model made them.

+What the AI was for
Classificationin the paper
uses embeddings from ProtT5 as input to train machine learning models to classify protein sequences into CATH superfamilieswhere the paper describes this · verbatim
+How it was taught
SupervisedSelf-supervisedin the paper
+Models named
ProtT5 · Off the shelfProtBERT · Off the shelfCATHe (ANN on ProtT5 embeddings) · Trained from scratchANN on ProtBERT embeddings · Trained from scratchANN on sequence lengths · Trained from scratchLogistic regression on ProtT5 embeddings · Trained from scratchLogistic regression on ProtBERT embeddings · Trained from scratchin the paper
+How results were checked
Held-out197 tested, 179 workedin the paper
a total of 179 domains out of 197 (90.86%) from Pfam-human that matched CATH superfamilies using CATHewhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
The code for the developed models is available on https://github.com/vam-sin/CATHewhere the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 9 items
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • Version of ProtT5Which version of the model was used is not stated.
  • Version of ProtBERTWhich version of the model was used is not stated.
  • Version of CATHe (ANN on ProtT5 embeddings)Which version of the model was used is not stated.
  • Version of ANN on ProtBERT embeddingsWhich version of the model was used is not stated.
  • Version of ANN on sequence lengthsWhich version of the model was used is not stated.
  • Version of Logistic regression on ProtT5 embeddingsWhich version of the model was used is not stated.
  • Version of Logistic regression on ProtBERT embeddingsWhich version of the model was used is not stated.

About this article

Record aix-00166, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error