~/aixsci
200 records · all checked

structural-biology/ai produced the result/Briefings in Bioinformatics 2025 · v2

Model labels which protein residues are active sites and what job they do

Researchers built M3Site, a model that reads a protein's sequence, its predicted three-dimensional shape and a written description of its function, then labels each building block as an active site of one of six functional kinds.

1. Collect and filter sequences, annotations and structures2. Cluster raw functional annotations into categories with an LLM3. Expert refinement, similarity clustering and dataset splitting4. Encode sequence, structure graph and function text into features5. Train M3Site and retrained baseline models6. Predict residue-level active site classes on held-out and unseen proteins7. Score predictions against annotations, baselines and ablations

spectrum · one line per step, placed by what the step does · bright lines used AI

M3Site: multiclass multimodal learning for protein active site identification and classification
Briefings in Bioinformatics, 2025

doi:10.1093/bib/bbaf590 · record aix-00140 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Classification
Model family
Transformer, Graph neural network, Protein language model, Large language model
Checked by
Held-out33 tested
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Proteins are long chains of building blocks called amino acid residues, folded into a shape. Most of the chain holds the structure together, but a small number of residues do the chemical work: they grip another molecule and change it. These are the active site. Knowing which residues they are, and what each one contributes, is how biologists explain what a protein does. Finding them in the laboratory is slow, so most proteins in the databases have no such annotation. The clues are scattered: the order of the residues, how close they sit in the folded shape, and what is already written about the protein's role.

The researchers set out to label residues not just as active site or not, but by the kind of chemical job they do. They gathered 25 883 proteins from Swiss-Prot, a hand-curated protein database, together with predicted shapes from the AlphaFold database, and sorted the existing free-text notes about active sites into six functional categories to use as labels.

Where AI came in

AI appears twice. First in building the dataset: the raw function notes in the database are written in loose prose, so OpenAI's o1 model was asked to group them into a handful of biologically sensible categories and to flag cases it found ambiguous. Human experts then refined these into the six final categories. Here the model stood in for the first pass of manual curation.

Then in the model itself. A protein language model, trained on sequences much as a text model is trained on words, turns the chain into numbers. A graph neural network reads the folded shape as a network of nearby residues. A language model trained on biomedical text reads the function description. The three readings are combined and a classifier assigns each residue a category, in place of laboratory measurement.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

M3Site is a model that labels each residue of a protein as a non-active site or as one of several active site functional categories, using three inputs at once: a sequence embedding from a protein language model, a graph of the predicted 3D structure encoded by an equivariant graph neural network, and a text description of the protein's function encoded by a biomedical language model. The authors assembled a dataset of 25 883 proteins from Swiss-Prot and the AlphaFold database, and used OpenAI's o1 model plus expert review to consolidate raw active site annotations into six functional categories. On the split clustered at 10% sequence similarity, the paper reports higher scores for M3Site than for the protein representation learning baselines it retrained, across the seven metrics it tabulates; on 33 proteins deposited after the dataset cut-off it reports an F1 of 0.912, and on a held-out Glycosyl-Hydrolase family split an F1 of 0.884. A Gradio application for prediction and visualisation accompanies the code and dataset.

How AI was used

An LLM (OpenAI o1, accessed via API) was prompted to cluster raw UniProt active site function annotations into 5–10 biologically meaningful groups and flag ambiguous cases; domain experts then refined these into six categories used as class labels, and MMSeqs2 sequence clustering at thresholds of 10% to 90% plus an 8:1:1 split produced the training, validation and test sets. Each protein was represented three ways: sequence embeddings from a pretrained protein language model (ESM-3 in the main configuration), a residue graph with edges between Cα atoms closer than 8 Å and 62-dimensional node features, encoded by a two-layer equivariant graph neural network, and a function description embedded with a biomedical language model (PubMedBERT-abs) and globally average pooled. A symmetric cross-attention module (FunICross) fuses the modalities in two branches, and an adaptive weighted fusion mechanism combines the fused and original features with a learned sigmoid coefficient. The classifier was trained end to end for 100 epochs with Adam at a learning rate of 5e-5, warm-up over 10% of steps and cosine annealing, under a loss combining weighted cross entropy with a class-centre loss and an inter-centre margin loss. Baselines, and variants substituting other protein and biomedical language models, were run by the authors with additional transformer layers and a classification head appended for comparable trainable parameter counts. The trained model was then run on the held-out split, on proteins deposited after the dataset cut-off date, on a Glycosyl-Hydrolase family hold-out, and inside a Gradio interface that takes a user's structure and function description.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONPREPARATIONREPRESENTATIONTRAININGINFERENCEVALIDATION1234567AIAIAIAICollect andfilter sequences,annotations and …Cluster rawfunctionalannotations into…Expertrefinement,similarity clust…Encode sequence,structure graphand function tex…Train M3Site andretrainedbaseline modelsPredictresidue-levelactive site clas…Score predictionsagainstannotations, bas…↤ manual curation↤ conventional algorithm↤ physical experiment
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Collect and filter sequences, annotations and structures

Obtaining raw data, whether by measurement, download or retrieval.

We began with collecting protein sequences from the Swiss-Prot in UniProt, deposited before 28 July 2024.where the paper describes this · verbatim
in the paper
2Preparation
AI

Cluster raw functional annotations into categories with an LLM

Cleaning, filtering, normalising or labelling data already obtained. The AI stood in for manual curation.

we employed OpenAI’s o1 model to cluster raw functional annotationswhere the paper describes this · verbatim
in the paper
3Preparation
no AI

Expert refinement, similarity clustering and dataset splitting

Cleaning, filtering, normalising or labelling data already obtained.

we employed MMSeqs2 to cluster sequences based on pairwise similaritywhere the paper describes this · verbatim
in the paper
4Representation
AI

Encode sequence, structure graph and function text into features

Encoding data into features, descriptors, embeddings or graphs. The AI stood in for conventional algorithm.

Sequence features were extracted using pretrained PLM, while functional text descriptions were encoded with biomedical language models (BLM).where the paper describes this · verbatim
in the paper
5Training
AI

Train M3Site and retrained baseline models

Fitting model parameters, including fine-tuning an existing model.

The model was trained for 100 epochs using the Adam optimizer with a learning rate of 5e-5.where the paper describes this · verbatim
in the paper
6Inference
AI

Predict residue-level active site classes on held-out and unseen proteins

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

we conducted a case study using unseen data collected from UniProt after 28 July 2024where the paper describes this · verbatim
in the paper
7Validation
no AI

Score predictions against annotations, baselines and ablations

Testing outputs against ground truth.

The evaluation metrics used include Precision, Recall, F1, AUROC, AUPRC, and MCCwhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's result is the performance of a learned multimodal model that assigns active-site classes to residues; every reported finding comes from running that model.

+What the AI was for
Classificationin the paper
a multimodal framework that integrates protein sequence embeddings, structural graph representations, and functional text annotationswhere the paper describes this · verbatim
+How it was taught
SupervisedTransfer / fine-tuningZero-shotin the paper
+Models named
M3Site · Trained from scratchESM-3 · Fine-tunedPubMedBERT-abs abs · Fine-tunedPubMedBERT-full full · Fine-tunedEGNN · Trained from scratchOpenAI o1 · Off the shelfESM-1b · Fine-tunedESM-1v · Fine-tunedESM-2-650M 650M · Fine-tunedProtT5-BFD BFD · Fine-tunedProtT5-UniRef UniRef · Fine-tunedProtBert · Fine-tunedProtAlbert · Fine-tunedProtXLNet · Fine-tunedProtElectra · Fine-tunedPETA · Fine-tunedS-PLM · Fine-tunedTAPE · Fine-tunedMIF · Fine-tunedPST · Fine-tunedin the paper
+How results were checked
Held-out33 testedin the paper
we identified 33 samples for evaluationwhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
The dataset, source code for experiments, and interactive application associated with this study are publicly availablewhere the paper describes this · verbatim
+Compute
NVIDIA GeForce RTX 4090 GPUs; PyTorch 2.2.2; 100 training epochsin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 17 items
  • Trained model weightsWhether the trained model is available is not stated.
  • Version of M3SiteWhich version of the model was used is not stated.
  • Version of ESM-3Which version of the model was used is not stated.
  • Version of EGNNWhich version of the model was used is not stated.
  • Version of OpenAI o1Which version of the model was used is not stated.
  • Version of ESM-1bWhich version of the model was used is not stated.
  • Version of ESM-1vWhich version of the model was used is not stated.
  • Version of ProtBertWhich version of the model was used is not stated.
  • Version of ProtAlbertWhich version of the model was used is not stated.
  • Version of ProtXLNetWhich version of the model was used is not stated.
  • Version of ProtElectraWhich version of the model was used is not stated.
  • Version of PETAWhich version of the model was used is not stated.
  • Version of S-PLMWhich version of the model was used is not stated.
  • Version of TAPEWhich version of the model was used is not stated.
  • Version of MIFWhich version of the model was used is not stated.
  • Version of PSTWhich version of the model was used is not stated.
  • What step 5 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00140, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error