~/aixsci
200 records · all checked

structural-biology/ai produced the result/Advanced Science 2026 · v2

Attention patterns inside a protein language model used to cut sequences into reusable units

Researchers split protein sequences into short segments they call protein words, using the internal attention patterns of the ESM2 language model, then trained a second model to link those words to molecular functions.

1. Assemble sequence and annotation datasets2. Extract ESM2 attention matrices and residue embeddings3. Binarize attention and cluster residues into raw words4. Compile degenerate UniRef50 and Pfam dictionaries5. Prioritize words for an analyte sequence by dictionary lookup6. Train Word2Function on word embeddings to predict GO terms7. Attribute words to GO terms and populate WordTableGO658. Evaluate coverage and GO prediction against baselines

spectrum · one line per step, placed by what the step does · bright lines used AI

Automatically Defining Protein Words for Diverse Functional Predictions Based on Attention Analysis of a Protein Language Model
Advanced Science, 2026

doi:10.1002/advs.202521970 · record aix-00072 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Segmentation, Classification
Model family
Protein language model, Transformer
Checked by
Benchmark182 tested
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Proteins are long chains of amino acids, and a chain's order determines what the protein does. Only some stretches matter for a given job: a few residues may grip a metal ion, bind DNA or form the pocket where a chemical reaction happens. Finding those stretches in a new sequence is hard. The classic approach is a dictionary of motifs, short patterns written down by curators from proteins already studied in the laboratory. That works when a new protein resembles an old one, but much of the protein world does not. Whole families are labelled domains of unknown function, meaning nobody has yet pinned a job to them.

The researchers set out to build such units automatically, without human curation. The idea was to divide a sequence into short segments, keep the ones that recur across many proteins, collect them into dictionaries, and then test whether those segments land on residues already known to matter and whether they can be mapped to functions.

Where AI came in

The segmentation came from a protein language model, ESM2-650M, which had been trained on large numbers of sequences to predict residues from their surroundings. Such a model builds attention matrices, internal tables recording which positions in a sequence it treats as related. Each sequence was passed through the model as supplied, without further tuning, yielding 660 such matrices and a 1280-number description of each residue. The matrices were then thresholded and cut into residue groups by a graph algorithm, Louvain, which finds clusters of densely connected points. Groups of 5 to 20 residues were kept as raw words. The model stood in for the human judgement behind curated motif dictionaries.

A second model, Word2Function, was trained from scratch on descriptions of these words, drawn from 56 299 sequences carrying 65 labels for molecular function, to predict those labels for a protein. An attribution method, Integrated Gradients, then scored how much each word contributed to a prediction, and the top-scoring words were written into a table keyed to the 65 functions. Performance was compared with the curated motif tools PROSITE and InterProScan and with other trained predictors.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The study parses protein sequences into units it calls 'protein words' by binarizing the 660 attention matrices that the ESM2-650M protein language model produces for a sequence and clustering residues with the Louvain community detection algorithm, then compiles dictionaries of high-occurrence words from 1 million UniRef50 sequences and from individual Pfam families (20 762 family dictionaries). A supervised transformer model, Word2Function, was trained on word embeddings from 56 299 sequences labelled with 65 GO molecular-function terms, and Integrated Gradients attributions were used to assign individual words to GO terms in a table called WordTableGO65. On 182 proteins from the ProteinGym deep mutational scanning set, median functional residue coverage was 0.900 with the combined Pfam and UniRef50 dictionaries, compared with 0.045 for the motif-based tool PROSITE and 0.017 for InterProScan; for whole-protein GO-term prediction on the PWNet dataset, mean functional MCC was 0.284 for Word2Function and 0.150 for PROSITE. The word table was also matched against 1 494 146 InterPro sequences annotated as domains of unknown function and against MHC peptides from the MHC Motif Atlas, where the UniRef50 dictionary matched 2832 MHC-I peptides versus 373 for a random k-mer dictionary.

How AI was used

A pre-trained protein language model, ESM2-650M, was run without fine-tuning on each input sequence (length up to 1024, canonical residues only) to obtain 660 attention matrices and 1280-dimensional per-residue embeddings. Each attention matrix was binarized with a dual cutoff (a fixed threshold of 0.1, plus a proportional threshold discarding the lowest 20% of scores in heads with diffuse attention), converted to a directed residue graph, and partitioned with the Louvain algorithm (resolution 1.0, modularity gain threshold 1e-7); communities of 5–20 residues, contiguous or with up to two gaps, were kept as raw words. Running this in multi-sequence mode over 1 million UniRef50 sequences (50 per each of 20 000 Pfam families) and over individual families produced a common dictionary and family dictionaries, with the 20 amino acids collapsed into 12 degenerate types and length-stratified occurrence thresholds; in single-sequence mode, raw words for an analyte sequence were degenerated and matched exactly against these dictionaries to prioritize words. For function mapping, word embeddings were formed by averaging the ESM2 residue embeddings inside each word and fed to transformer layers plus a fully connected layer (embedding and hidden dimension 1280, 20 attention heads) trained with binary cross-entropy, Adam, batch size 256, weight decay 1e-5, initial learning rate 1e-3, cosine decay and early stopping, to predict 65 GO terms over the ExpGO65 sequences. Integrated Gradients was then applied to the trained model to score each word's contribution, and words whose cumulative positive attribution reached a cutoff of half the total were written to the WordTableGO65 table, which was subsequently used for whole-protein annotation by exact word matching. Baselines were obtained by submitting sequences to ScanProsite and InterProScan, retraining ProtNote on the same ExpGO65 splits, and querying the GPSFun web server.

The shape of the work

Structural · the record, drawn

ACQUISITIONREPRESENTATIONINTERPRETATIONPREPARATIONSCREENINGTRAININGINTERPRETATIONVALIDATION12345678AIAIAIAIAssemble sequenceand annotationdatasetsExtract ESM2attentionmatrices and res…Binarizeattention andcluster residues…CompiledegenerateUniRef50 and Pfa…Prioritize wordsfor an analytesequence by dict…TrainWord2Function onword embeddings …Attribute wordsto GO terms andpopulate WordTab…Evaluate coverageand GO predictionagainst baselines↤ manual curation↤ manual curation
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Assemble sequence and annotation datasets

Obtaining raw data, whether by measurement, download or retrieval.

Protein sequences were downloaded from UniProt (2024/01)where the paper describes this · verbatim
in the paper
2Representation
AI

Extract ESM2 attention matrices and residue embeddings

Encoding data into features, descriptors, embeddings or graphs.

By inputting an analyte protein sequence into ESM2, we obtain 660 attention matrices from the corresponding attention heads.where the paper describes this · verbatim
in the paper
3Interpretation
no AI

Binarize attention and cluster residues into raw words

Extracting understanding from model behaviour.

We then use the Louvain algorithm, a graph‐based community detection algorithm, to segment the binary matrix into communities.where the paper describes this · verbatim
in the paper
4Preparation
no AI

Compile degenerate UniRef50 and Pfam dictionaries

Cleaning, filtering, normalising or labelling data already obtained.

We stratified the dictionary by length, retaining high occurrence raw words for each length.where the paper describes this · verbatim
in the paper
5Screening
no AI

Prioritize words for an analyte sequence by dictionary lookup

Reducing a candidate set by filtering or ranking, in a single pass.

we convert raw words into their degenerate forms and then search for exact matcheswhere the paper describes this · verbatim
in the paper
6Training
AI

Train Word2Function on word embeddings to predict GO terms

Fitting model parameters, including fine-tuning an existing model. The AI stood in for manual curation.

During training, we employed binary cross‐entropy loss to optimize the model, using a batch size of 256where the paper describes this · verbatim
in the paper
7Interpretation
AI

Attribute words to GO terms and populate WordTableGO65

Extracting understanding from model behaviour. The AI stood in for manual curation.

we use the Integrated Gradients algorithm to quantify the contribution of each word to the predictionwhere the paper describes this · verbatim
in the paper
8Validation
AI

Evaluate coverage and GO prediction against baselines

Testing outputs against ground truth.

The median and mean values for functional residue coverage using PROSITE were 0.045 and 0.393where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The object of the paper is the definition of 'protein words' obtained from ESM2 attention matrices and the GO-term mapping learned on those words; every reported result derives from the model's attention or from the trained Word2Function model.

+What the AI was for
The original ESM2‐650 M variant was run for all input sequences (without specific finetuning) to generate attention matriceswhere the paper describes this · verbatim
+Model families
+How it was taught
Self-supervisedSupervisedin the paper
+Models named
ESM2 ESM2-650M · Off the shelfWord2Function · Trained from scratchProtNote · Trained from scratchGPSFun · Off the shelfin the paper
+How results were checked
Benchmark182 testedin the paper
When applied to 182 proteins in the DMS dataset with sequence lengths between 50 and 1024 amino acidswhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
The data that support the findings of this study are openly available in ProteinWordwise atwhere the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 7 items
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • Version of Word2FunctionWhich version of the model was used is not stated.
  • Version of ProtNoteWhich version of the model was used is not stated.
  • Version of GPSFunWhich version of the model was used is not stated.
  • What step 2 replacedThe paper gives no basis for what the AI stood in for.
  • What step 8 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00072, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error