~/aixsci
200 records · all checked

structural-biology/ai produced the result/Bioinformatics Advances 2026 · v2

Neural networks trained to find the cutting site on serpin proteins

Serpins are proteins that disable enzymes, and the short loop that does the work is hard to locate from sequence alone. Researchers trained neural networks on expert-labelled serpins to mark that loop residue by residue.

1. Collect serpin and non-serpin sequences2. Expert annotation of RCL regions3. Encode sequences4. Train per-residue classifiers5. Predict RCL positions on held-out sequences6. Evaluate against UniProt annotations

spectrum · one line per step, placed by what the step does · bright lines used AI

U-Net-based reactive center loop-identifier for serpins
Bioinformatics Advances, 2026

doi:10.1093/bioadv/vbag132 · record aix-00152 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Segmentation
Model family
Convolutional neural network, Recurrent neural network, Protein language model
Checked by
Held-out78 tested
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Proteins are chains of amino acids, and serpins are a family of such chains that block enzymes which cut other proteins. The part that does the blocking is a short, exposed stretch called the reactive centre loop, or RCL. The enzyme bites the loop, and the serpin snaps shut around it. Knowing where that loop sits in a given serpin sequence matters for understanding how it works. The difficulty is that the loop's sequence varies greatly from one serpin to another. There is no shared pattern of letters to search for, so the usual trick of scoring a sequence against a matrix of expected amino acids at each position does not find it.

Labelling by hand is slow. Of the more than 48,000 serpins recorded in UniProt, a public catalogue of protein sequences, only 78 carry an RCL annotation. The researchers set out to see whether a model could learn to do the labelling instead. They gathered serpins from over 20 animal genomes, and trained experts marked the loop in each one by looking at three-dimensional structures, including high-confidence models from AlphaFold. That produced 1384 annotated loops, alongside 2048 non-serpin proteins of matching length as counter-examples.

Where AI came in

The task was framed as labelling every amino acid in a sequence as loop or not loop, a job much like marking out a region in an image. Each sequence was first turned into numbers three different ways: a plain code for each amino acid, a code based on the BLOSUM62 table of which amino acids substitute for which, and embeddings from ESM2, a language model trained on protein sequences and used here as released. On each of these the team trained three network types from scratch: a convolutional network, a U-Net, and a bidirectional LSTM, which reads the chain in both directions.

The trained models were then run on the 78 UniProt-annotated serpins, held back and never used in training, and their output compared with those existing annotations. The U-Net using ESM2 embeddings reached a residue-level F1 score of 0.9995 and accuracy of 0.9999, and both the U-Net and the LSTM got the whole loop exactly right in around 97% to 98% of sequences. The AI stood in for the structural inspection and hand-curation that had produced the training labels. Trimming the training set to 70% or 40% sequence similarity did not change accuracy much, and the BLOSUM-based U-Net sometimes flagged loop-like stretches in proteins that are not serpins.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The reactive center loop (RCL) of serpin proteins is highly variable and is annotated for only 78 of the more than 48,000 serpins in UniProt, so it cannot be located with position-matrix methods. The authors built an expert-annotated dataset of 1384 RCL regions plus 2048 length-matched non-serpins and trained CNN, U-Net and bidirectional LSTM models on one-hot, BLOSUM62 and ESM2 encodings to label each residue as RCL or non-RCL. On the 78 UniProt-annotated serpins used as an independent test set, the U-Net with ESM2 embeddings reached residue-level F1 = 0.9995, MCC = 0.9990 and accuracy = 0.9999, and per-sequence exact match was around 97%-98% for both U-Net and LSTM. Thresholding the training set to 70% or 40% sequence identity did not significantly change prediction accuracy, and applying the Blosum-Unet model to whole-genome protein sets sometimes flagged RCL-like regions in non-serpins.

How AI was used

Serpin sequences from over 20 animal genomes were downloaded from UniProt, and RCL regions were annotated by trained experts from publicly available 3D structures, including high-confidence AlphaFold models; serpins that already carried a UniProt RCL annotation were held back as an independent test set and length-matched non-serpin proteins were added as negatives. Sequences were truncated or zero-padded to 1024 residues and encoded three ways: one-hot, a BLOSUM62 substitution matrix encoding, and embeddings from the ESM2_650M protein language model used as released. For each encoding, three architectures were trained from scratch for per-residue two-class output: a four-block 1D CNN, a 1D U-Net with four DoubleConv encoder blocks, a bottleneck, a four-stage decoder with skip connections and attention gates, and a bidirectional LSTM. Training used Adam at learning rate 0.001 for up to 50 epochs with early stopping on validation F1, a masked binary cross-entropy loss that excluded padded positions, and batch sizes of 32 for one-hot or BLOSUM and 4 for ESM2, on one or two Nvidia B200 GPUs. The trained models were then run over the held-out sequences and their predicted RCL positions compared with the existing UniProt annotations at residue and sequence level; the training set was additionally thresholded at 70% and 40% sequence identity and the models retested.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONREPRESENTATIONTRAININGINFERENCEVALIDATION123456AIAIAICollect serpinand non-serpinsequencesExpert annotationof RCL regionsEncode sequencesTrain per-residueclassifiersPredict RCLpositions onheld-out sequenc…Evaluate againstUniProtannotations↤ conventional algorithm↤ manual curation↤ manual curation
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Collect serpin and non-serpin sequences

Obtaining raw data, whether by measurement, download or retrieval.

All serpin proteins encoded by these genomes were batch-downloaded from UniProt.where the paper describes this · verbatim
in the paper
2Preparation
no AI

Expert annotation of RCL regions

Cleaning, filtering, normalising or labelling data already obtained.

we meticulously annotated the RCL regions using publicly available 3D structural informationwhere the paper describes this · verbatim
in the paper
3Representation
AI

Encode sequences

Encoding data into features, descriptors, embeddings or graphs. The AI stood in for conventional algorithm.

Three embedding methods were implemented: one-hot encoding, BLOSUM62 substitution matrix-based encoding, or the ESM2 encoding with the ESM2_650M model.where the paper describes this · verbatim
in the paper
4Training
AI

Train per-residue classifiers

Fitting model parameters, including fine-tuning an existing model. The AI stood in for manual curation.

All models were trained for up to 50 epochs using the Adam optimizerwhere the paper describes this · verbatim
in the paper
5Inference
AI

Predict RCL positions on held-out sequences

Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.

Training and inference were conducted on the University of Florida HiPerGator supercomputerwhere the paper describes this · verbatim
in the paper
6Validation
no AI

Evaluate against UniProt annotations

Testing outputs against ground truth.

At the per-sequence level, the exact match rate is around 97%–98% for both U-Net and LSTMwhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's result is the trained RCL identifier itself and its per-residue/per-sequence annotation performance; there is no non-AI route to the reported annotations

+What the AI was for
Segmentationin the paper
We evaluated three neural network architectures—CNN, U-Net, and LSTM—to compare performance.where the paper describes this · verbatim
+How it was taught
Supervisedin the paper
+Models named
1D U-Net (onehot-Unet, Blosum-Unet, ESM2-Unet) · Trained from scratch1D CNN · Trained from scratchBidirectional LSTM · Trained from scratchESM2 ESM2_650M · Off the shelfin the paper
+How results were checked
Held-out78 testedin the paper
On the independent test dataset, U-Net-based models achieved ∼98% accuracy in identifying the RCL at the per-sequence level.where the paper describes this · verbatim
+Code · weights · data
code availableweights availabledata availablein the paper
Source code, training and testing datasets, and models are freely available at https://github.com/leizhou69/RCL-identifierwhere the paper describes this · verbatim
+Compute
University of Florida HiPerGator supercomputer, one or two Nvidia B200 GPUs; batch size reduced to 4 for ESM2 embeddings due to GPU memory constraintsin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 3 items
  • Version of 1D U-Net (onehot-Unet, Blosum-Unet, ESM2-Unet)Which version of the model was used is not stated.
  • Version of 1D CNNWhich version of the model was used is not stated.
  • Version of Bidirectional LSTMWhich version of the model was used is not stated.

About this article

Record aix-00152, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error