~/aixsci
200 records · all checked

structural-biology/ai produced the result/Antibodies 2024 · v2

Deep learning trained on robotic peptide mapping predicts where antibody proteins degrade

Researchers stressed 51 antibodies and measured chemical damage at thousands of sites with an automated robotic workflow. A protein language model turned those measurements into a predictor of which sites degrade, and by how much, from sequence alone.

1. Forced degradation of antibody panel2. Automated peptide mapping and LC-MS/MS measurement3. Curate and label deamidation dataset4. Encode antibody sequences as ESM-2 residue embeddings5. Train chimeric classifier and quantitative regression head6. Predict hot spots and extents on independent antibody set7. Compare predictions with peptide mapping ground truth8. Screen clone panel from sequence alone

spectrum · one line per step, placed by what the step does · bright lines used AI

The Accurate Prediction of Antibody Deamidations by Combining High-Throughput Automated Peptide Mapping and Protein Language Model-Based Deep Learning
Antibodies, 2024

doi:10.3390/antib13030074 · record aix-00110 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Classification, Property prediction
Model family
Protein language model, Transformer, Recurrent neural network, Multilayer perceptron
Checked by
Experimental312 tested
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Antibodies are large proteins used as medicines, and like all proteins they are chains of amino acid building blocks. Over time, two of those building blocks, asparagine and glutamine, can undergo a chemical change called deamidation, in which part of the side chain is altered. This can change how the antibody behaves, so drug developers want to know in advance which sites are likely to degrade. The trouble is that whether a given site changes depends on its surroundings in the folded protein, not just on the neighbouring letters in the chain. Finding out experimentally means storing samples under heat for weeks and then measuring them, which is slow.

The researchers set out to build a predictor that works from sequence alone. They first produced the measurements to learn from: a panel of 51 antibodies, including the reference material NISTmAb, was held under heat and sampled at zero, one, two, four and eight weeks. A robotic sample-preparation system chopped the proteins into peptides, and mass spectrometry measured how much deamidation had occurred at each site at each time. That gave 2,285 labelled asparagine and glutamine sites, of which 276 counted as hot spots and 2,009 as inactive.

Where AI came in

The AI supplied the model's sense of context. A pretrained protein language model, ESM-2, had already learned statistical patterns across large numbers of protein sequences; run over each full antibody sequence, it produced a numerical vector for every amino acid, including a 1,280-number vector for each asparagine or glutamine of interest. A second component read only a short window of letters around each candidate site, using a network suited to sequences. The two sets of outputs were joined and a further network was trained on the experimental labels to call each site a hot spot or not, with an extra output trained to predict the degree of deamidation at two, four and eight weeks.

On six held-out antibodies covering 312 possible sites, the model's calls were correct 95% of the time, identifying 28 hot spots with six false positives and eight sites missed; one flagged site was then checked with a further peptide mapping experiment and confirmed. Compared with published deamidation classifiers on the same data, it scored highest on accuracy and on a combined measure called MCC, while its precision was lower by 0.01. Fed only the sequences of 86 clones, it took several minutes and flagged eight predicted to deamidate below 5% after eight weeks, standing in for the weeks-long stress experiments.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors built an automated robotic peptide mapping and LC-MS/MS workflow, applied it to 255 thermally stressed samples from 51 antibodies, and used the resulting site-specific measurements to label a dataset of 2285 asparagine and glutamine sites, of which 276 were labelled deamidation hot spots and 2009 inactive. A chimeric deep learning model combining ESM-2 protein language model embeddings of the full antibody sequence with a word-embedding and bi-LSTM module over local sequence windows classified sites as hot spots and, with a regression head, predicted deamidation extents at 2, 4 and 8 weeks. On an independent set of six antibodies covering 312 potential sites the model reached 95% accuracy with an AUC of 0.986, correctly identifying 28 hot spots with six false positives and eight missed sites; against published deamidation classifiers applied to the same data it gave the highest MCC and accuracy but not the highest precision, which was lower by 0.01. Applied to 86 clones using sequence input only, the model flagged eight clones predicted to deamidate below 5% after 8 weeks of stress.

How AI was used

A pretrained ESM-2 model (esm2_t33_650M_UR50D, 33 layers, 650 million parameters) was run over full-length antibody sequences to produce per-residue contextual embeddings of dimension n×1280, giving a 1×1280 vector for each asparagine or glutamine site of interest. These embeddings fed a downstream deep neural network with two hidden layers and dropout as one base model. A second base model converted local sequence windows centred on each candidate site into supervised word embeddings and passed them through a bi-directional LSTM, with window size selected by fivefold cross-validation using MCC over windows from three to sixty-one residues. The two module output vectors were concatenated and a fully connected neural network meta-classifier was trained on the combined features, with architecture and hyperparameters chosen by fivefold stratified cross-validation against alternatives including logistic regression, random forest, ANN, 1D-CNN and RNN, and early stopping applied. A regression head with three output neurons was added and trained on measured deamidation extents at 2, 4 and 8 weeks. Labels came from automated peptide mapping, with a site called active when the measured deamidation increment from t0 to 1 week or 1 to 2 weeks exceeded 1.0%. The trained model was then run on an independent antibody set and on a panel of clones supplied as FASTA sequences alone, alongside published structure-based and sequence-based classifiers used for comparison.

The shape of the work

Structural · the record, drawn

EXPERIMENTEXPERIMENTPREPARATIONREPRESENTATIONTRAININGINFERENCEVALIDATIONSCREENING12345678AIAIAIAIForceddegradation ofantibody panelAutomated peptidemapping andLC-MS/MS measure…Curate and labeldeamidationdatasetEncode antibodysequences asESM-2 residue em…Train chimericclassifier andquantitative reg…Predict hot spotsand extents onindependent anti…Comparepredictions withpeptide mapping …Screen clonepanel fromsequence alone↤ conventional algorithm↤ physical experiment↤ physical experiment
AI stepNo AI↤ what the AI stood in for
1Experiment
no AI

Forced degradation of antibody panel

Physical execution, by hand or by robot.

A panel of 51 antibodies, including NISTmAb and 50 in-house antibodies of varied modalitieswhere the paper describes this · verbatim
in the paper
2Experiment
no AI

Automated peptide mapping and LC-MS/MS measurement

Physical execution, by hand or by robot.

A high-throughput fully automated peptide mapping sample preparation platform was developed by using the Lynx liquid handling robotic systemwhere the paper describes this · verbatim
in the paper
3Preparation
no AI

Curate and label deamidation dataset

Cleaning, filtering, normalising or labelling data already obtained.

each deamidation site was labeled by setting a fixed deamidation thresholdwhere the paper describes this · verbatim
in the paper
4Representation
AI

Encode antibody sequences as ESM-2 residue embeddings

Encoding data into features, descriptors, embeddings or graphs. The AI stood in for conventional algorithm.

The outputs from the ESM-2 consist of residue-level sequence embeddings with dimensions of n×1280where the paper describes this · verbatim
in the paper
5Training
AI

Train chimeric classifier and quantitative regression head

Fitting model parameters, including fine-tuning an existing model.

we concatenated the vectors from both sources and trained a fully connected (FC) neural network classification head as a meta-classifierwhere the paper describes this · verbatim
in the paper
6Inference
AI

Predict hot spots and extents on independent antibody set

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

Of the 312 total potential deamidation sites in this dataset involving N and Q residues, the chimeric model achieved an accuracy of 95%where the paper describes this · verbatim
in the paper
7Validation
no AI

Compare predictions with peptide mapping ground truth

Testing outputs against ground truth.

A follow-up LysC-based peptide mapping experiment confirmed this site as a true positivewhere the paper describes this · verbatim
in the paper
8Screening
AI

Screen clone panel from sequence alone

Reducing a candidate set by filtering or ranking, in a single pass. The AI stood in for physical experiment.

we ran a pilot study involving 86 clones from different transfection pools and fed only the FASTA sequences to the model frameworkwhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's reported result is the deamidation hot-spot and extent predictions produced by the chimeric deep learning model; the experimental peptide mapping supplies labels and ground truth for it.

+What the AI was for
the embeddings utilized in our work were derived from a pretrained ESM-2 modelwhere the paper describes this · verbatim
+How it was taught
SupervisedTransfer / fine-tuningin the paper
+Models named
ESM-2 (esm2_t33_650M_UR50D) esm2_t33_650m_UR50D · Off the shelfChimeric model (ESM-2 global module + local sequence module + fully connected meta-classifier, with classification and regression heads) · Trained from scratchESM-2 base model (DNN head on ESM-2 embeddings, two hidden layers with dropout) · Trained from scratchLocal sequence base model (supervised word embedding + bi-directional LSTM) · Trained from scratchNGOME · Off the shelfDecision tree classifier of Yan et al. · Trained from scratchRandom forest classifier of Jia et al. · Trained from scratchin the paper
+How results were checked
Experimental312 testedin the paper
Of the 312 total potential deamidation sites in this dataset involving N and Q residues, the chimeric model achieved an accuracy of 95%where the paper describes this · verbatim
+Code · weights · data
code not reportedweights not reporteddata not reportedin the paper
this screening process only took several minuteswhere the paper describes this · verbatim
+Compute
No hardware or accelerator details given; model-based screening of the clone panel is reported to take several minutes, and the automated peptide mapping plate takes 7 h.in the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 10 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • DataWhether the data are available is not stated.
  • Version of Chimeric model (ESM-2 global module + local sequence module + fully connected meta-classifier, with classification and regression heads)Which version of the model was used is not stated.
  • Version of ESM-2 base model (DNN head on ESM-2 embeddings, two hidden layers with dropout)Which version of the model was used is not stated.
  • Version of Local sequence base model (supervised word embedding + bi-directional LSTM)Which version of the model was used is not stated.
  • Version of NGOMEWhich version of the model was used is not stated.
  • Version of Decision tree classifier of Yan et al.Which version of the model was used is not stated.
  • Version of Random forest classifier of Jia et al.Which version of the model was used is not stated.
  • What step 5 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00110, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error