structural-biology/ai produced the result/Antibodies 2024 · v2
Deep learning trained on robotic peptide mapping predicts where antibody proteins degrade
Researchers stressed 51 antibodies and measured chemical damage at thousands of sites with an automated robotic workflow. A protein language model turned those measurements into a predictor of which sites degrade, and by how much, from sequence alone.
spectrum · one line per step, placed by what the step does · bright lines used AI
The Accurate Prediction of Antibody Deamidations by Combining High-Throughput Automated Peptide Mapping and Protein Language Model-Based Deep Learning
Antibodies, 2024
doi:10.3390/antib13030074 · record aix-00110 v2 · checked 2026-10-08
- AI was for
- Classification, Property prediction
- Model family
- Protein language model, Transformer, Recurrent neural network, Multilayer perceptron
- Checked by
- Experimental312 tested
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
Antibodies are large proteins used as medicines, and like all proteins they are chains of amino acid building blocks. Over time, two of those building blocks, asparagine and glutamine, can undergo a chemical change called deamidation, in which part of the side chain is altered. This can change how the antibody behaves, so drug developers want to know in advance which sites are likely to degrade. The trouble is that whether a given site changes depends on its surroundings in the folded protein, not just on the neighbouring letters in the chain. Finding out experimentally means storing samples under heat for weeks and then measuring them, which is slow.
The researchers set out to build a predictor that works from sequence alone. They first produced the measurements to learn from: a panel of 51 antibodies, including the reference material NISTmAb, was held under heat and sampled at zero, one, two, four and eight weeks. A robotic sample-preparation system chopped the proteins into peptides, and mass spectrometry measured how much deamidation had occurred at each site at each time. That gave 2,285 labelled asparagine and glutamine sites, of which 276 counted as hot spots and 2,009 as inactive.
Where AI came in
The AI supplied the model's sense of context. A pretrained protein language model, ESM-2, had already learned statistical patterns across large numbers of protein sequences; run over each full antibody sequence, it produced a numerical vector for every amino acid, including a 1,280-number vector for each asparagine or glutamine of interest. A second component read only a short window of letters around each candidate site, using a network suited to sequences. The two sets of outputs were joined and a further network was trained on the experimental labels to call each site a hot spot or not, with an extra output trained to predict the degree of deamidation at two, four and eight weeks.
On six held-out antibodies covering 312 possible sites, the model's calls were correct 95% of the time, identifying 28 hot spots with six false positives and eight sites missed; one flagged site was then checked with a further peptide mapping experiment and confirmed. Compared with published deamidation classifiers on the same data, it scored highest on accuracy and on a combined measure called MCC, while its precision was lower by 0.01. Fed only the sequences of 86 clones, it took several minutes and flagged eight predicted to deamidate below 5% after eight weeks, standing in for the weeks-long stress experiments.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors built an automated robotic peptide mapping and LC-MS/MS workflow, applied it to 255 thermally stressed samples from 51 antibodies, and used the resulting site-specific measurements to label a dataset of 2285 asparagine and glutamine sites, of which 276 were labelled deamidation hot spots and 2009 inactive. A chimeric deep learning model combining ESM-2 protein language model embeddings of the full antibody sequence with a word-embedding and bi-LSTM module over local sequence windows classified sites as hot spots and, with a regression head, predicted deamidation extents at 2, 4 and 8 weeks. On an independent set of six antibodies covering 312 potential sites the model reached 95% accuracy with an AUC of 0.986, correctly identifying 28 hot spots with six false positives and eight missed sites; against published deamidation classifiers applied to the same data it gave the highest MCC and accuracy but not the highest precision, which was lower by 0.01. Applied to 86 clones using sequence input only, the model flagged eight clones predicted to deamidate below 5% after 8 weeks of stress.
How AI was used
A pretrained ESM-2 model (esm2_t33_650M_UR50D, 33 layers, 650 million parameters) was run over full-length antibody sequences to produce per-residue contextual embeddings of dimension n×1280, giving a 1×1280 vector for each asparagine or glutamine site of interest. These embeddings fed a downstream deep neural network with two hidden layers and dropout as one base model. A second base model converted local sequence windows centred on each candidate site into supervised word embeddings and passed them through a bi-directional LSTM, with window size selected by fivefold cross-validation using MCC over windows from three to sixty-one residues. The two module output vectors were concatenated and a fully connected neural network meta-classifier was trained on the combined features, with architecture and hyperparameters chosen by fivefold stratified cross-validation against alternatives including logistic regression, random forest, ANN, 1D-CNN and RNN, and early stopping applied. A regression head with three output neurons was added and trained on measured deamidation extents at 2, 4 and 8 weeks. Labels came from automated peptide mapping, with a site called active when the measured deamidation increment from t0 to 1 week or 1 to 2 weeks exceeded 1.0%. The trained model was then run on an independent antibody set and on a panel of clones supplied as FASTA sequences alone, alongside published structure-based and sequence-based classifiers used for comparison.
The shape of the work
Structural · the record, drawn
no AI
Forced degradation of antibody panel
Physical execution, by hand or by robot.
A panel of 51 antibodies, including NISTmAb and 50 in-house antibodies of varied modalitieswhere the paper describes this · verbatim
no AI
Automated peptide mapping and LC-MS/MS measurement
Physical execution, by hand or by robot.
A high-throughput fully automated peptide mapping sample preparation platform was developed by using the Lynx liquid handling robotic systemwhere the paper describes this · verbatim
no AI
Curate and label deamidation dataset
Cleaning, filtering, normalising or labelling data already obtained.
each deamidation site was labeled by setting a fixed deamidation thresholdwhere the paper describes this · verbatim
AI
Encode antibody sequences as ESM-2 residue embeddings
Encoding data into features, descriptors, embeddings or graphs. The AI stood in for conventional algorithm.
The outputs from the ESM-2 consist of residue-level sequence embeddings with dimensions of n×1280where the paper describes this · verbatim
AI
Train chimeric classifier and quantitative regression head
Fitting model parameters, including fine-tuning an existing model.
we concatenated the vectors from both sources and trained a fully connected (FC) neural network classification head as a meta-classifierwhere the paper describes this · verbatim
AI
Predict hot spots and extents on independent antibody set
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
Of the 312 total potential deamidation sites in this dataset involving N and Q residues, the chimeric model achieved an accuracy of 95%where the paper describes this · verbatim
no AI
Compare predictions with peptide mapping ground truth
Testing outputs against ground truth.
A follow-up LysC-based peptide mapping experiment confirmed this site as a true positivewhere the paper describes this · verbatim
AI
Screen clone panel from sequence alone
Reducing a candidate set by filtering or ranking, in a single pass. The AI stood in for physical experiment.
we ran a pilot study involving 86 clones from different transfection pools and fed only the FASTA sequences to the model frameworkwhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's reported result is the deamidation hot-spot and extent predictions produced by the chimeric deep learning model; the experimental peptide mapping supplies labels and ground truth for it.
the embeddings utilized in our work were derived from a pretrained ESM-2 modelwhere the paper describes this · verbatim
Of the 312 total potential deamidation sites in this dataset involving N and Q residues, the chimeric model achieved an accuracy of 95%where the paper describes this · verbatim
this screening process only took several minuteswhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- DataWhether the data are available is not stated.
- Version of Chimeric model (ESM-2 global module + local sequence module + fully connected meta-classifier, with classification and regression heads)Which version of the model was used is not stated.
- Version of ESM-2 base model (DNN head on ESM-2 embeddings, two hidden layers with dropout)Which version of the model was used is not stated.
- Version of Local sequence base model (supervised word embedding + bi-directional LSTM)Which version of the model was used is not stated.
- Version of NGOMEWhich version of the model was used is not stated.
- Version of Decision tree classifier of Yan et al.Which version of the model was used is not stated.
- Version of Random forest classifier of Jia et al.Which version of the model was used is not stated.
- What step 5 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00110, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error