structural-biology/ai produced the result/Scientific Reports 2022 · v2
Protein language model embeddings used to predict succinylation sites in proteins
Researchers built LMSuccSite, a tool that predicts which lysines in a protein carry a succinyl tag, using only the protein's sequence. Numerical descriptions from a pre-trained protein language model replaced hand-designed sequence features.
spectrum · one line per step, placed by what the step does · bright lines used AI
Improving protein succinylation sites prediction using embeddings from protein language model
Scientific Reports, 2022
doi:10.1038/s41598-022-21366-2 · record aix-00181 v2 · checked 2026-10-09
- AI was for
- Classification
- Model family
- Protein language model, Transformer, Convolutional neural network, Multilayer perceptron, Random forest, Support vector machine, Gradient-boosted trees, Recurrent neural network
- Checked by
- Benchmark
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Proteins are chains of amino acids, and after a cell makes one it often decorates it with small chemical tags. Succinylation is one such tag, attached to the amino acid lysine, and it can change how a protein behaves. Finding which lysines carry the tag matters for understanding how proteins are regulated, but the usual route is laboratory work with mass spectrometry, which is slow and only ever catches part of the picture. That leaves many proteins with lysines whose status is simply unknown.
The alternative is to predict the answer from the sequence itself. The difficulty is that a lysine looks much like any other lysine on paper, and the clues lie in the surrounding stretch of amino acids. Traditionally researchers decided by hand which properties of that neighbourhood to measure. Here the researchers instead set out to let a model work out its own numerical description of each site, and to combine two such descriptions into a single predictor they call LMSuccSite.
Where AI came in
AI supplied the description of each candidate site and the judgement about it. A protein language model, ProtT5-XL-UniRef50, had already been trained on large numbers of protein sequences in the way a text model learns language, by predicting hidden pieces of its input. Feeding a whole protein through it returns a 1024-number vector for each amino acid, and the researchers kept the vector for the lysine in question. Separately, a short 33-residue window around each lysine was fed through a convolutional network that learned its own encoding. Between them these stood in for hand-crafted sequence features.
The two parts were then trained to classify sites as succinylated or not, and their internal outputs were joined and passed to a further neural network that made the final call. Random forests, support vector machines, gradient-boosted trees and other networks were trained on the same features for comparison, with architectures and settings chosen by cross-validation and grid search. The reported scores on a held-aside test set are the model's own predictions. The researchers also projected the learned features into two dimensions with t-SNE and retrained on smaller slices of the data.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The study builds LMSuccSite, a sequence-only predictor of lysine succinylation sites in proteins. It combines a supervised word embedding of a 33-residue window around each lysine, processed by a 2D convolutional network, with per-residue embeddings from the pre-trained protein language model ProtT5-XL-UniRef50, processed by a two-layer feed-forward network; features from the second-to-last layer of each module are concatenated and passed to a neural-network meta-classifier. On an independent test set the authors report MCC, sensitivity and specificity of 0.36, 0.79 and 0.79, and the highest MCC, sensitivity and g-mean among the predictors compared, while other methods scored higher on accuracy and specificity. t-SNE plots of the learned features and a training-size sensitivity analysis are also reported.
How AI was used
A previously published succinylation dataset was reused: sequences were redundancy-filtered, 33-residue windows were taken around each lysine, all non-annotated lysines in the same proteins were treated as negatives, and the negative training set was randomly under-sampled to match the positives. Two encodings were computed from sequence alone: a supervised word embedding learned in a Keras embedding layer over the window peptides (vocabulary size 21, output shape 33 x 21), and 1024-dimensional contextualised embeddings taken from the encoder of the pre-trained ProtT5-XL-UniRef50 language model applied to full-length sequences, with the vector for the central lysine retained. A 2D CNN was trained on the supervised embedding and an ANN with hidden layers of 256 and 128 units on the ProtT5 features; random forest, SVM, XGBoost, CNN1D and LSTM alternatives were also trained for comparison. Architectures and hyperparameters were chosen by tenfold cross-validation with grid search on the training set. The base modules were then frozen and their second-to-last-layer outputs concatenated into a 144-dimensional meta-feature (16 plus 128) used to train a feed-forward meta-classifier, optimised with Adam on binary cross-entropy with dropout and early stopping. The final model was applied to the held-aside independent test set, and learned features were projected with t-SNE (perplexity 30, learning rate 100), with models also retrained on 20-80% subsets of the training data.
The shape of the work
Structural · the record, drawn
no AI
Reuse curated succinylation site dataset
Obtaining raw data, whether by measurement, download or retrieval.
We used the dataset used during the development of DeepSuccinylSite to train and test our approach.where the paper describes this · verbatim
no AI
Build window peptides and balance negatives
Cleaning, filtering, normalising or labelling data already obtained.
we performed random under sampling on the negative training set to obtain the same number of negative sites (4750)where the paper describes this · verbatim
AI
Extract ProtT5 embeddings for lysine residues
Encoding data into features, descriptors, embeddings or graphs. The AI stood in for conventional algorithm.
This model takes the overall protein sequence as an input and returns an embedding vector of dimension 1024 for each amino acid.where the paper describes this · verbatim
AI
Train base modules and alternative ML/DL models
Fitting model parameters, including fine-tuning an existing model.
Initially, a 2D-CNN-based architecture was used for supervised embedding features while an artificial neural network (ANN)-based module was used for ProtT5 features.where the paper describes this · verbatim
AI
Select architectures and hyperparameters by cross-validated grid search
Iterative search over a space. Its result feeds back into an earlier step.
This architecture is chosen based on tenfold cross-validation on the training set using different architectures with different combinations of hyperparameters using grid search.where the paper describes this · verbatim
AI
Train stacked meta-classifier (LMSuccSite)
Fitting model parameters, including fine-tuning an existing model.
the embedding module and Prot-T5 module were combined using an ANN as a meta-classifierwhere the paper describes this · verbatim
AI
Score independent test set and compare with existing predictors
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
we used the independent test set described in Table 1 and computed parameters such as accuracy, MCC, sensitivity, specificity and g-meanwhere the paper describes this · verbatim
AI
Inspect learned features and training-size sensitivity
Extracting understanding from model behaviour.
we created four different training datasets by randomly selecting 20%, 40%, 60%, and 80% of the samples from our training setwhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is a trained predictor of succinylation sites; the reported scores are the model's own outputs, so the finding exists only through the model
we used the pretrained ProtT5 model to encode the featureswhere the paper describes this · verbatim
Importantly, the same training and test sets were used for all of the models tested.where the paper describes this · verbatim
The source code and trained models are publicly available in the GitHub repositorywhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of ProtT5-XL-UniRef50Which version of the model was used is not stated.
- Version of LMSuccSite (ANN meta-classifier)Which version of the model was used is not stated.
- Version of Embedding module (CNN2D on supervised word embedding)Which version of the model was used is not stated.
- Version of ProtT5 module (two-hidden-layer ANN)Which version of the model was used is not stated.
- Version of Random forest (ProtT5 features)Which version of the model was used is not stated.
- Version of Support vector machine (ProtT5 features)Which version of the model was used is not stated.
- Version of XGBoost (ProtT5 features)Which version of the model was used is not stated.
- Version of CNN1D (ProtT5 features)Which version of the model was used is not stated.
- Version of LSTM (supervised word embedding)Which version of the model was used is not stated.
- What step 4 replacedThe paper gives no basis for what the AI stood in for.
- What step 5 replacedThe paper gives no basis for what the AI stood in for.
- What step 6 replacedThe paper gives no basis for what the AI stood in for.
- What step 8 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00181, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error