structural-biology/ai produced the result/Briefings in Bioinformatics 2025 · v2
Protein language model trained to spot lactylated lysines in rice proteins
Researchers built PCBert-Kla, a tool that predicts which lysine residues in a protein carry a lactyl tag. A pretrained protein language model, ProtBert, supplied the sequence representation and was fine-tuned alongside the classifier.
spectrum · one line per step, placed by what the step does · bright lines used AI
PCBert-Kla: an efficient prediction method for lysine lactylation sites based on ProtBert and fusion of physicochemical features
Briefings in Bioinformatics, 2025
doi:10.1093/bib/bbaf615 · record aix-00185 v2 · checked 2026-10-09
- AI was for
- Classification
- Model family
- Transformer, Protein language model, Multilayer perceptron
- Checked by
- Benchmark
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Proteins are chains of amino acids, and after a cell has made one it often decorates it with small chemical tags. These tags change how the protein behaves. One such tag is lactylation: a lactyl group attached to the side chain of lysine, one of the twenty amino acids. Finding where these tags sit is slow laboratory work, so researchers try to predict the sites from sequence alone. The difficulty is that a lysine looks much like any other lysine on paper. Whether it gets tagged depends on the chemistry of the residues around it, in patterns that are not written down as simple rules.
The team set out to build a classifier that takes a short stretch of protein sequence centred on a lysine and decides whether that lysine is lactylated. They reused a published benchmark of rice sequences, built from windows of 51 residues, with a training set of 1720 lactylated and 1767 non-lactylated examples and a separate external test set of 177 of each. Keeping the dataset unchanged let them compare their model against two existing predictors, DeepKla and Auto-Kla.
Where AI came in
The central component is a protein language model. ProtBert is a transformer trained on large numbers of protein sequences, in the way text models are trained on text, so that it learns to represent a residue in the context of its neighbours. Here only the first four of its encoder layers were kept, and those layers were adjusted during training rather than held fixed. For each 51-residue window the model produced a vector of 1024 numbers summarising the sequence. This stood in for the hand-designed sequence encodings that such predictors have traditionally relied on.
That vector was joined to 27 physicochemical numbers computed with Biopython, covering molecular weight, isoelectric point, amino acid composition, secondary structure content, hydrophobicity and net charge, giving 1051 inputs to a small fully connected network with an attention step that weights the features. In five-fold cross-validation it averaged 0.9837 accuracy, and on the external test set 0.9497, against 0.9396 for a ProtBert-only variant, 0.9322 for DeepKla and 0.6893 for Auto-Kla. Shuffling each physicochemical feature in turn placed net charge top of the resulting ranking.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The study builds PCBert-Kla, a classifier that decides whether a lysine site in a protein sequence is lactylated. It fine-tunes the ProtBert protein language model over 51-residue sequence windows, concatenates the resulting 1024-dimensional representation with a 27-dimensional vector of physicochemical descriptors computed in Biopython, and feeds the combined vector to an attention-weighted fully connected classifier. On an existing rice Kla benchmark the model reached an average accuracy of 0.9837 in five-fold cross-validation, and on the benchmark's independent external test set it reached an accuracy of 0.9497, compared with 0.9396 for the authors' ProtBert-only variant, 0.9322 for DeepKla and 0.6893 for Auto-Kla. Permutation of each physicochemical feature placed net charge highest in the resulting importance ranking, and attention maps over positive training samples showed weight concentrated on residues adjacent to each position.
How AI was used
A pretrained protein language model, Rostlab/prot_bert, supplied the sequence representation: only the first four encoder layers were retained, and those layers were fine-tuned together with the rest of the network rather than frozen, yielding a 1024-dimensional [CLS] vector per 51-residue window. Separately, Biopython routines computed a 27-dimensional physicochemical vector (molecular weight, isoelectric point, amino acid composition, three-dimensional secondary structure content, Kyte–Doolittle hydrophobicity and net charge), min–max normalised; the two vectors were concatenated into 1051 dimensions and passed to a classifier of two hidden layers (32 and 8 neurons) with a feature-weighting attention mechanism in the first hidden layer, ReLU activations, dropout of 0.1 and 0.3, and a sigmoid output. Training used the published Lv et al. benchmark unchanged, 30 epochs, batch size 4, learning rate 0.003, stochastic gradient descent, binary cross-entropy loss, five-fold cross-validation and early stopping on validation accuracy, on an NVIDIA GeForce RTX 4090. The trained model was then run over the benchmark's independent external test set, with ablation variants (no physicochemical features, no attention, no fine-tuning, ESM2 in place of ProtBert) trained under the same conditions and DeepKla and Auto-Kla run through their public web services for comparison. Interpretation used permutation of each physicochemical feature column with the resulting AUC change summed across folds, and average ProtBert attention maps over positive training samples. The model is also served through a public web interface.
The shape of the work
Structural · the record, drawn
no AI
Reuse published Kla benchmark dataset
Obtaining raw data, whether by measurement, download or retrieval.
We directly used the benchmark dataset proposed by Lv et al.where the paper describes this · verbatim
no AI
Compute physicochemical descriptors
Encoding data into features, descriptors, embeddings or graphs.
we utilized traditional feature extraction techniques to derive a 27-dimensional feature vector, which was then normalized using min–max scalingwhere the paper describes this · verbatim
AI
Encode sequences with ProtBert
Encoding data into features, descriptors, embeddings or graphs. The AI stood in for conventional algorithm.
The ProtBert model generates 1024-dimensional feature vectors for each protein sequence subsequently used to predict Kla sites.where the paper describes this · verbatim
AI
Fuse features and fine-tune the classifier
Fitting model parameters, including fine-tuning an existing model.
five-fold cross-validation was used, and early stopping was applied based on the highest validation accuracywhere the paper describes this · verbatim
AI
Predict Kla sites on the independent external set
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
we evaluated them using an independent external datasetwhere the paper describes this · verbatim
AI
Compare against ablations and published predictors
Testing outputs against ground truth.
The results were obtained through their publicly available web services.where the paper describes this · verbatim
AI
Interpret model via permutation importance and attention maps
Extracting understanding from model behaviour.
we computed four average attention maps based on all positive samples from the training set using the ProtBert modulewhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is the trained predictor itself; every reported finding is a property of the model's output on the benchmark data.
The retained layers were not frozen; instead, the entire architecture was fine-tuned during training.where the paper describes this · verbatim
we compared the performance of our models with the advanced algorithmic models DeepKla and Auto-Kla using data from an independent test setwhere the paper describes this · verbatim
The source code has been uploaded to GitHub and can be accessed at:where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of ProtBert (Rostlab/prot_bert)Which version of the model was used is not stated.
- Version of PCBert-Kla attention-based fully connected classifierWhich version of the model was used is not stated.
- Version of Bert-Kla (ProtBert-only ablation variant)Which version of the model was used is not stated.
- Version of ESM2 (substituted for ProtBert in a variant)Which version of the model was used is not stated.
- Version of DeepKlaWhich version of the model was used is not stated.
- Version of Auto-KlaWhich version of the model was used is not stated.
- What step 4 replacedThe paper gives no basis for what the AI stood in for.
- What step 6 replacedThe paper gives no basis for what the AI stood in for.
- What step 7 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00185, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error