structural-biology/ai produced the result/PLoS ONE 2026 · v2
Machine learning sorts antibody heavy chains by the germ they target
Researchers turned 1111 antibody heavy chain sequences into numerical descriptors and trained tree-based classifiers, an ensemble and a Transformer to predict which of five antigens each antibody targets, in place of laboratory binding tests.
spectrum · one line per step, placed by what the step does · bright lines used AI
Computational models for the classification of antibody specificity using heavy chain features
PLoS ONE, 2026
doi:10.1371/journal.pone.0349143 · record aix-00188 v2 · checked 2026-10-09
- AI was for
- Classification
- Model family
- Gradient-boosted trees, Random forest, Transformer, Support vector machine, Linear model
- Checked by
- Held-out
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Antibodies are proteins the immune system makes to latch onto a specific invader. Each one is built from two kinds of chain, and the heavier of the two carries much of the region that does the gripping. Which target an antibody binds is normally established in the laboratory, by testing it against candidate molecules. That is slow and costly, and databases now hold far more antibody sequences than anyone could test. So a question arises: does the amino acid sequence of the heavy chain alone carry enough signal to say what the antibody was raised against? The chains that bind very different germs are not obviously different to the eye.
The authors gathered antibody heavy chain sequences from a public protein database for five targets: the dengue virus, influenza, tetanus, SARS-CoV-2 and the tuberculosis bacterium. After removing near-duplicates, 1111 sequences remained. Each was summarised as a list of 81 numbers describing its amino acid composition, the order of its residues, and chemical properties such as charge and bulk. The task was then to see whether a computer could read those numbers back to the correct antigen class, and which features mattered.
Where AI came in
The machine learning did the classifying. Several kinds of model were trained from scratch on the numerical descriptors, with the correct antigen supplied as the answer during training: five tree-based methods including CatBoost and XGBoost, a stacked arrangement in which simpler models feed a further one, and a Transformer, the architecture behind recent language models, here reading feature lists rather than text. Settings were tuned by repeatedly holding back portions of the training data. The models were then asked about a fifth of the sequences they had never seen, and about a separate published dataset.
Those predictions are the result. On the held-out sequences the stacked model was correct 0.7803 of the time and CatBoost 0.7713, while the Transformer reached 0.7399. On the external data the stacked model fell to 0.7240 and XGBoost led at 0.7729. A method called SHAP, which apportions credit among the inputs, ranked sequence-order and evolutionary descriptors highest, with the amino acid cysteine contributing most. The models stand in for the binding experiments that would otherwise assign each antibody to a target.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors assembled 1111 low-redundancy antibody heavy chain sequences from the NCBI Protein database across five antigen classes (anti-dengue, anti-influenza, anti-tetanus, anti-SARS-CoV-2 and anti-Mycobacterium tuberculosis) and encoded each sequence as an 81-dimensional vector of evolutionary, sequence-order and physicochemical descriptors. Tree-based classifiers, a stacking ensemble and a feature-based Transformer were trained to predict the antigen class from these features; on the held-out test set CatBoost reached an accuracy of 0.7713, the stacking ensemble 0.7803, and the Transformer an accuracy of 0.7399 with an F1-score of 0.6761. On an external dataset of translated heavy chain sequences the stacking model's accuracy fell to 0.7240 while XGBoost reached the highest accuracy of 0.7729. SHAP attribution ranked sequence-order (PseAAC) and evolutionary (AAC-PSSM) features highest, with cysteine the largest single contributor.
How AI was used
Heavy chain amino acid sequences were retrieved from NCBI with per-antigen search strings, clustered with CD-HIT at a 40% identity threshold, length-standardised and mapped to standard residues, then converted by web servers (POSSUM, the PseAAC server and iFeature) into an 81-dimensional vector of 20 AAC-PSSM, 22 PseAAC and 39 CTDC descriptors. These vectors were z-score normalised and used to fit supervised multiclass classifiers: XGBoost, LightGBM, Random Forest, CatBoost and AdaBoost as single models, a scikit-learn StackingClassifier with logistic regression, SVM and KNN base learners feeding a Random Forest meta-learner trained on out-of-fold predictions, and a feature-based Transformer with eight-head self-attention, a 256-dimensional feedforward block, six encoder layers, L2 regularisation and dropout. Hyperparameters were chosen by Bayesian optimisation with 5-fold cross-validation on an 80/20 stratified split, the trained models were then run over the held-out split and over an independently published, translated heavy chain dataset processed through the same pipeline, and SHAP was applied to the CatBoost, Stacking and Transformer models to rank feature contributions overall and per antibody class.
The shape of the work
Structural · the record, drawn
no AI
Retrieve antigen-specific heavy chain sequences
Obtaining raw data, whether by measurement, download or retrieval.
We collected antigen-specific immunoglobulin sequences from the National Center for Biotechnology Information (NCBI) Protein databasewhere the paper describes this · verbatim
no AI
Preprocess and deduplicate sequences
Cleaning, filtering, normalising or labelling data already obtained.
we utilized CD-HIT(Cluster Database at High Identity with Tolerance) to cluster sequences at a 40% identity thresholdwhere the paper describes this · verbatim
no AI
Encode sequences as engineered feature vectors
Encoding data into features, descriptors, embeddings or graphs.
an 81-dimensional feature vector was generated for each antibody sequence using three complementary encoding strategieswhere the paper describes this · verbatim
AI
Train classifiers and tune hyperparameters
Fitting model parameters, including fine-tuning an existing model.
Hyperparameter optimization was performed using Bayesian optimization with 5-fold cross-validation.where the paper describes this · verbatim
AI
Classify held-out test sequences
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
The dataset was randomly split into training (80%) and test (20%) sets, maintaining class distribution.where the paper describes this · verbatim
AI
External validation on an independent dataset
Testing outputs against ground truth. The AI stood in for physical experiment.
we conducted rigorous external validation using an independent dataset from Wang et al. (2022)where the paper describes this · verbatim
no AI
Attribute predictions to sequence features
Extracting understanding from model behaviour.
We utilized SHAP (SHapley Additive exPlanations) to explain our model outputs by assigning importance values to features.where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is the classification performance of the trained models themselves; every reported finding, including the SHAP feature rankings, comes from model output
Our approach begins with five robust tree-based algorithms as primary classifierswhere the paper describes this · verbatim
To assess the generalizability of our models, we conducted rigorous external validation using an independent datasetwhere the paper describes this · verbatim
The source code is available on GitHub (https://github.com/LJxp22/AbClass-Classifier).where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of CatBoostWhich version of the model was used is not stated.
- Version of XGBoostWhich version of the model was used is not stated.
- Version of LightGBMWhich version of the model was used is not stated.
- Version of Random ForestWhich version of the model was used is not stated.
- Version of AdaBoostWhich version of the model was used is not stated.
- Version of Stacking ensemble (logistic regression, SVM, KNN base learners with Random Forest meta-learner)Which version of the model was used is not stated.
- Version of Feature-Based TransformerWhich version of the model was used is not stated.
- What step 4 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00188, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error