structural-biology/ai produced the result/bioRxiv (Cold Spring Harbor Laboratory) 2023 · v2
Benchmark measures how well AlphaFold predicts antibody-antigen complex structures
Researchers ran AlphaFold-Multimer on 429 antibody-antigen complexes released after the model's training cut-off, then scored its predicted shapes against the experimentally determined ones. The AI did the predicting; the scoring was conventional software.
spectrum · one line per step, placed by what the step does · bright lines used AI
Evaluation of AlphaFold Antibody-Antigen Modeling with Implications for Improving Predictive Accuracy
bioRxiv (Cold Spring Harbor Laboratory), 2023
doi:10.1101/2023.07.05.547832 · record aix-00052 v2 · checked 2026-10-08
- AI was for
- Structure determination
- Model family
- Transformer
- Checked by
- Held-out429 tested
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
Antibodies are immune proteins that latch onto a target, called an antigen, and the grip depends on shape. The two molecules fit together at a small patch of contact, and knowing the exact arrangement of atoms there tells researchers how the antibody recognises its target. Working that arrangement out experimentally is slow and demanding work. Predicting it from the sequence of amino acids alone is harder still for antibodies than for most proteins, because the loops that do the binding vary enormously between antibodies and have few close relatives in the evolutionary record for a method to learn from.
The researchers set out to measure how often AlphaFold, a deep neural network that predicts protein structures from sequence, gets these pairings right. They assembled a non-redundant set of 429 complexes from a public antibody structure database, choosing cases whose structures were released after the model's training data stopped, so the answers could not already be known to it. They also compared the results against an older, physics-style docking approach, and examined whether the model's own confidence numbers flag which predictions to trust.
Where AI came in
AlphaFold did the prediction itself. Given antibody and antigen sequences, two released versions of AlphaFold-Multimer, run off the shelf on a local cluster of GPUs, produced ranked three-dimensional models of each complex. The same sequences were also run through ColabFold, a wrapper around the model, adjusted to save a prediction at every internal refinement pass. AlphaFold additionally modelled the antibody and antigen separately, and those single-molecule models were fed into the ZDOCK docking program as the comparison baseline. In the record's terms, the model stood in for a physical experiment: the structures it output take the place of ones that would otherwise be determined in a laboratory.
Everything else was conventional software. The predicted shapes were compared with the experimentally determined structures using DockQ and the CAPRI accuracy classes, which are standard yardsticks for how close a predicted complex is to the real one. The model's own per-atom confidence values, averaged over the residues at the contact patch, were turned into a score for judging which predictions were likely to be accurate. Top-ranked predictions reached Acceptable or better accuracy in 25 per cent of the 429 cases and Medium or better in 18 per cent, while the docking baseline reached Medium or better in 1 per cent of the 390 cases it was run on.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The study benchmarked AlphaFold's prediction of antibody-antigen complex structures on 429 nonredundant complexes released after the model's training cutoff, scoring predictions against the experimentally determined structures using CAPRI criteria and DockQ. Top-ranked predictions reached Acceptable or higher accuracy for 25% of cases and Medium or higher accuracy for 18%, while a rigid-body docking baseline (ZDOCK with IRAD reranking) using AlphaFold-modelled unbound inputs reached Medium or higher accuracy top-ranked predictions for 1% of the 390 cases it was run on. An interface pLDDT score computed over antibody-antigen interface residues separated Incorrect from Medium-or-higher models with an AUC of 0.93, versus 0.88 for AlphaFold's model confidence score. On a subset of 39 complexes modelled by both versions, AlphaFold v2.3 produced Medium or higher accuracy top-ranked predictions for 36% of cases versus 23% for v2.2, and supplying experimentally determined chains as templates raised near-native top-ranked success from 18% to 52% on a 100-case subset.
How AI was used
Antibody-antigen complex structures were predicted directly from sequence by AlphaFold-Multimer, run off-the-shelf in two released versions (v2.2.0 and v2.3.0) on a local cluster, with antibody sequences trimmed to variable domains by ANARCI and template date cutoffs set to precede each version's training cutoff. The same models were also run through ColabFold 1.3.0, modified to emit predictions at each recycling iteration and to produce 25 predictions per complex using different random seeds. AlphaFold was additionally used in Monomer or Multimer setting to model unbound antibody and antigen subunits, which served both as inputs to the ZDOCK rigid-body docking baseline and as custom templates. The pipeline was further modified to accept specified per-chain PDB templates and to restrict MSA features to the query sequence for single-sequence runs. Model-derived pLDDT values were averaged over interface residues to form an interface confidence score, and all predictions were scored against experimental structures with DockQ, CAPRI criteria, TM-score, CDR loop RMSDs and Rosetta interface energies.
The shape of the work
Structural · the record, drawn
no AI
Obtain candidate antibody-antigen structures
Obtaining raw data, whether by measurement, download or retrieval.
we downloaded the full SAbDab antibody structure dataset in January 2022where the paper describes this · verbatim
no AI
Filter to nonredundant benchmark and prepare input sequences
Cleaning, filtering, normalising or labelling data already obtained.
Structural nonredundancy criteria were then applied to the set.where the paper describes this · verbatim
AI
Predict complexes with AlphaFold v2.2 and v2.3
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
Both versions of AlphaFold were installed on a local computing cluster.where the paper describes this · verbatim
AI
Predict complexes with ColabFold and output each recycling iteration
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
AlphaFold modeling in ColabFold was performed with ColabFold version 1.3.0where the paper describes this · verbatim
AI
Model unbound antibody and antigen subunits
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
Unbound antibody and antigen input structures for the ZDOCK algorithm were generated by AlphaFoldwhere the paper describes this · verbatim
no AI
Rigid-body docking baseline with ZDOCK and IRAD reranking
Producing candidate objects that did not previously exist.
ZDOCK version 3.0.2 was used to generate antibody-antigen docking models using unbound or bound input structureswhere the paper describes this · verbatim
AI
Re-run AlphaFold with custom templates and without MSAs
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
Modifications were made to the AlphaFold pipeline to optionally input specific selected PDB templates for each chain.where the paper describes this · verbatim
no AI
Score models against experimental structures and assess confidence metrics
Testing outputs against ground truth.
We assessed antibody-antigen complex model accuracy using DockQwhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's findings are measurements of AlphaFold's own predictions; every reported result is produced by running the model, so the claims cannot exist without it
end-to-end deep neural network designed to predict protein structures from sequencewhere the paper describes this · verbatim
AlphaFold generated Acceptable or higher accuracy models as top-ranked predictions for 25% of the 429 test caseswhere the paper describes this · verbatim
AlphaFold and ColabFold modeling runs were performed using NVIDIA Titan RTX and Quadro 6000 GPUs.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
About this article
Record aix-00052, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error