structural-biology/ai produced the result/Protein Science 2025 · v2
Five structure-prediction models tested on 509 protein–peptide pairs under altered inputs
Researchers ran AlphaFold2, AlphaFold2-Multimer, AlphaFold3, Boltz-1 and Chai-1 over 509 known protein–peptide complexes, then fed the same models deliberately degraded inputs to see what their predictions actually depend on.
spectrum · one line per step, placed by what the step does · bright lines used AI
Training bias and sequence alignments shape protein–peptide docking by AlphaFold and related methods
Protein Science, 2025
doi:10.1002/pro.70331 · record aix-00164 v2 · checked 2026-10-09
- AI was for
- Structure determination
- Model family
- Transformer, Diffusion model
- Checked by
- Held-out509 tested
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
Many jobs inside a cell are done by a short stretch of protein — a peptide — settling into a groove on the surface of a larger protein. Working out the shape such a pair makes, and where exactly the peptide sits, normally means growing crystals or using other laboratory methods on one complex at a time. Software that predicts these shapes from sequence alone would be much quicker. But a peptide is short and floppy, with few of its own structural clues, so a prediction that comes out right does not by itself tell you what the program used to get there.
The authors assembled a set of 509 protein–peptide complexes whose shapes had already been determined experimentally, filtering entries from the public structure database by peptide length, the presence of helper molecules and crystal contacts near the binding site. They then asked a narrower question than how often prediction software is correct: what information it leans on when it is.
Where AI came in
Five released prediction models — AlphaFold2, AlphaFold2-Multimer, AlphaFold3, Boltz-1 and Chai-1 — did the structure prediction. None was trained or adjusted here; each was run as supplied, taking the top-ranked of five predictions, and the results were scored against the experimental structures. In this step the software stood in for the laboratory work that would otherwise have resolved each shape.
The same models were then re-run on altered inputs. These programs normally read a stack of related sequences from other species, an alignment, alongside the sequence of interest. The authors supplied alignments with the species pairings shuffled, alignments with the peptide's own removed, peptide sequences replaced by blanks, scrambled residues or plain glycine, and experimental structures as templates. For one model they also blocked parts of its internal attention between chains. Separately they counted how often each binding site already appeared in a model's training data.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors assembled a non-redundant set of 509 experimentally resolved protein-peptide complexes and predicted each one with AlphaFold2, AlphaFold2-Multimer, AlphaFold3, Boltz-1 and Chai-1, scoring the predictions against the experimental structures with DockQ. Comparing a 68-structure post-training-cutoff set with the 68 latest pre-cutoff structures, and separately counting how often each binding site appeared in each model's training data, they found lower accuracy for AlphaFold3, Boltz-1 and Chai-1 on post-cutoff structures and for binding sites with no match in the training set. A series of alignment manipulations showed little accuracy difference between paired and species-shuffled paired alignments, while removing the peptide alignment improved 17%-26% of predictions when present; of complexes predicted successfully with the peptide sequence supplied, 40%-51% remained successful when that sequence was replaced by mask tokens.
How AI was used
Five released structure-prediction models were run as inference engines over a curated test set of protein-peptide complexes, with no model fitted or fine-tuned in this study. Predictions were made from sequence with unpaired multiple-sequence alignments and no templates, taking the top-ranked of five predictions, and compared to the experimental structures using DockQ, backbone and all-atom RMSD, native-contact recovery, TM-score and DSSP-assigned secondary structure, alongside the models' own confidence outputs. The same models were then re-run under systematically altered inputs: no alignments, deeper paired alignments built from UniProt parent sequences with 50 or 100 residues of flanking context, those paired alignments with protein-peptide pairings shuffled, peptide alignments removed, native structures supplied as templates in place of alignments, peptide sequences replaced by unknown tokens, poly-glycine or scrambled residues, and Boltz-1 pocket residues or a Chai-1 distance restraint supplied. For AlphaFold2-Multimer, row- and column-wise self-attention between rows and columns belonging to different chains was masked inside the Evoformer, and masked-peptide predictions were read out from the distogram head using a pairwise-distance comparison metric. Model behaviour was related to training-set overlap by searching for homologous proteins by sequence or structure among structures released before each model's training cutoff and checking whether they bound a structurally similar partner, and to alignment statistics via inter-chain mutual information, Jensen-Shannon conservation and interface hydrophobicity.
The shape of the work
Structural · the record, drawn
no AI
Curate non-redundant protein-peptide test set
Cleaning, filtering, normalising or labelling data already obtained.
we filtered structures from the PDB by peptide length, presence of cofactors, and crystal contacts near the peptide binding sitewhere the paper describes this · verbatim
no AI
Construct sequence alignments and alternative inputs
Encoding data into features, descriptors, embeddings or graphs.
we constructed deeper paired MSAs for cases where the peptide could be mapped to a canonical UniProt entrywhere the paper describes this · verbatim
AI
Predict complex structures with five models
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
We made predictions using unpaired MSAs and no templates, took the top‐ranked structure out of five predictionswhere the paper describes this · verbatim
no AI
Score predictions against experimental structures
Testing outputs against ground truth.
calculated DockQ values between predicted and experimental structureswhere the paper describes this · verbatim
no AI
Search training data for binding-site matches
Extracting understanding from model behaviour.
we searched for homologous proteins (by sequence or structure) in the training setwhere the paper describes this · verbatim
AI
Re-run models under input and attention ablations
Running a trained model over new data to predict, classify or score.
We compared prediction performance using the paired MSAs to predictions made using MSAs containing the same sequences but with protein‐peptide pairings randomizedwhere the paper describes this · verbatim
no AI
Analyse what drives prediction accuracy
Extracting understanding from model behaviour.
We calculated inter‐chain mutual information (MI) for the original and shuffled alignmentswhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The study's findings are entirely about the behaviour of structure-prediction models; every reported result is derived from model predictions the authors ran
Applying filters and clustering resulted in a non‐redundant test set of 509 protein‐peptide complexeswhere the paper describes this · verbatim
The data that support the findings of this study are openly availablewhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- Version of AlphaFold2Which version of the model was used is not stated.
- Version of AlphaFold2-MultimerWhich version of the model was used is not stated.
- Version of AlphaFold3Which version of the model was used is not stated.
- Version of Boltz-1Which version of the model was used is not stated.
- Version of Chai-1Which version of the model was used is not stated.
- What step 6 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00164, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error