structural-biology/ai produced the result/PLoS ONE 2026 · v2
Benchmarking AI tools that predict and design peptides for cell-surface receptors
Researchers tested six deep-learning tools on G protein-coupled receptors: three predicted how known peptides sit in 113 receptor complexes, three designed new peptides for three receptors. Every result came from the software itself.
spectrum · one line per step, placed by what the step does · bright lines used AI
Assessment of generative de novo peptide design methods for G protein-coupled receptors
PLoS ONE, 2026
doi:10.1371/journal.pone.0355549 · record aix-00189 v2 · checked 2026-10-09
- AI was for
- Structure determination, Candidate generation
- Model family
- Transformer, Diffusion model, Graph neural network
- Checked by
- Held-out113 tested
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
G protein-coupled receptors sit in the outer membrane of cells and pass messages inward. Many are switched on by peptides, short chains of amino acids that slot into a pocket on the receptor's outer face. Knowing exactly how a peptide sits in that pocket matters for anyone hoping to design a new one. But working this out by experiment is slow: it means crystallising the receptor with its partner and reading off the atomic positions. Software that predicts such arrangements from sequence alone would be quicker, and software that invents new peptides to fit a chosen pocket quicker still. Whether either works reliably on these receptors is a separate question from whether it works on proteins generally.
The researchers set out to measure both. They gathered crystal structures of receptor-peptide pairs from public databases and filtered them to 113 non-redundant pairs. Each pair was then handed to three prediction tools to see whether they could reproduce the known arrangement, each run 50 times with a different random starting seed. Separately, three generative tools were asked to invent peptides for three chosen receptors, and the resulting designs were checked for whether they sat in the pocket and whether their atoms collided with the receptor.
Where AI came in
The artificial intelligence here is the object of study as well as the instrument. AlphaFold2 Initial Guess, Boltz-2 and RosettaFold3 produced the predicted structures; BindCraft, BoltzGen and RFdiffusion3 each generated 10,000 candidate peptides per receptor; ProteinMPNN wrote fresh sequences for 90 selected designs. All were run with their released weights and near-default settings, with no further training. Every reported finding is an output of one of these tools, measured by conventional scoring software afterwards.
The generative step stands in for the judgement of an expert designer choosing which sequence might bind. The final step, in which designed peptides were fed back through the prediction tools and scored, stands in for laboratory testing: the designs were never made or measured in the lab.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The study benchmarks deep-learning structure prediction and generative peptide design on G protein-coupled receptor targets. A curated set of 113 unique receptor-peptide dimers was predicted 50 times each with AlphaFold2 Initial Guess, Boltz-2 and RosettaFold3, giving median DockQ scores of 0.03, 0.56 and 0.41 respectively, with PAE-derived confidence metrics often assigning high confidence to misplaced peptides and accuracy dropping for complexes deposited after each method's training cutoff. In the second part, 10000 peptides were generated per target by BindCraft, BoltzGen and RFdiffusion3 for three receptors; all methods placed peptides inside the orthosteric pocket, BoltzGen stayed closest to the reference peptides, and a large share of designs showed steric clashes. Regenerating sequences for 90 selected backbones with ProteinMPNN shifted reprediction DockQ scores towards the acceptable and medium quality categories.
How AI was used
Crystal structures of GPCR-peptide and GPCR-protein complexes were collected from GPCRdb and RCSB and filtered to non-redundant dimers with gapless, canonical peptides. Each dimer was predicted 50 times with different seeds by AlphaFold2 Initial Guess, Boltz-2 and RosettaFold3, with the receptor supplied as template or initial guess and no multiple sequence alignment for the peptide; predictions were compared with the crystal reference using DockQ, iRMSD and fnat, and pLDDT and PAE-derived confidence values were correlated with structural deviation. For three receptors, BindCraft, BoltzGen and RFdiffusion3 each generated 10000 peptides restricted to the native peptide length and guided by four to six hotspot residues; designs were characterised by hotspot distance, Ca-RMSD to the native peptide and Rosetta clash metrics. Ninety non-clashing designs inside the binding pocket were repredicted with the three prediction methods, both with their original sequences and with one ProteinMPNN sequence generated per backbone at temperature 0.05, and scored with DockQ against the initial design. All models were run with released weights and near-default or author-recommended settings.
The shape of the work
Structural · the record, drawn
no AI
Collect and filter GPCR-peptide complexes
Cleaning, filtering, normalising or labelling data already obtained.
duplicate dimers were filtered for the best crystal structure resolution, resulting in 113 receptor-peptide dimerswhere the paper describes this · verbatim
AI
Predict known complexes with three methods
Running a trained model over new data to predict, classify or score.
Each receptor-peptide dimer was consequently predicted 50 times using a different seed for each run.where the paper describes this · verbatim
no AI
Score predictions against crystal references
Testing outputs against ground truth.
the DockQ score was calculated for the predicted peptide to the reference peptidewhere the paper describes this · verbatim
AI
Generate de novo peptides for three targets
Producing candidate objects that did not previously exist. The AI stood in for expert judgement.
With each method 10000 designs were generated per targetwhere the paper describes this · verbatim
no AI
Assess placement and select non-clashing designs
Reducing a candidate set by filtering or ranking, in a single pass.
Lastly, we selected 90 (10 per method and receptor) sequences of non-clashing peptides inside the binding pocketwhere the paper describes this · verbatim
AI
Redesign sequences for selected backbones
Producing candidate objects that did not previously exist.
we generated one ProteinMPNN sequence with the default temperature of 0.05 for each of the 90 backboneswhere the paper describes this · verbatim
AI
Repredict designed peptides
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
validated them with AF2IG, Boltz-2 and RF3 (50 predictions with different seeds)where the paper describes this · verbatim
no AI
Score repredictions against designs
Testing outputs against ground truth.
used DockQ to assess their placement compared to the initial designwhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
Every reported result is an output of the deep-learning prediction or generation tools being benchmarked; no non-AI measurement produces the findings
we generated 10000 putative peptides with BindCraft, BoltzGen and RFdiffusion3 each with the goal of mimicking the native peptidewhere the paper describes this · verbatim
Collectively, over all 50 predictions for all 113 dimers, AF2IG achieved a median DockQ score of 0.03where the paper describes this · verbatim
All data underlying the results of the study can be found on Zenodowhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- Version of BindCraftWhich version of the model was used is not stated.
- What step 2 replacedThe paper gives no basis for what the AI stood in for.
- What step 6 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00189, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error