~/aixsci
200 records · all checked

structural-biology/ai produced the result/PLoS ONE 2026 · v2

Benchmarking AI tools that predict and design peptides for cell-surface receptors

Researchers tested six deep-learning tools on G protein-coupled receptors: three predicted how known peptides sit in 113 receptor complexes, three designed new peptides for three receptors. Every result came from the software itself.

1. Collect and filter GPCR-peptide complexes2. Predict known complexes with three methods3. Score predictions against crystal references4. Generate de novo peptides for three targets5. Assess placement and select non-clashing designs6. Redesign sequences for selected backbones7. Repredict designed peptides8. Score repredictions against designs

spectrum · one line per step, placed by what the step does · bright lines used AI

Assessment of generative de novo peptide design methods for G protein-coupled receptors
PLoS ONE, 2026

doi:10.1371/journal.pone.0355549 · record aix-00189 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Structure determination, Candidate generation
Model family
Transformer, Diffusion model, Graph neural network
Checked by
Held-out113 tested
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

G protein-coupled receptors sit in the outer membrane of cells and pass messages inward. Many are switched on by peptides, short chains of amino acids that slot into a pocket on the receptor's outer face. Knowing exactly how a peptide sits in that pocket matters for anyone hoping to design a new one. But working this out by experiment is slow: it means crystallising the receptor with its partner and reading off the atomic positions. Software that predicts such arrangements from sequence alone would be quicker, and software that invents new peptides to fit a chosen pocket quicker still. Whether either works reliably on these receptors is a separate question from whether it works on proteins generally.

The researchers set out to measure both. They gathered crystal structures of receptor-peptide pairs from public databases and filtered them to 113 non-redundant pairs. Each pair was then handed to three prediction tools to see whether they could reproduce the known arrangement, each run 50 times with a different random starting seed. Separately, three generative tools were asked to invent peptides for three chosen receptors, and the resulting designs were checked for whether they sat in the pocket and whether their atoms collided with the receptor.

Where AI came in

The artificial intelligence here is the object of study as well as the instrument. AlphaFold2 Initial Guess, Boltz-2 and RosettaFold3 produced the predicted structures; BindCraft, BoltzGen and RFdiffusion3 each generated 10,000 candidate peptides per receptor; ProteinMPNN wrote fresh sequences for 90 selected designs. All were run with their released weights and near-default settings, with no further training. Every reported finding is an output of one of these tools, measured by conventional scoring software afterwards.

The generative step stands in for the judgement of an expert designer choosing which sequence might bind. The final step, in which designed peptides were fed back through the prediction tools and scored, stands in for laboratory testing: the designs were never made or measured in the lab.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The study benchmarks deep-learning structure prediction and generative peptide design on G protein-coupled receptor targets. A curated set of 113 unique receptor-peptide dimers was predicted 50 times each with AlphaFold2 Initial Guess, Boltz-2 and RosettaFold3, giving median DockQ scores of 0.03, 0.56 and 0.41 respectively, with PAE-derived confidence metrics often assigning high confidence to misplaced peptides and accuracy dropping for complexes deposited after each method's training cutoff. In the second part, 10000 peptides were generated per target by BindCraft, BoltzGen and RFdiffusion3 for three receptors; all methods placed peptides inside the orthosteric pocket, BoltzGen stayed closest to the reference peptides, and a large share of designs showed steric clashes. Regenerating sequences for 90 selected backbones with ProteinMPNN shifted reprediction DockQ scores towards the acceptable and medium quality categories.

How AI was used

Crystal structures of GPCR-peptide and GPCR-protein complexes were collected from GPCRdb and RCSB and filtered to non-redundant dimers with gapless, canonical peptides. Each dimer was predicted 50 times with different seeds by AlphaFold2 Initial Guess, Boltz-2 and RosettaFold3, with the receptor supplied as template or initial guess and no multiple sequence alignment for the peptide; predictions were compared with the crystal reference using DockQ, iRMSD and fnat, and pLDDT and PAE-derived confidence values were correlated with structural deviation. For three receptors, BindCraft, BoltzGen and RFdiffusion3 each generated 10000 peptides restricted to the native peptide length and guided by four to six hotspot residues; designs were characterised by hotspot distance, Ca-RMSD to the native peptide and Rosetta clash metrics. Ninety non-clashing designs inside the binding pocket were repredicted with the three prediction methods, both with their original sequences and with one ProteinMPNN sequence generated per backbone at temperature 0.05, and scored with DockQ against the initial design. All models were run with released weights and near-default or author-recommended settings.

The shape of the work

Structural · the record, drawn

PREPARATIONINFERENCEVALIDATIONGENERATIONSCREENINGGENERATIONINFERENCEVALIDATION12345678AIAIAIAICollect andfilterGPCR-peptide com…Predict knowncomplexes withthree methodsScore predictionsagainst crystalreferencesGenerate de novopeptides forthree targetsAssess placementand selectnon-clashing des…Redesignsequences forselected backbon…Repredictdesigned peptidesScorerepredictionsagainst designs↤ expert judgement↤ physical experiment
AI stepNo AI↤ what the AI stood in for
1Preparation
no AI

Collect and filter GPCR-peptide complexes

Cleaning, filtering, normalising or labelling data already obtained.

duplicate dimers were filtered for the best crystal structure resolution, resulting in 113 receptor-peptide dimerswhere the paper describes this · verbatim
in the paper
2Inference
AI

Predict known complexes with three methods

Running a trained model over new data to predict, classify or score.

Each receptor-peptide dimer was consequently predicted 50 times using a different seed for each run.where the paper describes this · verbatim
in the paper
3Validation
no AI

Score predictions against crystal references

Testing outputs against ground truth.

the DockQ score was calculated for the predicted peptide to the reference peptidewhere the paper describes this · verbatim
in the paper
4Generation
AI

Generate de novo peptides for three targets

Producing candidate objects that did not previously exist. The AI stood in for expert judgement.

With each method 10000 designs were generated per targetwhere the paper describes this · verbatim
in the paper
5Screening
no AI

Assess placement and select non-clashing designs

Reducing a candidate set by filtering or ranking, in a single pass.

Lastly, we selected 90 (10 per method and receptor) sequences of non-clashing peptides inside the binding pocketwhere the paper describes this · verbatim
in the paper
6Generation
AI

Redesign sequences for selected backbones

Producing candidate objects that did not previously exist.

we generated one ProteinMPNN sequence with the default temperature of 0.05 for each of the 90 backboneswhere the paper describes this · verbatim
in the paper
7Inference
AI

Repredict designed peptides

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

validated them with AF2IG, Boltz-2 and RF3 (50 predictions with different seeds)where the paper describes this · verbatim
in the paper
8Validation
no AI

Score repredictions against designs

Testing outputs against ground truth.

used DockQ to assess their placement compared to the initial designwhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

Every reported result is an output of the deep-learning prediction or generation tools being benchmarked; no non-AI measurement produces the findings

+What the AI was for
we generated 10000 putative peptides with BindCraft, BoltzGen and RFdiffusion3 each with the goal of mimicking the native peptidewhere the paper describes this · verbatim
+How it was taught
Zero-shotin the paper
+Models named
AlphaFold2 Initial Guess model_2_ptm.npz weights · Off the shelfBoltz-2 weights downloaded 06/20/25 · Off the shelfRosettaFold3 rf3_latest.pt obtained 08/15/25 · Off the shelfBindCraft · Off the shelfBoltzGen weights cached 12/04/25 · Off the shelfRFdiffusion3 rfd3_latest.ckpt obtained 12/03/25 · Off the shelfProteinMPNN protein_mpnn_v48_030.pt · Off the shelfin the paper
+How results were checked
Held-out113 testedin the paper
Collectively, over all 50 predictions for all 113 dimers, AF2IG achieved a median DockQ score of 0.03where the paper describes this · verbatim
+Code · weights · data
code not reportedweights not reporteddata availablein the paper
All data underlying the results of the study can be found on Zenodowhere the paper describes this · verbatim
+Compute
Leipzig University Computing Cluster and the high-performance computer at the NHR Center of TU Dresdenin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 5 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • Version of BindCraftWhich version of the model was used is not stated.
  • What step 2 replacedThe paper gives no basis for what the AI stood in for.
  • What step 6 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00189, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error