~/aixsci
200 records · all checked

structural-biology/ai produced the result/Bioinformatics 2022 · v2

Neural networks pick amino acid sequences for short peptides in protein binding sites

Researchers trained two neural networks, PepSeP1 and PepSeP6, to choose the amino acids of six-residue peptide fragments lodged in a protein's binding site. The networks produced the sequences; Rosetta software then refined and scored them.

1. Extract peptide–binding site complexes from crystal structures2. Perturb backbones, glycine-mutate and split datasets3. Encode complexes as distance maps and sequence features4. Train PepSeP1 and PepSeP6 networks5. Design peptide sequences for held-out and perturbed backbones6. Rosetta FastDesign redesign and baseline design7. Refine complexes and compute binding energies8. Score sequence and hot-spot recovery against native sequences

spectrum · one line per step, placed by what the step does · bright lines used AI

Deep learning of protein sequence design of protein–protein interactions
Bioinformatics, 2022

doi:10.1093/bioinformatics/btac733 · record aix-00153 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Candidate generation
Model family
Convolutional neural network, Recurrent neural network
Checked by
Held-out1245 tested
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Proteins do much of their work by sticking to one another. A short stretch of one protein settles into a groove on the surface of another, and the fit depends on which amino acids — the twenty chemical building blocks of proteins — sit at each position along that stretch. Designing such a binder means answering a hard question: given the shape of the groove and the path the peptide's backbone takes through it, which amino acids should fill the positions? The number of possible combinations is enormous, and the usual approach is to search through them with physics-based scoring software, which is slow and does not always land on sequences resembling those nature uses.

The authors set out to have a neural network make that choice directly. They assembled peptide–binding site pairs from 9002 co-crystal structures in the Protein Data Bank, a public archive of experimentally determined protein shapes. In each pair, the peptide was a six-residue fragment and the binding site a patch on the partner protein. They stripped the peptides back to plain backbones, nudged them away from their natural shapes, and asked whether a model could recover the original amino acids from geometry alone.

Where AI came in

The AI is where the sequence comes from. Each complex was turned into maps of the distances between backbone atoms, within the peptide and between peptide and binding site, plus a description of the binding site's amino acids and shape. A convolutional network — the kind used for reading images — compressed that into a summary, and a second network with an attention mechanism read the summary off one residue at a time, much as image-captioning systems write a sentence describing a picture. PepSeP1 gives one sequence per complex; PepSeP6 reuses PepSeP1's compressed summary and emits six, each conditioned on the one before.

The networks stood in for the physics-based search that normally picks the amino acids. The authors ran Rosetta's FastDesign on the same bare backbones as a comparison, and also used PepSeP1's output to guide Rosetta redesigns. On a held-out test set of 1245 cases, PepSeP1 matched the natural amino acid at 40.83% of positions and at 45.95% of positions flagged as binding hot-spots; PepSeP6 averaged 38.67% across its outputs. Matching was lower on interfaces between different proteins, at 27.71%, and lower again on the antibody–antigen sets, at 26.00% and 16.78%. Code and example data are published.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors trained two attention-based encoder–decoder networks, PepSeP1 and PepSeP6, to assign amino acid sequences to 6-residue peptide fragments sitting in a protein binding site, using peptide–binding site complexes extracted from 9002 co-crystal structures in the Protein Data Bank. On the independent test set, PepSeP1 recovered 40.83% of native peptide residues overall and 45.95% of residues identified as binding hot-spots, while PepSeP6, which emits six sequences per complex, averaged 38.67% across its outputs. Recovery was lower on hetero-oligomeric interfaces (27.71%) and on the antibody–antigen benchmark subsets (26.00% and 16.78%). Compared with Rosetta's FastDesign applied to the same all-glycine peptides, PepSeP1 recovered more native residues overall, though FastDesign gave higher hot-spot recovery on the T-he and B-ag/ab subsets.

How AI was used

Peptide–binding site complexes were extracted from multichain PDB co-crystal structures, with the peptide defined as a 6-residue fragment of one partner and the binding site as a 24–48 residue patch of the other; peptides were mutated to all-glycine and perturbed from their native conformation, and the complexes were split into training, validation, test and a held-out antibody–antigen benchmark set. Each complex was encoded as intramolecular and intermolecular backbone distance maps plus binding-site secondary structure, binding-site amino acid types and a homo- or hetero-oligomeric flag. These features were passed to a convolutional encoder of two blocks (8 and 4 layers) whose concatenated feature vectors fed a bidirectional Bahdanau-attention LSTM decoder, trained with categorical cross-entropy in three learning-rate stages and repeated 20 times, with the run selected on recovery over hetero-oligomeric subsets. PepSeP6 reused the PepSeP1 encoder with frozen weights and ran its decoder five times, each conditioned on the previous prediction's final hidden state, with the PepSeP1 output as the sixth sequence. Designed sequences were threaded back onto the perturbed backbones, refined with Rosetta FastRelax under harmonic constraints, and scored with InterfaceAnalyzerMover under ref15; Rosetta FastDesign was run by the authors as a baseline both on all-glycine backbones directly and as a redesign constrained by PepSeP1 position-specific scoring matrices in three schemes (RD3, RD5, RD20). Recovery was measured against native sequences overall and at alanine-scanning hot-spot positions, and the models were additionally applied to highly perturbed backbones generated by the iNNterfaceDesign method in antibody case studies.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONREPRESENTATIONTRAININGGENERATIONOPTIMISATIONSIMULATIONVALIDATION12345678AIAIExtractpeptide–bindingsite complexes f…Perturbbackbones,glycine-mutate a…Encode complexesas distance mapsand sequence fea…Train PepSeP1 andPepSeP6 networksDesign peptidesequences forheld-out and per…RosettaFastDesignredesign and bas…Refine complexesand computebinding energiesScore sequenceand hot-spotrecovery against…↤ conventional algorithm↤ conventional algorithm
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Extract peptide–binding site complexes from crystal structures

Obtaining raw data, whether by measurement, download or retrieval.

The complexes were extracted from 9002 co-crystal structures.where the paper describes this · verbatim
in the paper
2Preparation
no AI

Perturb backbones, glycine-mutate and split datasets

Cleaning, filtering, normalising or labelling data already obtained.

Peptides were perturbed up to 1.07 Å root-mean-square deviation (RMSD) of their native conformationwhere the paper describes this · verbatim
in the paper
3Representation
no AI

Encode complexes as distance maps and sequence features

Encoding data into features, descriptors, embeddings or graphs.

We use two types of distance maps as the main geometrical descriptors of the structureswhere the paper describes this · verbatim
in the paper
4Training
AI

Train PepSeP1 and PepSeP6 networks

Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.

Training of the PepSeP1 model was conducted 20 times in three stages: 5 epochs with a learning rate of 0.001where the paper describes this · verbatim
in the paper
5Generation
AI

Design peptide sequences for held-out and perturbed backbones

Producing candidate objects that did not previously exist. The AI stood in for conventional algorithm.

Five outputs are generated by passing feature vectors produced by the encoder into the decoder of PepSeP6 five timeswhere the paper describes this · verbatim
in the paper
6Optimisation
no AI

Rosetta FastDesign redesign and baseline design

Iterative search over a space.

Peptide ligands designed by the PepSeP1 method underwent an additional design step after refinement using the FastDesign protocolwhere the paper describes this · verbatim
in the paper
7Simulation
no AI

Refine complexes and compute binding energies

Numerical or physics simulation, including where a learned surrogate replaces it.

Binding free energies of the complexes were estimated using InterfaceAnalyzerMover with repacking chains after separation.where the paper describes this · verbatim
in the paper
8Validation
no AI

Score sequence and hot-spot recovery against native sequences

Testing outputs against ground truth.

To evaluate the performance of PepSeP1, we utilized sequence recovery.where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's result is the designed peptide sequences themselves, which are produced by the trained networks PepSeP1 and PepSeP6; the reported recovery rates and binding energies are measurements of those model outputs.

+What the AI was for
We developed an attention-based deep learning model inspired by algorithms used for image-caption assignments to design peptideswhere the paper describes this · verbatim
+How it was taught
SupervisedTransfer / fine-tuningin the paper
+Models named
PepSeP1 · Trained from scratchPepSeP6 · Trained from scratchin the paper
+How results were checked
Held-out1245 testedin the paper
The native sequence recovery rate of PepSeP1 is 40.83% on the independent test setwhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
All the code and example data are available at https://github.com/strauchlab/iNNterfaceDesignwhere the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 4 items
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • Version of PepSeP1Which version of the model was used is not stated.
  • Version of PepSeP6Which version of the model was used is not stated.

About this article

Record aix-00153, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error