~/aixsci
200 records · all checked

structural-biology/ai produced the result/Science 2022 · v3

Neural network writes amino acid sequences to fit fixed protein backbones

Researchers trained a graph neural network, ProteinMPNN, to choose amino acid sequences for a given protein shape. The network wrote every sequence that was then made in bacteria and tested in the laboratory.

1. Assemble and cluster PDB structure sets2. Encode backbones as graph edge features3. Train ProteinMPNN with backbone noise4. Generate sequences for fixed backbones with tied positions5. Rank sequences by model log probability6. Predict structures of designed sequences with AlphaFold7. Express and biochemically characterise designs8. Solve structures and compare with design models

spectrum · one line per step, placed by what the step does · bright lines used AI

Robust deep learning based protein sequence design using ProteinMPNN
Science, 2022

doi:10.1126/science.add2187 · record aix-00023 v3 · checked 2026-10-08

ai-resultrole of AI
AI was for
Candidate generation, Structure determination
Model family
Graph neural network
Checked by
Experimental96 tested, 73 worked
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

A protein is a chain of amino acids that folds into a particular three-dimensional shape, and that shape largely determines what the protein does. Designers often work backwards: they draw the shape they want, down to the path traced by the chain's backbone, and then must decide which amino acid to place at each position so the chain actually folds that way. The number of possible sequences is vast, and the usual approach has been to score candidates with physics-based energy calculations, which is slow and often yields sequences that fail when made in the laboratory.

The researchers set out to train a network to do this sequence-writing step instead. They wanted it to handle not only single chains but assemblies of several chains, and to respect symmetry, where the same position in each repeated copy of a unit must carry the same amino acid.

Where AI came in

Protein structures from the public databank were grouped by similarity and split into training, validation and test sets. Each backbone was described to the network only as a graph: distances between the backbone atoms of residue pairs, how far apart they sit along the chain, and whether they belong to the same chain. A three-layer encoder and three-layer decoder were trained in PyTorch on a single graphics processor to predict the real amino acid at each position, with small random jitter added to the coordinates. Because the order of decoding was shuffled during training, positions could later be fixed or tied together.

The trained network then wrote sequences for chosen backbones, including ones whose earlier sequences, produced by the physics-based Rosetta method or by AlphaFold-based design, had failed in experiments. Candidates were ranked by the network's own confidence. A separate program, AlphaFold, predicted the folded shape of each designed sequence so it could be compared with the intended backbone before anything was made, standing in for an early round of laboratory work. Of 96 designs made in bacteria, 73 dissolved properly, and selected ones were examined by crystallography and electron microscopy.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors trained a message passing graph neural network, ProteinMPNN, to write amino acid sequences for a given protein backbone, including backbones made of multiple chains and backbones with symmetry constraints that tie residue identities across positions. On a test set of 402 monomer backbones it recovered 52.4% of native sequence positions against 32.9% for Rosetta fixed-backbone design, and took 1.2 seconds rather than 4.3 minutes per 100 residues. Sequences were then generated for backbones whose earlier Rosetta or AlphaFold-derived sequences had failed experimentally: of 96 designs produced in E. coli, 73 expressed solubly and 50 had the target monomeric or oligomeric state by size exclusion chromatography, while 13 of 76 tetrahedral nanoparticle sequences formed assemblies of the expected mass. Crystal, cryoEM and negative stain EM structures of selected designs were compared against the design models.

How AI was used

Protein assemblies from the PDB solved by X-ray crystallography or cryoEM to better than 3.5 A and under 10,000 residues were clustered at 30% sequence identity with mmseqs2 into 25,361 clusters, split into training, validation and test groups. Backbones were encoded only as graph edge features: radial basis function distances between N, Ca, C, O and a virtual Cb atom, a relative positional encoding capped at plus or minus 32 residues, and a binary same-chain or different-chain indicator. A three-layer encoder with edge updates and a three-layer decoder, each with 128 hidden dimensions, were trained in PyTorch on a single A100 GPU with negative log likelihood and label smoothing, using a decoding order sampled at random from all permutations so that positions can be fixed or tied; Gaussian noise was added to backbone coordinates during training at several levels, and a Ca-only variant was trained separately. At inference, sequences were sampled at a chosen temperature for fixed backbones, with residue identities at symmetry-equivalent positions tied by averaging logits, and ranked by the model's averaged log probability of the sequence given the structure. Designed sequences were passed as single sequences to five AlphaFold ptm models with three recycles, taking the highest average pLDDT model, and compared with the target backbones by lDDT-Ca, pLDDT and inter-chain PAE before selected designs were encoded in synthetic genes, expressed in E. coli and characterised.

The shape of the work

Structural · the record, drawn

PREPARATIONREPRESENTATIONTRAININGGENERATIONSCREENINGVALIDATIONEXPERIMENTVALIDATION12345678AIAIAIAssemble andcluster PDBstructure setsEncode backbonesas graph edgefeaturesTrain ProteinMPNNwith backbonenoiseGeneratesequences forfixed backbones …Rank sequences bymodel logprobabilityPredictstructures ofdesigned sequenc…Express andbiochemicallycharacterise des…Solve structuresand compare withdesign models↤ conventional algorithm↤ physical experiment
AI stepNo AI↤ what the AI stood in for
1Preparation
no AI

Assemble and cluster PDB structure sets

Cleaning, filtering, normalising or labelling data already obtained.

Sequences were clustered at 30% sequence identity cutoff using mmseqs2 (10) resulting in 25,361 clusters.where the paper describes this · verbatim
in the paper
2Representation
no AI

Encode backbones as graph edge features

Encoding data into features, descriptors, embeddings or graphs.

The edge features consisted of distances between residues in Euclidean space and distances between residues in the primary sequence spacewhere the paper describes this · verbatim
in the paper
3Training
AI

Train ProteinMPNN with backbone noise

Fitting model parameters, including fine-tuning an existing model.

Models were trained using pytorch (24), batch size of 10k tokens, automatic mixed precision, and gradient checkpointing on a single NVIDIA A100 GPU.where the paper describes this · verbatim
in the paper
4Generation
AI

Generate sequences for fixed backbones with tied positions

Producing candidate objects that did not previously exist. The AI stood in for conventional algorithm.

we kept the backbones of the original designs fixed but discarded the original sequences and generated new ones using ProteinMPNNwhere the paper describes this · verbatim
in the paper
5Screening
no AI

Rank sequences by model log probability

Reducing a candidate set by filtering or ranking, in a single pass.

ranked them using the log probability of the model (confidence)where the paper describes this · verbatim
in the paper
6Validation
AI

Predict structures of designed sequences with AlphaFold

Testing outputs against ground truth. The AI stood in for physical experiment.

We ran all 5 AlphaFold ptm models with 3 recycles and selected the model with the highest average pLDDTwhere the paper describes this · verbatim
in the paper
7Experiment
no AI

Express and biochemically characterise designs

Physical execution, by hand or by robot.

Synthetic genes encoding the designs were obtained, and the proteins expressed in E. coli and characterized biochemically and structurally.where the paper describes this · verbatim
in the paper
8Validation
no AI

Solve structures and compare with design models

Testing outputs against ground truth.

The design model was used as the search model for molecular replacement with Phaser 2.8.where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

All sequences that were synthesised and characterised were produced by the trained network; the paper's claims are about those sequences

+What the AI was for
We used encoder-decoder message passing neural networks for this taskwhere the paper describes this · verbatim
+Model families
+How it was taught
Supervisedin the paper
+Models named
ProteinMPNN · Trained from scratchCa-only ProteinMPNN · Trained from scratchMPNN baseline model (Table 1 baseline) · Trained from scratchAlphaFold (5 ptm models) · Off the shelfin the paper
+How results were checked
Experimental96 tested, 73 workedin the paper
of 96 designs produced in E. coli, 73 were expressed solublywhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
ProteinMPNN code is available at https://github.com/dauparas/ProteinMPNN.where the paper describes this · verbatim
+Compute
Trained on a single NVIDIA A100 GPU; validation loss converged after about 150k optimizer steps (~100 epochs). Inference ~1.2 s per 100-residue protein versus 258.8 s for Rosetta on a single CPU.in the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 6 items
  • Trained model weightsWhether the trained model is available is not stated.
  • Version of ProteinMPNNWhich version of the model was used is not stated.
  • Version of Ca-only ProteinMPNNWhich version of the model was used is not stated.
  • Version of MPNN baseline model (Table 1 baseline)Which version of the model was used is not stated.
  • Version of AlphaFold (5 ptm models)Which version of the model was used is not stated.
  • What step 3 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00023, version 3, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY-4.0; quotations are at most 25 words. How we work · Report an error