structural-biology/ai produced the result/Nature Communications 2024 · v2
Neural network predicts protein pair binding strength from simplified molecular models
Researchers built MCGLPPI, a framework that turns protein complexes into coarse-grained graphs. Graph neural networks learned from these graphs to predict binding strength and to tell real biological interfaces from artefacts of crystal packing.
spectrum · one line per step, placed by what the step does · bright lines used AI
Integration of molecular coarse-grained model into geometric representation learning framework for protein-protein complex property prediction
Nature Communications, 2024
doi:10.1038/s41467-024-53583-w · record aix-00179 v2 · checked 2026-10-09
- AI was for
- Property prediction, Classification
- Model family
- Graph neural network, Diffusion model, Multilayer perceptron
- Checked by
- Benchmark161 tested
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Proteins rarely work alone. They latch onto one another, and how tightly they hold determines much of what happens inside a cell. That grip strength is usually written as a binding free energy, often shortened to delta G. Measuring it in the laboratory is slow, so there is long-standing interest in predicting it from a complex's three-dimensional shape instead. The difficulty is scale. A pair of proteins can contain tens of thousands of atoms, and tracking every one is costly. Separately, structures solved by crystallography contain contacts between protein chains that are real in the crystal but not in the living cell, and telling the two apart is its own problem.
One way to cut the cost is coarse-graining: lumping several atoms into a single bead, so the model keeps the overall shape while carrying far fewer points. The MARTINI force field is a widely used recipe for doing this, and it also supplies the bead types and the bonds, angles and twists between them. The authors set out to feed that coarse-grained description directly into a learning system, and to test the result on binding strength prediction and on interface classification.
Where AI came in
The coarse-grained beads became the nodes of a graph, with MARTINI bond types and nearby contacts as seven kinds of connecting edge, and the graph was cropped down to the region where the two proteins meet. A graph neural network, GearNet-Edge, read these graphs and condensed each complex into a single numerical summary. A three-layer network then turned that summary into either a predicted delta G value or a judgement about the interface type. The learned model stands in here for laboratory measurement of binding strength.
The encoder was also pre-trained without labels, in a denoising scheme: coordinates and sequences of domain-domain interaction graphs were deliberately scrambled with noise, and the network learned to undo it, before being fine-tuned on each task. Atom-scale and residue-scale versions of the same encoder, plus GVP-GNN, were trained as comparisons. The coarse-grained model used roughly five times and three times less GPU memory than those two, and ran three times faster than each.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
MCGLPPI is a graph neural network framework that represents protein-protein complexes at the MARTINI coarse-grained scale, mapping CG beads to graph nodes and MARTINI bonded parameters to edges and node features, and predicts overall complex properties from the cropped interaction region. The authors evaluated it with tenfold cross-validation on a curated 1270-sample PDBbind strict-dimer set and a 531-structure ATLAS TCR-pMHC set for binding affinity regression, and on the MANY/DC sets for biological-versus-crystal interface classification. Against atom- and residue-scale versions of the same GearNet-Edge encoder on the 915-sample PDBbind subset, the CG model reduced GPU consumption by approximately 5x and 3x and total elapsed time by 3x and 3x at the same batch size. Self-supervised denoising pre-training on 3DID domain-domain interaction structures raised Pearson correlation for MCGLPPI-M2 from 0.597 to 0.606 on PDBbind and 0.825 to 0.830 on ATLAS, while AUPR on MANY/DC fell from 0.880 to 0.866.
How AI was used
Atomistic complex structures were repaired with pdbfixer and converted to MARTINI22 (martinize.py) or MARTINI3 (Martinize2/Vermouth) coarse-grained structures and force field parameters. Each complex was then encoded as a multi-relational graph in which beads are nodes, MARTINI bond types plus intra- and inter-residue contact edges within 5 Angstroms form seven edge types, and bead types, angles and dihedrals are node and edge features; a backbone-distance rule cropped the graph to a core interaction region within 8.5 Angstroms plus adjacent residues within 10 Angstroms. A multi-relational heterogeneous GNN encoder (GearNet-Edge) operating on these CG graphs produced a graph-level representation that a three-layer MLP mapped to either a binding affinity value or an interface class. The encoder was additionally pre-trained in a self-supervised diffusion denoising scheme that adds noise to CG bead coordinates and sequences of 3DID domain-domain interaction graphs, then fine-tuned on each downstream task. Atom- and residue-scale GearNet-Edge models and GVP-GNN were trained by the authors under matched cropping and hyper-parameter settings as baselines, all with PyTorch and TorchDrug, Adam at learning rate 0.0001, and a single A100 GPU.
The shape of the work
Structural · the record, drawn
no AI
Curate downstream PPI benchmark datasets
Obtaining raw data, whether by measurement, download or retrieval.
we obtained 1270 dimer samples with binding affinity labels △G, referred to as the PDBbind-strict-dimer datasetwhere the paper describes this · verbatim
no AI
Curate 3DID domain-domain interaction pre-training set
Obtaining raw data, whether by measurement, download or retrieval.
we obtained a pre-training dataset which provides 41,663 DDI structure samples in totalwhere the paper describes this · verbatim
no AI
Generate MARTINI coarse-grained structures and force field parameters
Cleaning, filtering, normalising or labelling data already obtained.
was used to generate MARTINI22-based CG structure and force field parameters for each protein complexwhere the paper describes this · verbatim
no AI
Build and crop CG-scale multi-relational complex graphs
Encoding data into features, descriptors, embeddings or graphs.
First, an edge will be wired if any two bead nodes have the Euclidean distance smaller than 5Åwhere the paper describes this · verbatim
AI
Self-supervised denoising pre-training of the CG graph encoder
Fitting model parameters, including fine-tuning an existing model.
a CG-scale complex pre-training technique was developed, which adds noise with changing magnitudes into 3D coordinates and sequences of MARTINI-based CG bead nodeswhere the paper describes this · verbatim
AI
Train from scratch and fine-tune task models
Fitting model parameters, including fine-tuning an existing model.
we fine-tuned the CG graph encoder that had undergone pre-training for each respective downstream taskwhere the paper describes this · verbatim
AI
Predict complex binding affinity and interface type
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
the generated representation was further learnt by a three-layer task-specific multi-layer perception (MLP) to give the final property prediction resultswhere the paper describes this · verbatim
no AI
Evaluate accuracy and computational cost against baselines
Testing outputs against ground truth.
The standard tenfold cross-validation (CV) strategy was used to evaluate the modelwhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is the predictive framework itself and the property predictions it produces; every reported finding comes from training and running learned graph neural network models.
we introduce MCGLPPI, a geometric representation learning framework that combines graph neural networks (GNNs) with MARTINI molecular coarse-grained (CG) modelswhere the paper describes this · verbatim
the MANY and DC datasets were utilized, containing 5739 and 161 dimers respectivelywhere the paper describes this · verbatim
The source code of MCGLPPI (Version 1.0) can be downloaded from https://github.com/arantir123/MCGLPPI.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- Version of MCGLPPI-M2 (GearNet-Edge encoder on MARTINI22 CG graphs)Which version of the model was used is not stated.
- Version of MCGLPPI-M3 (GearNet-Edge encoder on MARTINI3 CG graphs)Which version of the model was used is not stated.
- Version of GearNet-Edge (atom-scale and residue-scale baselines)Which version of the model was used is not stated.
- Version of GVP-GNNWhich version of the model was used is not stated.
- What step 5 replacedThe paper gives no basis for what the AI stood in for.
- What step 6 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00179, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error