structural-biology/ai produced the result/Nature Communications 2026 · v2
Peptides designed by simulation and machine learning sit at condensate surfaces
Researchers built a pipeline that designed short peptides to gather at the edge of protein droplets inside cells. A neural network learned to predict simulation results, and an optimiser searched sequences through it.
spectrum · one line per step, placed by what the step does · bright lines used AI
De novo design of peptides localizing at the interface of biomolecular condensates
Nature Communications, 2026
doi:10.1038/s41467-026-73099-9 · record aix-00163 v2 · checked 2026-10-09
- AI was for
- Simulation surrogate, Property prediction, Candidate generation, Experimental design
- Model family
- Multilayer perceptron, Linear model, Support vector machine, Gradient-boosted trees
- Checked by
- Experimental3 tested, 3 worked
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Cells are not simply bags of evenly mixed molecules. Some proteins gather together into dense liquid-like droplets, called biomolecular condensates, which behave a little like oil drops in water. Many of the proteins that do this have floppy, shapeless stretches rather than a fixed folded form, which makes their behaviour hard to predict from structure alone. The surface where a droplet meets the surrounding fluid is a distinct place with its own chemistry, and a molecule that prefers to sit there is not the same as one that dissolves into the droplet's interior. Designing a short protein chain to stay at that boundary means searching an enormous number of possible sequences.
The researchers set out to design 30-residue peptides, short chains of amino acids, that would collect at the interface of condensates formed by the disordered regions of three proteins: hnRNPA1, LAF-1 and DDX4. They used simplified molecular simulations to measure, for a given peptide, how strongly it favoured the interface and how strongly it stuck to copies of itself. Candidate peptides were also filtered with sequence-based tools that flag chains likely to clump. One peptide per target, plus a control, was then made chemically and imaged.
Where AI came in
The simulations that scored each peptide were slow, so they could not be run for every sequence worth considering. A neural network with two hidden layers of 50 neurons was trained on the simulation results, learning to predict the interface and self-interaction values directly from 44 numerical descriptors of a sequence, such as its composition, charge and the arrangement of its residues. In that role the network stood in for the simulations as a fast stand-in, or surrogate.
Because the network's maths is piecewise linear, it could be rewritten as a set of algebraic constraints and handed to a mathematical optimiser, which searched sequence space for the best trade-offs rather than testing candidates one by one. Promising sequences went back into simulation, the network was retrained, and the cycle repeated. The sequences that were synthesised and imaged came out of this loop; microscopy showed all three designed peptides at the interface, while the controls spread through the droplet interior.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
A computational pipeline combined coarse-grained molecular dynamics, a neural-network surrogate model and mixed-integer linear programming to design 30-residue peptides that partition at the interface of condensates formed by the disordered regions of hnRNPA1, LAF-1 and DDX4. Simulations quantified each peptide's interface partitioning (pint*) and self-interaction (B2*), a multi-output network was trained on those values, and the network embedded in an MILP produced Pareto-optimal sequences for the next simulation round. Three designed peptides, one per condensate target, were synthesized with a Cy5 label, and confocal microscopy showed interfacial localization in all three cases, while control peptides distributed uniformly through the dense phase. The designs had surfactant-like architectures, with an aromatic-rich tail inserting into the condensate and an excluded tail whose composition varied with the net charge of the condensate-forming protein.
How AI was used
After randomized 30-residue peptides were filtered with the sequence-based aggregation predictors Waltz, TANGO and AGGRESCAN, coarse-grained Mpipi simulations with adaptive biasing force were run for each peptide to obtain interface-partitioning free energies and a second virial coefficient. A multi-output neural network with two fully connected hidden layers of 50 neurons and ReLU activations was trained in PyTorch on these simulation outputs, taking as input 44 engineered descriptors that are linear transformations of the one-hot encoded sequence. Because the features are linear and ReLU is piecewise linear, the trained network was embedded via OMLT into a Pyomo model and reformulated with big-M constraints as a mixed-integer linear program, with the AGGRESCAN predictor added as a constraint and equation (1) linearized by supporting hyperplanes. Gurobi solved this bi-objective program to global optimality, using the ε-constrained method for exploitation sequences and a weighted-sum formulation with randomly fixed sequence positions for exploration sequences; the selected sequences were simulated and the surrogate retrained, iterating until the Pareto-front hypervolume stagnated. Final sequences were passed through the Waltz and TANGO filters before chemical synthesis, Cy5 labelling and confocal imaging, and the selected peptides were also simulated with the CALVADOS 2 and Martini3-IDP force fields.
The shape of the work
Structural · the record, drawn
no AI
Define condensate target and build simulation slab
Obtaining raw data, whether by measurement, download or retrieval.
modeling the condensate as a minimalistic slab containing 16 protein copieswhere the paper describes this · verbatim
no AI
Initialize random peptide library and apply aggregation filters
Reducing a candidate set by filtering or ranking, in a single pass.
we applied three different sequence-based aggregation predictors: Waltz, TANGO, and AGGRESCANwhere the paper describes this · verbatim
no AI
Coarse-grained simulations of interface partitioning and homotypic interaction
Numerical or physics simulation, including where a learned surrogate replaces it.
We performed two separate coarse-grained simulations for each peptidewhere the paper describes this · verbatim
no AI
Encode sequences as engineered linear features
Encoding data into features, descriptors, embeddings or graphs.
We engineered a set of 44 features based on overall compositionwhere the paper describes this · verbatim
AI
Train multi-output surrogate neural network
Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.
a multi-output neural network with two fully connected hidden layers of 50 neurons each was trained on ΔG1, ΔG2, and B2* with PyTorchwhere the paper describes this · verbatim
AI
Solve MILP with embedded network for Pareto-optimal sequences
Iterative search over a space. The AI stood in for exhaustive search. Its result feeds back into an earlier step.
we leveraged MILP to identify globally optimal sequences, leading to new peptides to simulatewhere the paper describes this · verbatim
no AI
Apply final aggregation filters and select peptides
Reducing a candidate set by filtering or ranking, in a single pass.
Following convergence of the cycle, we applied final aggregation filters and selected a sequence for experimental validationwhere the paper describes this · verbatim
no AI
Synthesize labelled peptides and image condensates
Physical execution, by hand or by robot.
We conjugated the synthesized peptides with the fluorescent dye Cy5, before incubation with hnRNPA1-LCD condensateswhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The sequences that were synthesized and imaged came out of the neural-network surrogate embedded in the MILP; the designed peptides are the result the paper reports
we trained a surrogate model (multi-output neural network) on the simulation resultswhere the paper describes this · verbatim
Interfacial accumulation of the designed peptides was experimentally confirmed in vitro for all three caseswhere the paper describes this · verbatim
The developed code, molecular dynamics input files, and final configurations are available on Zenodowhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- Version of multi-output neural network (two hidden layers of 50 neurons, ReLU)Which version of the model was used is not stated.
- Version of elastic netWhich version of the model was used is not stated.
- Version of support vector machineWhich version of the model was used is not stated.
- Version of gradient-boosted decision treeWhich version of the model was used is not stated.
About this article
Record aix-00163, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error