structural-biology/ai produced the result/Nature Communications 2024 · v2
Teaching protein language models to rank mutants from a few dozen lab measurements
Researchers built a training strategy, FSFP, that adapts large protein language models using only tens of measured mutants. The models then ranked mutants across a public benchmark and picked candidates for a DNA-copying enzyme tested in the lab.
spectrum · one line per step, placed by what the step does · bright lines used AI
Enhancing efficiency of protein language models with minimal wet-lab data through few-shot learning
Nature Communications, 2024
doi:10.1038/s41467-024-49798-6 · record aix-00180 v2 · checked 2026-10-09
- AI was for
- Property prediction, Experimental design
- Model family
- Protein language model, Transformer, Linear model
- Checked by
- Experimental20 tested
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Proteins are chains of amino acids, and changing a single one can make the protein more stable, less stable, or stop it working. Biologists call the useful measure of how well a variant performs its fitness. Working out which single changes help is hard: a modest protein has thousands of possible single swaps, and measuring each one in the laboratory costs time and money. Methods that learn from measurements usually need many of them, which is exactly what is in short supply when a new protein is being engineered.
The authors set out to make large models of protein sequence useful when only a handful of laboratory measurements exist. Their strategy, FSFP, adapts an already-trained model using tens of measured single-amino-acid variants of the protein in hand, rather than hundreds or thousands.
Where AI came in
Three pre-trained protein language models — ESM-1v, ESM-2 and SaProt, programs trained on large collections of protein sequences much as text models are trained on text — did the predicting. The model first compared the target protein with proteins in the ProteinGym collection of mutation measurements and borrowed the two most similar sets as practice material. A third practice set came from GEMME, a method based on alignments of related sequences, which supplied stand-in labels. The model was then primed on these practice tasks and finally tuned on the target protein's own labelled mutants, learning to put them in order rather than predict exact values.
The trained models were scored on held-out mutants across the 87 benchmark datasets, and compared with their untrained versions and with a ridge regression baseline. For Phi29 DNA polymerase, an enzyme that copies DNA, ESM-1v chose the top 20 single-amino-acid variants for laboratory melting-temperature measurements, standing in for screening every possible variant at the bench; those measurements then fed a further round of training.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors built FSFP, a training strategy that adapts pre-trained protein language models to predict the fitness of protein mutants using only tens of labelled single-site mutants from the target protein. It combines retrieval of similar mutational-scanning datasets, pseudo labels from the alignment-based method GEMME, meta-training with MAML, low-rank adaptation, and a listwise ranking loss. Across the 87 deep mutational scanning datasets of the ProteinGym substitution benchmark, FSFP-trained ESM-1v, ESM-2 and SaProt scored higher on average Spearman correlation than their zero-shot versions and than the ridge regression baseline at every training size tested. Applied to Phi29 DNA polymerase, the top 20 single-site mutants predicted by ESM-1v after FSFP training had an average melting temperature more than 1 °C higher and a positive rate 25% higher than the top 20 from the zero-shot model.
How AI was used
Three pre-trained protein language models — ESM-1v, ESM-2 and SaProt, each at 650 M parameters — were adapted to rank mutant fitness. The model to be trained first embedded wild-type sequences (or structures, for SaProt) of the target protein and of the proteins in ProteinGym, and the two datasets with highest cosine similarity became auxiliary meta-learning tasks; a third task was built by scoring candidate mutants of the target protein with the alignment-based method GEMME to produce pseudo labels. MAML was then used to meta-train the model on these tasks, with updates confined to LoRA rank-decomposition matrices injected into the self-attention and feed-forward weights (rank 16) while pre-trained weights stayed frozen, and a first-order approximation used for the outer gradient. The meta-trained initialisation was fine-tuned on the target protein's labelled mutants with a ListMLE listwise ranking loss, mutant scores being computed from the model's residue probabilities relative to the wild type; Monte Carlo cross-validation on the training data set the number of steps and early-stopped meta-training. Trained models were then run over held-out mutants of the benchmark and, for Phi29 DNA polymerase, over saturated single-site mutants to pick the top 20 for wet-lab melting-temperature assays, whose labels fed a further round of training. A ridge regression approach on one-hot plus density features, GEMME, and ridge-augmented variants were run as baselines.
The shape of the work
Structural · the record, drawn
no AI
Assemble benchmark and few-shot splits
Cleaning, filtering, normalising or labelling data already obtained.
we first randomly sample 20 single-site mutants as an initial training setwhere the paper describes this · verbatim
AI
Embed wild-type proteins and retrieve similar datasets
Encoding data into features, descriptors, embeddings or graphs. The AI stood in for conventional algorithm.
the relevance between the target protein and a candidate protein is measured by the cosine similarity between their embeddings p and qwhere the paper describes this · verbatim
AI
Generate MSA-based pseudo labels
Running a trained model over new data to predict, classify or score.
Utilizing GEMME, we score the candidate mutated sequences of the target protein and build the dataset of the third task.where the paper describes this · verbatim
AI
Meta-train PLM on auxiliary tasks with LoRA
Fitting model parameters, including fine-tuning an existing model.
We apply MAML, a state-of-the-art meta-learning algorithm, to enable PLMs to better utilize the few-shot training data.where the paper describes this · verbatim
AI
Fine-tune on target data with listwise ranking loss
Fitting model parameters, including fine-tuning an existing model. The AI stood in for statistical model.
We use a listwise LTR approach, namely ListMLE to train PLMs on few-shot training data.where the paper describes this · verbatim
AI
Score held-out mutants and benchmark against baselines
Testing outputs against ground truth.
the predictive performance is measured by two metrics: Spearman rank correlation and normalized discounted cumulative gain (NDCG) with the fitness labels as ground truthwhere the paper describes this · verbatim
AI
Select Phi29 single-site mutants for testing
Reducing a candidate set by filtering or ranking, in a single pass. The AI stood in for physical experiment.
Then, 20 mutants with the highest predicted scores are chosen to measure experimental Tm values.where the paper describes this · verbatim
no AI
Express, purify and measure Tm of mutants
Physical execution, by hand or by robot. Its result feeds back into an earlier step.
The Tm values are determined by differential scanning fluorimetry (DSF) method using the Protein Thermal Shift Dye Kit (Thermo Fisher).where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is the trained-model strategy itself and the mutants it selected for wet-lab testing; both the benchmark finding and the Phi29 engineering outcome depend entirely on the models
By combining meta-transfer learning, learning to rank, and parameter-efficient fine-tuning, FSFP can significantly boost the performance of various protein language modelswhere the paper describes this · verbatim
The top 20 predictions from the trained model for single-site mutants are selected for the next iteration of wet-lab experiments.where the paper describes this · verbatim
The source code of FSFP is available at https://github.com/ai4protein/FSFP.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- Version of GEMMEWhich version of the model was used is not stated.
- What step 3 replacedThe paper gives no basis for what the AI stood in for.
- What step 4 replacedThe paper gives no basis for what the AI stood in for.
- What step 6 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00180, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error