structural-biology/ai produced the result/Protein Science 2025 · v2
Adding a term for the unfolded protein sharpens AI stability predictions
Researchers added a fitted correction for the unfolded state of a mutated protein to scores from two machine-learning predictors, ESM-IF1 and Pythia. The models supplied the folded-state scores; a small linear model learned the correction.
spectrum · one line per step, placed by what the step does · bright lines used AI
Mass balance approximation of unfolding boosts potential‐based protein stability predictions
Protein Science, 2025
doi:10.1002/pro.70134 · record aix-00063 v2 · checked 2026-10-08
- AI was for
- Property prediction
- Model family
- Protein language model, Graph neural network, Linear model
- Checked by
- Benchmark
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Proteins are chains of amino acids that fold into particular shapes. A folded protein is only marginally more stable than the loose, unfolded chain, and changing a single amino acid can tip that balance. Biologists measure the shift as ΔΔG, the change in the free energy of folding caused by a mutation. Predicting it from a structure alone is hard, partly because stability is a difference between two states. Most scoring methods look closely at the folded shape, where the mutated residue sits among its neighbours, and treat the unfolded chain as though the mutation barely mattered there.
The authors set out to supply that missing half. They added a term standing for the free-energy difference between the unfolded states of the original and the mutated protein, and asked whether bolting it onto existing predictors changed how well those predictors matched experiment.
Where AI came in
Two learned models produced the folded-state scores. ESM-IF1 is a protein language model that reads backbone atom coordinates and judges how likely a sequence is given that shape. Pythia is a graph neural network trained on proteins without labelled stability data, so it can score mutations it has never seen. Both were run over original and mutant structures, including shapes predicted by AlphaFold, alongside FoldX, a non-learned empirical potential whose published scores were reused.
The correction itself was learned too, but by a much simpler model. Each mutation was written as a twenty-element vector marking which amino acid left and which arrived, and ridge regression fitted coefficients on a training set of 3322 mutations. On the S461 test set Pythia's Pearson correlation with measured values went from 0.41 to 0.56. Methods already carrying such information changed little or got worse.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors added a mass‐balance correction term, representing the free‐energy difference between the unfolded states of wild‐type and mutant, to ΔΔG scores produced by potential‐like stability predictors including the protein language model ESM‐IF1, the graph neural network Pythia and the empirical potential FoldX. The correction was a linear term fitted by ridge regression on a 3322‐mutation training set, or a two‐parameter combination with the experimentally derived Rose solvation scale, and required no refitting of the base methods. Pearson correlation on the S461 test set increased for the corrected methods, with Pythia reported as going from 0.41 to 0.56, while methods that already encode mass‐balance information, such as Stability Oracle and DDGun3D, changed little or decreased. On an independent mega‐scale dataset the corrected Pythia score improved Pearson correlation by 0.07 over the original score, reaching a correlation close to 0.70 and an RMSE of 1.43 kcal/mol.
How AI was used
Learned models supplied the folded‐state term of the stability change: ESM‐IF1, a structure‐conditioned protein language model, and Pythia, a self‐supervised graph neural network for zero‐shot ΔΔG prediction, were run over wild‐type and mutant structures (both experimental PDB structures and AlphaFold models) to produce per‐mutation scores, alongside the non‐learned FoldX potential whose published scores were reused. Each mutation was separately encoded as a 20‐element occurrence vector (−1 for the wild‐type residue, +1 for the substitution), and a linear model over that encoding plus the original method score was fitted by ridge regression in Scikit‐learn on the VBS3322 training set, giving 21 coefficients; a two‐parameter variant replaced the encoding with the difference of Rose‐scale values. The fitted coefficients were then applied to method scores on the S461 test set, on an independent mega‐scale dataset, and to previously published predictions of 48 methods, with coefficients for that last comparison fitted on reported Ssym predictions. Pearson correlation and RMSE against experimental ΔΔG were computed, and the fitted residue coefficients were correlated with the Kyte–Doolittle and Rose scales.
The shape of the work
Structural · the record, drawn
no AI
Assemble mutation datasets and repair structures
Cleaning, filtering, normalising or labelling data already obtained.
The main training set used in this work, namely VBS3322, consists of 3322 mutationswhere the paper describes this · verbatim
no AI
Encode each mutation as an occurrence vector
Encoding data into features, descriptors, embeddings or graphs.
We encode the mutation in the sequence as a 20‐element array, one element for each of the natural amino acidswhere the paper describes this · verbatim
AI
Score mutations with potential‐like methods
Running a trained model over new data to predict, classify or score. The AI stood in for simulation.
a large protein‐language model (PLM) trained to predict a protein sequence likelihood from its backbone atom coordinateswhere the paper describes this · verbatim
AI
Fit mass‐balance correction coefficients
Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.
by fitting it to the training set using ridge regression implemented in Scikit‐learnwhere the paper describes this · verbatim
AI
Apply corrected scores to test sets and to other methods
Running a trained model over new data to predict, classify or score.
we computed the Pythia/MBC(dd) and Pythia/MBC(Rose) scores using the parameters derived from our VBS3322 training setwhere the paper describes this · verbatim
no AI
Evaluate against experimental ΔΔG and baselines
Testing outputs against ground truth.
The standard scoring values calculated in our assessment are the Pearson correlation coefficients (PCC) and the root mean square error (RMSE)where the paper describes this · verbatim
no AI
Compare fitted coefficients with hydrophobicity and solvation scales
Extracting understanding from model behaviour.
We performed a Pearson correlation analysis among the residue‐specific parameters fitted using the VBS3322 datasetwhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The reported finding is a change in predictive accuracy of learned ΔΔG predictors (ESM‐IF1, Pythia) once a fitted mass‐balance term is added, so the result exists only through the models' outputs
Pythia, a self‐supervised graph neural network tailored for zero‐shot ∆∆G predictionswhere the paper describes this · verbatim
We used the S461 dataset (Hernández et al., ) as the test set to perform comparisonswhere the paper describes this · verbatim
The python codes and the data used in this study can be downloaded from Githubwhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of ESM‐IF1Which version of the model was used is not stated.
- Version of PythiaWhich version of the model was used is not stated.
- Version of Stability OracleWhich version of the model was used is not stated.
- Version of MBC(dd) ridge‐regression linear model (ddMBC)Which version of the model was used is not stated.
- Version of MBC(Rose) two‐parameter linear modelWhich version of the model was used is not stated.
- Version of AlphaFoldWhich version of the model was used is not stated.
- What step 5 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00063, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error