structural-biology/ai produced the result/Frontiers in Molecular Biosciences 2022 · v2
Machine learning tested against three definitions of protein shape-shifting
Researchers trained random forest classifiers on sequence-derived biophysical predictions to tell ordered, disordered and 'ambiguous' protein residues apart. The models' scores and rules were the evidence used to compare rival definitions of protein order.
spectrum · one line per step, placed by what the step does · bright lines used AI
Challenges in describing the conformation and dynamics of proteins with ambiguous behavior
Frontiers in Molecular Biosciences, 2022
doi:10.3389/fmolb.2022.959956 · record aix-00172 v2 · checked 2026-10-09
- AI was for
- Classification, Property prediction
- Model family
- Random forest
- Checked by
- Held-out
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
Proteins are chains of amino acids, and much of biology textbook teaching assumes each chain folds into one fixed shape. Many do not. Some stay floppy until they meet a partner molecule, then settle into a shape only once bound. Others flip between two different folds. Both behaviours sit awkwardly between the neat categories of 'ordered' and 'disordered'. That matters because the categories are not just labels: they feed the curated databases and prediction tools that researchers use to reason about what a protein does. If a residue's behaviour is ambiguous, it is not obvious which bin it belongs in, or whether the two kinds of ambiguity belong in the same bin at all.
The authors assembled residue-by-residue labels from curated structural resources: disorder-to-order transitions, conformationally stable residues, secondary-structure switches annotated from pairs of structures, and a set of mutually folding proteins kept aside. They then asked how separable these classes really are from sequence alone, and how the answer changes when folding-upon-binding and fold-switching are merged into a single ambiguous class.
Where AI came in
Seven per-residue properties were predicted from each sequence by existing tools — backbone and side-chain dynamics, helix, sheet and coil propensities, early folding propensity and disorder. Those predictions alone, with the amino acid identities withheld, were the input to random forest classifiers: one trained on the folding-upon-binding labels, one on the fold-switching labels, one on the merged set. A random forest is a crowd of simple decision trees that vote on an answer. The models were built for interpretability rather than top scores, so the authors read off which features mattered most and summarised each forest as a set of plain if-then rules.
The classifiers stood in for the judgement a curator would otherwise make residue by residue. Their performance became the measurement: where a class scored poorly, that was taken as evidence about the definition rather than only about the model. The merged model was then applied to the held-out mutually folding set, and its per-residue calls across the human proteome were compared with AlphaFold2's per-residue confidence scores, with sites of chemical modification, and with disease-linked and benign mutations.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors assembled residue-level datasets of three kinds of protein conformational behaviour — ordered, disordered, and 'ambiguous' residues that either fold upon binding or switch secondary structure — and trained random forest classifiers on seven sequence-based biophysical predictions to separate these classes. Sequence-predicted disorder, early folding propensity and backbone dynamics were the most important features; the model trained on the folding-upon-binding set had its lowest F1 score for the disorder class, and the fold-switching model reached an F1 of 0.36 for converting residues, with a recall of 0.26. Merging the folding-upon-binding and fold-switching categories into one combined model lowered F1 performance for the merged ordered class relative to the separate ordered and same classes. Applying the combined model to a held-out MFIB set of mutually folding proteins predicted 79.6% of residues as ordered, 20% as ambiguous and under 1% as disordered, and comparisons with AlphaFold2 pLDDT values placed ambiguous residues between the ordered and disordered classes.
How AI was used
Residues were labelled from curated structural resources: disorder-to-order transitions from DisProt, conformationally stable residues from CoDNaS clusters filtered at a 2 A maximum pairwise RMSD, secondary-structure switches from DSSP annotations of fold-switcher structure pairs, and a redundancy-filtered MFIB set held out of training. For every sequence, seven per-residue features were predicted with the b2bTools suite — backbone and side-chain dynamics and conformational propensities from DynaMine, early folding propensity from EFoldMine, and disorder from DisoMine — and these features alone, with no amino acid codes, were the inputs to random forest classifiers built with scikit-learn: one on the DisProt/CoDNaS labels split 90/10 into train and test, one on the fold-switching labels, and one on the merged combined set split 70/30, each with hyperparameters chosen by 3-fold cross-validation. The stated aim was interpretability rather than best performance, so feature importances were extracted and each forest was additionally summarised by a surrogate rule model induced with the Weka implementation of the Ripper algorithm over the forest's predictions. The combined model was then run over the MFIB set, and the per-residue predictions for the human proteome were cross-tabulated against AlphaFold2 pLDDT values and DSSP categories from downloaded AlphaFold2 models, against PTM sites compiled from four databases, and against deleterious, benign, somatic and germline missense variants.
The shape of the work
Structural · the record, drawn
no AI
Assemble labelled residue datasets
Obtaining raw data, whether by measurement, download or retrieval.
we downloaded a custom set of human proteins with manually curated disorder-to-order structural transitions, resulting in 138 different proteinswhere the paper describes this · verbatim
no AI
Assemble proteome, PTM and mutation annotation sets
Obtaining raw data, whether by measurement, download or retrieval.
AlphaFold 2’s mmCIF files for the human proteome were downloaded on 2 September 2021, from the AlphaFold protein structure database.where the paper describes this · verbatim
AI
Predict per-residue biophysical features from sequence
Encoding data into features, descriptors, embeddings or graphs.
seven biophysical features were predicted at the residue level using the following methods: backbone dynamics (DynaMine)where the paper describes this · verbatim
AI
Train and test random forest residue classifiers
Fitting model parameters, including fine-tuning an existing model. The AI stood in for manual curation.
We used a combination of these datasets (disprot_codnas_set) to train a random forest (RF) predictorwhere the paper describes this · verbatim
AI
Classify MFIB residues with the combined model
Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.
the disordered class, without ambiguous folding propensity, was shown to be depleted in the output of the combined_RF predictorwhere the paper describes this · verbatim
AI
Derive surrogate rule models for interpretation
Extracting understanding from model behaviour. The AI stood in for expert judgement.
The RF models were interpreted using a surrogate model trained over the predictions for each of the models.where the paper describes this · verbatim
no AI
Relate classes to pLDDT, PTMs and variants
Extracting understanding from model behaviour.
related the key biophysical predictions of the selected_human_set with the respective pLDDT values of the AlphaFold2 modelswhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's comparisons between definitions of order, disorder and ambiguity are read off the performance, feature importance and surrogate rules of random forest classifiers the authors trained, and off sequence-based predictor outputs; without those models there is no result.
The classification model was trained using seven predicted biophysical features at the residue levelwhere the paper describes this · verbatim
The RF model is trained using those hyperparameters and finally tested on the remaining 10% of the data (test set)where the paper describes this · verbatim
The complete dataset is available at https://bitbucket.org/bio2byte/protein_ambiguity/.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of folding_upon_binding_RF (random forest)Which version of the model was used is not stated.
- Version of fold_switching_RF (random forest)Which version of the model was used is not stated.
- Version of combined_RF (random forest)Which version of the model was used is not stated.
- Version of Ripper (Weka JRip) surrogate rule modelWhich version of the model was used is not stated.
- Version of DynaMineWhich version of the model was used is not stated.
- Version of DisoMineWhich version of the model was used is not stated.
- Version of EFoldMineWhich version of the model was used is not stated.
- Version of AlphaFold2Which version of the model was used is not stated.
- What step 3 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00172, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error