structural-biology/ai produced the result/Frontiers in Molecular Biosciences 2025 · v2
AlphaFold's confidence scores used to predict which disordered protein segments bind LC8
Researchers asked a structure-prediction program to model a hub protein alongside short stretches of its partners, then used the program's own confidence numbers to sort binders from non-binders and to flag candidate sites in 24 proteins.
spectrum · one line per step, placed by what the step does · bright lines used AI
Successful prediction of LC8 binding to intrinsically disordered proteins sheds light on AlphaFold’s black box
Frontiers in Molecular Biosciences, 2025
doi:10.3389/fmolb.2025.1531793 · record aix-00169 v2 · checked 2026-10-09
- AI was for
- Classification, Structure determination, Property prediction
- Model family
- Transformer, Linear model
- Checked by
- Held-out
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Many proteins have no fixed shape. These intrinsically disordered regions stay floppy until they dock onto a partner, and a short stretch of the chain then settles into place. One such partner is LC8, a small protein that works as a pair and acts as a hub, gripping short motifs in dozens of other proteins. Working out which stretch of a long, shapeless chain actually binds is hard. The motifs are short, their sequences vary, and the usual way to settle the question is laboratory work on one candidate at a time. Simple sequence patterns catch some sites and miss others.
The researchers set out to test whether a general-purpose structure predictor could answer the narrower question of whether a given sequence binds LC8 at all, and if so how tightly, and then to use it to propose new sites.
Where AI came in
AlphaFold-Multimer, run locally through ColabFold, was given short client segments together with one or two copies of LC8 and asked to model the complex. The point was not the models themselves but the confidence numbers AlphaFold reports about them: an overall confidence score, how well-ordered the contact surface looks, and three measures of how sure the program is about the relative placement of the parts. A genetic algorithm then fitted cut-offs across ten of these numbers at once, giving a stricter setting that reports few false binders and a looser one that misses few real ones.
Those numbers took the place of binding measurements at the bench. Scores were also compared against affinities already measured in the laboratory, which let the authors write simple linear relations and turn scores into rough strength estimates. Applying the cut-offs to 24 proteins, scanned in overlapping 16-residue windows, produced lists of candidate sites, which were filtered by estimated strength and compared with a dedicated sequence-based predictor and with sites already known from the literature. The candidate sites remain untested.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
AlphaFold-Multimer, run locally through ColabFold, was used to co-predict structures of the LC8 dimer together with client peptide sequences, and AlphaFold's own confidence metrics (confidence score, interface pLDDT, and three PAE measures, in 1-client and 2-client contexts) were used to decide whether a sequence binds LC8. A genetic algorithm fitted a composite threshold over 10 of these score parameters; the paper reports an AUROC of 0.9274 for the composite threshold against 0.9104 for the best single parameter, an exclusive threshold with a 5.5% false-positive and 28.1% false-negative rate, and an inclusive threshold with a 27.9% false-positive and 7.5% false-negative rate. Scores were then compared with 62 experimentally measured binding-site affinities using linear regression and adjusted mutual information, and the authors use qualitative Bayesian reasoning to argue that AlphaFold has learned an energy function that matches the natural one near minima but is inaccurate elsewhere. Applying the thresholds to 24 proteins known to bind LC8, parsed into 16-residue segments, gave 247 predicted sites under the exclusive threshold and 776 under the inclusive threshold, reduced to 133 and 213 after filtering by predicted affinity; the predicted sites are untested.
How AI was used
Binder and nonbinder client sequences were collected from the literature and from AAA anchor mutants, and a Python script parsed protein sequences into subsequences and wrote FASTA inputs pairing each client with one or two LC8 copies. ColabFold, installed on a local server, ran AlphaFold-Multimer on these inputs in both 1-client and 2-client configurations, returning mostly 25 structures per prediction (a few early runs returned five, some later runs 100); approximately 7,500 predictions were run, including 641 nonbinding and 333 binding sequences with clients 10 to 71 residues long, plus more than 6,000 predictions scanning 24 proteins in 16-residue segments with 8-residue overlap. Python scripts extracted five score types per run — confidence score, average interface pLDDT, LtoP PAE, PtoL PAE and LC8 dimer interface PAE — assigned binding status, and linked 2-client to 1-client results. A genetic algorithm optimised the AUROC of a composite threshold over all 10 resulting parameters, with learning curves generated by training on different fractions of the data and 10 iterations per fraction; the reported thresholds came from the highest-AUROC run trained on 90% of the data, and were reported as an exclusive and an inclusive set. Scores for predictions matching 62 sequences with known affinities were regressed against affinity in Excel using LINEST, and adjusted mutual information was computed with scikit-learn under three affinity binning schemes and three score binning schemes. The strongest linear relations for TQT and non-TQT anchors were written as equations to estimate affinities for newly predicted sites, which were filtered at 40 µM and screened against structured-domain assignments from the AlphaFold Database, then compared with LC8Pred predictions. ChatGPT was used in writing the analysis scripts.
The shape of the work
Structural · the record, drawn
no AI
Assemble binder and nonbinder sequence sets
Obtaining raw data, whether by measurement, download or retrieval.
The information on LC8 binders was collected from the literature, which included our database LC8hub, previously assembled libraries, and other reports on binders.where the paper describes this · verbatim
no AI
Parse sequences into segments and build prediction inputs
Cleaning, filtering, normalising or labelling data already obtained.
a Python script was employed that reads a full protein sequence, extracts a subsequence, adds the necessary LC8 sequences, and generates a new FASTA filewhere the paper describes this · verbatim
AI
Co-predict LC8–client complex structures with AlphaFold
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
All sequences of interest were run using ColabFold, locally installed on a server hosted by the Vanegas laboratorywhere the paper describes this · verbatim
no AI
Extract and assemble AlphaFold score parameters
Encoding data into features, descriptors, embeddings or graphs.
processed with a set of three Python scripts to extract the AlphaFold scores from each predictionwhere the paper describes this · verbatim
AI
Train composite score threshold classifier
Fitting model parameters, including fine-tuning an existing model. The AI stood in for exhaustive search.
In training the classifier, a genetic algorithm was used to optimize the AUROC.where the paper describes this · verbatim
no AI
Relate scores to measured affinities and reason about the learned energy function
Extracting understanding from model behaviour.
Experimental binding affinities were plotted against 21 scoring metricswhere the paper describes this · verbatim
no AI
Screen 24 proteins for new binding sites
Reducing a candidate set by filtering or ranking, in a single pass.
The exclusive threshold resulted in 247 predicted binding sites among the 24 proteins, while the inclusive threshold yielded 776 predicted binding sites.where the paper describes this · verbatim
no AI
Compare predictions with LC8Pred and known binding sites
Testing outputs against ground truth.
AlphaFold identifies the known binding site in all of these except MARK3.where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
AlphaFold's own confidence scores are the measurement from which every binder/nonbinder call, the affinity correlations and the 24-protein site predictions are derived; the paper's conclusions do not exist without them
we probed the ability of a general structure predictor, AlphaFold, to predict whether a given sequence binds to LC8where the paper describes this · verbatim
by training on different fractions of the available data and testing on the restwhere the paper describes this · verbatim
The scripts used here are available at https://github.com/DouglasRWalker/AlphaFold_BindingPrediction.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of AlphaFold-Multimer (run via ColabFold)Which version of the model was used is not stated.
- Version of Composite 10-parameter threshold classifier (genetic-algorithm optimised)Which version of the model was used is not stated.
- Version of LC8PredWhich version of the model was used is not stated.
- Version of ChatGPTWhich version of the model was used is not stated.
About this article
Record aix-00169, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error