~/aixsci
200 records · all checked

structural-biology/ai produced the result/Frontiers in Molecular Biosciences 2025 · v2

AlphaFold's confidence scores used to predict which disordered protein segments bind LC8

Researchers asked a structure-prediction program to model a hub protein alongside short stretches of its partners, then used the program's own confidence numbers to sort binders from non-binders and to flag candidate sites in 24 proteins.

1. Assemble binder and nonbinder sequence sets2. Parse sequences into segments and build prediction inputs3. Co-predict LC8–client complex structures with AlphaFold4. Extract and assemble AlphaFold score parameters5. Train composite score threshold classifier6. Relate scores to measured affinities and reason about the learned energy function7. Screen 24 proteins for new binding sites8. Compare predictions with LC8Pred and known binding sites

spectrum · one line per step, placed by what the step does · bright lines used AI

Successful prediction of LC8 binding to intrinsically disordered proteins sheds light on AlphaFold’s black box
Frontiers in Molecular Biosciences, 2025

doi:10.3389/fmolb.2025.1531793 · record aix-00169 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Classification, Structure determination, Property prediction
Model family
Transformer, Linear model
Checked by
Held-out
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Many proteins have no fixed shape. These intrinsically disordered regions stay floppy until they dock onto a partner, and a short stretch of the chain then settles into place. One such partner is LC8, a small protein that works as a pair and acts as a hub, gripping short motifs in dozens of other proteins. Working out which stretch of a long, shapeless chain actually binds is hard. The motifs are short, their sequences vary, and the usual way to settle the question is laboratory work on one candidate at a time. Simple sequence patterns catch some sites and miss others.

The researchers set out to test whether a general-purpose structure predictor could answer the narrower question of whether a given sequence binds LC8 at all, and if so how tightly, and then to use it to propose new sites.

Where AI came in

AlphaFold-Multimer, run locally through ColabFold, was given short client segments together with one or two copies of LC8 and asked to model the complex. The point was not the models themselves but the confidence numbers AlphaFold reports about them: an overall confidence score, how well-ordered the contact surface looks, and three measures of how sure the program is about the relative placement of the parts. A genetic algorithm then fitted cut-offs across ten of these numbers at once, giving a stricter setting that reports few false binders and a looser one that misses few real ones.

Those numbers took the place of binding measurements at the bench. Scores were also compared against affinities already measured in the laboratory, which let the authors write simple linear relations and turn scores into rough strength estimates. Applying the cut-offs to 24 proteins, scanned in overlapping 16-residue windows, produced lists of candidate sites, which were filtered by estimated strength and compared with a dedicated sequence-based predictor and with sites already known from the literature. The candidate sites remain untested.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

AlphaFold-Multimer, run locally through ColabFold, was used to co-predict structures of the LC8 dimer together with client peptide sequences, and AlphaFold's own confidence metrics (confidence score, interface pLDDT, and three PAE measures, in 1-client and 2-client contexts) were used to decide whether a sequence binds LC8. A genetic algorithm fitted a composite threshold over 10 of these score parameters; the paper reports an AUROC of 0.9274 for the composite threshold against 0.9104 for the best single parameter, an exclusive threshold with a 5.5% false-positive and 28.1% false-negative rate, and an inclusive threshold with a 27.9% false-positive and 7.5% false-negative rate. Scores were then compared with 62 experimentally measured binding-site affinities using linear regression and adjusted mutual information, and the authors use qualitative Bayesian reasoning to argue that AlphaFold has learned an energy function that matches the natural one near minima but is inaccurate elsewhere. Applying the thresholds to 24 proteins known to bind LC8, parsed into 16-residue segments, gave 247 predicted sites under the exclusive threshold and 776 under the inclusive threshold, reduced to 133 and 213 after filtering by predicted affinity; the predicted sites are untested.

How AI was used

Binder and nonbinder client sequences were collected from the literature and from AAA anchor mutants, and a Python script parsed protein sequences into subsequences and wrote FASTA inputs pairing each client with one or two LC8 copies. ColabFold, installed on a local server, ran AlphaFold-Multimer on these inputs in both 1-client and 2-client configurations, returning mostly 25 structures per prediction (a few early runs returned five, some later runs 100); approximately 7,500 predictions were run, including 641 nonbinding and 333 binding sequences with clients 10 to 71 residues long, plus more than 6,000 predictions scanning 24 proteins in 16-residue segments with 8-residue overlap. Python scripts extracted five score types per run — confidence score, average interface pLDDT, LtoP PAE, PtoL PAE and LC8 dimer interface PAE — assigned binding status, and linked 2-client to 1-client results. A genetic algorithm optimised the AUROC of a composite threshold over all 10 resulting parameters, with learning curves generated by training on different fractions of the data and 10 iterations per fraction; the reported thresholds came from the highest-AUROC run trained on 90% of the data, and were reported as an exclusive and an inclusive set. Scores for predictions matching 62 sequences with known affinities were regressed against affinity in Excel using LINEST, and adjusted mutual information was computed with scikit-learn under three affinity binning schemes and three score binning schemes. The strongest linear relations for TQT and non-TQT anchors were written as equations to estimate affinities for newly predicted sites, which were filtered at 40 µM and screened against structured-domain assignments from the AlphaFold Database, then compared with LC8Pred predictions. ChatGPT was used in writing the analysis scripts.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONINFERENCEREPRESENTATIONTRAININGINTERPRETATIONSCREENINGVALIDATION12345678AIAIAssemble binderand nonbindersequence setsParse sequencesinto segments andbuild prediction…Co-predictLC8–clientcomplex structur…Extract andassembleAlphaFold score …Train compositescore thresholdclassifierRelate scores tomeasuredaffinities and r…Screen 24proteins for newbinding sitesComparepredictions withLC8Pred and know…↤ physical experiment↤ exhaustive search
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Assemble binder and nonbinder sequence sets

Obtaining raw data, whether by measurement, download or retrieval.

The information on LC8 binders was collected from the literature, which included our database LC8hub, previously assembled libraries, and other reports on binders.where the paper describes this · verbatim
in the paper
2Preparation
no AI

Parse sequences into segments and build prediction inputs

Cleaning, filtering, normalising or labelling data already obtained.

a Python script was employed that reads a full protein sequence, extracts a subsequence, adds the necessary LC8 sequences, and generates a new FASTA filewhere the paper describes this · verbatim
in the paper
3Inference
AI

Co-predict LC8–client complex structures with AlphaFold

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

All sequences of interest were run using ColabFold, locally installed on a server hosted by the Vanegas laboratorywhere the paper describes this · verbatim
in the paper
4Representation
no AI

Extract and assemble AlphaFold score parameters

Encoding data into features, descriptors, embeddings or graphs.

processed with a set of three Python scripts to extract the AlphaFold scores from each predictionwhere the paper describes this · verbatim
in the paper
5Training
AI

Train composite score threshold classifier

Fitting model parameters, including fine-tuning an existing model. The AI stood in for exhaustive search.

In training the classifier, a genetic algorithm was used to optimize the AUROC.where the paper describes this · verbatim
in the paper
6Interpretation
no AI

Relate scores to measured affinities and reason about the learned energy function

Extracting understanding from model behaviour.

Experimental binding affinities were plotted against 21 scoring metricswhere the paper describes this · verbatim
in the paper
7Screening
no AI

Screen 24 proteins for new binding sites

Reducing a candidate set by filtering or ranking, in a single pass.

The exclusive threshold resulted in 247 predicted binding sites among the 24 proteins, while the inclusive threshold yielded 776 predicted binding sites.where the paper describes this · verbatim
in the paper
8Validation
no AI

Compare predictions with LC8Pred and known binding sites

Testing outputs against ground truth.

AlphaFold identifies the known binding site in all of these except MARK3.where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

AlphaFold's own confidence scores are the measurement from which every binder/nonbinder call, the affinity correlations and the 24-protein site predictions are derived; the paper's conclusions do not exist without them

~What the AI was for
we probed the ability of a general structure predictor, AlphaFold, to predict whether a given sequence binds to LC8where the paper describes this · verbatim
~Model families
~How it was taught
Zero-shotSupervisedour reading
~Models named
AlphaFold-Multimer (run via ColabFold) · Off the shelfComposite 10-parameter threshold classifier (genetic-algorithm optimised) · Trained from scratchLC8Pred · Off the shelfChatGPT · Off the shelfour reading
+How results were checked
Held-outin the paper
by training on different fractions of the available data and testing on the restwhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
The scripts used here are available at https://github.com/DouglasRWalker/AlphaFold_BindingPrediction.where the paper describes this · verbatim
+Compute
ColabFold installed locally on a server hosted by the Vanegas laboratory; no accelerator type, run time or total compute is givenin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 6 items
  • Trained model weightsWhether the trained model is available is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of AlphaFold-Multimer (run via ColabFold)Which version of the model was used is not stated.
  • Version of Composite 10-parameter threshold classifier (genetic-algorithm optimised)Which version of the model was used is not stated.
  • Version of LC8PredWhich version of the model was used is not stated.
  • Version of ChatGPTWhich version of the model was used is not stated.

About this article

Record aix-00169, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error