~/aixsci
200 records · all checked

structural-biology/ai produced the result/Journal of Proteins and Proteomics 2025 · v2

Software predictors compare protein floppiness across retinal disease protein sets

Researchers assembled six sets of retinal proteins, one healthy control and five tied to eye disease, then used off-the-shelf prediction tools to estimate how disordered each protein is and how readily it might form liquid droplets. No laboratory measurements were made.

1. Assemble control and disease retinal gene/protein sets2. Map genes to reviewed UniProt entries and retrieve sequences3. Predict per-residue intrinsic disorder4. Classify proteins by disorder level and CH–CDF quadrant5. Predict liquid–liquid phase separation propensity6. Statistically compare proteomes, including unique versus overlapping proteins

spectrum · one line per step, placed by what the step does · bright lines used AI

The effects of retinal disease on intrinsic protein disorder and liquid–liquid-phase separation
Journal of Proteins and Proteomics, 2025

doi:10.1007/s42485-025-00188-6 · record aix-00124 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Property prediction
Model family
Gradient-boosted trees
Checked by
None stated
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Most proteins are pictured as folding into a fixed shape, but many contain stretches that never settle into one. These floppy stretches are called intrinsically disordered regions. They matter because they help proteins stick together loosely and, in some cases, gather into droplets inside the cell, rather like oil separating in water. Biologists call that liquid–liquid phase separation. Measuring either property in the laboratory is slow, and doing it for thousands of proteins at once is not practical. So researchers often turn to software that guesses these properties from the amino acid sequence alone.

The team looked at the retina, the light-sensing tissue at the back of the eye. They drew up a control list of retinal proteins from the Human Protein Atlas, and five disease lists: inherited retinal disease genes from the RetNet database, plus proteins reported in published studies of age-related macular degeneration, glaucoma, and diabetic retinopathy with and without gliosis. Gene names were matched to reviewed database entries to obtain each protein's amino acid sequence. The aim was to compare predicted disorder and predicted droplet-forming tendency across the six groups.

Where AI came in

Every number about disorder and droplet formation in this work came from trained prediction models, used as released, with nothing fitted to the study's own data. The sequences were fed to a platform called RIDAO, which runs a set of established disorder predictors, returning a score for each amino acid and summary figures per protein. The same sequences went to PSPredictor, a model built from gradient-boosted decision trees — a method that combines many simple decision rules — and to ParSe V2, which flags disordered regions likely to drive droplet formation.

The models stood in for physical experiments. Rather than purifying each protein and testing its structure or its droplet behaviour at the bench, the researchers took the predicted values as their measurements, then compared them between groups using standard statistical tests. The paper states that it offers no experimental confirmation of these predictions, so the findings rest on how well the models generalise. The control set was predicted to have the largest share of disordered residues, at 43.660%, and the macular degeneration set the smallest, at 33.260%.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

Sequence-based predictors were run over six retinal protein sets — a Human Protein Atlas retinal control plus RetNet, age-related macular degeneration, glaucoma and diabetic retinopathy with and without gliosis — to estimate intrinsic protein disorder and the propensity to undergo liquid–liquid phase separation. The HPA control proteome showed the highest percentage of predicted disordered residues at 43.660%, while the AMD proteome showed the lowest at 33.260%. Average PSPredictor phase-separation scores were 0.325 for HPA and 0.329 for RetNet, compared with 0.168 for AMD and 0.148 for glaucoma. The paper reports no experimental validation of these predictions.

How AI was used

Gene and protein lists were drawn from the Human Protein Atlas retinal transcriptome, the RetNet database and three published disease proteomic studies, then mapped to reviewed Swiss-Prot entries to obtain canonical FASTA sequences. Those sequences were submitted to the RIDAO platform, which runs the PONDR VSL2, VL3, VLXT and FIT predictors together with IUPred short and long, yielding per-residue disorder scores summarised as average disorder score and percentage of predicted disordered residues; PONDR VSL2B output was used for the main analysis and for charge-hydropathy and cumulative distribution function quadrant assignment. The same FASTA files were submitted to PSPredictor, a gradient boosting decision tree model over protein embeddings, and to ParSe V2, which identifies phase-separating intrinsically disordered regions and produces recall plots against a reference human proteome. All predictors were used as released, with no model fitted to this study's data; the predicted scores were then compared across proteomes with ANOVA, pairwise t-tests, Tukey HSD and chi-squared tests, including a comparison of proteins unique to each disease set against those overlapping the control set.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONINFERENCEINTERPRETATIONINFERENCEINTERPRETATION123456AIAIAssemble controland diseaseretinal gene/pro…Map genes toreviewed UniProtentries and retr…Predictper-residueintrinsic disord…Classify proteinsby disorder leveland CH–CDF quadr…Predictliquid–liquidphase separation…Statisticallycompareproteomes, inclu…↤ physical experiment↤ physical experiment
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Assemble control and disease retinal gene/protein sets

Obtaining raw data, whether by measurement, download or retrieval.

Our control group utilizes the defined retinal proteome from the Human Protein Atlas (HPA), a Swedish-based open-access repository started in 2003.where the paper describes this · verbatim
in the paper
2Preparation
no AI

Map genes to reviewed UniProt entries and retrieve sequences

Cleaning, filtering, normalising or labelling data already obtained.

the ID mapping tool was used to return a list of unique UniProt IDs for each proteomic groupwhere the paper describes this · verbatim
in the paper
3Inference
AI

Predict per-residue intrinsic disorder

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

we quantified intrinsic disorder at the level of individual residues using the Rapid Intrinsic Disorder Analysis Online (RIDAO) platformwhere the paper describes this · verbatim
in the paper
4Interpretation
no AI

Classify proteins by disorder level and CH–CDF quadrant

Extracting understanding from model behaviour.

Proteins were characterized as highly ordered if their percentage of predicted disordered residues (PPDR) was less than 10%where the paper describes this · verbatim
in the paper
5Inference
AI

Predict liquid–liquid phase separation propensity

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

To evaluate our six proteomes for their corresponding likelihood to undergo LLPS, we inputted each post-UniProt mapped FASTA file into these two predictors.where the paper describes this · verbatim
in the paper
6Interpretation
no AI

Statistically compare proteomes, including unique versus overlapping proteins

Extracting understanding from model behaviour.

we performed analysis of variance (ANOVA) tests to test for statistical significance of variations in ADS and PPDR across our protein groupswhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

All reported disorder and LLPS quantities are outputs of sequence-based predictors; no experimental measurement is reported, so the findings rest on the models' predictions.

~What the AI was for
applies a Gradient Boosting Decision Tree (GBDT) algorithm to predict the likelihood of proteins undergoing LLPSwhere the paper describes this · verbatim
~Model families
~How it was taught
Supervisedour reading
~Models named
PSPredictor · Off the shelfParSe V2 V2 · Off the shelfPONDR VSL2 (VSL2B output used) · Off the shelfPONDR VL3 · Off the shelfPONDR VLXT · Off the shelfPONDR FIT · Off the shelfIUPred (short and long) · Off the shelfour reading
+How results were checked
None statedin the paper
our analysis is computational and bioinformatics-based, lacking experimental validationwhere the paper describes this · verbatim
+Code · weights · data
code not reportedweights not reporteddata availablein the paper
Data is provided within the manuscript or supplementary information files.where the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 10 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • ValidationNo validation of the AI is described.
  • Version of PSPredictorWhich version of the model was used is not stated.
  • Version of PONDR VSL2 (VSL2B output used)Which version of the model was used is not stated.
  • Version of PONDR VL3Which version of the model was used is not stated.
  • Version of PONDR VLXTWhich version of the model was used is not stated.
  • Version of PONDR FITWhich version of the model was used is not stated.
  • Version of IUPred (short and long)Which version of the model was used is not stated.

About this article

Record aix-00124, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error