~/aixsci
200 records · all checked

structural-biology/ai produced the result/Scientific Reports 2022 · v2

Microscopy and a trained classifier map which E. coli proteins clump together

Researchers overexpressed 2577 E. coli proteins tagged with a fluorescent marker and imaged the cells. A supervised neural network sorted each cell's glow pattern into three categories, which set each protein's aggregation class.

1. Overexpress GFP-fusion protein library2. Acquire high-content microscopy images3. Pre-process images and extract single-cell features4. Classify single-cell fluorescence phenotypes5. Assign each protein an aggregation class6. Validate classification by western blot and external dataset overlap7. Compute sequence- and structure-derived protein features8. Test features for discrimination between aggregation classes

spectrum · one line per step, placed by what the step does · bright lines used AI

Proteome-wide landscape of solubility limits in a bacterial cell
Scientific Reports, 2022

doi:10.1038/s41598-022-10427-1 · record aix-00157 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Classification, Property prediction
Model family
Multilayer perceptron, Linear model
Checked by
Experimental46 tested
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Proteins are chains of amino acids that must fold into a particular shape to work. If a cell makes too much of one, or if the chain folds badly, copies can stick to each other and fall out of solution as clumps, called aggregates. The amount of a protein a cell can hold before this happens is its solubility limit. Measuring that limit across a whole proteome is awkward: there are thousands of different proteins, each needs to be made in quantity and then checked, and clumping inside a living cell is hard to see directly. Tagging each protein with green fluorescent protein, a jellyfish protein that glows, gives a visible readout of where the copies end up.

The authors took a library of E. coli genes, each fused to that glowing tag, pushed the cells to make large amounts of one protein at a time, and photographed them under an automated microscope. A clump often shows up as a bright spot at one end of the cell, while a soluble protein spreads evenly. Some clumps do not glow at all. The task was to sort thousands of strains by which pattern dominated, then ask which properties of a protein's sequence and predicted structure go with each outcome.

Where AI came in

Standard image software first corrected uneven lighting, outlined individual cells and measured each one's brightness, texture and shape. A supervised artificial neural network, run through the Advanced Cell Classifier software built on Weka, then read those per-cell measurements and assigned each cell to one of three patterns the researchers had defined by hand: no glow, even glow, or a concentrated spot at the cell's pole. Cells fitting none of the three were discarded. The network took the place of a person looking at each cell and judging it, at a scale of roughly a thousand cells per strain. A simple majority rule, not learning, then turned those cell labels into one class per protein: soluble, dark aggregate or fluorescent aggregate.

Software also supplied the protein properties used afterwards. Sequence- and structure-based predictors, including PONDR VSL2B, Espritz, DisEMBL and FOLD-RATE, estimated things like how much of a protein is disordered rather than neatly folded, and how fast it folds, yielding 115 features in all. Logistic regression, a simple statistical model, was then fitted to the study's own labels to see which of those features separated the classes, with tenfold cross-validation used to score each one. Separately, western blots on 46 overexpressions checked the image-based assignments against a direct measurement of what was soluble.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors overexpressed 2577 cytosolic E. coli proteins as C-terminal GFP fusions and imaged them by high-content microscopy, then used a supervised neural-network classifier to assign the fluorescence phenotype of about 1000 cells per strain into three categories: no signal, homogeneous signal, or polar fluorescent foci. Each protein was then labelled soluble, dark aggregate or fluorescent aggregate according to its predominant cellular phenotype, giving 37.8% soluble, 18.6% dark aggregates and 43.6% fluorescent aggregates. Logistic regressions over 115 precomputed protein features, for 1631 cytoplasmic monomers, identified contact order as the most predictive feature separating dark from fluorescent aggregates, and linked high disorder content and low surface stickiness to solubility. Western blot analysis of selected overexpressions was used to check the image-based assignments.

How AI was used

Microscopy images were first corrected for uneven illumination and segmented with conventional thresholding and watershed steps in CellProfiler, which extracted per-cell intensity, texture and morphology features. A supervised artificial neural network, run through the Advanced Cell Classifier software on top of Weka and trained on the three predefined fluorescence phenotypes, then labelled each segmented cell; cells not fitting the three categories were discarded. A non-learned majority rule converted the per-cell labels into one aggregation class per overexpressed protein. Separately, learned and non-learned sequence- and structure-based predictors (including PONDR VSL2B, Espritz, DisEMBL, IUpred, FOLD-RATE and AggreScan) were run over protein sequences and retrieved predicted structures to compute disorder content, folding rate and related features, and logistic regression models were fitted to the study's own labels, with tenfold cross-validated AUC used to measure each feature's discriminative power.

The shape of the work

Structural · the record, drawn

EXPERIMENTACQUISITIONPREPARATIONINFERENCEPREPARATIONVALIDATIONREPRESENTATIONTRAINING12345678AIAIAIOverexpressGFP-fusionprotein libraryAcquirehigh-contentmicroscopy imagesPre-processimages andextract single-c…Classifysingle-cellfluorescence phe…Assign eachprotein anaggregation classValidateclassification bywestern blot and…Compute sequence-andstructure-derive…Test features fordiscriminationbetween aggregat…↤ manual curation
AI stepNo AI↤ what the AI stood in for
1Experiment
no AI

Overexpress GFP-fusion protein library

Physical execution, by hand or by robot.

the C-terminal GFP fusion version of the E. coli K-12 Open Reading Frame Archive library was grown in the original host strainwhere the paper describes this · verbatim
in the paper
2Acquisition
no AI

Acquire high-content microscopy images

Obtaining raw data, whether by measurement, download or retrieval.

Images of two channels (DAPI and GFP) were collected using a 60 × high-NA objectivewhere the paper describes this · verbatim
in the paper
3Preparation
no AI

Pre-process images and extract single-cell features

Cleaning, filtering, normalising or labelling data already obtained.

cells were identified on the DAPI signal using Otsu adaptive threshold and a Watershed algorithm to split touching cellswhere the paper describes this · verbatim
in the paper
4Inference
AI

Classify single-cell fluorescence phenotypes

Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.

Supervised classification of cells into predefined groups was performed using the Advanced Cell Classifier softwarewhere the paper describes this · verbatim
in the paper
5Preparation
no AI

Assign each protein an aggregation class

Cleaning, filtering, normalising or labelling data already obtained.

the proteins were assigned to one of the three classes, depending on which phenotype was predominant in the cell populationwhere the paper describes this · verbatim
in the paper
6Validation
no AI

Validate classification by western blot and external dataset overlap

Testing outputs against ground truth.

100% and 88.9% of proteins classified as fluorescent and dark aggregates, respectively, are indeed predominantly present in their aggregated formswhere the paper describes this · verbatim
in the paper
7Representation
AI

Compute sequence- and structure-derived protein features

Encoding data into features, descriptors, embeddings or graphs.

Protein disorder data was retrieved from the MobiDB database and calculated with several different predictorswhere the paper describes this · verbatim
in the paper
8Training
AI

Test features for discrimination between aggregation classes

Fitting model parameters, including fine-tuning an existing model.

We used logistic regressions to test the statistical association between aggregation class and each protein feature individuallywhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The proteome-wide aggregation classification that the paper reports is produced by the supervised single-cell phenotype classifier; the downstream feature analysis rests on those labels

+What the AI was for
an artificial neural network method was used based on the Weka softwarewhere the paper describes this · verbatim
+Model families
+How it was taught
Supervisedin the paper
+Models named
Weka artificial neural network (via Advanced Cell Classifier) · Trained from scratchLogistic regression models of protein features · Trained from scratchPONDR VSL2B · Off the shelfEspritz · Off the shelfDisEMBL · Off the shelfFOLD-RATE · Off the shelfin the paper
+How results were checked
Experimental46 testedin the paper
we next carried out western blot analysis after separating the soluble and insoluble fractions of 46 protein overexpressionswhere the paper describes this · verbatim
+Code · weights · data
code not reportedweights not reporteddata availablein the paper
The raw microscopy data can be accessed atwhere the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 11 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • Version of Weka artificial neural network (via Advanced Cell Classifier)Which version of the model was used is not stated.
  • Version of Logistic regression models of protein featuresWhich version of the model was used is not stated.
  • Version of PONDR VSL2BWhich version of the model was used is not stated.
  • Version of EspritzWhich version of the model was used is not stated.
  • Version of DisEMBLWhich version of the model was used is not stated.
  • Version of FOLD-RATEWhich version of the model was used is not stated.
  • What step 7 replacedThe paper gives no basis for what the AI stood in for.
  • What step 8 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00157, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error