structural-biology/ai produced the result/Scientific Reports 2022 · v2
Microscopy and a trained classifier map which E. coli proteins clump together
Researchers overexpressed 2577 E. coli proteins tagged with a fluorescent marker and imaged the cells. A supervised neural network sorted each cell's glow pattern into three categories, which set each protein's aggregation class.
spectrum · one line per step, placed by what the step does · bright lines used AI
Proteome-wide landscape of solubility limits in a bacterial cell
Scientific Reports, 2022
doi:10.1038/s41598-022-10427-1 · record aix-00157 v2 · checked 2026-10-09
- AI was for
- Classification, Property prediction
- Model family
- Multilayer perceptron, Linear model
- Checked by
- Experimental46 tested
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
Proteins are chains of amino acids that must fold into a particular shape to work. If a cell makes too much of one, or if the chain folds badly, copies can stick to each other and fall out of solution as clumps, called aggregates. The amount of a protein a cell can hold before this happens is its solubility limit. Measuring that limit across a whole proteome is awkward: there are thousands of different proteins, each needs to be made in quantity and then checked, and clumping inside a living cell is hard to see directly. Tagging each protein with green fluorescent protein, a jellyfish protein that glows, gives a visible readout of where the copies end up.
The authors took a library of E. coli genes, each fused to that glowing tag, pushed the cells to make large amounts of one protein at a time, and photographed them under an automated microscope. A clump often shows up as a bright spot at one end of the cell, while a soluble protein spreads evenly. Some clumps do not glow at all. The task was to sort thousands of strains by which pattern dominated, then ask which properties of a protein's sequence and predicted structure go with each outcome.
Where AI came in
Standard image software first corrected uneven lighting, outlined individual cells and measured each one's brightness, texture and shape. A supervised artificial neural network, run through the Advanced Cell Classifier software built on Weka, then read those per-cell measurements and assigned each cell to one of three patterns the researchers had defined by hand: no glow, even glow, or a concentrated spot at the cell's pole. Cells fitting none of the three were discarded. The network took the place of a person looking at each cell and judging it, at a scale of roughly a thousand cells per strain. A simple majority rule, not learning, then turned those cell labels into one class per protein: soluble, dark aggregate or fluorescent aggregate.
Software also supplied the protein properties used afterwards. Sequence- and structure-based predictors, including PONDR VSL2B, Espritz, DisEMBL and FOLD-RATE, estimated things like how much of a protein is disordered rather than neatly folded, and how fast it folds, yielding 115 features in all. Logistic regression, a simple statistical model, was then fitted to the study's own labels to see which of those features separated the classes, with tenfold cross-validation used to score each one. Separately, western blots on 46 overexpressions checked the image-based assignments against a direct measurement of what was soluble.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors overexpressed 2577 cytosolic E. coli proteins as C-terminal GFP fusions and imaged them by high-content microscopy, then used a supervised neural-network classifier to assign the fluorescence phenotype of about 1000 cells per strain into three categories: no signal, homogeneous signal, or polar fluorescent foci. Each protein was then labelled soluble, dark aggregate or fluorescent aggregate according to its predominant cellular phenotype, giving 37.8% soluble, 18.6% dark aggregates and 43.6% fluorescent aggregates. Logistic regressions over 115 precomputed protein features, for 1631 cytoplasmic monomers, identified contact order as the most predictive feature separating dark from fluorescent aggregates, and linked high disorder content and low surface stickiness to solubility. Western blot analysis of selected overexpressions was used to check the image-based assignments.
How AI was used
Microscopy images were first corrected for uneven illumination and segmented with conventional thresholding and watershed steps in CellProfiler, which extracted per-cell intensity, texture and morphology features. A supervised artificial neural network, run through the Advanced Cell Classifier software on top of Weka and trained on the three predefined fluorescence phenotypes, then labelled each segmented cell; cells not fitting the three categories were discarded. A non-learned majority rule converted the per-cell labels into one aggregation class per overexpressed protein. Separately, learned and non-learned sequence- and structure-based predictors (including PONDR VSL2B, Espritz, DisEMBL, IUpred, FOLD-RATE and AggreScan) were run over protein sequences and retrieved predicted structures to compute disorder content, folding rate and related features, and logistic regression models were fitted to the study's own labels, with tenfold cross-validated AUC used to measure each feature's discriminative power.
The shape of the work
Structural · the record, drawn
no AI
Overexpress GFP-fusion protein library
Physical execution, by hand or by robot.
the C-terminal GFP fusion version of the E. coli K-12 Open Reading Frame Archive library was grown in the original host strainwhere the paper describes this · verbatim
no AI
Acquire high-content microscopy images
Obtaining raw data, whether by measurement, download or retrieval.
Images of two channels (DAPI and GFP) were collected using a 60 × high-NA objectivewhere the paper describes this · verbatim
no AI
Pre-process images and extract single-cell features
Cleaning, filtering, normalising or labelling data already obtained.
cells were identified on the DAPI signal using Otsu adaptive threshold and a Watershed algorithm to split touching cellswhere the paper describes this · verbatim
AI
Classify single-cell fluorescence phenotypes
Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.
Supervised classification of cells into predefined groups was performed using the Advanced Cell Classifier softwarewhere the paper describes this · verbatim
no AI
Assign each protein an aggregation class
Cleaning, filtering, normalising or labelling data already obtained.
the proteins were assigned to one of the three classes, depending on which phenotype was predominant in the cell populationwhere the paper describes this · verbatim
no AI
Validate classification by western blot and external dataset overlap
Testing outputs against ground truth.
100% and 88.9% of proteins classified as fluorescent and dark aggregates, respectively, are indeed predominantly present in their aggregated formswhere the paper describes this · verbatim
AI
Compute sequence- and structure-derived protein features
Encoding data into features, descriptors, embeddings or graphs.
Protein disorder data was retrieved from the MobiDB database and calculated with several different predictorswhere the paper describes this · verbatim
AI
Test features for discrimination between aggregation classes
Fitting model parameters, including fine-tuning an existing model.
We used logistic regressions to test the statistical association between aggregation class and each protein feature individuallywhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The proteome-wide aggregation classification that the paper reports is produced by the supervised single-cell phenotype classifier; the downstream feature analysis rests on those labels
an artificial neural network method was used based on the Weka softwarewhere the paper describes this · verbatim
we next carried out western blot analysis after separating the soluble and insoluble fractions of 46 protein overexpressionswhere the paper describes this · verbatim
The raw microscopy data can be accessed atwhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- Version of Weka artificial neural network (via Advanced Cell Classifier)Which version of the model was used is not stated.
- Version of Logistic regression models of protein featuresWhich version of the model was used is not stated.
- Version of PONDR VSL2BWhich version of the model was used is not stated.
- Version of EspritzWhich version of the model was used is not stated.
- Version of DisEMBLWhich version of the model was used is not stated.
- Version of FOLD-RATEWhich version of the model was used is not stated.
- What step 7 replacedThe paper gives no basis for what the AI stood in for.
- What step 8 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00157, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error