~/aixsci
200 records · all checked

structural-biology/ai produced the result/Nature Communications 2021 · v2

Machine learning picked which engineered enzymes to build for fatty alcohol production

Researchers stitched three bacterial enzymes into a library of 4,374 hybrids. Gaussian process models chose which hybrids to build and test across ten rounds, ending with a variant that made 54 ± 11 mg/L of fatty alcohols in E. coli.

1. Design chimeric ATR library by SCHEMA recombination2. Select informative seed set by greedy entropy maximisation3. Assemble genes and measure in vivo fatty alcohol titers4. Train activity classifier and titer regression model5. Predict untested chimeras and pick next batch by UCB6. Characterise expression and in vitro kinetics of top enzymes7. Model block contributions to activity8. Dock ACP and compute interface net charge

spectrum · one line per step, placed by what the step does · bright lines used AI

Machine learning-guided acyl-ACP reductase engineering for improved in vivo fatty alcohol production
Nature Communications, 2021

doi:10.1038/s41467-021-25831-w · record aix-00024 v2 · checked 2026-10-07

ai-resultrole of AI
AI was for
Property prediction, Classification, Experimental design
Model family
Gaussian process, Probabilistic graphical model
Checked by
Experimental
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Some bacteria carry enzymes that turn fatty acid building blocks into fatty alcohols, the kind of molecule used in detergents and cosmetics. Engineers would like E. coli to make them instead. One way to improve an enzyme is recombination: take several related natural versions, chop each into blocks, and swap the blocks between them to make hybrids, or chimeras. The trouble is arithmetic. Eight blocks drawn from three parents give thousands of possible combinations, and here the library held 4,374 of them. Measuring what each one produces means growing cells and running their contents through a gas chromatograph, a slow instrument that separates a mixture into its components one sample at a time.

So the full library could not be tested. The researchers set out to find high-producing chimeras by building and assaying only a small fraction of them, letting a statistical model decide which fraction that should be.

Where AI came in

The learning sat in the choosing. Twenty starting sequences were picked to be as informative as possible about the whole library, then built and measured. Their titres trained two models: a Gaussian Naïve Bayes classifier, which sorts sequences into active and inactive, and Gaussian process regression, which predicts a number and also how uncertain that prediction is. Both were then applied to every untested chimera. Predicted duds were dropped, and a rule favouring sequences that were either predicted high or still poorly understood selected ten to twelve to build next. The new measurements retrained the models, and the cycle ran ten times.

In place of testing everything, the models stood in for the untested majority of the library. A Gaussian process model was later fitted to all the collected data to estimate how much each sequence block contributed to output. The enzyme purification, the kinetic measurements and the docking calculations that followed did not involve learning.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

Three natural alcohol-forming fatty acyl reductases were recombined by SCHEMA into a library of 4374 chimeric acyl-thioester reductase domains, and a Gaussian process model with an upper-confidence-bound criterion chose which chimeras to build and assay over ten design-test-learn rounds, beginning from 20 seed sequences. The top chimera, ATR-83, produced a total titer of 54 ± 11 mg/L fatty alcohols in E. coli, 4.9-fold the titer of MA-ACR and about twofold that of the best natural sequence, MB-ACR. In vitro kinetics on palmitoyl-ACP showed a larger turnover number for ATR-83 than for parent B, with similar KM and no significant differences in enzyme expression. A Gaussian process model fitted to the sequence-function data was used to estimate each block's contribution, and net charge near the docked ACP interface correlated with titer.

How AI was used

Machine learning was used to search a combinatorial chimera space that could not be assayed exhaustively with the low-throughput gas chromatography readout. A greedy Gaussian-entropy criterion first picked 20 seed sequences that maximised mutual information with the full 4374-sequence library; these were assembled by Golden Gate, expressed in E. coli and measured. A Gaussian Naïve Bayes classifier on one-hot encoded sequences was then trained to separate active from inactive variants, and Gaussian process regression with a linear kernel over Hamming or contact-pair encodings was trained on the active sequences' titers, with the variance hyperparameter chosen by leave-one-out cross-validation. In each of ten rounds, both models were applied to all untested chimeras, inactive-predicted sequences were excluded, and a batch-mode upper-confidence-bound rule (mean plus one standard deviation, with the chosen sequence's predicted titer fed back as pseudo-data before reselecting) picked 10–12 sequences to construct and assay; the new titers retrained the models for the following round. After optimisation, a Gaussian process regression model fitted to the collected sequence-function data was used to estimate the contribution of each sequence block, alongside non-learned RosettaDock docking and interface charge calculation.

The shape of the work

Structural · the record, drawn

GENERATIONSCREENINGEXPERIMENTTRAININGOPTIMISATIONVALIDATIONINTERPRETATIONSIMULATION12345678AIAIAIDesign chimericATR library bySCHEMA recombina…Selectinformative seedset by greedy en…Assemble genesand measure invivo fatty alcoh…Train activityclassifier andtiter regression…Predict untestedchimeras and picknext batch by UCBCharacteriseexpression and invitro kinetics o…Model blockcontributions toactivityDock ACP andcompute interfacenet charge↤ exhaustive searchloops back · 10 rounds
AI stepNo AI↤ what the AI stood in for
1Generation
no AI

Design chimeric ATR library by SCHEMA recombination

Producing candidate objects that did not previously exist.

we used SCHEMA-RASPP to determine 7 additional crossover locations within the ATR domainwhere the paper describes this · verbatim
in the paper
2Screening
no AI

Select informative seed set by greedy entropy maximisation

Reducing a candidate set by filtering or ranking, in a single pass.

We sought to identify the set of 20 chimera sequences that is maximally informative of the full chimera landscape.where the paper describes this · verbatim
in the paper
3Experiment
no AI

Assemble genes and measure in vivo fatty alcohol titers

Physical execution, by hand or by robot.

We then constructed these sequences and experimentally measured their fatty alcohol titers in three E. coli strainswhere the paper describes this · verbatim
in the paper
4Training
AI

Train activity classifier and titer regression model

Fitting model parameters, including fine-tuning an existing model.

The fatty alcohol titer data from these 20 initial sequences was used to train Gaussian process (GP) sequence-function modelswhere the paper describes this · verbatim
in the paper
5Optimisation
AI

Predict untested chimeras and pick next batch by UCB

Iterative search over a space. The AI stood in for exhaustive search. Its result feeds back into an earlier step.

We then applied the GNB and GP models to make functional predictions over all untested chimeras.where the paper describes this · verbatim
in the paper
6Validation
no AI

Characterise expression and in vitro kinetics of top enzymes

Testing outputs against ground truth.

Next, we purified the enzymes and measured their kinetic properties on palmitoyl-ACPwhere the paper describes this · verbatim
in the paper
7Interpretation
AI

Model block contributions to activity

Extracting understanding from model behaviour.

We trained a GP regression model to predict fatty alcohol titers from sequence.where the paper describes this · verbatim
in the paper
8Simulation
no AI

Dock ACP and compute interface net charge

Numerical or physics simulation, including where a learned surrogate replaces it.

We ran 1000 docking simulations and selected a model based on minimizing the total energy and the interface score.where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The improved enzymes the paper reports are the sequences chosen by the GP/UCB search; every tested variant after the seed round was selected by the model.

+What the AI was for
a Gaussian Naïve Bayes (GNB) classifier to distinguish inactive versus active sequences and Gaussian process (GP) regression to model a sequence’s fatty alcohol titerwhere the paper describes this · verbatim
+How it was taught
SupervisedActive learningin the paper
+Models named
Gaussian process regression (linear/Hamming and structure kernels) · Trained from scratchGaussian Naïve Bayes active/inactive classifier (scikit-learn) · Trained from scratchin the paper
+How results were checked
Experimentalin the paper
measured each strain’s fatty alcohol titer using gas chromatographywhere the paper describes this · verbatim
+Code · weights · data
code not reportedweights not reporteddata availablein the paper
Source data underlying Fig. 2c are provided as a Source Data file.where the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 8 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of Gaussian process regression (linear/Hamming and structure kernels)Which version of the model was used is not stated.
  • Version of Gaussian Naïve Bayes active/inactive classifier (scikit-learn)Which version of the model was used is not stated.
  • What step 4 replacedThe paper gives no basis for what the AI stood in for.
  • What step 7 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00024, version 2, checked by a person on 2026-10-07. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY-4.0; quotations are at most 25 words. How we work · Report an error