~/aixsci
200 records · all checked

structural-biology/ai produced the result/Frontiers in Plant Science 2022 · v2

Machine learning sorts plant protein sequences into disease-resistance proteins or not

Researchers built StackRPred, a classifier that labels a plant protein sequence as a resistance protein or not. Six machine learning models feed a seventh, which makes the final call from features derived from a residue energy matrix.

1. Assemble plant R protein and non-R protein sequences2. Remove redundancy and split into training and test sets3. Encode sequences as residue pairwise energy features4. Select feature subset with SVM-RFE + CBR5. Train two-layer stacking classifier6. Evaluate by cross-validation and independent test

spectrum · one line per step, placed by what the step does · bright lines used AI

Prediction of Plant Resistance Proteins Based on Pairwise Energy Content and Stacking Framework
Frontiers in Plant Science, 2022

doi:10.3389/fpls.2022.912599 · record aix-00107 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Classification
Model family
Support vector machine, Random forest, Gradient-boosted trees
Checked by
Held-out92 tested
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Plants fight off disease with the help of so-called resistance proteins, or R proteins. These act rather like sentries inside the plant cell, recognising signs of an invading pathogen and setting off a defence response. Knowing which of a plant's many thousands of proteins are R proteins matters to anyone trying to breed hardier crops. But finding them is awkward. A protein is, at bottom, a long string of amino acids, and its job is not written plainly in that string. The usual approach is to compare a new sequence against known R proteins and look for family resemblance, which works less well when the resemblance is faint.

The authors set out to make that judgement from sequence alone, without relying on alignment to known examples. They gathered R proteins from a curated database alongside ordinary plant proteins from a public sequence archive, stripped out near-duplicates so that no two non-R sequences were more than 30 per cent similar, and split what remained into a training set and a separate test set of 92 sequences held back for the final check.

Where AI came in

The machine learning is not a tool used along the way here; it is the result. Each sequence was first turned into numbers by way of a published table of pairwise energies between amino acid types, treating the resulting grid as a kind of signal and summarising it with a wavelet transform, a standard way of describing a pattern at several scales at once. A further procedure, which repeatedly trains support vector machines and discards the least useful measurements, trimmed the description down to 112 numbers per protein.

Those numbers trained a two-layer arrangement called stacking. Six classifiers of different kinds sit in the first layer, among them random forests and gradient-boosted trees, and their verdicts become the input to a seventh, a support vector machine, which issues the final label. Settings were chosen by grid search. The result was checked by five-fold cross-validation on the training data and then on the held-back 92 sequences, where it reached an accuracy of 0.967 and an AUC of 0.997. Two earlier predictors were compared using figures from their own papers; one of them, prPred, reported higher precision.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

StackRPred classifies plant protein sequences as resistance (R) proteins or not. Each sequence was encoded through a residue pairwise energy content matrix, from which discrete wavelet transform and PseRECM features were extracted and then reduced to 112 dimensions by SVM-RFE + CBR. The features were used to train a two-layer stacking classifier whose base layer holds six classifiers (KNN, GBDT, SVM, XGBoost, LightGBM, RF) and whose meta-layer is an SVM. On the 92-sample independent test set the authors report accuracy 0.967, precision 0.980, recall 0.968, F1-score 0.980 and AUC 0.997; the compared prPred model is reported with accuracy 0.935 and a precision of 1, higher than StackRPred's precision.

How AI was used

Machine learning supplied the predictor itself. Protein sequences from a curated set of plant R proteins and non-R proteins, deduplicated with CD-HIT at 30% similarity and split 8:2 into training and independent test partitions, were converted into numeric features by treating a published 20x20 residue pairwise energy content matrix as the encoding, extracting five-level discrete wavelet transform statistics and discrete cosine coefficients from the resulting two-dimensional signal, and computing PseRECM descriptors analogous to PsePSSM. An SVM-RFE + CBR procedure, which repeatedly fits SVMs to rank and eliminate features, reduced the feature space to 112 dimensions. These features trained a stacking ensemble: a base layer of KNN, GBDT, SVM, XGBoost, LightGBM and random forest, whose outputs became the input to an SVM meta-classifier. Grid search tuned the SVM cost and RBF kernel parameters, tree counts for XGBoost and random forest, and n_estimators, max_depth and learning_rate for LightGBM; KNN and GBDT used default parameters. The model was assessed by five-fold cross-validation on the training set and by scoring the independent test partition, with metrics compared against values published for prPred and prPred-DRLF.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONREPRESENTATIONPREPARATIONTRAININGVALIDATION123456AIAIAIAssemble plant Rprotein and non-Rprotein sequencesRemove redundancyand split intotraining and tes…Encode sequencesas residuepairwise energy …Select featuresubset withSVM-RFE + CBRTrain two-layerstackingclassifierEvaluate bycross-validationand independent …↤ conventional algorithm
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Assemble plant R protein and non-R protein sequences

Obtaining raw data, whether by measurement, download or retrieval.

R proteins of 35 plant species were obtained from the PRGdb databasewhere the paper describes this · verbatim
in the paper
2Preparation
no AI

Remove redundancy and split into training and test sets

Cleaning, filtering, normalising or labelling data already obtained.

proteins with sequence similarity greater than 30% were excluded from the non-R protein dataset using CD-HITwhere the paper describes this · verbatim
in the paper
3Representation
no AI

Encode sequences as residue pairwise energy features

Encoding data into features, descriptors, embeddings or graphs.

extract the PsePSSM and DWT characteristics of each Plant R protein based on the RECM matrixwhere the paper describes this · verbatim
in the paper
4Preparation
AI

Select feature subset with SVM-RFE + CBR

Cleaning, filtering, normalising or labelling data already obtained.

we employ the SVM-RFE + CBR algorithm to select the best feature subsetwhere the paper describes this · verbatim
in the paper
5Training
AI

Train two-layer stacking classifier

Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.

the 112-dimensional feature information was fed into the constructed stacking model for trainingwhere the paper describes this · verbatim
in the paper
6Validation
AI

Evaluate by cross-validation and independent test

Testing outputs against ground truth.

To further compare the performance of our proposed method with other methods in independent testswhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's result is the trained classifier itself; the reported finding is the predictor's performance on plant R protein identification

+What the AI was for
Classificationin the paper
we selected six classification algorithms as the base classifier for the first layerwhere the paper describes this · verbatim
+How it was taught
Supervisedin the paper
+Models named
XGBoost · Trained from scratchLightGBM · Trained from scratchGradient Boosting Decision Tree (GBDT) · Trained from scratchRandom Forest · Trained from scratchK-Nearest Neighbor · Trained from scratchSupport Vector Machine (RBF kernel; base classifier and meta-classifier) · Trained from scratchin the paper
+How results were checked
Held-out92 testedin the paper
the number of training samples is 364, and the number of independent test samples is 92where the paper describes this · verbatim
−Code · weights · data
code not reportedweights not reporteddata not reportednot reported
−Compute
not reportednot reported

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 12 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • DataWhether the data are available is not stated.
  • ComputeThe hardware or time used is not stated.
  • Version of XGBoostWhich version of the model was used is not stated.
  • Version of LightGBMWhich version of the model was used is not stated.
  • Version of Gradient Boosting Decision Tree (GBDT)Which version of the model was used is not stated.
  • Version of Random ForestWhich version of the model was used is not stated.
  • Version of K-Nearest NeighborWhich version of the model was used is not stated.
  • Version of Support Vector Machine (RBF kernel; base classifier and meta-classifier)Which version of the model was used is not stated.
  • What step 4 replacedThe paper gives no basis for what the AI stood in for.
  • What step 6 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00107, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error