~/aixsci
200 records · all checked

structural-biology/ai produced the result/arXiv 2025 · v2

Phage proteins drawn as pictures, then sorted by image-recognition networks

Researchers turned viral protein sequences into colour-coded images and fine-tuned three off-the-shelf image-recognition networks to say which proteins form part of a virus's physical shell, then measured how confident those networks were.

1. Curate PVP and non-PVP sequence dataset2. Encode sequences as images with Knight Encoding3. Fine-tune pre-trained CNNs on encoded images4. Classify held-out encoded sequences5. Score metrics and compare with published tools6. Quantify prediction uncertainty by Monte Carlo Dropout

spectrum · one line per step, placed by what the step does · bright lines used AI

ProteoKnight: Convolution-based Phage Virion Protein Classification and Uncertainty Analysis
arXiv, 2025

doi:10.48550/arxiv.2508.07345 · record aix-00128 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Classification
Model family
Convolutional neural network
Checked by
Benchmark
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Diagram showing the structural constituents of a bacteriophage.
Illustration of the structural constituents of a phage.Figure 1 from Neha et al., arXiv 2025 · source · CC BY · resized

Bacteriophages are viruses that infect bacteria. Like other viruses, each one is built from protein parts. Some of those proteins make up the physical body of the virus particle: the outer shell, the tail, the fibres it uses to grip a host cell. These are called virion proteins. Others do jobs inside the infected cell and never become part of the particle itself. Telling the two groups apart matters for understanding how a phage attaches to and enters its host. The traditional way to find out is laboratory work, such as mass spectrometry or protein arrays, which weigh and probe the actual molecules. That is slow, and genome sequencing now produces far more phage protein sequences than anyone can test by hand.

A protein sequence is just a long string of letters, each standing for one of twenty amino acids. The researchers set out to classify such strings without laboratory work, by first turning each one into a picture. Their scheme, which they call Knight Encoding, treats the twenty amino acids as twenty directions spaced evenly around a circle, eighteen degrees apart. Starting from a point, it steps a fixed distance in the direction given by each amino acid in turn and marks a coloured dot, returning to the centre whenever the walk reaches the edge. The result is an image whose pattern reflects the order of the sequence.

Where AI came in

The classifying was done by convolutional neural networks, the kind of software used for ordinary image recognition. Three of them were taken ready-made from a public library, already trained on general images: GoogLeNet, EfficientNet v2 and MobileNet. Each was then fine-tuned, meaning further trained, on the encoded protein pictures for twenty-five rounds, using a public dataset of 35,213 virion and 46,883 non-virion sequences. On held-out sequences the GoogLeNet-based version reached about 89 per cent accuracy, with sensitivity of 91.1 per cent and specificity of 92.7 per cent; EfficientNet V2 reached about 91.2 per cent. Sorting virion proteins into their eight finer groups was less accurate, between about 72 and 78 per cent.

The networks stand in for the laboratory step: instead of measuring a protein to learn what it does, the models assign a label from its sequence alone, and the scores were tabulated against existing software tools for the same task. The researchers also probed how sure the models were, using a technique called Monte Carlo Dropout. This switches off random parts of the network at the moment of prediction and repeats the prediction many times, so the spread of answers shows how stable it is. The spread was narrower for non-virion proteins than for virion ones, and narrower for shorter sequences than longer ones.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

Diagram showing the structural constituents of a bacteriophage.
Illustration of the structural constituents of a phage.Figure 1 from Neha et al., arXiv 2025 · source · CC BY · resized

The study introduces Knight Encoding, a walk-based scheme that turns phage protein sequences into colour-coded images, and classifies those images with pre-trained convolutional networks fine-tuned on a public phage protein dataset of 35,213 PVP and 46,883 non-PVP sequences. On the test set, the fine-tuned GoogLeNet reached about 89% accuracy with sensitivity 91.1% and specificity 92.7%, EfficientNet V2 reached about 91.2% accuracy, and multi-class accuracy fell in a 72% to 78% range. Monte Carlo Dropout over the GoogLeNet model showed lower prediction variance and entropy for non-PVP sequences than PVP sequences, and for shorter sequences than longer ones.

How AI was used

Protein sequences from a curated public phage dataset were converted into M x M images by a rule-based polar-coordinate walk, in which each of the 20 amino acids is assigned an 18-degree angle on an icosagon and drawn as a coloured point at a fixed radius from the previous residue's position, restarting from the image centre on reaching a boundary. These images were used to fine-tune several pre-trained CNN architectures taken from the PyTorch library, including GoogLeNet, EfficientNet v2 and MobileNet, for 25 epochs with batch size 32 under binary cross-entropy loss, for both binary PVP/non-PVP and multi-class PVP classification. The fine-tuned models were then run over held-out encoded images, and standard metrics were computed and tabulated against previously published PVP tools. For uncertainty quantification, the selected GoogLeNet model was run with dropout active at test time, with 100 randomly chosen sequences from each of four categories, defined by class and by a sequence-length split, each predicted 100 times; the variance and entropy of the resulting softmax distributions were compared across categories, and the pattern was checked at dropout rates of 0.1, 0.2 and 0.3.

The shape of the work

Structural · the record, drawn

PREPARATIONREPRESENTATIONTRAININGINFERENCEVALIDATIONINTERPRETATION123456AIAIAICurate PVP andnon-PVP sequencedatasetEncode sequencesas images withKnight EncodingFine-tunepre-trained CNNson encoded imagesClassify held-outencoded sequencesScore metrics andcompare withpublished toolsQuantifypredictionuncertainty by M…↤ physical experiment
AI stepNo AI↤ what the AI stood in for
1Preparation
no AI

Curate PVP and non-PVP sequence dataset

Cleaning, filtering, normalising or labelling data already obtained.

The CD-hit algorithm was then applied to identify clusters of sequences with a similarity threshold of 90%.where the paper describes this · verbatim
in the paper
2Representation
no AI

Encode sequences as images with Knight Encoding

Encoding data into features, descriptors, embeddings or graphs.

we introduce a novel walk-based encoding technique specifically for protein sequenceswhere the paper describes this · verbatim
in the paper
3Training
AI

Fine-tune pre-trained CNNs on encoded images

Fitting model parameters, including fine-tuning an existing model.

each was fine-tuned on the encoded image dataset for 25 epochs, with a batch size of 32where the paper describes this · verbatim
in the paper
4Inference
AI

Classify held-out encoded sequences

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

Evaluating on the test set, our GoogLeNet-based method achieved an accuracy of approximately 89%where the paper describes this · verbatim
in the paper
5Validation
no AI

Score metrics and compare with published tools

Testing outputs against ground truth.

Table 5 provides a comprehensive comparison of the results of our GoogLeNet-based approach with existing state-of-the-art toolswhere the paper describes this · verbatim
in the paper
6Interpretation
AI

Quantify prediction uncertainty by Monte Carlo Dropout

Extracting understanding from model behaviour.

we activate the model’s dropout layer with a 0.2% dropout rate during the test phasewhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's reported result is the classification performance of fine-tuned CNNs on the proposed image encoding, plus an uncertainty analysis of those same model predictions; without the models there is no finding.

+What the AI was for
Classificationin the paper
each was fine-tuned on the encoded image dataset for 25 epochswhere the paper describes this · verbatim
+Model families
+How it was taught
SupervisedTransfer / fine-tuningin the paper
+Models named
GoogLeNet · Fine-tunedEfficientNet v2 v2 · Fine-tunedMobileNet · Fine-tunedin the paper
+How results were checked
Benchmarkin the paper
This study utilizes the benchmark dataset established by Shang et al.where the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
publicly available on GitHub at https://github.com/eniac00/ProteoKnightwhere the paper describes this · verbatim
+Compute
AMD Ryzen 7 3700X 8-core CPU, 16GB DDR4 RAM, NVIDIA GeForce RTX 3060 Ti GPU; Windows 11 Pro, Python 3.11in the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 6 items
  • Trained model weightsWhether the trained model is available is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of GoogLeNetWhich version of the model was used is not stated.
  • Version of MobileNetWhich version of the model was used is not stated.
  • What step 3 replacedThe paper gives no basis for what the AI stood in for.
  • What step 6 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00128, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error