~/aixsci
200 records · all checked

structural-biology/ai produced the result/Bioinformatics Advances 2024 · v2

Changing the training penalty improves how a network spots beta-sheets in blurry protein maps

Researchers trained a 3D U-Net to label every point in a medium-resolution cryo-electron-microscopy map as helix, beta-sheet or background, and compared five ways of penalising its mistakes during training. The combination of focal and Dice loss scored highest for beta-sheets.

1. Collect medium-resolution map and structure pairs2. Remove redundancy and screen pair quality3. Box-crop maps and label voxels by DSSP assignment4. Train one U-Net per loss function and pick hyperparameters5. Segment held-out test maps6. Score voxel- and residue-level F1 against atomic structures

spectrum · one line per step, placed by what the step does · bright lines used AI

The combined focal loss and dice loss function improves the segmentation of beta-sheets in medium-resolution cryo-electron-microscopy density maps
Bioinformatics Advances, 2024

doi:10.1093/bioadv/vbae169 · record aix-00079 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Segmentation
Model family
Convolutional neural network
Checked by
Held-out62 tested
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Cryo-electron microscopy freezes copies of a protein and images them, producing a three-dimensional map of how densely matter sits at each point in space. At high resolution, individual atoms can be placed in such a map. At medium resolution the picture is blurrier, and researchers must instead recognise shapes. Proteins fold into a few recurring local forms, chiefly the corkscrew-shaped alpha helix and the beta-sheet, where strands lie side by side in a flat or twisted layer. Helices show up as fat rods and are comparatively easy to pick out. Sheets are thinner and their strands sit close together, so in a blurry map they are harder to tell from surrounding density.

The researchers set out to label each small cube of a map, or voxel, as helix, beta-sheet or neither. Their question was narrower than the architecture: given the same network and the same data, how much does the choice of training penalty, the loss function, change what gets found? Because sheet voxels are far outnumbered by background, a penalty that rewards overall accuracy can quietly ignore them.

Where AI came in

The work rests on a convolutional neural network, a 3D U-Net, trained from scratch to do the labelling. Maps deposited in the public EMDB archive were paired with the matching atomic structures from the PDB, filtered for redundancy and quality, and cropped into boxes around a single protein chain. Voxels near a backbone atom took the helix or sheet label assigned to that residue by the standard DSSP method; the rest became background. Those structure-derived labels were the teaching signal, so the network learned to infer from density alone what would otherwise require an atomic model.

Five networks were trained, identical but for the loss function: cross-entropy, focal loss, Dice loss, and the pairings of cross-entropy with Dice and focal with Dice. Each was then run over 62 held-out test maps it had never seen, and its voxel labels compared with the structure-derived ones using F1 scores, a measure balancing missed features against false ones. Focal plus Dice reached the highest weighted-average F1 for beta-sheet voxels, 48.7% against 39.9% for cross-entropy, while helix scores stayed close across all five. That model was packaged as DeepSSETracer 1.1, which runs inside the ChimeraX viewer.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The study trained a 3D U-Net segmentation network (the DeepSSETracer architecture) to label each voxel of a medium-resolution cryo-EM density map as helix, β-sheet or background, and compared five training loss functions: cross-entropy, focal loss, Dice loss, and the two combinations cross-entropy plus Dice and focal plus Dice. A dataset of 1355 box-cropped map/structure pairs was split into 1246 training, 47 validation and 62 test cases. On the 62-case test set, the focal-plus-Dice loss gave the highest weighted-average F1 score for β-sheet voxels (48.7%, against 39.9% for cross-entropy), while helix F1 scores were close across the five losses (62.9% for focal loss and 62.0% for focal plus Dice). The model trained with focal plus Dice was packaged as DeepSSETracer 1.1, which runs inside ChimeraX.

How AI was used

Cryo-EM maps from EMDB were paired with PDB atomic structures, filtered for redundancy and match quality, and cropped into rectangular subregions around a central chain with 34 Å of padding; voxels within 3 Å of a Cα atom were labelled helix or β-sheet according to the DSSP assignment of that residue, and all remaining voxels background. A five-layer 3D U-Net convolutional network was then trained from scratch on the training split, once per loss function, with three learning rates and four focal γ values explored and the best setting for each loss chosen on the validation split; the loss was computed with the padding excluded. Each trained network was run over the held-out box-cropped test maps to assign one of the three classes to every voxel, and the predicted labels in the unpadded centre box were compared with the structure-derived labels to compute weighted voxel-level F1 scores, with residue-level scores obtained by voting over the voxels within 3 Å of each Cα atom. Four further random splits, constrained by central-chain sequence identity, were used to repeat the comparison.

The shape of the work

Structural · the record, drawn

ACQUISITIONSCREENINGPREPARATIONTRAININGINFERENCEVALIDATION123456AIAICollectmedium-resolutionmap and structur…Remove redundancyand screen pairqualityBox-crop maps andlabel voxels byDSSP assignmentTrain one U-Netper loss functionand pick hyperpa…Segment held-outtest mapsScore voxel- andresidue-level F1against atomic s…↤ conventional algorithm↤ conventional algorithm
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Collect medium-resolution map and structure pairs

Obtaining raw data, whether by measurement, download or retrieval.

we developed a method of selecting protein chains from EMDB-deposited medium-resolution cryo-EM mapswhere the paper describes this · verbatim
in the paper
2Screening
no AI

Remove redundancy and screen pair quality

Reducing a candidate set by filtering or ranking, in a single pass.

chains that shared over 70% sequence identity in the same PDB entry were removedwhere the paper describes this · verbatim
in the paper
3Preparation
no AI

Box-crop maps and label voxels by DSSP assignment

Cleaning, filtering, normalising or labelling data already obtained.

the density map region was cropped around the specified central chain using ChimeraXwhere the paper describes this · verbatim
in the paper
4Training
AI

Train one U-Net per loss function and pick hyperparameters

Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.

For each loss function, training was conducted using the same training dataset.where the paper describes this · verbatim
in the paper
5Inference
AI

Segment held-out test maps

Running a trained model over new data to predict, classify or score. The AI stood in for conventional algorithm.

Each voxel of the map had a predicted label indicating one of the following three classes: helix, β-sheet, and other or background.where the paper describes this · verbatim
in the paper
6Validation
no AI

Score voxel- and residue-level F1 against atomic structures

Testing outputs against ground truth.

the center box without the 34 Å padding was evaluated for the F1 score in each test casewhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The reported result is the per-voxel helix/β-sheet segmentation produced by the trained U-Net under five loss functions; there is no non-AI route to the finding.

+What the AI was for
Segmentationin the paper
It was based on the 3D U-Net model and used 4D tensors with different sizes as inputswhere the paper describes this · verbatim
+Model families
+How it was taught
Supervisedin the paper
+Models named
DeepSSETracer 3D U-Net secondary-structure segmentation CNN 1.1 · Trained from scratchin the paper
+How results were checked
Held-out62 testedin the paper
We used the test dataset of 62 box-cropped cryo-EM density maps to investigate the performance of five modelswhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
DeepSSETracer 1.1 is downloadable at https://www.cs.odu.edu/∼bioinfo/B2I_Tools/.where the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 2 items
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.

About this article

Record aix-00079, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error