structural-biology/ai produced the result/Database 2023 · v2
AlphaFold models used to map membrane-crossing segments in human proteins
Researchers built a database of the stretches of human proteins that pass through cell membranes. They took predicted three-dimensional structures from AlphaFold and used a physics-based program to settle each model into a membrane.
spectrum · one line per step, placed by what the step does · bright lines used AI
AFTM: a database of transmembrane regions in the human proteome predicted by AlphaFold
Database, 2023
doi:10.1093/database/baad008 · record aix-00154 v2 · checked 2026-10-09
- AI was for
- Structure determination, Property prediction
- Model family
- Transformer
- Checked by
- Benchmark601 tested
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
Cells are wrapped in membranes, thin sheets of fat that keep the inside separate from the outside. Many proteins sit in those sheets, with parts of the chain threading right through them. Knowing exactly which stretches of amino acids cross the membrane matters for understanding how a protein works, yet those stretches are hard to pin down. Membrane proteins are notoriously awkward to study by the usual experimental methods, which is why many human entries carry annotations based on indirect reasoning rather than a measured structure. Different databases often disagree about where a crossing starts and ends, or whether one is there at all.
The researchers set out to annotate these segments across the human proteome by working from predicted structures rather than sequence alone. They gathered 5491 candidate human membrane proteins using existing annotations from the UniProt and HTP databases, then worked out, protein by protein, which parts of each chain lie inside the membrane.
Where AI came in
The AI's part was supplying the shapes. AlphaFold predicts a protein's folded three-dimensional structure from its sequence, a job that otherwise needs laboratory structure determination, and the authors used its models as the raw material for everything that followed. AlphaFold also reports how sure it is: a per-residue confidence score and a matrix estimating the error between pairs of positions. Low-confidence positions were stripped out, and the error matrix guided an in-house script that split each model into separate ordered domains.
The membrane placement itself was not done by a learned model. A program called PPM3, which works from energy calculations rather than training data, settled each model into a simulated membrane and reported which segments crossed it; fixed geometric rules then tidied the output. Separately, the website shows off-the-shelf sequence predictions of local structure and of floppy, disordered regions. No model was trained or fine-tuned for this work.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The authors assembled 5491 candidate human transmembrane proteins from UniProt and HTP annotations, then located transmembrane segments by placing AlphaFold structural models in a membrane with the PPM3 program, after trimming low-confidence and peptide regions and splitting models into domains. Rule-based corrections separated re-entrant regions, merged broken helices and added segments PPM3 had skipped. Against transmembrane regions mapped from PDBTM experimental structures for 601 proteins, 66 PDBTM-defined segments were not predicted by AFTM, compared with 107 for TmAlphaFold, 197 for UniProt and 177 for HTP; on 2322 Membranome single-pass proteins, AFTM reported no segment for 309, compared with 69 for UniProt and 141 for HTP. Results from AFTM, UniProt, HTP, TmAlphaFold, PDBTM and Membranome are published together in an online database.
How AI was used
AlphaFold structural models of human proteins, together with their per-residue pLDDT scores and predicted aligned error matrices, were the substrate for the whole annotation procedure: pLDDT was used to strip low-confidence positions from full-length models, and PAE densities were used by an in-house splitting script to partition models into ordered and disordered domains. Both the trimmed full-length models and the individual domains were then passed to PPM3, a non-learned free-energy-based membrane placement program, whose reported segments were post-processed by fixed geometric rules into transmembrane versus re-entrant regions, with merging of segments of the same orientation separated by fewer than ten residues and insertion of segments PPM3 failed to report. Segment sets were compared with UniProt, HTP, TmAlphaFold, PDBTM-derived and Membranome annotations, using DIAMOND BLAST to map external segments onto human sequences. The web pages additionally display off-the-shelf per-residue predictions of secondary structure (PSIPRED, SPIDER3) and disorder (SPOT-Disorder, IUPRED2A). No model was trained or fine-tuned in this study.
The shape of the work
Structural · the record, drawn
no AI
Compile candidate human transmembrane protein set
Obtaining raw data, whether by measurement, download or retrieval.
This dataset includes 5491 potential human TMPs compiled from a combination of reviewed UniProt entrieswhere the paper describes this · verbatim
AI
Obtain AlphaFold structural models with pLDDT and PAE
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
In addition to the predicted 3D structure, AlphaFold provides the predicted aligned error (PAE) for each residue pair in a proteinwhere the paper describes this · verbatim
no AI
Strip peptides and low-confidence regions from models
Cleaning, filtering, normalising or labelling data already obtained.
positions with low pLDDT scores (<0.5) were removed from the full-length AlphaFold modelwhere the paper describes this · verbatim
no AI
Partition models into ordered and disordered domains
Cleaning, filtering, normalising or labelling data already obtained.
We wrote an in-house script to iterate the following procedure to split any AlphaFold model into segments (domains)where the paper describes this · verbatim
no AI
Position models in membrane with PPM3
Numerical or physics simulation, including where a learned surrogate replaces it.
We used the positioning of proteins in membranes, version 3 (PPM3) method to predict the localizations of TMSswhere the paper describes this · verbatim
no AI
Correct and label segments into final AFTM annotations
Cleaning, filtering, normalising or labelling data already obtained.
we merged the consecutive PPM3 segments if they have the same orientation and are separated by <10 residueswhere the paper describes this · verbatim
no AI
Benchmark against experimental and curated annotations
Testing outputs against ground truth.
we compared their TMS predictions to those derived from experimental structures in the PDBTM databasewhere the paper describes this · verbatim
AI
Add sequence-feature predictions to protein web pages
Running a trained model over new data to predict, classify or score.
secondary structure predictions by PSI-blast based secondary structure PREDiction (PSIPRED) and SPIDER3where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
Every transmembrane segment reported by AFTM is derived from AlphaFold structural models; without the predicted structures there is no result, although the segment assignment itself is made by the non-learned PPM3 program
The PPM3 program was used to predict the TMSs in AlphaFold models of human proteinswhere the paper describes this · verbatim
A total of 601 human proteins in our dataset were mapped to at least one entrywhere the paper describes this · verbatim
The AFTM database is available at http://conglab.swmed.edu/AFTMwhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- Version of AlphaFoldWhich version of the model was used is not stated.
- Version of PSIPREDWhich version of the model was used is not stated.
- Version of SPIDER3Which version of the model was used is not stated.
- Version of SPOT-DisorderWhich version of the model was used is not stated.
- What step 8 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00154, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error