⮝ Full datasets listing

PXD062269

PXD062269 is an original dataset announced via ProteomeXchange.

Dataset Summary
TitleAn integrated landscape of mRNA and protein isoforms
DescriptionCellular processes like alternative splicing (AS) and proteolytic processing generate various protein isoforms from the same protein-coding gene. Studies using high-throughput sequencing indicate that about 90% of multi-exon genes in humans undergo AS. These AS events often occur in a tissue and developmental stage specific manner. Proteolytic processing is important for regulation of protein function and involved in processes like cell cycle regulation, apoptosis and protein degradation. Standard bottom-up proteomics involves digesting proteins into peptides. Since many peptide sequences match to several proteoforms, the information to which proteoform a given peptide belongs is lost at this step. Furthermore, there is an evolutionarily conserved preferential usage of lysine and arginine at splicing junctions. In combination with tryptic digestion, this impedes detection of splice junction-spanning peptides. Hence, the contribution of AS to protein diversity remains understudied. In this study we combined full-length mRNA sequencing (Iso-Seq) with proteomic analyses to obtain an integrated landscape of mRNA and protein isoforms in human RPE-1 cells. To overcome the limitations of bottom proteomics to study protein isoforms we resolved proteins by extensive protein-level fractionation by SDS-PAGE using the GELFREE 8100 fractionation system. After digestion of individual fractions and TMT labeling, the samples were combined and analyzed using conventional LC-MS/MS. A newly developed computational framework enabled us to automatically differentiate proteoforms from bottom up proteomics at a higher scale. As a result we established a reference set of ~45,000 full-length transcripts, ~32,000 ORFs and ~16,000 protein isoforms. Our data reveals multiple protein isoforms for many genes and provides an integrated landscape of mRNA and protein isoforms to reveal how transcriptional, translational and post-translational processes contribute to proteome complexity.
HostingRepositoryPRIDE
AnnounceDate2026-07-07
AnnouncementXMLSubmission_2026-07-07_03:58:01.336.xml
DigitalObjectIdentifier
ReviewLevelPeer-reviewed dataset
DatasetOriginOriginal dataset
RepositorySupportUnsupported dataset by repository
PrimarySubmitterHenrik Zauber
SpeciesList scientific name: Homo sapiens (Human); NCBI TaxID: NEWT:9606;
ModificationListacetylated residue; monohydroxylated residue; iodoacetamide derivatized residue
InstrumentQ Exactive HF; Orbitrap Exploris 480
Dataset History
RevisionDatetimeStatusChangeLog Entry
02025-03-26 13:47:54ID requested
12026-07-07 03:58:01announced
Publication List
Dataset with its publication pending
Keyword List
submitter keyword: PACBIO, Isoforms,Proteogenomics,SDS-Gel, Protein Processing,TMT, Alternative Splicing
Contact List
Matthias Selbach
contact affiliationMax Delbrück Center for Molecular Medicine in the Helmholtz Association
contact emailmatthias.selbach@mdc-berlin.de
lab head
Henrik Zauber
contact affiliationMDC Berlin-Buch
contact emailhenrik.zauber@mdc-berlin.de
dataset submitter
Full Dataset Link List
Dataset FTP location
NOTE: Most web browsers have now discontinued native support for FTP access within the browser window. But you can usually install another FTP app (we recommend FileZilla) and configure your browser to launch the external application when you click on this FTP link. Or otherwise, launch an app that supports FTP (like FileZilla) and use this address: ftp://ftp.pride.ebi.ac.uk/pride/data/archive/2026/07/PXD062269
PRIDE project URI
Repository Record List
[ + ]