As global proteomics continues to advance, the number of identifiable proteins has increased substantially. However, this does not inherently ensure optimal quantitative performance. While targeted assays using isotope-labeled peptides can be developed, label-free strategies remain an attractive and cost-efficient option for methodological validation. Yet, systematic evaluations of data analysis workflows for label-free targeted proteomics, particularly those incorporating AI based tools, are still limited. Therefore, this study aimed to benchmark multiple data analysis approaches for label-free prmPASEF proteomics as a validation framework for results obtained from global analyses.
A benchmarking dataset was prepared using a constant human proteome background (30 proteins measured) and a spike-in of 3 yeast proteins at varying concentrations. After in-solution digestion, a label-free prmPASEF analysis was performed using Bruker timsTOF Pro system. Results were imported into Skyline-daily and checked for integrity. Separate peptide fragment areas were exported into the R environment for subsequent analysis.
Calibration curves were plotted, and only measurements within the determined LOQ - LOL range were included in workflow comparisons. Missing-data imputation (MDI) strategies, including no MDI, k-nearest neighbors, data-driven, and AI-based methods, were evaluated, alongside consolidation and testing frameworks such as mathematical summation, best-peak selection, AI-based scaling with univariate statistics, p-value integration, and multivariate testing. Data-driven MDI combined with p-value integration consistently showed the strongest performance across accuracy, precision, specificity, and false-discovery rate, outperforming all other strategies.
These findings demonstrate that careful selection of data analysis workflows can yield substantially improved quantitative outcomes compared with commonly used approaches such as simple mathematical summation. Although our conclusions are based on a controlled benchmarking dataset comprising three yeast proteins spiked into a constant human background using only the prmPASEF method, they can be generally applied as a default workflow for biomarker validation using the label-free approach.