Integrated data management and validation platform for phosphorylated tandem mass spectrometry data
Author(s)
Lahesmaa-Korpinen, Anna-Maria; Carlson, Scott M.; White, Forest M.; Hautaniemi, Sampsa
DownloadWhite_Integrated data.pdf (1002.Kb)
OPEN_ACCESS_POLICY
Open Access Policy
Creative Commons Attribution-Noncommercial-Share Alike
Terms of use
Metadata
Show full item recordAbstract
MS/MS is a widely used method for proteome-wide analysis of protein expression and PTMs. The thousands of MS/MS spectra produced from a single experiment pose a major challenge for downstream analysis. Standard programs, such as MASCOT, provide peptide assignments for many of the spectra, including identification of PTM sites, but these results are plagued by false-positive identifications. In phosphoproteomic experiments, only a single peptide assignment is typically available to support identification of each phosphorylation site, and hence minimizing false positives is critical. Thus, tedious manual validation is often required to increase confidence in the spectral assignments. We have developed phoMSVal, an open-source platform for managing MS/MS data and automatically validating identified phosphopeptides. We tested five classification algorithms with 17 extracted features to separate correct peptide assignments from incorrect ones using over 2600 manually curated spectra. The naïve Bayes algorithm was among the best classifiers with an AUC value of 97% and PPV of 97% for phosphotyrosine data. This classifier required only three features to achieve a 76% decrease in false positives as compared with MASCOT while retaining 97% of true positives. This algorithm was able to classify an independent phosphoserine/threonine data set with AUC value of 93% and PPV of 91%, demonstrating the applicability of this method for all types of phospho-MS/MS data. PhoMSVal is available at http://csbi.ltdk.helsinki.fi/phomsval.
Date issued
2010-09Department
Massachusetts Institute of Technology. Department of Biological Engineering; Koch Institute for Integrative Cancer Research at MITJournal
PROTEOMICS
Publisher
Wiley Blackwell
Citation
Lahesmaa-Korpinen, Anna-Maria et al. “Integrated Data Management and Validation Platform for Phosphorylated Tandem Mass Spectrometry Data.” PROTEOMICS 10.19 (2010): 3515–3524.
Version: Author's final manuscript
ISSN
1615-9853
1615-9861