OpenMS
Loading...
Searching...
No Matches
ProteinQuantifier

Compute peptide and protein abundances from annotated feature/consensus maps or from identification results.

potential predecessor tools → ProteinQuantifier → potential successor tools
IDMapper external tools
e.g. for statistical analysis
FeatureLinkerUnlabeled
(or another feature grouping tool)

Reference:
Weisser et al.: An automated pipeline for high-throughput label-free quantitative proteomics (J. Proteome Res., 2013, PMID: 23391308).

Input: featureXML or consensusXML

Quantification is based on the intensity values of the features in the input files. Feature intensities are first accumulated to peptide abundances, according to the peptide identifications annotated to the features/feature groups. Then, abundances of the peptides of a protein are aggregated to compute the protein abundance.

The peptide-to-protein step uses the (e.g. 3) most abundant proteotypic peptides per protein to compute the protein abundances. This is a general version of the "top 3 approach" (but only for relative quantification) described in:
Silva et al.: Absolute quantification of proteins by LCMSE: a virtue of parallel MS acquisition (Mol. Cell. Proteomics, 2006, PMID: 16219938).

Only features/feature groups with unambiguous peptide annotation are used for peptide quantification. It is possible to resolve ambiguities before applying ProteinQuantifier using one of several equivalent mechanisms in OpenMS: IDConflictResolver, ConsensusID (algorithm best), or FileFilter (option id:keep_best_score_id).

Similarly, only proteotypic peptides (i.e. those matching to exactly one protein) are used for protein quantification by default. Peptide/protein IDs from multiple identification runs can be handled, but will not be differentiated (i.e. protein accessions for a peptide will be accumulated over all identification runs). See section "Optional input: Protein inference/grouping results" below for exceptions to this.

Peptides with the same sequence, but with different modifications are quantified separately on the peptide level, but treated as one peptide for the protein quantification (i.e. the contributions of differently-modified variants of the same peptide are accumulated).

Output granularity: assays vs. files and channels

By default one protein and peptide abundance is reported per assay. An assay is the experimental-design pair (fraction_group, label): it spans every fraction file of that fraction group at that label, and its reported value aggregates over those files. Columns are named abundance_fgroupF_labelL, where F is the design's Fraction_Group and L its Label. The SampleSection remains metadata and may group several assays as technical or biological replicates; ProteinQuantifier does not sum those replicates.

How the fractions of a group are combined into its assay values is controlled by fractions:aggregate. The default sum adds them up, treating them as the parts of one separated sample that they are. best instead keeps one fraction per peptide and fraction group and discards the rest: the fraction with the most labels at a positive abundance wins, ties are broken by the total of those abundances and then by the lower fraction number. The choice is made once for the whole fraction group and never per label - taking one channel from one fraction and another channel from a different fraction would mix physical aliquots and destroy the reporter-ion ratios that isobaric quantification consists of. Note that best reports a fraction of the material rather than all of it, which matters for LFQ, where the value is an absolute intensity, more than for isobaric data, where quantification is relative within a run and any single fraction preserves the ratios.

With file_and_channel_level_output the protein abundances are instead reported per (file, channel) cell. These cells are computed with the same peptide-level policy as the assay values (all peptidoforms are accumulated into one peptide; all charge states contribute by default, or only each peptidoform's selected charge with best_charge), but the peptide selection and the aggregation are applied per file. Two consequences are worth knowing:

  • The cells only decompose the assay value exactly for top:N 0 together with top:aggregate sum. Top-N selection, median, mean and weighted_mean do not commute with aggregation across fractions, so for those settings the cells of an assay neither sum nor average to the assay value.
  • The top:N requirement ("at least N peptides") is likewise enforced per file, not per assay. In a fractionated experiment this is considerably stricter than the assay-level rule: a protein can easily have N peptides in an assay while no individual fraction contains N of them, in which case the protein is quantified at the assay level but all of its (file, channel) cells are reported as 0. Use top:N 0 (optionally with top:aggregate sum) or top:include_all if per-file values are wanted for such data.
  • fractions:aggregate does not apply to them. A (file, channel) cell is one fraction by definition, so every file is always reported, even under best where the assay value comes from a single fraction. The two granularities then describe the data at different completeness on purpose.

With best_charge, one charge is selected globally for each modified peptide. Charges are ranked first by the number of distinct assays with a positive abundance and then, on a tie, by total abundance across all assays (an exact tie deterministically keeps the lower charge). Every observation of the selected charge is retained and then combined over the fractions of an assay according to fractions:aggregate; the detailed peptide output still reports all observed fraction and charge combinations. The same selected-charge policy is used for assay and file/channel protein quantities.

The detailed peptide_out table that this flag produces has one row per (fraction, charge) and one abundance column per (file, channel) covering every file of the experimental design, so a row reports 0.0 for the files and channels its fraction does not cover. Without the flag, peptide_out instead writes one column per assay and one row per peptide, with fraction reported as "all".

Input: idXML

Quantification based on identification results uses spectral counting, i.e. the abundance of each peptide is the number of times that peptide was identified from an MS2 spectrum (considering only the best hit per spectrum). Different identification runs in the input become distinct inferred assays; this makes it possible to quantify several related runs at once by merging the corresponding idXML files with IDMerger. Depending on the presence of multiple runs, output format and applicable parameters are the same as for featureXML and consensusXML, respectively.

The notes above regarding quantification on the protein level and the treatment of modifications also apply to idXML input. In particular, this means that the settings top 0 and aggregate sum should be used to get the "classical" spectral counting quantification on the protein level (where all identifications of all peptides of a protein are summed up).

Optional input: Protein inference/grouping results

By default only proteotypic peptides (i.e. those matching to exactly one protein) are used for protein quantification. However, this limitation can be overcome: Protein inference results for the complete data set can be supplied with the protein_groups option (or included in a featureXML input). In that case, the peptide-to-protein references from that file are used (rather than those from in), and groups of indistinguishable proteins will be quantified. Each reported protein quantity then refers to the total for the respective group.

In order for everything to work correctly, it is important that the protein inference results come from the same identifications that were used to annotate the quantitative data. We suggest to use the OpenMS tool ProteinInference ProteinInference.

More information below the parameter specification.

Optional output: QPX Parquet (out_qpx)

out_qpx writes the quantification as a QPX collection - quantms.feature.parquet, quantms.psm.parquet and quantms.pg.parquet - for consensusXML input. QPX is an interchange format with a strict value contract, and OpenMS refuses to write a table it cannot represent rather than emit one that will not join. A refusal aborts the tool and leaves no files behind, including any view already written.

Consensus maps produced by ProteomicsLFQ and IsobaricWorkflow satisfy the contract by construction; those two are the supported producers. A map assembled by a different pipeline may not. What the contract requires, and how to satisfy it, is documented in one place: on OpenMS::QPXValueValidation, the class that enforces it.

Note
Currently mzIdentML (mzid) is not directly supported as an input/output format of this tool. Convert mzid files to/from idXML using IDFileConverter if necessary.

The command line parameters of this tool are:

ProteinQuantifier -- Compute peptide and protein abundances
Full documentation: http://www.openms.de/doxygen/nightly/html/TOPP_ProteinQuantifier.html
Version: 3.6.0-pre-nightly-2026-09-29 Sep 30 2026, 01:45:35, Revision: 55f7bdb
To cite OpenMS:
 + Pfeuffer, J., Bielow, C., Wein, S. et al.. OpenMS 3 enables reproducible analysis of large-scale mass spec
   trometry data. Nat Methods (2024). doi:10.1038/s41592-024-02197-7.

Usage:
  ProteinQuantifier <options>

Options (mandatory options marked with '*'):
  -in <file>*                              Input file (valid formats: 'featureXML', 'consensusXML', 'idXML')
  -protein_groups <file>                   Protein inference results for the identification runs that were 
                                           used to annotate the input (e.g. via the ProteinInference tool).
                                           Information about indistinguishable proteins will be used for prot
                                           ein quantification. (valid formats: 'idXML')
  -design <file>                           Input file containing the experimental design (valid formats: 'tsv
                                           ')
  -out <file>                              Output file for protein abundances (valid formats: 'csv')
  -peptide_out <file>                      Output file for peptide abundances (valid formats: 'csv')
  -mztab <file>                            Output file (mzTab) (valid formats: 'mzTab')
  -out_qpx <directory>                     Output directory for QPX Parquet files (quantms.feature.parquet, 
                                           quantms.psm.parquet, quantms.pg.parquet). Only supported for conse
                                           nsusXML input.
                                           QPX has a strict value contract; input that does not meet it is 
                                           refused outright and no files are written. Maps produced by Proteo
                                           micsLFQ or IsobaricWorkflow satisfy it by construction, other pipe
                                           lines may not. The contract is documented on the OpenMS::QPXValueV
                                           alidation class, which enforces it.
                                           
  -method <choice>                         - top - quantify based on three most abundant peptides (number 
                                           can be changed in 'top').
                                           - iBAQ (intensity based absolute quantification), calculate the 
                                           sum of all peptide peak intensities divided by the number of theor
                                           etically observable tryptic peptides (https://rdcu.be/cND1J). Warn
                                           ing: only consensusXML or featureXML input is allowed! (default: 
                                           'top') (valid: 'top', 'iBAQ')
  -best_charge                             Distinguish between fraction and charge states in detailed peptide
                                            output. For protein quantification, select one charge per modifie
                                           d peptide globally: maximize the number of (fraction group, label)
                                            assays with a positive abundance, then break ties by total abunda
                                           nce; retain that charge's values in every assay.
                                           By default, protein abundances are summed over all charge states. 
                                           How the retained values of several fractions are combined is gover
                                           ned by 'fractions:aggregate', not by this flag.

Additional options for custom quantification using top N peptides.:
  -top:N <number>                          Calculate protein abundance from this number of proteotypic peptid
                                           es (most abundant first; '0' for all) (default: '3') (min: '0')
  -top:aggregate <choice>                  Aggregation method used to compute protein abundances from peptide
                                            abundances (default: 'median') (valid: 'median', 'mean', 'weighte
                                           d_mean', 'sum')
  -top:include_all                         Include results for proteins with fewer proteotypic peptides than 
                                           indicated by 'N' (no effect if 'N' is 0 or 1)

Options for combining the fractions of a fraction group.:
  -fractions:aggregate <choice>            How the fractions of one fraction group are combined into that 
                                           group's (fraction group, label) assay values.
                                           - sum - add up every fraction, i.e. treat them as the parts of 
                                           one separated sample that they are.
                                           - best - keep a single fraction per peptide and fraction group 
                                           and discard the others. The fraction is chosen ONCE per peptide, 
                                           ranked by the number of labels in which it has a positive abundanc
                                           e and then by the total of those abundances (an exact tie keeps 
                                           ...
                                            are per fraction by definition and always report every file. (def
                                           ault: 'sum') (valid: 'sum', 'best')

Additional options for consensus maps (and identification results comprising multiple runs):
  -consensus:normalize                     Scale peptide abundances so that the median of each (fraction grou
                                           p, label) assay matches the overall median.
                                           Abundances of zero count as 'not detected' and are left out of 
                                           the medians; an assay without any positive abundance takes no part
                                            in the normalization.
  -consensus:fix_peptides                  Use the same peptides for protein quantification across all (fract
                                           ion group, label) assays.
                                           With 'N 0',all peptides that occur in every assay are considered.
                                           Otherwise ('N'), the N peptides that occur in the most assays (ind
                                           ependently of each other) are selected,
                                           breaking ties by total abundance (there is no guarantee that the 
                                           best co-ocurring peptides are chosen!).
                                           A peptide counts as occurring in an assay only where its abundance
                                           ...
                                           measurement of absence.

  -greedy_group_resolution <choice>        Pre-process identifications with greedy resolution of shared pepti
                                           des based on the protein group probabilities. (Only works with an 
                                           idXML file given as protein_groups parameter). (default: 'false') 
                                           (valid: 'true', 'false')
  -file_and_channel_level_output <choice>  Output protein abundances with detailed file+channel level headers
                                            (similar to detailed peptide output). When enabled, protein outpu
                                           t will show abundance_filename_channel columns instead of assay 
                                           columns.
                                           Note that peptide selection and aggregation are then applied per 
                                           file, not per assay: 'top:N' requires N peptides in that single 
                                           file (much stricter than the assay-level rule for fractionated 
                                           data, where all cells of a quantified protein can end up 0), and 
                                           the cells only decompose the assay-level value for 'top:N' 0 with 
                                           'top:aggregate' sum. (default: 'false') (valid: 'true', 'false')

Output formatting options:
  -format:separator <sep>                  Character(s) used to separate fields; by default, the 'tab' charac
                                           ter is used
  -format:quoting <method>                 Method for quoting of strings: 'none' for no quoting, 'double' 
                                           for quoting with doubling of embedded quotes,
                                           'escape' for quoting with backslash-escaping of embedded quotes 
                                           (default: 'double') (valid: 'none', 'double', 'escape')
  -format:replacement <x>                  If 'quoting' is 'none', used to replace occurrences of the separat
                                           or in strings before writing (default: '_')

                                           
Common TOPP options:
  -ini <file>                              Use the given TOPP INI file
  -threads <n>                             Sets the number of threads allowed to be used by the TOPP tool (0 
                                           = all available cores) (default: '1')
  -write_ini <file>                        Writes the default configuration file
  --help                                   Shows options
  --helphelp                               Shows all options (including advanced)

INI file documentation of this tool:

Legend:
required parameter
advanced parameter

This section lists all parameters supported by the tool. Parameters are organized into hierarchical subsections that group related settings together. Subsections may contain further subsections or individual parameters.

Each parameter entry contains the following information:

  • Name The identifier used in configuration files and on the command line.
  • Default value The value used if the parameter is not explicitly specified.
  • Description A short explanation describing the purpose and behavior of the parameter.
  • Tags Additional metadata associated with the parameter.
  • Restrictions Allowed value ranges for numeric parameters or valid options for string parameters.

Parameter tags provide additional information about how a parameter is used. Some tags indicate whether a parameter is required or intended for advanced configuration, while others may be used internally by OpenMS or workflow tools.

Parameters highlighted as required must be specified for the tool to run successfully. Parameters marked as advanced allow fine-tuning of algorithm behavior and are typically not needed for standard workflows.

+ProteinQuantifierCompute peptide and protein abundances
version3.6.0-pre-nightly-2026-09-29 Version of the tool that generated this parameters file.
++1Instance '1' section for 'ProteinQuantifier'
in Input fileinput file*.featureXML, *.consensusXML, *.idXML
protein_groups Protein inference results for the identification runs that were used to annotate the input (e.g. via the ProteinInference tool).
Information about indistinguishable proteins will be used for protein quantification.
input file*.idXML
design input file containing the experimental designinput file*.tsv
out Output file for protein abundancesoutput file*.csv
peptide_out Output file for peptide abundancesoutput file*.csv
mztab Output file (mzTab)output file*.mzTab
out_qpx Output directory for QPX Parquet files (quantms.feature.parquet, quantms.psm.parquet, quantms.pg.parquet). Only supported for consensusXML input.
QPX has a strict value contract; input that does not meet it is refused outright and no files are written. Maps produced by ProteomicsLFQ or IsobaricWorkflow satisfy it by construction, other pipelines may not. The contract is documented on the OpenMS::QPXValueValidation class, which enforces it.
output dir
methodtop - top - quantify based on three most abundant peptides (number can be changed in 'top').
- iBAQ (intensity based absolute quantification), calculate the sum of all peptide peak intensities divided by the number of theoretically observable tryptic peptides (https://rdcu.be/cND1J). Warning: only consensusXML or featureXML input is allowed!
top, iBAQ
best_chargefalse Distinguish between fraction and charge states in detailed peptide output. For protein quantification, select one charge per modified peptide globally: maximize the number of (fraction group, label) assays with a positive abundance, then break ties by total abundance; retain that charge's values in every assay.
By default, protein abundances are summed over all charge states. How the retained values of several fractions are combined is governed by 'fractions:aggregate', not by this flag.
true, false
greedy_group_resolutionfalse Pre-process identifications with greedy resolution of shared peptides based on the protein group probabilities. (Only works with an idXML file given as protein_groups parameter).true, false
file_and_channel_level_outputfalse Output protein abundances with detailed file+channel level headers (similar to detailed peptide output). When enabled, protein output will show abundance_filename_channel columns instead of assay columns.
Note that peptide selection and aggregation are then applied per file, not per assay: 'top:N' requires N peptides in that single file (much stricter than the assay-level rule for fractionated data, where all cells of a quantified protein can end up 0), and the cells only decompose the assay-level value for 'top:N' 0 with 'top:aggregate' sum.
true, false
log Name of log file (created only when specified)
debug0 Sets the debug level
threads1 Sets the number of threads allowed to be used by the TOPP tool (0 = all available cores)
no_progressfalse Disables progress logging to command linetrue, false
forcefalse Overrides tool-specific checkstrue, false
testfalse Enables the test mode (needed for internal use only)true, false
+++topAdditional options for custom quantification using top N peptides.
N3 Calculate protein abundance from this number of proteotypic peptides (most abundant first; '0' for all)0:∞
aggregatemedian Aggregation method used to compute protein abundances from peptide abundancesmedian, mean, weighted_mean, sum
include_allfalse Include results for proteins with fewer proteotypic peptides than indicated by 'N' (no effect if 'N' is 0 or 1)true, false
+++fractionsOptions for combining the fractions of a fraction group.
aggregatesum How the fractions of one fraction group are combined into that group's (fraction group, label) assay values.
- sum - add up every fraction, i.e. treat them as the parts of one separated sample that they are.
- best - keep a single fraction per peptide and fraction group and discard the others. The fraction is chosen ONCE per peptide, ranked by the number of labels in which it has a positive abundance and then by the total of those abundances (an exact tie keeps the lowest fraction number), and ALL of its labels are then taken from it. The choice is deliberately not made per label: taking one channel from one fraction and another channel from a different fraction would mix physical aliquots and destroy the reporter-ion ratios that isobaric quantification consists of.
Only the assay values are affected. Per-(file, channel) quantities are per fraction by definition and always report every file.
sum, best
+++consensusAdditional options for consensus maps (and identification results comprising multiple runs)
normalizefalse Scale peptide abundances so that the median of each (fraction group, label) assay matches the overall median.
Abundances of zero count as 'not detected' and are left out of the medians; an assay without any positive abundance takes no part in the normalization.
true, false
fix_peptidesfalse Use the same peptides for protein quantification across all (fraction group, label) assays.
With 'N 0',all peptides that occur in every assay are considered.
Otherwise ('N'), the N peptides that occur in the most assays (independently of each other) are selected,
breaking ties by total abundance (there is no guarantee that the best co-ocurring peptides are chosen!).
A peptide counts as occurring in an assay only where its abundance is positive: an abundance stored as zero means 'not detected' (e.g. an isobaric reporter below 'min_reporter_intensity'), not a measurement of absence.
true, false
+++formatOutput formatting options
separator Character(s) used to separate fields; by default, the 'tab' character is used
quotingdouble Method for quoting of strings: 'none' for no quoting, 'double' for quoting with doubling of embedded quotes,
'escape' for quoting with backslash-escaping of embedded quotes
none, double, escape
replacement_ If 'quoting' is 'none', used to replace occurrences of the separator in strings before writing

Output format

The output files produced by this tool have a table format, with columns as described below:

Protein output (one protein/set of indistinguishable proteins per line):

  • protein: Protein accession(s) (as in the annotations in the input file; separated by "/" if more than one).
  • n_proteins: Number of indistinguishable proteins quantified (usually "1").
  • protein_score: Protein score, e.g. ProteinProphet probability (if available).
  • n_peptides: Number of proteotypic peptides observed for this protein (or group of indistinguishable proteins) across all assays. Note that not necessarily all of these peptides contribute to the protein abundance (depending on parameter top).
  • abundance_fgroupF_labelL: Computed protein abundance for assay (F, L). There is one self-describing column per assay in the experimental design.

Peptide output (one peptide or - if best_charge is set - one charge state and fraction of a peptide per line):

  • peptide: Peptide sequence. Only peptides that occur in unambiguous annotations of features are reported.
  • protein: Protein accession(s) for the peptide (separated by "/" if more than one).
  • n_proteins: Number of proteins this peptide maps to. (Same as the number of accessions in the previous column.)
  • charge: Charge state quantified in this line. "0" (for "all charges") unless best_charge was set.
  • abundance_fgroupF_labelL: Computed peptide abundance for assay (F, L). If the charge in the preceding column is 0, this is the total abundance over all charge states; otherwise, it is only the abundance observed for the indicated charge (in this case, the detailed table uses file/channel columns instead). For consensusXML input, the reported values are already normalized if consensus:normalize was set.

Protein quantification examples

While quantification on the peptide level is fairly straight-forward, a number of options influence quantification on the protein level - especially for consensusXML input. The three parameters top:N, top:include_all and consensus:fix_peptides determine which peptides are used to quantify proteins in different assays.

As an example, consider a protein with four proteotypic peptides. Each peptide is detected in a subset of three assays, as indicated in the table below. The peptides are ranked by abundance (1: highest, 4: lowest; assuming for simplicity that the order is the same in all assays).

assay 1 assay 2 assay 3
peptide 1 X X
peptide 2 X X
peptide 3 X X X
peptide 4 X X

Different parameter combinations lead to different quantification scenarios, as shown here:

parameters
"*": no effect in this case
peptides used for quantification
"(...)": not quantified here because ...
explanation
top include_all c.:fix_peptides assay 1 assay 2 assay 3
0 * no 1, 2, 3, 4 2, 3, 4 1, 3 all peptides
1 * no 1 2 1 single most abundant peptide
2 * no 1, 2 2, 3 1, 3 two most abundant peptides
3 no no 1, 2, 3 2, 3, 4 (too few peptides) three most abundant peptides
3 yes no 1, 2, 3 2, 3, 4 1, 3 three or fewer most abundant peptides
4 no * 1, 2, 3, 4 (too few peptides) (too few peptides) four most abundant peptides
4 yes * 1, 2, 3, 4 2, 3, 4 1, 3 four or fewer most abundant peptides
0 * yes 3 3 3 all peptides present in every assay
1 * yes 3 3 3 single peptide present in most assays
2 no yes 1, 3 (peptide 1 missing) 1, 3 two peptides present in most assays
2 yes yes 1, 3 3 1, 3 two or fewer peptides present in most assays
3 no yes 1, 2, 3 (peptide 1 missing) (peptide 2 missing) three peptides present in most assays
3 yes yes 1, 2, 3 2, 3 1, 3 three or fewer peptides present in most assays

Further considerations for parameter selection

With best_charge and the protein aggregation settings, there is a trade-off between comparability of protein abundances within an assay and of abundances for the same protein across different assays.
Setting best_charge may increase reproducibility between assays, but will distort the proportions of protein abundances within an assay. The reason is that ionization properties vary between peptides, but should remain constant across assays. Filtering by charge state can help to reduce the impact of feature detection differences between assays.
For aggregate, there is a qualitative difference between (intensity weighted) mean/median and sum in the effect that missing peptide abundances have (only if include_all is set or top is 0): (intensity weighted) mean and median ignore missing cases, averaging only present values. If low-abundant peptides are not detected in some assays, the computed protein abundances for those assays may thus be too optimistic. sum implicitly treats missing values as zero, so this problem does not occur and comparability across assays is ensured. However, with sum the total number of peptides ("summands") available for a protein may affect the abundances computed for it (depending on top), so results within an assay may become unproportional.