OpenMS
Loading...
Searching...
No Matches
ProSEAlgorithm::CandidatePoolStats_ Struct Reference

Running summary of the complete candidate pool of one spectrum. More...

#include <OpenMS/ANALYSIS/ID/ProSEAlgorithm.h>

Public Member Functions

void add (double score)
 Fold one freshly scored candidate into the summary.
 
double deltaScore () const
 Best minus runner-up over the full pool.
 
double zScore () const
 How much of an outlier the best score is within its own candidate pool.
 

Public Attributes

double sum = 0.0
 sum of all candidate scores
 
double sumsq = 0.0
 sum of squared candidate scores
 
double best = 0.0
 best candidate score (HyperScore is non-negative, so 0 doubles as "none seen")
 
double second_best = 0.0
 runner-up candidate score
 
Size count = 0
 number of candidates scored
 

Detailed Description

Running summary of the complete candidate pool of one spectrum.

scoreSpectraAgainstIndex_() prunes each spectrum to max(report_top_hits_, 2) candidates as soon as it has scored them, so by the time postProcessHits_() runs the surviving hits are no longer a sample of the search space: with the default report:top_hits=1 only two candidates remain, and derived features such as ln_num_candidates or hyperscore_zscore would degenerate into a binary flag and a rescaled delta score respectively.

add() is therefore called for every candidate the moment it is scored – before pruning and before zero-scoring candidates are dropped – so the summary reflects the full pool. Instances accumulate across chunks in the chunked search paths, where each chunk contributes its own candidates for the same spectrum.

Member Function Documentation

◆ add()

void add ( double  score)
inline

Fold one freshly scored candidate into the summary.

◆ deltaScore()

double deltaScore ( ) const
inline

Best minus runner-up over the full pool.

◆ zScore()

double zScore ( ) const

How much of an outlier the best score is within its own candidate pool.

Standard score of best against the mean and (population) SD of all count candidates – a lightweight significance proxy, since HyperScore (unlike e.g. MS-GF+'s SpecEValue) has no closed-form e-value.

Standardising against the whole pool rather than against the count - 1 non-best candidates keeps the statistic bounded: by Samuelson's inequality a pool member cannot deviate from its own mean by more than sqrt(count - 1) SDs, so the feature tops out at sqrt(scoring:max_candidates_per_spectrum - 1) (7 at the default cap of 50). The leave-one-out form has no such bound – with two candidates the SD of the single remaining score is 0 by construction, and the ratio diverges. That mattered in practice: on a 70k-PSM Orbitrap run the leave-one-out form produced values up to 6.5e+07, which is enough to dominate Percolator's feature standardisation.

Returns 0 when fewer than two candidates were scored, and when every candidate scored the same – in both cases the best candidate does not stand out from the pool.

Member Data Documentation

◆ best

double best = 0.0

best candidate score (HyperScore is non-negative, so 0 doubles as "none seen")

◆ count

Size count = 0

number of candidates scored

◆ second_best

double second_best = 0.0

runner-up candidate score

◆ sum

double sum = 0.0

sum of all candidate scores

◆ sumsq

double sumsq = 0.0

sum of squared candidate scores