![]() |
OpenMS
|
Converts a UniProtKB XML protein database to PEFF 1.0 with per-entry modification, processing, variant and disulfide-bond annotations.
This is an OpenMS-native port of the standalone C# UniPEFF tool (David L. Tabb, UMC Groningen). For each entry it emits a PEFF descriptor line with \PName \GName \NcbiTaxId \TaxName \Length \SV \EV \PE \ID \AltAC \ModResPsi \ModResUnimod \ModRes \VariantSimple \VariantComplex \Processed and \DisulfideBond (the latter unless -omit_amino_acid_modifications is set).
UniProt's <feature type="disulfide bond"> entries are translated into half-cystine modifications (PSI-MOD:00798) which are then merged into the main modified-residue list in position-sorted order. Disulfide connectivity is reported as \DisulfideBond=(bond:idA,idB) tuples referencing id: prefixes on the half-cystine tuples inside \ModResPsi, for every bond whose two half-cystines UniProt locates in this sequence (<begin>/<end>). These are mostly intrachain bonds, but also bonds between chains cleaved from the same precursor (e.g. insulin's "Interchain (between B and A chains)"). A bond given as a single <position> (partner cysteine in another molecule) or with an endpoint beyond the sequence end gets no connectivity; its in-range half-cystines are written as plain modifications (see below for the rest). By default only the reported bonds and their half-cystines are labeled, in bond order: of K reported bonds, bond k (counting from 0) labels its two half-cystines 2k and 2k+1 and is itself labeled 2K+k, so the ids 0..3K-1 are unique within the entry, as PEFF requires (see issue 9829). With -annotation_identifiers (PEFF "Option B") every annotation tuple instead carries a global sequential id (0, 1, 2, ...) and \DisulfideBond references those ids. Each database block with ids declares # HasAnnotationIdentifiers=true.
Annotations positioned beyond the end of the sequence (malformed input), i.e. modifications, variants and processed regions, are omitted with a warning, since PEFF treats such a position as an error in the file.
Features whose <location sequence="..."> names another isoform carry coordinates of that isoform, not of the entry's canonical sequence. They are never applied to the canonical entry. If that isoform is written as an entry of its own (see below), they become annotations of that entry, in its coordinates; otherwise they are not applied. Summary log lines count both kinds among the features that would otherwise have become annotations. A location naming the "displayed" isoform refers to the canonical sequence.
Modification accession lookup uses UniProt's ptmlist.txt (a snapshot is bundled under share/OpenMS/CHEMISTRY/UniProt_ptmlist.txt); override with -ptmlist. Canonical OBO names come from PSI-MOD.obo (bundled) and an optional unimod.obo; without them, names fall back to the UniProt ptmlist.txt ID and a warning is printed.
UniProt alternative products (isoforms) are expanded into their own PEFF entries: every isoform whose sequence is "described" by <feature type="splice variant"> records is reconstructed from the canonical sequence (a referenced region without a replacement is deleted, one with a <variation> is substituted) and emitted directly after its parent entry, e.g. >sp:P02768-2 \PName=(Isoform 2 of Albumin) ..., with gene/taxonomy/mnemonic metadata inherited from the parent (mirroring UniProt's isoform FASTA convention, which carries no per-isoform PE/SV). The isoform flagged "displayed" is identical to the canonical sequence and is not emitted again; isoforms typed "external" or "not described" have no reconstructable sequence and are reported in the log. An isoform entry carries the features UniProt annotates on that isoform (<location sequence="...">, see above), with ids and \DisulfideBond connectivity as for canonical entries; the canonical entry's annotations are not mapped onto the spliced sequence. Disable isoform expansion with -omit_isoforms.
Both plain .xml and .xml.gz UniProt inputs are accepted (gzip is auto-detected by the underlying parser).
The command line parameters of this tool are:
UniPEFF -- Convert a UniProtKB XML protein database to PEFF 1.0 with rich annotations.
Full documentation: http://www.openms.de/doxygen/nightly/html/TOPP_UniPEFF.html
Version: 3.6.0-pre-nightly-2026-09-29 Sep 30 2026, 01:45:35, Revision: 55f7bdb
To cite OpenMS:
+ Pfeuffer, J., Bielow, C., Wein, S. et al.. OpenMS 3 enables reproducible analysis of large-scale mass spec
trometry data. Nat Methods (2024). doi:10.1038/s41592-024-02197-7.
Usage:
UniPEFF <options>
Options (mandatory options marked with '*'):
-in <file>* Input UniProtKB XML file (plain or gzip; gzip is auto-detected). (valid
formats: 'xml')
-out <file>* Output PEFF 1.0 file. (valid formats: 'peff')
-ptmlist <file> UniProt ptmlist.txt; defaults to the bundled snapshot. (valid formats:
'txt')
-psimod_obo <file> PSI-MOD.obo for canonical modification names; defaults to the bundled Open
MS PSI-MOD.obo. (valid formats: 'obo')
-unimod_obo <file> Optional unimod.obo for canonical Unimod names; if absent, names fall back
to the UniProt ptmlist IDs. (valid formats: 'obo')
-prefix <string> Force a single PEFF prefix for every entry (e.g. 'sp'); if empty, sp/tr
is derived from the UniProt dataset.
-dbversion <string> Value for the mandatory '# DbVersion=' PEFF header line. (default: 'unknow
n')
-annotation_identifiers Emit PEFF Option B: assign a global sequential id: prefix to every annotat
ion tuple, referenced by \DisulfideBond. By default only the \DisulfideBon
d tuples and the half-cystines they reference get ids: of K bonds, bond k
(counting from 0) labels its half-cystines 2k and 2k+1 and is itself label
ed 2K+k.
-omit_molecular_processing Skip the \Processed annotations (initiator methionine, signal/transit pept
ide, propeptide, chain).
-omit_amino_acid_modifications Skip \ModResPsi / \ModResUnimod / \ModRes and \DisulfideBond; ptmlist is
not read.
-omit_sequence_variations Skip \VariantSimple and \VariantComplex annotations.
-omit_isoforms Skip the additional PEFF entries for UniProt isoforms (alternative product
s whose sequences are reconstructed from splice-variant features).
Common TOPP options:
-ini <file> Use the given TOPP INI file
-threads <n> Sets the number of threads allowed to be used by the TOPP tool (0 = all
available cores) (default: '1')
-write_ini <file> Writes the default configuration file
--help Shows options
--helphelp Shows all options (including advanced)
INI file documentation of this tool:
This section lists all parameters supported by the tool. Parameters are organized into hierarchical subsections that group related settings together. Subsections may contain further subsections or individual parameters.
Each parameter entry contains the following information:
Parameter tags provide additional information about how a parameter is used. Some tags indicate whether a parameter is required or intended for advanced configuration, while others may be used internally by OpenMS or workflow tools.
Parameters highlighted as required must be specified for the tool to run successfully. Parameters marked as advanced allow fine-tuning of algorithm behavior and are typically not needed for standard workflows.