PepShop Output Documentation




A. Query Prohormone Database


The PepShop database can be searched by UniProt accession number, gene symbol, organism name (7 species), exact amino acid sequence, and peptide monoisotopic mass with adjustable mass tolerance level. The output of this search is a list of prohormones meeting the search criteria.

Description of the results columns:

Match number:
The number of prohormones matching the search criteria.
Symbol:
The gene symbol of the matched prohormone.
Species:
Common english name of the organism.
UniProt Accn:
The primary accession number of prohormone in UniProt database.
NCBI Gene:
The primary accession number of prohormone in NCBI Gene database.
Number of peptides:
The total number of peptides in the searched prohormone sequence.
Experimental peptides:
The source of the reported peptides in prohormone sequences. The 'SwePep' denotes experimentally verified peptides in SwePep database. All other peptides are labelled as 'Unknown'.
Details:
The further details of the listed prohormone including sequence, peptides, cross-references to public databases including UniProt and NCBI, and calculated isotopic masses and isoelectric point of prohormones.
Select:
PepShop allow user to select one or more prohormones and perform cleavage site prediction (NeuroPred), pairwise sequence alignment (blastp) and mutiple sequence alignment (Muscle) to gain additional insight on prohormones and neuropeptides.

1. Prohormone Details


The Details column in above result table provides further information for each prohormone. The information includes: sequence of prohormone, cross-references to public databases, prohormone properties such as calculated isotopic masses and isoelectric point, and list of peptides.

Description for the prohormone details page:

Prohormone Information:
The entry name, prohormone symbol, gene symbol, gene name and common organism name to identify each prohormone.
Cross-References:
Each prohormone is further linked to UniPort, NCBI-Gene, NCBI-Protein, UniGene, Ensembl, and Genome Databases using primary accession number of prohormone in each database.
Protein sequence:
The sequence of each prohormone.
Protein properties:
The sequence length, calculated isotopic masses, isoelectric point and number of peptides in each prohormone.
Peptides:
The location in prohormone, sequence, sequence length, calculated isotopic masses, isoelectric point and source for all peptides in prohormone. Each peptide is linked to peptide database in PepShop.

2. Peptide Details


Each peptide of prohormone is further linked to peptide details page. The peptides page provides peptide name, sequence, source, and spectrum information (for SwePep peptides).

Description for the entries on peptide page:

Peptide Information:
The peptide information section provides information about prohormone name, species of origin, prohormone symbol, and location of peptide in prohormone.
Source:
The source of listed peptide. The spectrum information of each peptide can be viewed in the SwePep database.
Sequence:
The sequence of each peptide.
Properties:
The sequence length, calculated isotopic masses, isoelectric point, and number of occurrences of each peptide in prohormone.

3. SwePep Information


The spectrum information of each SwePep peptide can be viewed using view spectra link on peptides detail page. The SwePep page provides list of all spectra reported for the selected peptide. The SwePep information of TKN1_MOUSE[58-64] is shown below.
Screenshot of SwePep input

Description of columns on SwePep page:

log(e):
The log-transformed expectation value of X!Tandem.
m+h:
The observed mass of peptide.
delta:
The difference between observed and calculated mass of each peptide.
z:
The peptide ion charge state.
peptide:
The sequence of peptide and its location in prohormone sequence.


The peptide column links to spectrum page that provides information about prohormone, plot of m/z vs. intensity of each MS/MS peak, and a table of observed b and y ions. An example is shown below. The spectrum of TKN1_MOUSE[58-64] peptide is shown below.
Screenshot of the output of a SwePep search

D. NeuroPred


1. Cleavage Prediction Diagram


The cleavage prediction diagram is provided for all Output Selection Tasks except for Print Probabilities of Basic Sites only task. Each sequence entered is automatically converted to upper case and split into groups of a maximum of 50 amino acids. Each group is presented in sequence order as follows: The first column, Sequence, denotes that first line is the sequence and contains up to five blocks in which each block holds a maximum of ten amino acids. Immediately below this line is another line for each selected model, and the final line consists of a consensus report of all models. This is repeated until the sequence is completely shown. For the model and consensus lines, a series of "s" is provided to indicate the signal sequence determined by either the global default value or sequence specific values. For each model, sites where the cleavage probability for that model exceeded the threshold probability are denoted with the letter "C" below the site while non-cleaved sites are designated by a period ".". The consensus line is defined for each site as "C" if at least one model predicted cleavage or "." if all models did not predict cleavage. By default, the rules of Amare et al. 2006 and Southey et al. 2008 are implemented and the resulting redundant sites are denoted by an 'r'. This symbol will not appear when the Ignore processing rules option in the Advanced Options Interface is set to No.

2. Mass of Predicted Peptides


When the Obtain Mass of Predicted peptides task is selected, the original sequence is cleaved based on predicted values for each model. The resulting peptides are extended by an order of 2. Only peptides where the selected post-translational modifications have been successfully applied are reported using small blue brackets (e.g. PQQFFGLM[Amide] denotes amidation of the peptide PQQFFGLM). The average and monoiostopic masses of the predicted peptides are calculated at all stages such that masses are available for every combination of PTM to address the possibility that some PTMs may be absent. Note that the standard mass or molecular weight, not the MH+ or M+H mass, is calculated. The results are presented in the Mass of Predicted Peptides table:

Description of the columns for the Mass of Predicted Peptides table

Abb. Peptide:
Abbreviated Peptide where the start and the end of the peptide is provided by a single letter amino acid code and location within the sequence.
NCut:
The model or models that predicted cleavage at this site that corresponds to the start or N-terminal region of the peptide. Signal Peptidase or Sequence Start are used to denote the start of the peptide whether a signal peptide is indicated or not.
CCut:
The model or models that predicted cleavage at this site that corresponds to the end or C-terminal region of the peptide. Sequence End is used to denote the end of the sequence.
PTM Applied:
Provides which PTMs have been applied to the peptide. Most PTM designations are self-explanatory; however, "Cleaved" denotes that the peptide has only been cleaved, "Extended" denotes that adjacent peptides have been joined, and "TrimKR" indicates that all the C-terminal K and R amino acids have been removed after cleavage. When a PTM is applied multiple times, only the mass with all occurrences is reported and the frequency is also reported in the "PTM Applied" column by indicating the number of PTM occurrences, followed by the times symbol ("x"), (e.g., Acetylation 3x).
Predicted Aver. Mass:
Predicted average mass of the peptide including any PTMs.
Predicted Mono. Mass:
Predicted monoisotopic mass of the peptide including any PTMs.
Peptide Sequence:
Complete sequence of the peptide. Any post-translational modifications that have been successfully applied are reported using small blue brackets (e.g. PQQFFGLM[Amide] denotes amidation of the peptide PQQFFGLM).

E. Multiple Sequence Alignment Using Muscle


The Multiple sequence alignment displays CLUSTALW format with the peotein name in the left column and aligned sequence in the right column.


F. Local Pairwise Sequence Alignment Using BLASTP


Blastp program searches protein databases using protein query. The matched protein sequences from the prohormone database are arragned in ascending order of there E-value with more significant hits are at the top of the list.
NIDA Logo
Bioinformatic Portal is supported by NIH/NIDA Grants: P30 DA 018310 and R21 DA027548
NIH Logo