Evidence Infrastructure
How to Use Peptide Sequence Databases Without Losing Context
A peptide sequence database can connect a string to an organism, precursor, processing annotation, stable identifier, and supporting evidence. The strength of that connection depends on the underlying record.
Educational content only. Not medical advice.
A sequence string needs an organism and molecular context
Short amino-acid strings can occur in many proteins or organisms. A reliable search records the sequence alphabet, length, organism or taxon, expected precursor, modification state, and whether ambiguity codes are present. Exact matching can connect a peptide to candidate proteins, but it does not prove which biological source produced a measured signal. Shorter sequences have more chance matches and require stronger contextual evidence.
Canonical protein, isoform, and mature chain are different records
UniProt's canonical sequence is the coordinate reference for an entry, while alternative isoforms may add or remove regions. Secreted peptide hormones can be annotated as mature chains cut from a larger precursor. NCBI RefSeq and GenBank records may track transcripts, genomic annotations, and translated proteins with versioned accessions. Copying the entire precursor when a study names a mature peptide creates a sequence-identity error.
Stable identifiers are more durable than display names
Gene symbols, protein names, and peptide aliases can change or collide. An accession plus version, organism, database, retrieval date, and feature coordinates creates a reproducible citation. Cross-references connect UniProt, RefSeq, Ensembl, PDB, and literature, but records may be curated or predicted at different levels. A database link should preserve its evidence status rather than flatten every entry into confirmed biology.
Database search supports identification but does not finish it
A match can generate hypotheses for mass-spectrometry spectra, antibody targets, precursor processing, or comparative biology. Experimental identity also depends on sample provenance, measurement quality, modifications, sequence coverage, and appropriate controls. Databases are updated, so versioned exports or access dates matter. This page explains evidence infrastructure and does not validate a vendor label or material from a sequence match alone.
Evidence limits
- Database records differ in curation, prediction status, versioning, and feature annotation.
- A sequence match does not prove sample origin, expression, processing, or chemical identity.
- This workflow cannot authenticate a commercial or personal-use material.
Sources and further reading
These sources ground the definitions and evidence boundaries on this page. A citation is a route for verification, not an endorsement of a product or personal use.
UniProt Consortium
Peptide Search
Official guidance for matching peptide strings against UniProtKB with organism restrictions and programmatic access.
Open sourceNational Center for Biotechnology Information
NCBI Protein
Official sequence collection integrating RefSeq, GenBank translations, Swiss-Prot, PDB, and other sources.
Open sourceCommon questions
What is a canonical protein sequence?
It is the database's main coordinate reference for an entry, not proof that it is the only biological isoform or mature product.
Why include an accession version?
Versions preserve which sequence record was used when databases later correct or update annotations.
Can a short peptide match multiple proteins?
Yes. Short strings can occur by chance or in conserved regions, so organism and precursor context are essential.
Continue with context
Continue in the evidence workspace
Explore the complete PeptideSchool research workspace to organize sources, compare evidence layers, and follow related peptide science. Premium tools remain educational and do not provide individualized medical guidance.