Research use only. All Vinnix products are for laboratory research only. Not for human or veterinary use.

Research Library: Peptide Fundamentals

Understanding Peptide Sequences

How to read a peptide sequence: one-letter and three-letter codes, N- and C-termini, and the notation used for modified residues.

Quick answerA peptide sequence is the exact order of amino acids in a peptide chain, written from the N-terminus to the C-terminus using standard one-letter or three-letter codes. The peptide sequence is the primary identifier of a peptide: two chains with the same residues in a different order are different compounds.
Vinnix Research TeamUpdated October 6, 20264 min read3 references
Understanding Peptide Sequences illustration

Key facts

Key facts
Reading direction
N-terminus (left) to C-terminus (right)
Code systems
Three-letter (Gly) and one-letter (G), defined by IUPAC-IUB
Standard residues
20, plus selenocysteine (U) and pyrrolysine (O)
C-terminal amide
Written -NH2
N-terminal acetyl
Written Ac-
Example
BPC-157: GEPPPGKPADDAGLV (15 residues)

What Is a Peptide Sequence?

A peptide sequence is the ordered list of amino acids in a peptide, first residue to last. Chemists also call it the primary structure: the first level of description, before folding enters the picture [2].

If you could know only one thing about a peptide, the sequence would be the one to pick. With any modifications, it fixes the molecular formula and so the expected molecular weight, and that weight is exactly what a mass spectrometer checks when it confirms identity. The sequence also sets the chain's chemistry. It decides which residues can oxidize, which positions carry charge, and where enzymes or acids are likely to cut.

One-Letter and Three-Letter Codes

Each standard amino acid has a three-letter abbreviation and a one-letter code, both set by the IUPAC-IUB Joint Commission on Biochemical Nomenclature [1].

Three-letter codes are easier to read and are joined by hyphens: Gly-His-Lys. One-letter codes are compact and run together without separators: GHK. Most one-letter codes are simply the first letter of the name. Where names clashed, the committee picked other letters, so lysine is K, tryptophan W, tyrosine Y and glutamic acid E.

Amino acid codes (IUPAC-IUB)
Amino acid Three-letter One-letter
Alanine Ala A
Arginine Arg R
Asparagine Asn N
Aspartic acid Asp D
Cysteine Cys C
Glutamine Gln Q
Glutamic acid Glu E
Glycine Gly G
Histidine His H
Isoleucine Ile I
Leucine Leu L
Lysine Lys K
Methionine Met M
Phenylalanine Phe F
Proline Pro P
Serine Ser S
Threonine Thr T
Tryptophan Trp W
Tyrosine Tyr Y
Valine Val V
Selenocysteine Sec U
Pyrrolysine Pyl O

Databases add a few extra symbols. B stands for Asp or Asn when the two can't be told apart, Z for Glu or Gln, and X for an unknown or unspecified residue. Our peptide glossary lists them with short definitions.

N-Terminus and C-Terminus

Peptide sequences are written from the N-terminus, the end with a free amino group, to the C-terminus, the end with a free carboxyl group [1].

The convention comes from the chemistry of the peptide bond. Every bond in the chain points the same way, so the chain has a built-in direction. Reading MOTS-c as MRWQEMGYIFYPRKLR puts methionine at the N-terminus and arginine at the C-terminus [3]. Reverse the string and you get RLKRPYFIYGMEQWRM: a different peptide, yet with the same composition and the same mass. That has a practical consequence. Mass spectrometry alone can't tell a sequence from its reverse, or from any other rearrangement with identical composition, so fragmentation (MS/MS) or other evidence is needed when order has to be confirmed.

Positions are counted from the N-terminus, so "position 2" always means the second residue from the amino end. Fragment names like GHRH (1-29) follow the same numbering. Sermorelin, for instance, corresponds to residues 1 to 29 of growth hormone-releasing hormone, with a C-terminal amide.

How Are Peptide Modifications Written?

Modifications go into the sequence as standard prefixes, suffixes and abbreviations. Each one changes the molecule's formula and mass, so leaving it out gives the wrong reference value.

Common notation for modified peptides
Notation Meaning Example from the Vinnix library
-NH2 at the C-terminus C-terminal amide instead of a free acid Sermorelin: ...DIMSR-NH2
-OH at the C-terminus Free carboxylic acid, stated explicitly PT-141: ...Lys]-OH
Ac- at the N-terminus N-terminal acetylation PT-141: Ac-Nle-cyclo[...]
D- prefix (or lowercase one-letter code) D-amino acid, the mirror image of the natural L form SS-31: D-Arg-Dmt-Lys-Phe-NH2
Aib 2-aminoisobutyric acid, a non-standard residue Ipamorelin: Aib-His-D-2-Nal-D-Phe-Lys-NH2
cyclo[...] Ring closed between the bracketed residues PT-141 (lactam ring)
Named N-terminal group Acyl group added to the N-terminus Tesamorelin: trans-3-hexenoyl group on residue 1

Two cautions. Lowercase one-letter codes for D-residues are a common convention, not a universal rule, so a careful record spells out D- forms in three-letter notation. And any residue outside the standard 20, such as Dmt (2',6'-dimethyltyrosine) in SS-31 or Nle (norleucine) in PT-141, should be defined next to the sequence. Longer modified peptides get a backbone sequence plus a list of modifications; retatrutide, with Aib at positions 2 and 13 and a fatty-acid side chain on lysine 20, is a good example.

Why Sequence Matters

Sequence matters because it defines molecular identity. Two peptides of the same length, or even the same composition, are different compounds once the order of their residues differs.

  • Identity. GHK and KHG contain the same three amino acids, yet they are different molecules. A product record that gives only a name or a length is incomplete without the sequence.
  • Chemistry. Sequence decides which degradation routes are open. MOTS-c has two methionines that can oxidize, and the oxidized forms can show up as related peaks in HPLC testing.
  • Formula and mass. The sequence plus modifications gives the expected molecular formula, which is the reference value for identity testing by mass spectrometry.
  • Naming. Different compounds can share similar names or origins. Never treat a short fragment, an analogue and the full-length peptide as interchangeable.

Reading a sequence on a product page

Check the direction (N to C), the code system, any terminal groups (Ac-, -NH2, -OH) and definitions for non-standard residues. Then compare the stated formula and molecular weight with the mass reported on the certificate of analysis.

Each compound page in the Vinnix compound library lists the sequence alongside the formula, molecular weight and CAS number, so you can check the identity information against the batch documentation.

FAQFrequently asked questions

What is a peptide sequence?

A peptide sequence is the order of amino acids in a peptide chain, written from the N-terminus to the C-terminus. It's the peptide's primary structure and its main identifier. Add any modifications and the sequence gives you the molecular formula and expected mass, which is what analytical labs use to confirm identity.

How do you read a one-letter peptide sequence?

Read it left to right, starting at the N-terminus, and translate each capital letter with the IUPAC-IUB code. G is glycine, E is glutamic acid, P is proline and K is lysine, for example. So GEPPPGKPADDAGLV, the BPC-157 sequence, starts with glycine and ends with valine.

Why is lysine K and not L?

Because L was already taken by leucine. Where names clashed, the code committee picked other letters: lysine became K, tryptophan W, tyrosine Y, glutamic acid E, glutamine Q, aspartic acid D, asparagine N, arginine R and phenylalanine F. These assignments come from the IUPAC-IUB nomenclature recommendations.

What does -NH2 mean at the end of a peptide sequence?

An -NH2 after the last residue means the C-terminus is an amide, not a free carboxylic acid. That lowers the mass by about 1 Da compared with the free acid and removes a negative charge. Sermorelin, ipamorelin and SS-31 are all C-terminally amidated peptides.

Can two peptides have the same mass but different sequences?

Yes. Any rearrangement of the same residues, the reversed sequence included, has the same formula and the same intact mass. Intact-mass measurement therefore confirms composition only. To confirm order you need tandem mass spectrometry fragmentation, or solid knowledge of how the material was synthesized.

REFScientific references

  1. IUPAC-IUB Joint Commission on Biochemical Nomenclature (JCBN). Nomenclature and symbolism for amino acids and peptides. Recommendations 1983. Eur J Biochem. 1984;138(1):9-37. PubMed 6692818
    nomenclature standard
  2. Alberts B, Johnson A, Lewis J, Raff M, Roberts K, Walter P. Molecular Biology of the Cell. 4th ed. New York: Garland Science; 2002. The Shape and Structure of Proteins. NCBI Bookshelf NBK26830. Source
    textbook (NCBI Bookshelf)
  3. Lee C, Zeng J, Drew BG, et al. The mitochondrial-derived peptide MOTS-c promotes metabolic homeostasis and reduces obesity and insulin resistance. Cell Metab. 2015;21(3):443-454. PubMed 25738459
    in vitro and animal study (cited here for sequence origin only)

Research use only. Vinnix products are supplied for laboratory, analytical and scientific research. They are not for human or veterinary use, consumption, diagnosis or treatment. Information on this page is educational and is not a claim about any effect of any product. See the Product & Research Information Disclosure.

Keep exploring

Related resources

Research Library: Peptide Fundamentals

What Are Peptides?

What are peptides, and how do amino acids and peptide bonds put them together? This plain-English guide also covers how…

Research Library: Peptide Fundamentals

What Is a Peptide Bond?

How a peptide bond forms, why it's flat and rigid, and why it holds peptide chains together so dependably.

Research Library

What Is Mass Spectrometry?

Mass spectrometry turns peptides into ions, measures their mass-to-charge ratio and tells you whether a sample is the…

Reference

Peptide Glossary

A peptide glossary of plain-English definitions, A to Z, for the terms you'll meet in peptide chemistry, analytical…

Vinnix BPC-157 Peptide research vialPeptides

BPC-157 Peptide

BPC-157, a synthetic 15-amino-acid peptide: its sequence, molecular data, published research sorted by study type, and…

Compound Library

Ipamorelin

Ipamorelin is a synthetic pentapeptide that Novo Nordisk scientists first described in 1998. This entry covers its…

Get Vinnix research updates

New compounds, batch documentation and research-library articles. No spam, unsubscribe any time.