Key facts
- Three-letter codes
- Mostly first three letters (Gly, Ala); hyphen = peptide bond
- One-letter codes
- Single capitals, no separators (GHK)
- Standardized by
- IUPAC-IUB: one-letter rules 1968, full recommendations 1983
- Irregular three-letter
- Asn, Gln, Ile, Trp
- Ambiguity symbols
- B (Asx), Z (Glx), X (unknown)
- Reading direction
- N-terminus to C-terminus
Two Codes for the Same Amino Acids
The two codes are interchangeable. Gly-His-Lys and GHK name exactly the same tripeptide.
The IUPAC-IUB biochemical nomenclature commissions standardized both. One-letter notation arrived as tentative rules in 1968, when protein sequences were getting too long to print and compare comfortably in three-letter form [1]. The 1983 recommendations then fixed the three-letter symbols along with conventions for writing peptides, modifications and termini [2]. You'll find the full code table in our peptide sequences guide. Here the focus is on how the two systems differ and when to use which.
How Three-Letter Codes Work
Most three-letter codes are just the first three letters of the amino acid's name, capitalized at the start and linked by hyphens that stand for peptide bonds.
Sixteen of the twenty standard codes follow the first-three-letters pattern: Ala, Arg, Asp, Cys, Glu, Gly, His, Leu, Lys, Met, Phe, Pro, Ser, Thr, Tyr, Val and so on. Four break it to avoid clashes or ambiguity: Asn (asparagine), Gln (glutamine), Ile (isoleucine) and Trp (tryptophan) [2]. Three-letter symbols are easy to extend, so non-standard residues fit in without fuss: Aib for 2-aminoisobutyric acid, Nle for norleucine, pGlu for pyroglutamic acid, and a D- prefix for D-amino acids.
How Were the One-Letter Codes Chosen?
Where an initial is unique, it's used. Where initials clash, the code borrows a phonetic or nearby letter.
| Amino acid | Three-letter | One-letter | Reason commonly given |
|---|---|---|---|
| Arginine | Arg | R | Sounds like 'aRginine'; A is alanine |
| Asparagine | Asn | N | Contains N ('asparagiNe') |
| Aspartic acid | Asp | D | 'asparDic'; close to B |
| Glutamic acid | Glu | E | 'glutEmic'; G is glycine |
| Glutamine | Gln | Q | 'Q-tamine' (sounds like cute) |
| Lysine | Lys | K | Near L in the alphabet; L is leucine |
| Phenylalanine | Phe | F | Sounds like 'Fenylalanine' |
| Tryptophan | Trp | W | Double ring, 'tWyptophan' |
| Tyrosine | Tyr | Y | 'tYrosine'; T is threonine |
The other standard residues get their initials: A, C, G, H, I, L, M, P, S, T and V. The 1968 rules also set aside B for aspartic acid or asparagine when the two can't be told apart, Z for glutamic acid or glutamine, and X for an unknown or unspecified residue [1]. Sequence databases later added U for selenocysteine and O for pyrrolysine. The mnemonics in the table are memory aids, not part of the rules.
One-Letter vs Three-Letter: When to Use Each
Pick three-letter codes when readability and modifications matter. Pick the one-letter code when length and machine processing matter.
| Situation | Better choice | Why |
|---|---|---|
| Short peptide in text | Three-letter | Readable; hyphens show each bond |
| Long sequence (20+ residues) | One-letter | Compact; easier to align and compare |
| Database or search query | One-letter | Standard input format for sequence tools |
| Non-standard or D-residues | Three-letter | Extends cleanly (Aib, D-Phe, Nle) |
| Modified termini | Either, with suffixes | Ac-, H-, -OH, -NH2 work with both |
| Certificate of Analysis | Often both | One-letter for compactness, three-letter for clarity |
Wherever possible, pages in the Vinnix Compound Library give both. Take BPC-157: Gly-Glu-Pro-Pro-Pro-Gly-Lys-Pro-Ala-Asp-Asp-Ala-Gly-Leu-Val in three-letter code, GEPPPGKPADDAGLV in one-letter code.
Where One-Letter Codes Fall Short
Most non-standard residues, D-amino acids and many modifications have no direct one-letter form. That gap is why synthetic peptides are still written in three-letter notation.
Residues with no standard one-letter code
Ipamorelin is Aib-His-D-2-Nal-D-Phe-Lys-NH2. Neither Aib nor D-2-naphthylalanine has a standard one-letter code, so a one-letter version would need made-up symbols. SS-31 (D-Arg-Dmt-Lys-Phe-NH2) combines a D-amino acid with 2',6'-dimethyltyrosine, and PT-141 is cyclic. Some authors put D-residues in lowercase (r for D-arginine, say), but the convention is informal and needs defining wherever it appears.
Turning a modified sequence into plain one-letter code can quietly drop D-configuration, terminal amides or non-standard residues. Carry terminal groups and modifications into both notations, every time.
How Do You Convert Between the Two Codes?
Conversion is a residue-by-residue swap. The catch is that terminal groups and modifications have to be carried across by hand.
- Write the sequence out from the N-terminus, one residue at a time.
- Swap each three-letter symbol for its one-letter code, or the other way round, using the IUPAC-IUB table [1][2].
- Remove the hyphens going to one-letter code; put them in going to three-letter code.
- Keep terminal groups as suffixes or prefixes. KPV stays KPV, but a C-terminal amide has to stay -NH2 in both forms.
- Flag any residue that lacks a standard one-letter code instead of forcing a substitute.
- Count residues in both versions to make sure none went missing.
Lys-Pro-Val goes straight to KPV, and Gly-His-Lys to GHK. Sermorelin, YADAIFTNSYRKVLGQLSARKLLQDIMSR-NH2 in one-letter code, keeps its -NH2 when you expand it to three-letter code.
Common Reading Errors
The usual mistake is assuming a letter stands for the amino acid whose name starts with it.
- K is lysine, not potassium or anything starting with K.
- D is aspartic acid and E is glutamic acid, while N and Q are their amide forms, asparagine and glutamine.
- W is tryptophan and Y is tyrosine, not tyrosine and tryptophan.
- F is phenylalanine, not a placeholder.
- Reading right to left reverses the peptide. Both notations run from the N-terminus on the left to the C-terminus on the right [2].
A reversed or mistranslated sequence can share the composition and mass of the intended peptide, so a mass check by itself may miss it. Fragmentation in mass spectrometry will catch it, because fragment ions are read from each end of the chain [3].
FAQFrequently asked questions
Why is lysine K in one-letter code?
Leucine had already claimed L, so lysine got K, the letter right before L in the alphabet. The one-letter system, issued as IUPAC-IUB tentative rules in 1968, used initials where they were unique and fell back on phonetic or nearby letters when several amino acid names started with the same letter.
Which is better, one-letter or three-letter code?
Neither, in general. Three-letter code reads more easily and handles non-standard residues and modifications cleanly, which makes it a good fit for short synthetic peptides. One-letter code is compact and is the standard for databases and long sequences. Plenty of documents give both so readers can use whichever suits them.
What do B, Z and X mean in a protein sequence?
B means aspartic acid or asparagine when the analysis can't separate them, Z means glutamic acid or glutamine, and X means an unknown or unspecified residue. These ambiguity symbols date from the 1968 one-letter rules and still turn up in sequence databases, particularly for older or partial sequences.
Do the hyphens in three-letter code mean anything?
Yes. Each hyphen between residue symbols stands for a peptide bond, so Gly-His-Lys shows three residues and two bonds. Hyphens also attach terminal groups, as in H- for a free amino terminus or -NH2 for a C-terminal amide. Leave them out and that information is lost.
How do you write D-amino acids in sequence notation?
In three-letter notation, put D- in front, as in D-Phe for D-phenylalanine. Anything unmarked is assumed to be L. There's no official one-letter symbol for D-residues. Some authors use lowercase letters, but if you do that, define the convention in the document so nobody misreads it.
Which direction are amino acid sequences read?
Left to right, N-terminus to C-terminus, in both one-letter and three-letter notation. That matches the direction in which ribosomes build proteins. Read a sequence backwards and you're describing a different peptide, even though it contains exactly the same amino acids.
REFScientific references
-
IUPAC-IUB Commission on Biochemical Nomenclature. A one-letter notation for amino acid sequences. Tentative rules. J Biol Chem. 1968;243(13):3557-3559. PubMed 5658538
nomenclature standard -
IUPAC-IUB Joint Commission on Biochemical Nomenclature (JCBN). Nomenclature and symbolism for amino acids and peptides. Recommendations 1983. Eur J Biochem. 1984;138(1):9-37. PubMed 6692818
nomenclature standard -
Roepstorff P, Fohlman J. Proposal for a common nomenclature for sequence ions in mass spectra of peptides. Biomed Mass Spectrom. 1984;11(11):601. PubMed 6525415
nomenclature proposal (mass spectrometry)

Research Library: Peptide Fundamentals