Research use only. This is an independent literature notebook about laboratory peptide chemistry. It contains no dosing guidance, no purchasing information and no advice of any kind for humans or animals.

Peptide and Protein Structure: Amino Acids, Peptide Bonds and Beyond

Written by Marion Kessler · Reviewed by Douglas ReyesIndependent compilation of public literature; no institutional affiliation.
Corrections welcome. If a summary here misreads a paper, a note with the citation gets it fixed in the next pass. The notebook answers questions about its own sources only - never about sourcing, dosing or anything clinical.

Ask a room of first-year students what the structural unit of peptides and proteins is and you get the same answer almost every time: the amino acid. I gave that answer myself for years, and I still think it is roughly right. What I no longer think is that it is sufficient. This page is my attempt to write down the version I wish I had been given, with the measured numbers attached and with the exceptions named in the text rather than deferred to a footnote nobody reads.

A note on scope before anything else. This is a structural chemistry page, not a product page. Everything below concerns laboratory research material and the primary literature that describes it; none of it is a set of instructions for handling in humans or animals, and I write as one compiler of papers rather than as anyone's instructor. Where I give a number, I give the kind of number that appears in a crystallography or NMR table. Where I give an opinion, I mark it as mine.

I keep this page beside my what is ps r3 peptide pillar note, because the residue-and-backbone vocabulary set out here is the vocabulary that pillar assumes. Readers arriving from a search for the structural unit of peptides and proteins usually want the definition first, so the opening section gives one plainly. The later sections are where I argue with that definition, and I think that is the part worth staying for.

The amino acid residue as the structural unit of peptides and proteins

Start with the plain version, because it is the one most people need and it is not wrong. The structural unit of peptides and proteins is the amino acid residue: a central carbon, the alpha carbon, carrying an amino group, a carboxyl group, a hydrogen atom and a side chain. Twenty side chains are encoded in the standard genetic code, and they differ in size, charge, hydrogen-bonding capacity and hydrophobicity. Once residues join into a chain each has lost the elements of water, so the careful word is residue rather than amino acid. The chain has a direction, written from the free amino end to the free carboxyl end, and every sequence database follows that convention.

The distinction between residue and free amino acid sounds pedantic until you start drawing structures. What repeats along a backbone is a three-atom unit, nitrogen, alpha carbon and carbonyl carbon, with the side chain hanging off the middle atom. Among the twenty, two are outliers I always flag. Glycine has hydrogen as its side chain, so it has no beta carbon and far more conformational freedom than anything else. Proline is the harder case: its side chain loops back and bonds to the backbone nitrogen, making it a cyclic imino acid rather than a standard amine, removing the amide hydrogen, and locking the preceding torsion angle.

The count of twenty is itself a teaching approximation, and I think it should be labelled as one. Selenocysteine and pyrrolysine are genetically encoded in some organisms. Many natural peptides carry D-configured residues or N-methylated backbone nitrogens installed enzymatically rather than by the ribosome. None of that overturns the residue model, but it does mean the model describes ribosomal proteins better than it describes peptides in general, which matters here because most of the material I read about is synthetic or non-ribosomal. Everything on this page concerns laboratory research material only.

  • What the residue model gives me: a way to write any sequence as a string, a reproducible molecular weight, and a predictable set of backbone atoms.
  • What it hides: side-chain modification, backbone N-methylation, D-configuration and backbone cyclisation, none of which change the residue letter.
  • What it cannot predict: whether the chain folds into one arrangement, into an ensemble, or not at all.

The peptide bond is planar, and resonance is the reason

The most useful number in this note is a bond length. The carbon-nitrogen bond of a peptide group measures about 1.32 angstrom, which sits between a carbon-nitrogen single bond near 1.47 angstrom and a carbon-nitrogen double bond near 1.27 angstrom. That intermediate value is the physical evidence for what every textbook asserts: the amide nitrogen lone pair is delocalised into the carbonyl, so the bond carries partial double-bond character. The consequence is geometric. The carbonyl carbon, the oxygen, the nitrogen and its hydrogen lie in one plane, and rotation about that bond costs roughly 20 kilocalories per mole.

Planarity is what makes the omega angle worth naming. Omega describes rotation about the peptide bond itself and takes one of two values: trans near 180 degrees, with the alpha carbons on opposite sides, or cis near 0 degrees. For every pair except those involving proline, trans is favoured by roughly a thousand to one, because cis stacks the adjacent side chains on top of each other. Proline narrows that gap, since its five-membered ring makes trans only modestly better than cis. cis X-Pro bonds therefore appear at a few percent, and interconverting them is slow enough that dedicated isomerase enzymes exist to catalyse it.

Two things follow that I did not absorb for years. First, a residue is not a free rotor: with omega in practice locked, each residue contributes only two rotatable backbone angles, phi and psi, and the chain is already stiff before any side chain is considered. Second, the locked geometry fixes hydrogen-bonding directions, with the NH as donor and the carbonyl oxygen as acceptor at constrained angles, and most of secondary structure falls out of that constraint. I read these as measured crystallographic and NMR descriptions of the bond, not as assumptions about any particular molecule.

Peptide bond properties I keep at hand when reading a structure
QuantityTypical valueWhat it tells me
Peptide C-N bond lengthabout 1.32 angstromPartial double-bond character from amide resonance
C-N single / C=N double referenceabout 1.47 / 1.27 angstromThe two values the measured length sits between
Rotation barrier about C-Nroughly 20 kcal per moleThe bond is locked near room temperature
omega, trans / cis180 degrees / 0 degreesTwo discrete states rather than a continuum
cis fraction of peptide bondswell under 1 percent; a few percent for X-ProProline is where cis geometry becomes common

Primary, secondary, tertiary, quaternary: the ladder and where it breaks

Once the structural unit of peptides and proteins is repeated into a chain, textbooks introduce a four-level ladder. Primary is the covalent sequence written N to C. Secondary is the local pattern of backbone hydrogen bonds, alpha helix and beta sheet above all. Tertiary is how one chain folds as a whole. Quaternary is how two or more folded chains arrange themselves. I use the ladder constantly, because it tells me which experiment can see what: sequencing gives primary, far-UV circular dichroism estimates secondary content, crystallography or cryo-electron microscopy gives tertiary and quaternary.

What I try to remember is that the ladder is an index, not a description of nature, and several rungs are optional. Quaternary structure simply does not exist for a large number of single-chain proteins. Tertiary structure, in the sense of one stable folded arrangement, is not what an intrinsically disordered region has, and such regions are now common enough that calling them exceptions is hard to defend. Disulfide bonds are covalent and therefore primary-level facts, yet a sequence string alone does not say which cysteines are paired. Metal centres and glycosylation sit outside all four levels while shaping most of the chemistry.

The non-ribosomal peptides I read about break the ladder more thoroughly still. A backbone-cyclised peptide has no N or C terminus to write a sequence from, so primary structure has to be given as a graph rather than a string, and N-methylation changes hydrogen-bonding capacity without changing the residue letter at all. My working rule is narrow: the ladder tells me what a structural claim is about, and nothing about whether that claim is complete. This is also how I read the papers behind my mt2 peptide pill note, where one residue substitution changes the whole conformational picture.

Level of structure, what each level defines, and what it does not tell you
Level of structureWhat is definedWhat it does not tell you
PrimaryThe covalent residue order from N to C, including disulfide connectivityWhich side chains are chemically modified, and whether the chain folds at all
SecondaryLocal backbone hydrogen-bond patterns: helix, sheet, turnHow those elements pack against each other, or how long they persist
TertiaryThe three-dimensional arrangement of one chainWhether that arrangement is a single state or an ensemble; it is silent on disordered regions
QuaternaryHow two or more chains are arranged relative to each otherNothing at all for the many proteins that function as a single chain

Backbone dihedral angles and the intuition behind a Ramachandran plot

Two rotatable angles per residue is a small enough number to plot, and that is the whole idea. Phi is rotation about the bond from the amide nitrogen to the alpha carbon; psi is rotation about the bond from the alpha carbon to the carbonyl carbon. Ramachandran and colleagues computed in 1963 which combinations are possible when atoms are treated as hard spheres that cannot overlap, and their map still looks essentially the same in modern validation reports. Each residue in a structure becomes one point, and where the points fall tells you at a glance what kind of conformation you are looking at.

The intuition I carry is that the map is mostly a steric map rather than an energetic one. The large empty region in the upper right is empty for one reason: the atoms would clash. A right-handed alpha helix sits near minus 60, minus 45. Beta strands sit in the upper left, around minus 130, plus 130. There is a small mirror-image island near plus 60, plus 45, the left-handed helix, which almost nothing reaches except glycine. Glycine's allowed area is large precisely because it has no beta carbon to collide with anything, and proline's is tiny because the ring holds phi near minus 60.

In practice I use the plot as a first-pass check rather than an oracle. Clusters of outliers in a deposited model usually point to a tracing or refinement problem, while a single outlier in a strained binding site is often real and interesting. Favoured, allowed and outlier percentages are reported routinely now, which is useful, but I keep two limits in mind. The map says nothing about side-chain conformation, and nothing about whether a permitted conformation is actually populated in solution. For that, solution NMR and the disordered-protein literature are the better guide.

  • Where the points cluster tells me the secondary structure faster than any ribbon diagram.
  • Where a single point sits outside the allowed region, I ask whether it is strain or a modelling error before I believe either.
  • Where the map is silent, on side chains and on dynamics, I go to a different experiment rather than pushing the plot further.

Why I think the structural unit phrasing is taught too simply

Here is the opinion I have come to hold, flagged as an opinion. The phrase the structural unit of peptides and proteins invites a one-word answer, and the one-word answer quietly teaches three wrong things: that the unit is invariant, that it is freely jointed, and that it behaves the same wherever it appears. All three are false in ways that matter to anyone reading peptide literature, and I think the simplification survives mostly because it is easy to examine rather than because it is true.

On invariance: the twenty units are not interchangeable parts. Proline has no amide hydrogen, so it cannot donate the hydrogen bond every other residue donates, which is why it interrupts helices. Glycine is so flexible that it is often the residue letting a tight turn close. N-methylated backbones, D-configured residues and backbone cyclisation each remove a capability the standard model quietly assumes. A description of the structural unit of peptides and proteins that cannot represent proline without a caveat is not a description I would hand to someone starting in a laboratory.

On flexibility: if the repeating unit were freely jointed, three torsion angles per residue would give a random coil. In reality resonance locks omega, so two angles per residue carry all the freedom, and the chain is stiff and directional before any side chain is considered. The beads-on-a-string picture most of us were shown is the wrong mental image, and it makes the persistence of secondary structure look more mysterious than it is. I would rather teach the planar amide on day one and let helix and sheet fall out of it.

On context: the same residue is not the same chemical object in two positions, since a buried glutamate and a solvent-exposed one differ in protonation behaviour, and a cysteine in a reducing compartment is not the cysteine in an oxidising one. Add the ensemble problem, that many sequences occupy a distribution of conformations rather than one, and the unit framing captures very little of what the chain actually does. I still use the phrase, and I still give the one-word answer when asked, but with an asterisk attached. The applied side of this vocabulary appears in my best bone growth peptide note. mt2 is a trade-style name used in vendor catalogues; I cite it for structural comparison only and imply no endorsement or affiliation. Everything above concerns laboratory research material only.

Sources & further reading

Search links into public bibliographic databases; the notebook quotes no paywalled full text.

Related notes