AccueilGlossairePrimary Structure

Primary Structure

Definition

The primary structure of a peptide or protein designates the first level of structural organization: the linear amino acid sequence linked by covalent peptide bonds. It is the most fundamental backbone, entirely defined by the N-terminus → C-terminus sequence without considering three-dimensional folding. It constitutes the peptide's identity charter, and all structural information at higher levels (secondary, tertiary, quaternary) stems from this primary sequence.

Four hierarchical levels organize protein structure. The primary level, subject of this article, corresponds to the linear sequence text — for example Gly-Glu-Pro-Pro-Pro-Gly-Lys-Pro-Ala-Asp-Asp-Ala-Gly-Leu-Val for BPC-157. The secondary level groups regular motifs (alpha-helix, beta-sheet, beta-turn) formed by local hydrogen bonds between nearby residues in the sequence. The tertiary level describes the peptide's global spatial folding, stabilized by long-range interactions (hydrophobic, ionic, disulfide bridges). The quaternary level, reserved for multimeric proteins, assembles multiple distinct chains into a functional complex.

The primary structure carries critical information about peptide properties. Net charge at physiological pH (7.4) is deduced from counts of basic residues (Lys, Arg, partially His) and acidic ones (Asp, Glu). This charge influences solubility in aqueous solution, isoelectric point (pI), and electrophoretic behavior. Global hydrophobicity (GRAVY score) predicts solubility in organic solvents and precipitation risk. Presence of oxidizable (Met, Cys, Trp), deamidable (Asn, Gln), or labile residues (Asp-Gly cycloisomerization) identifies chemical vulnerability points to monitor during storage.

Beyond chemistry, the primary structure directly encodes biological activity. An RGD motif in the sequence signals a potential ligand for membrane integrins. A C-terminal -GKKK-NH2 motif suggests a hormonal amidated peptide. The "YGGF" sequence identifies endogenous opioid enkephalins. Modern bioinformatics tools (BLAST, Pfam, InterPro) compare an unknown sequence against millions of annotated sequences and deduce within seconds the probable functional family, conserved domains, and structural homologs.

For the researcher, mastering the primary structure is an absolute prerequisite. Before any experimentation, the sequence must be confirmed on the COA (observed mass = theoretical mass, LC-MS/MS sequencing). Any reconstitution, storage protocol, or result interpretation depends on it. A misread sequence (N/C terminal inversion, Leu/Ile confusion, omission of a post-translational modification such as C-terminal amidation) can invalidate a week of experiments. Always check that the sequence and mass stated on the COA match the expected primary structure.