- Terminal modifications
- Free N-terminus; C-terminal amide at Ser39
- Structural modifications
- Fatty-acid moiety linked to the peptide backbone
- Stored three-letter notation
- Tyr · Aib · Glu · Gly · Thr · Phe · Thr · Ser · Asp · Tyr · Ser · Ile · Aib · Leu · Asp · Lys · Ile · Ala · Gln · Lys · Ala · Phe · Val · Gln · Trp · Leu · Ile · Ala · Gly · Gly · Pro · Ser · Ser · Gly · Ala · Pro · Pro · Pro · Ser · NH2
- Sequence notes
- Thirty-nine residues numbered from the N-terminus, with the C-terminal amide carried as the trailing NH2 token. NO one-letter string is published here: positions 2 and 13 are both Aib, 2-aminoisobutyric acid, and no standard letter names that residue. Residue 1 is Tyr, the GIP N-terminus rather than GLP-1's His. There are two lysines, at 16 and 20, and the acylation site is Lys20. The backbone is not the whole molecule: that Lys20 side chain is acylated through two AEEA spacers and a gamma-L-glutamyl unit to a C20 dicarboxylic fatty acid. Summing the thirty-nine residues with Aib at 2 and 13, removing thirty-eight waters, applying the Ser39 amide, then adding two AEEA units, one glutamyl unit and eicosanedioic acid less the four waters of the four amide bonds, gives C225H348N48O68, reproducing the recorded formula exactly; the derived average mass is 4813.53 Da against a recorded 4813, which is a truncation of that value rather than a rounding of it. Residues 29 to 39 are Gly-Gly-Pro-Ser-Ser-Gly-Ala-Pro-Pro-Pro-Ser; that tail is shared with other 39-residue incretin peptides and is not by itself evidence of ancestry, which the structural notes take up. That arithmetic is the identity check this record rests on; the automated sequence verifier cannot reach it, because Aib and the acyl side chain are outside its derivation model.