Chromatography Trace Illustration

Reversed-phase HPLC is the standard method for assessing peptide purity. The chromatogram is a visual representation of detector response over retention time, and purity is calculated as the area percent of the main peak.

What the trace shows

The target compound appears as the main peak; earlier and later smaller peaks represent related impurities. Reviewers evaluate peak shape, retention time, resolution from neighboring peaks, and the overall impurity profile. Method parameters — column, mobile phase, gradient, flow rate, and detection wavelength — provide the context needed to interpret the trace.

Purity documentation

HPLC is widely used because it is robust and reproducible. Quantifying purity by area percent and reporting it with full traceability is central to a defensible COA.

Research use only. Products discussed are intended strictly for in-vitro laboratory research and are not for human or veterinary use.

The short version

A purity by area figure is a ratio computed inside one chromatogram, not a measurement of what fraction of the vial is peptide. The software integrates every peak the detector produced, divides the main peak area by the total, and reports the result. Everything outside that sum, anything that does not absorb at the chosen wavelength, anything that never left the column inside the run window, anything smaller than the integration threshold, is silently absent from the denominator. That makes the number sensitive to choices an analyst makes before the sample is ever loaded: wavelength, gradient slope, column chemistry, run length, and then to processing choices made afterward on data that is already fixed. The sections below work through each of those levers, the peak-shape vocabulary that tells you which one is in play, and how to read a supplied trace when the method behind it is not stated.

How the percentage is built and why it is only a ratio

Integration software does two things in sequence. It first decides where each peak begins and ends and where the baseline sits underneath it, then it computes the area enclosed between the curve and that baseline. The reported purity is one of those areas divided by the sum of all of them, multiplied by a hundred. The word to hold onto is sum: the denominator is built only from peaks the software chose to integrate in this particular run.

Three consequences follow, and all three tend to get lost when the figure is copied onto a certificate.

The first is that area percent assumes every component responds to the detector in proportion to how much of it is present, on the same scale. That is false in general. Ultraviolet absorbance depends on what is absorbing, so an impurity with weaker absorbance at the monitored wavelength contributes less area per unit mass than the main compound, and one with stronger absorbance contributes more. A related impurity that has lost a tryptophan residue does not simply produce a smaller peak, it produces a peak that understates its own abundance if the run is monitored where tryptophan dominates the signal. Area percent is an approximation to mass percent, not a synonym for it.

The second is that the figure is normalized. It always sums to a hundred across the peaks present. If a component never produces a peak, the remaining peaks do not shrink to make room for it; they simply divide the same hundred percent between themselves. Adding an invisible contaminant to a vial cannot lower an area-percent purity figure. It can only lower a mass-based measurement.

The third is that a purity percentage says nothing about how much target material was in the vial in the first place. Water retained after lyophilization, the counterion paired with basic residues, residual salts from cleavage and workup, and trapped solvent all contribute to the weight on the label without contributing any area. That is why a separate content or assay figure exists at all, and why the two numbers on a certificate are usually different, with content the lower of the two.

What the denominator contains, and what quietly sits outside it

Sample componentAppears in the area sum?Consequence
Target peptideYes, as the main peakSupplies the numerator and part of the denominator
Related peptide impuritiesYes, if resolved and above thresholdDeletion, truncation, oxidation and dimer species that absorb much like the target
Counterion, commonly trifluoroacetate or acetateNo useful peakReal dry mass with no area; pulls content down, leaves purity untouched
Residual water from lyophilizationNoWeighs, does not absorb, invisible to the ratio
Inorganic salts and small polar speciesUsually notElute at or near the solvent front, where peaks are commonly excluded
Features below the area reject settingNoRemoved from the sum before the division, so the main peak looks larger
Species retained past the end of the runNoLate hydrophobic material never elutes inside the window at all

Read the ratio as a statement about the visible population. It answers one question: of everything this method could see, in this window, above this threshold, what share was the main peak. That is a genuinely useful question and reversed-phase separation answers it well. It is not the same question as what fraction of the powder is the named compound, and it is not usually the question a receiving lab thinks it is asking. Keeping the two apart is most of the difference between using a certificate and merely quoting one.

What the detection wavelength decides for you

Ultraviolet detection is selective whether or not anyone intends it to be. The detector reports absorbance at a wavelength the analyst picked, and every component in the sample is weighted by how strongly it happens to absorb there.

Near 214 nm the dominant absorbing feature is the peptide bond itself. Every amide linkage in the backbone contributes, so response scales roughly with chain length and is largely independent of which side chains are present. That is why 214 nm, or 215 and 220 nm as near equivalents, is the conventional choice for purity work. It is the closest thing to a universal peptide detector, and impurities that are themselves peptides show up whether or not they carry aromatic residues.

Near 280 nm the picture changes completely. Absorbance there comes from aromatic side chains, principally tryptophan and tyrosine, with a smaller contribution from phenylalanine and from disulfide bonds. A sequence containing neither tryptophan nor tyrosine barely absorbs at 280 nm at all. Running a purity trace at that wavelength on such a peptide produces either a nearly flat line or, worse, a chromatogram in which only the aromatic-containing impurities are visible while the main compound is faint. More often the failure runs the other way: the main peak is aromatic, several impurities are not, those impurities vanish, and the purity figure rises without any change in the material.

Both settings share a blind spot that no ultraviolet wavelength closes. A contaminant with no chromophore in the accessible range, an inorganic salt, a sugar, glycerol, a silicone leached from tubing or a stopper, produces no peak anywhere. It is not a small peak, it is nothing. The chromatogram is silent about it and the area percent is entirely unaffected by however much of it is present. Only a detector working on a different principle, evaporative, charged-aerosol, refractive index or mass spectrometric, will register it.

There is also a practical artifact of the far ultraviolet. Trifluoroacetic acid, the usual ion-pairing modifier, absorbs meaningfully below about 220 nm, and because the aqueous and organic mobile phases rarely match exactly in modifier content, a gradient run monitored at 214 nm often shows a rising or falling baseline. Where the software puts the baseline under that drift feeds directly into the area of every peak on the trace.

Detection settings and what each one can and cannot see

SettingWhat produces the signalWhat it misses
214 to 220 nmThe backbone amide bond, in every residueAnything with no ultraviolet chromophore; also sits where the acid modifier absorbs
254 nmA legacy general-purpose setting, weak for most peptidesMost non-aromatic impurities and much of the main peak signal too
280 nmTryptophan and tyrosine side chainsAny species lacking those residues, which may include the target itself
Diode array, full spectrum storedAbsorbance at every wavelength at every time pointStill blind to non-absorbing species, but allows peak-spectrum comparison
Evaporative or charged-aerosol detectionAnything less volatile than the mobile phaseVolatile impurities; response is not linear in the way ultraviolet response is
Mass spectrometry in lineAnything that ionizes under the source conditionsPoorly ionizing species, and isomers that share the target mass

The useful habit is to check the stated wavelength before reading the percentage, and to ask whether it was the right one for this sequence. A trace at 280 nm on a peptide carrying a single tyrosine is not wrong, but it is a narrow view, and it will flatter the material relative to a 214 nm run on the same vial. Where a diode-array detector was used, a report can show the same separation at two wavelengths at once. When the two numbers differ noticeably, that difference is information rather than an error.

Conditions that move the figure without touching the material

The separation is the measurement. Change the separation and the number changes, even though nobody has touched the contents of the vial.

Gradient slope is the largest single lever. A shallow gradient, a small change in organic content per minute, spreads the elution window and pulls closely related species away from the main peak. Each one that resolves becomes its own integrated area and leaves the main peak smaller. A steep gradient compresses everything toward the same retention time, impurities that were separate merge into the main peak, and their area is counted as target. The steeper method reports the higher purity, reliably, on the same sample. This is the most common way an unremarkable lot acquires an impressive number.

Run length works the same way at the back end. Strongly retained species, aggregates, hydrophobic side products, fragments that never fully deprotected, elute late and sometimes very late. If the gradient ends and the run stops before a hold at high organic content, that material stays on the column, does not appear, and is not in the denominator. It also tends to reappear in the next chromatogram, which is one origin of ghost peaks. A method that ends with a wash step and a stated re-equilibration is a method that has at least looked for late material.

Column chemistry sets selectivity, not just retention. C18 is the default; C8 and C4 retain less and suit larger or more hydrophobic sequences; phenyl and biphenyl phases add aromatic interaction and often separate pairs that co-elute on straight alkyl chains. Two impurities that sit under the main peak on one phase can be baseline resolved on another. Nothing about the sample changed between the two runs.

Particle size, column length and temperature act on efficiency rather than selectivity, but the effect on the number is the same, since narrower peaks resolve more neighbors. Raising column temperature sharpens peaks and can also collapse or reveal the splitting caused by slow rotation around proline residues.

Loading matters as well. Overload the column and the main peak broadens and often fronts, burying small impurities inside its own footprint. Load too little and impurities sitting near the integration threshold fall below it and disappear from the sum.

How each method variable typically moves the reported percentage

VariableChange madeUsual direction of the number
Gradient slopeA steep fast ramp instead of a shallow oneHigher, because near neighbors merge into the main peak
Run length and washStopping at the end of the ramp with no high-organic holdHigher, because late hydrophobic material never elutes
Column chemistryC18 swapped for phenyl, C8 or a different bondingEither direction; selectivity decides which pairs separate
Particle size and bed lengthSmaller particles, longer column, core-shell packingUsually lower, as efficiency exposes shoulders as separate peaks
Column temperatureRaised from ambient to 40 to 60 degrees CEither direction; sharpens peaks, can merge or split rotamer pairs
Mobile-phase modifierTrifluoroacetic acid replaced by formic acid or a bufferOften broader peaks and a different impurity profile at the same nominal purity
Amount on columnA larger load to make small impurities visibleHigher, as the broadened main peak absorbs its small neighbors

None of this makes a fast method dishonest. Short runs have a real place in process control, and a supplier screening incoming lots at speed is doing something reasonable. What matters is that the method appears on the report. A purity figure with no column, no gradient table and no wavelength is an assertion rather than a measurement, because there is no way to judge whether the separation was capable of finding anything. When two certificates for the same compound disagree, comparing the two method sections usually explains the gap before anything else does.

Integration choices: one trace, several defensible answers

Everything above happens before the data exists. The next set of choices happens afterward, on data that is already fixed, and it can move the reported figure by a surprising amount.

Baseline placement is the foundation. Peak area is the region between the curve and a baseline the software constructs, so where that baseline is drawn determines how much area exists at all. On a drifting trace, common at 214 nm in a gradient run, a baseline drawn horizontally from the peak start gives a different area than one drawn along the drift. Analysts can and do reintegrate by hand, and manual reintegration is legitimate practice; it becomes a problem only when it is invisible in the report.

Then there is how partly resolved peaks are divided. The drop-line, or perpendicular drop, splits the pair with a vertical line at the valley between them, assigning everything to the left to one peak and everything to the right to the other. Tangent skimming instead treats the smaller peak as a rider sitting on the tail of the larger one, drawing a curved baseline underneath it, which assigns far less area to the small peak. On a shoulder or a tail rider the two conventions produce visibly different impurity percentages from identical raw data. Neither is wrong in principle, each suits a different physical situation, and a report that does not say which was used has left something out.

Exclusion windows are the next lever. Almost every method excludes the region at the very start of the run, where unretained material, salts and the sample solvent all arrive together as the solvent front, along with the small positive or negative disturbance that sample introduction itself produces on the detector. There is a good reason for that exclusion: nothing in the void volume has been separated, so nothing there can be quantified. But a window extending well past the void volume starts discarding real, if polar, impurities.

Finally there are thresholds. Integration parameters include a minimum area or height below which a feature is not treated as a peak at all. Setting that threshold loosely removes a scatter of small impurities from the denominator before the division happens. Reporting conventions often disregard peaks below a stated limit, which is sound practice when the limit is stated and misleading when it is not.

Processing choices and their effect on the same raw data

Processing choiceWhat it doesEffect on the main-peak percentage
Baseline placement on a drifting traceSets the floor under every peakEither direction; small peaks are affected proportionally more
Drop-line at the valleySplits an unresolved pair with a vertical lineLower than skimming, because the small peak keeps more of the shared area
Tangent skim of a rider peakTreats a shoulder as sitting on the main peak tailHigher, because the shoulder is assigned much less area
Solvent-front exclusion windowDiscards everything before a chosen timeHigher, correctly so at the void volume, questionably so beyond it
Area or height reject thresholdIgnores features below a set sizeHigher, by removing many small peaks from the sum entirely
Blank or gradient-blank subtractionRemoves features seen in a solvent-only runHigher; correct when the feature is from the system, wrong if it is carryover
Manual reintegrationAn analyst redraws peak start and end limitsEither direction; acceptable practice, but it should be disclosed

The point is not that integration is arbitrary. It is that a single trace supports a range of defensible numbers, usually within a few tenths of a percent for a clean sample and considerably wider for a messy one. That range is why a printed percentage carries much less information than the picture it came from. Given the chromatogram itself, with axes and a peak table, a reviewer can see the baseline, see the valleys, and form an independent view. Given only the number, there is nothing to check.

Peak shape and artifacts, and what each usually indicates

Shape is diagnostic. Before reading any percentage it is worth spending a minute on what the peaks actually look like, because several shapes signal that the number attached to them is unreliable.

Tailing is the most common defect: the peak rises quickly and returns to baseline slowly, dragging a long asymmetric edge behind it. In peptide work the usual cause is secondary interaction between basic residues, lysine, arginine and histidine, and residual free silanols on the silica surface, which is exactly what the ion-pairing modifier is there to suppress. Aging columns, voids in the bed and excessive tubing volume produce the same signature. The practical harm is that a long tail conceals small later-eluting impurities inside it, and it forces the integrator to choose between a drop-line and a skim.

Fronting is the mirror image and usually means overload: more material than the stationary phase can hold in the linear range, so the leading edge runs ahead of the band. It also appears when the sample was dissolved in a solvent stronger than the starting mobile phase, so the band travels some distance through the head of the column before it focuses.

A shoulder is a partly resolved peak, and treating it as noise is the most consequential mistake in this whole exercise. Shoulders on peptide main peaks are commonly a deamidation product, an oxidized methionine, a diastereomer from racemization during coupling, or a sequence missing a single residue. All of those elute close to the parent because chemically they differ from it very little. A method with more resolving power will pull them apart; a faster method merges them entirely.

A split peak, one that dips toward baseline and rises again, is usually mechanical: a void or channel at the column inlet, a partly blocked frit, or a sample solvent that is too strong. It can also be genuine, when cis and trans proline rotamers interconvert slowly enough on the timescale of the separation to travel as two bands.

Ghost peaks appear where no sample was loaded and belong to the system rather than to the material: contaminated mobile phase, leached plasticizer, or material accumulated on the column from earlier runs and released by the gradient. Carryover is the related case, where a peak from the previous sample turns up in the following chromatogram because the wash was inadequate.

Shape features and the usual reading

FeatureUsual causeEffect on interpretation
Tailing main peakSilanol interaction with basic residues, aging column, extra tubing volumeHides small late impurities in the tail; integration becomes a judgment call
Fronting main peakColumn overload, or sample solvent stronger than the starting eluentThe broadened peak covers its neighbors and purity reads high
Shoulder on the main peakDeamidation, oxidation, a diastereomer or a deletion sequenceA real impurity that a faster method would simply count as target
Split main peakVoid or blocked frit at the column head, strong sample solvent, proline rotamersEither an instrument fault or two genuine conformers; the run should be repeated
Ghost peaks in a blank runMobile-phase contamination, leachables, material released from the columnNot sample impurities, but they must be identified before they are subtracted
Carryover from the previous sampleInadequate needle or column wash between runsAdds area that belongs to a different vial
Rising or falling baselineModifier absorbance mismatch between the two mobile phases at low wavelengthShifts every peak area through the baseline the software draws

A trace with sharp, symmetrical, well-separated peaks is evidence that the method was working, and that is a precondition for the number meaning anything. A trace showing a broad tailing main peak with a figure of ninety nine and change attached is a document that has reported a result the separation was not in a position to support. In that situation the honest conclusion is not that the material is bad. It is that the chromatogram has not established anything yet and the run should be repeated on a column in better condition.

A single peak is not proof of a single compound

Co-elution is the limit case of everything above. Two species that share a retention time under one set of conditions produce one peak, and one peak integrates as one component. The chromatogram cannot distinguish a pure band from a perfectly overlapped pair, because absorbance plotted against time contains no information about composition.

Some overlaps are far more likely than others. Species differing from the target by a small, hydrophobically neutral change tend to sit almost exactly on top of it: a deamidated residue, an epimerized center where one amino acid has racemized, an isoaspartate rearrangement, sometimes a conservative substitution introduced by impure starting material. Several of those also share the target mass, so even a mass measurement on the collected peak will not separate them.

There are three ways to look underneath a peak. A diode-array detector allows a peak purity assessment: spectra are collected across the peak and compared, and a change in composition on the upslope, apex and downslope shows up as a spectral mismatch. This only works when the two components genuinely have different ultraviolet spectra, which for two closely related peptide sequences is often not the case. Mass spectrometry coupled to the separation is stronger, because it reports mass continuously across the peak and will show a second component of different mass immediately. It remains blind to isomers.

The third approach is orthogonal separation: run the same sample on a mechanism that separates by something other than hydrophobicity. Ion-exchange separates by charge and readily resolves deamidation products that reversed-phase merges. A substantially different mobile-phase pH changes the ionization state of side chains and shifts selectivity as well. For stereochemistry the honest answer is that neither separation by hydrophobicity nor a mass measurement closes the gap; it takes hydrolysis followed by chiral analysis of the released amino acids.

That is the real shape of the limitation. A chromatogram establishes that the material is dominated by one species under one separation mechanism. It never establishes what that species is, and on its own it never establishes that the species is single. Identity confirmation, and where it matters an orthogonal separation, are what turn a high area percent into a claim that can be defended.

Reading a supplied chromatogram: what to check before accepting the number

What to look forA report you can work withA report you cannot
AxesBoth axes labeled with units, time in minutes, response in absorbance unitsUnlabeled axes, or an image with the axes cropped away
Run windowThe full run shown, from the solvent front through the high-organic washA view zoomed to a few minutes around the main peak only
Vertical scaleScale stated, or an inset at higher gain showing the baseline regionOne tall main peak with the baseline flattened by the chosen scale
Peak tableRetention time, area and area percent listed for every integrated peakA single percentage printed underneath the picture
Method statementColumn and dimensions, mobile phases, gradient, flow, temperature, wavelengthNo method at all, or a partial one with the gradient omitted
TraceabilityLot number, sample identifier, acquisition date, instrument and analystA generic image reused across lots with no lot number on it
Integration disclosureThreshold and any exclusion window stated, baseline visible under the peaksIntegration marks removed, or the trace redrawn as a clean graphic

Two quiet checks are worth running on any peak table. First, add the area percentages up. They should come to a hundred, and if they do not, something has been excluded that the report did not mention. Second, compare the main peak retention time against the total run length. A main peak eluting in the last tenth of the run leaves no room to see anything more hydrophobic, and a main peak arriving in the first minute may be sitting in the void volume, where nothing has been separated from anything. Neither is fatal, but both change what the percentage can support.

Questions this guide gets asked

Why do two labs report different purity for the same lot?

Almost always because they ran different methods, not because one of them is wrong. Gradient slope, run length, column chemistry, temperature, amount on column and detection wavelength each shift the figure, and integration settings shift it again afterward. A steep short run on a less efficient column reports a higher number than a shallow run on a modern column, using the same vial. Before treating a discrepancy as a dispute, put the two method sections side by side. If the numbers still differ once the separations look comparable, the question becomes whether the material itself changed between the two samplings, which storage and shipping history can sometimes explain.

Which detection wavelength should a purity chromatogram use?

For general peptide purity work, somewhere in the 210 to 220 nm range, because the signal there comes from the backbone amide bond and every peptide species responds regardless of its side chains. A 280 nm trace is useful as a confirmatory second channel when the sequence contains tryptophan or tyrosine, and it is close to meaningless when it does not. If a certificate reports purity from a 280 nm run for a sequence with no aromatic residues, that is worth querying directly. A diode-array acquisition that stores the whole spectrum lets a reviewer look at both views without rerunning the sample at all.

Is a symmetrical single peak proof of a single compound?

No. It is proof that under that one separation mechanism nothing else eluted at a resolvably different time. Species differing from the target by deamidation, by the position of an oxidation, or by a racemized center often co-elute almost exactly, and some of them also match the target mass, so mass detection alone does not always separate them either. Confidence comes from stacking mechanisms: a diode-array peak purity check, a mass measurement taken across the peak rather than at the apex only, and where it matters a second separation on a different mechanism such as ion exchange or a substantially different mobile-phase pH.

Why are peaks near the solvent front usually excluded?

Because nothing there has been separated. Unretained material passes through the column in the void volume, and everything that does not interact with the stationary phase, salts, the sample solvent itself, small polar species, arrives together at that same time. A peak in that region is a mixture by definition, so integrating it as though it were a component would be misleading. The legitimate exclusion window is the void volume and a little either side of it. A window extending well past that starts discarding polar impurities that were in fact separating, and a good report states where the window ends.

What does a shoulder on the main peak usually mean?

It means a second species is present and the method is only just failing to resolve it. Common candidates in peptide work are a deamidation product, an oxidized methionine, a diastereomer produced by racemization during coupling, and a sequence missing one residue. All of these are chemically close to the target and therefore elute close to it. Whether the shoulder is counted as an impurity depends on integration, since a drop-line assigns it real area while a tangent skim assigns it much less. The informative response is to rerun with a shallower gradient or a different stationary phase and see whether it separates.

What should be asked for when only a cropped image is supplied?

Ask for the full acquisition report rather than a picture: the whole run from the solvent front to the end of the wash, both axes labeled, the peak table with retention time and area for every integrated peak, and the method section with column, gradient, flow, temperature and wavelength. Ask whether the lot number on the trace matches the vial in hand. If what comes back is the same image again, that is itself informative. An image with no axes, no peak table and no method cannot be checked by anyone, and the percentage on it should be treated as a claim rather than as data.

Where to read next

All materials referenced here are supplied strictly for laboratory research use. They are not drugs, foods, cosmetics or medical devices, and they are not for human or veterinary use, diagnostic use, or any form of consumption. Nothing in this guide is guidance for use outside a controlled research setting. The analytical descriptions are general explanations of common chromatographic practice and do not replace a qualified analyst reviewing a specific certificate and the data behind it.

Leave a Reply

Your email address will not be published. Required fields are marked *

0