Skip to content
Pureza

Research concepts

Residue Numbering: Precursors and Fragments

A residue number is a coordinate, not a sufficient identity. The same region can receive different numbers depending on the reference sequence or an article's convention. Resolve this by documenting the coordinate system before comparing mutations, fragments or modified sites.

Source editorial review:

First identify the reference sequence

UniProt states that its positional annotations refer to the canonical sequence displayed by default. This choice organizes data within an entry but does not mean that every publication uses exactly the same numbering.

A reproducible description needs the sequence identifier and the specific form being compared. Several sequences may share a protein name. If an article studies an alternative form, its coordinates must be interpreted on that form rather than transferred because of a similar name.

The precursor and mature chain are related maps

Processing annotations distinguish signal peptide, propeptide, chain and peptide. These categories describe different regions: not every removed segment is a signal peptide, and not every mature chain corresponds to the complete sequence initially presented.

Local numbering starting at the mature chain can differ from precursor positions. Relating them requires knowing the annotated boundaries. Subtracting a number remembered from another protein or species is not a valid correspondence.

A fragment needs two ends and an origin

Writing only a range leaves questions open: which sequence it belongs to, whether the ends are complete and whether the material includes additional changes. The range locates a region but does not, by itself, document every feature of a synthetic sample.

UniProt can also represent uncertain boundaries or regions extending beyond a stated position. That uncertainty must be retained. Turning an approximate annotation into an exact endpoint creates apparent precision the source does not provide.

A structure can use another numbering system

RCSB PDB documents two schemes: sequential numbering assigned to the polymer and numbering supplied by the authors. Their fields are label_seq_id and auth_seq_id. Depending on the viewer or file consulted, either may be displayed.

The chain also matters when citing a structural site. A number without a chain or scheme can indicate more than one location. Correspondence between structure and sequence must be checked explicitly, especially when the structural material represents only part of the protein.

What should accompany a published position

A useful reference brings together identifier, selected sequence or form, range and numbering convention. For structures, it adds the chain and scheme used. This distinguishes a genuine sequence discrepancy from a purely documentary difference.

The aim is not to make all sources renumber alike, but to preserve a verifiable correspondence. Two different positions can represent the same site, while two identical numbers can belong to entirely different regions.

Questions and answers

Is residue 1 always the beginning of the mature chain?

No. It depends on the sequence and numbering scheme used.

Is a fragment's range sufficient to identify it?

No. The source reference and any modifications distinguishing the material are also needed.

Sources

  1. UniProt: Sequences
  2. UniProt: Sequence annotation (Features)
  3. RCSB PDB: Identifiers in PDB