What is Genetic Variation?

Humans share approximately 99.9% of their DNA. The remaining fraction consists of millions of SNVs and INDELs, as well as tens of thousands of larger structural variants. These differences drive biological diversity, influence disease risk, and affect treatment response [1,2]. Variants differ in size (from a single base to megabases), timing (germline versus somatic), and function (including coding, regulatory, structural, and copy-number). Understanding this variation is essential to decoding life and the basis of disease.

Fundamental Definitions

The landscape of genetic variation encompasses several distinct classes, each with unique characteristics and functional implications. These variant types differ in how they form, when they appear (whether inherited or acquired), and how they affect biology, depending on their location and function in the genome. 

A DNA sequence with base pairs highlighted: green for adenine (A), red for guanine (G), yellow for cytosine (C), and blue for thymine (T). The sequence includes several marked small variants.

1. Small Variants 

  • Single Nucleotide Variants (SNVs): change in a single base pair (e.g., A→T) 

  • Insertions and Deletions (INDELs): The insertion or deletion of a small number of bases. 

Illustration of structural variant types represented on chromosomes, including deletions, duplications, inversions, insertions, and translocations.

2. Structural Variants (SVs)

  • Balanced: Rearrangements without a net gain / loss of DNA. 

  • Unbalanced: Copy number changes (deletions, duplications).

  • Complex: Multiple breakpoints and combinations of SV types.

Small Variants 

Illustration of a DNA double helix with a colored base representing a single nucleotide genomic variation.

SNVs

The simplest form of genetic variation, SNVs involve the substitution of one nucleotide for another [2,3]. These variants occur throughout the genome at a frequency of approximately 1 in 1,000 base pairs and can be categorized into two types: 

Transitions (ex: A ↔ G or C ↔ T)

Transversions (purine ↔ pyrimidine conversions). 

Transitions tend to be less disruptive compared to transversion [2].

Depending on the type of change and where it occurs, SNVs can significantly impact gene function. Changes to encoded amino acids through nonsynonymous substitutions can impact protein folding and/or function.

Alternatively, base changes in non-coding regions can influence gene regulation. This may occur either by directly affecting promoter activity or by altering regulatory elements such as enhancers and silencers that modulate gene expression. In addition, mRNA conformation and structure may be affected, influencing mRNA half-life, function, or cellular localization [2].

Illustration of a DNA double helix with a colored DNA segment representing a genomic variation.

INDELs

For this review, we are defining a gain or loss of DNA from 1 to 500 base pairs as an INDEL, with their size and location significantly impacting any functional consequences. INDELs that are not multiples of three bases can cause reading frame shifts, often disrupting protein coding and resulting in loss of function.


Structural Variants 

Translocations

These consist of movement of DNA segments to new chromosomal positions, either intra- or inter-chromosomal. These can disrupt genes or create oncogenic fusions [4,5]. Robertsonian translocations (where two acrocentric chromosomes fuse at the centromere) are the most common rearrangement in humans.

Illustration of two chromosomes showing a large structural rearrangement in which one chromosome has lost an arm that has been added to another chromosome.

Inversions

These consist of movement of DNA segments to new chromosomal positions, either intra- or inter-chromosomal. These can disrupt genes or create oncogenic fusions [4,5]. Robertsonian translocations (where two acrocentric chromosomes fuse at the centromere) are the most common rearrangement in humans.

Illustration of a structural variant showing a genomic inversion, where a DNA segment is reversed in orientation.

Insertions

Large additions of DNA segments may originate from other genomic locations or external sources. Insertions can disrupt coding regions, regulatory elements, or generate fusion genes.

Illustration of a structural variant showing a genomic insertion, where a DNA segment has been inserted into a gene region.

Copy Number Variations (CNVs)

CNVs alter the number of copies of a genomic segment. These include deletions (loss of DNA) and duplications/amplifications (gain of DNA), which can affect gene dosage, epigenetic state, and regulation [3,4,6,7]. CNVs may appear as tandem or displaced duplications, and can be intrachromosomal, interchromosomal, or even extrachromosomal. They are also key drivers of genomic evolution [6].

Deletion

Illustration of a structural variant showing a genomic deletion, where a DNA segment is missing.

Duplication

Illustration of a structural variant showing a genomic duplication, where an additional DNA segment has been duplicated.

Complex Structural Variants

Many of the variants described above can occur in parallel, leading to more complex rearrangements that extend beyond single events. These complex SVs are illustrated in Case Study 1 and discussed in greater depth in our downloadable Genetic Variation Ebook.

Balanced Vs. Unbalanced Structural Variants

Variant Type Balanced UnBalanced
Translocation
Inversion
Insertion
Copy Number Variant

Since SVs can lead to either a gain or loss of DNA, distinguishing them as either balanced or unbalanced provides a useful framework for classification. A balanced SV does not involve a net gain or loss of genetic material, though the DNA is rearranged. An unbalanced SV involves either a loss or gain of DNA sequence. Certain classes, such as translocations and insertions, include variants that can be either balanced or unbalanced, depending on the specific event. The table below summarizes these categories.

Germline Vs. Somatic Variants

Illustration of genetic inheritance across a family.

Germline variants are inherited and present in every cell of the body. They drive inherited risk for a range of conditions, including breast and ovarian cancer (e.g., BRCA1), cystic fibrosis (CFTR), and fragile X syndrome (FMR1)[8,9,10,11]. These variants are stable across tissues and generations, making them key targets for population screening and family-based genetic studies.

Clipboard with genetic DNA symbols and checkmarks.

Somatic variants arise post-zygotically due to replication errors, environmental exposures, or cellular aging [5,8,12,13]. Unlike germline mutations, they are not inherited but accumulate over time in specific tissues. Mutation rates vary significantly depending on cell turnover and environmental stress. Rapidly dividing epithelial tissues, like lung, colon, and bladder, accumulate more somatic mutations than quiescent tissues like neurons [8,12]. Cancer reflects the clinically dominant consequence of somatic variation. 

Clinical Impact & Disease Relevance 

Genetic variation plays a critical role in human health, from rare Mendelian disorders to complex diseases. 

SVs, while less numerous than SNVs and INDELs, affect larger genomic regions and drive major changes in gene function and regulation. Only a minority of detected variants are pathogenic, underscoring the need for accurate interpretation and contributing to missed diagnoses [4,6]. 

The Pan-Cancer Analysis of Whole Genomes (PCAWG) analyses suggest SVs account for approximately 55% of cancer drivers, via gene dosage changes, fusions, enhancer hijacking, and chromatin domain disruption [5,14]. We will describe these mechanisms in more detail in a later chapter. 

*all reference information can be found in the Reference section of the eBook