Characterizing nucleotide variation and expansion dynamics in human-specific variable number tandem repeats

2021 
Abstract There are over 55,000 variable number tandem repeats (VNTRs) in the human genome, notable for both their striking polymorphism and mutability. Despite their role in human evolution and genomic variation, they have yet to be studied collectively and in detail, partially due to their large size, variability, and predominant location in non-coding regions. Here, we examine 467 VNTRs that are human-specific expansions, unique to one location in the genome, and not associated with retrotransposons. We leverage publicly available long-read genomes – including from the Human Genome Structural Variant Consortium – to ascertain the exact nucleotide composition of these VNTRs, and compare their composition of alleles. We then confirm repeat unit composition in over 3000 short-read samples from the 1000 Genomes Project. Our analysis reveals that these VNTRs contain remarkably structured repeat motif organization, modified by frequent deletion and duplication events. While overall VNTR compositions tend to remain similar between 1000 Genomes Project super-populations, we describe a notable exception with substantial differences in repeat composition (in PCBP3), as well as several VNTRs that are significantly different in length between super-populations (in ART1, PROP1, WDR60, and LOC102723906). We also observe that most of these VNTRs are expanded in archaic human genomes, yet remain stable in length between single generations. Collectively, our findings indicate that repeat motif variability, repeat composition, and repeat length are all informative modalities to consider when characterizing VNTRs and their contribution to genomic variation.
    • Correction
    • Source
    • Cite
    • Save
    • Machine Reading By IdeaReader
    22
    References
    0
    Citations
    NaN
    KQI
    []