A multi-institution consortium coordinated through the National Human Genome Research Institute showed that roughly 8 percent of the genome, mostly repetitive regions long considered too difficult to read, had actually been left out of that 2003 draft.
The Telomere-to-Telomere Consortium, including lead author Sergey Nurk, used newer long-read sequencing to assemble a genome gapless on every chromosome except Y, filling in nearly 200 million base pairs earlier methods could not sequence reliably. The gap-filled regions were repetitive stretches near chromosome centers and tips that standard techniques had failed on for two decades.
The newly resolved sequence revealed 1,956 additional predicted genes, of which 99 are believed to code for proteins, none of which were visible in the supposedly complete 2003 genome. The assembly used DNA from a single cell line rather than a standard human sample with two different sets of chromosomes, and later projects have begun applying the same method to other individuals to capture human genetic variation more fully.

Read the original study
Nurk et al., Science, 2022 · doi.org/10.1126/science.abj6987
Eight percent of the map had been blank.
Source
- Nurk, S., Koren, S., Rhie, A., Rautiainen, M., Bzikadze, A. V., Mikheenko, A., Vollger, M. R., Altemose, N., Uralsky, L., Gershman, A., Aganezov, S., Hoyt, S. J., Diekhans, M., Logsdon, G. A., Alonge, M., Antonarakis, S. E., Borchers, M., Bouffard, G. G., Brooks, S. Y., . . . Phillippy, A. M. (2022). The complete sequence of a human genome. Science, 376(6588), 44–53. https://doi.org/10.1126/science.abj6987