For years, genomics has had a dirty little secret: even our best reference genomes have gaps. Now, a large international team of the National Human Genome Research Institute (NHGRI) at the NIH, along with the National Institute of Standards and Technology (NIST), has closed most of those gaps for good. Their new study introduces what they call a “genome benchmark,” a complete, error-checked, two-copy (diploid) version of a well-known human genome sample called HG002. The work brings together dozens of labs, including groups from Johns Hopkins, UC Santa Cruz, the University of Washington, Baylor College of Medicine, and several international institutions, all working under something nicknamed the “Q100 Project.”
Why This Matters
When researchers sequence someone’s DNA, they usually compare the results to a “reference genome,” a kind of master template. The problem is that the standard reference has always been incomplete. Certain repetitive, tangled-up stretches of DNA (think of them as the genome’s junk drawers) have historically been too messy to piece together, so they got left out of comparisons. That meant roughly 12% of a person’s genetic code, including entire chromosomes in some cases, simply wasn’t being checked for accuracy.
This isn’t a small technical footnote. Many of these skipped regions contain genes linked to real health conditions, such as immune function, cancer risk, and neurological disorders. If your sequencing test can’t reliably “see” that part of your DNA, it can’t tell you what’s happening there either.
What the Team Actually Built
Instead of comparing DNA fragments to an imperfect reference and cataloging the differences (the traditional approach, called variant benchmarking), the team flipped the whole idea on its head. They assembled the entire genome of HG002, both the copy inherited from mom and the copy from dad, from telomere to telomere, meaning literally from one end of each chromosome to the other, with almost nothing missing.
To pull this off, they combined multiple sequencing technologies (including PacBio and Oxford Nanopore long-read platforms). They ran the assembly through several rounds of careful, partly crowdsourced correction, fixing tens of thousands of small errors along the way. The result, version 1.1 of the T2T-HG002 genome, is free of detectable mistakes across 99.4% of its length. To put that in perspective, it adds back over 700 million base pairs of previously invisible DNA on the regular chromosomes, plus both sex chromosomes in full, which together make up about 15% of the entire genome that older benchmarks simply couldn’t account for.

The team also built and released a companion piece of software called the Genome Quality Checker, or GQC, which lets other researchers test their own sequencing data, assemblies, or genetic variant calls against this new, far more complete gold standard.
A More Honest Way to Grade Genomic Tools
One of the more striking findings is how much this new approach reveals about the limits of older methods. When the researchers compared modern DNA assembly techniques against the new benchmark, they found that building a genome from scratch (de novo assembly) is now roughly ten times more accurate than the traditional method of calling variants against a reference, something that simply wasn’t visible before because there was no complete ground truth to measure against. They also found that popular quality-scoring tools had been overestimating the accuracy of certain sequencing technologies while underestimating others, largely because those tools couldn’t “see” missing parts of the genome.
The team also produced a full gene annotation for both haplotypes, cataloguing over 39,000 protein-coding genes and turning up interesting biological findings along the way, including genes that exist on one parental copy but not the other, differences in gene copy number between the maternal and paternal chromosomes, and a well-known case involving the gene DUSP22, which shows up in a completely different genomic neighborhood on one haplotype than the other, something no prior benchmark could have caught.
Where This Leads
The authors are careful to note that even this benchmark isn’t 100% flawless; a small sliver of highly repetitive ribosomal DNA remains genuinely difficult to resolve with current technology, and some natural cell-to-cell variation is expected. But by building genomic quality control around a real, complete, personal genome rather than a patchwork reference, this work sets a new bar for what “accurate” sequencing actually means.
Practically speaking, this benchmark should help sequencing companies, clinical labs, and researchers develop better tools for reading the full genome, not just the easy 85% of it. As the field moves toward analyzing complete, personalized genomes rather than genomes riddled with blind spots, resources like this one are likely to become the new standard against which every future method gets measured.
Article Source: Reference Paper
Disclaimer:
The research discussed in this article was conducted and published by the authors of the referenced paper. CBIRT has no involvement in the research itself. This article is intended solely to raise awareness about recent developments and does not claim authorship or endorsement of the research.
Follow Us!
Learn More:
Anchal is a consulting scientific writing intern at CBIRT with a passion for bioinformatics and its miracles. She is pursuing an MTech in Bioinformatics from Delhi Technological University, Delhi. Through engaging prose, she invites readers to explore the captivating world of bioinformatics, showcasing its groundbreaking contributions to understanding the mysteries of life. Besides science, she enjoys reading and painting.












