A complete diploid human genome benchmark for personalized genomics
TL;DR
The authors present a near-perfect telomere-to-telomere diploid genome benchmark for HG002 that adds 15.3 percent of previously unmapped genomic sequence and shows de novo assembly outperforms standard variant calling by an order of magnitude.
Problem / question
Standard genome resequencing relies on mapping reads to a reference genome, which introduces biases and fails to map complex, duplicated, or structurally polymorphic regions, meaning existing variant benchmarks cannot assess these difficult areas.
Methods
The researchers generated a complete telomere-to-telomere diploid genome benchmark for the human genome HG002. They annotated genes, transposable elements, segmental duplications, and satellite repeats, and developed custom tools to measure the accuracy of sequencing reads, phased variant calls, and genome assemblies against this diploid reference.
Key findings
The new benchmark achieves near-perfect accuracy across 99.4 percent of the diploid HG002 genome. It adds 701.4 Mb of autosomal sequence and 216.8 Mb of sex chromosome sequence, representing 15.3 percent of the genome missing from previous benchmarks, and annotates 39,144 protein-coding genes across both haplotypes. Using this benchmark, the authors found that state-of-the-art de novo assembly methods resolve 2 to 7 percent more sequence than standard variant calling and achieve an error rate of just one per 100 kb across 99.9 percent of the benchmark regions.
Why it matters
Providing a complete, genome-wide benchmark enables the accurate evaluation of sequencing and assembly methods in previously inaccessible complex regions, paving the way for comprehensive personalized genomics and better genomic medicine.
Limitations
The provided text does not specify any limitations or caveats of the benchmark or methods.
Takeaway
A new complete diploid benchmark for the HG002 genome includes 15.3 percent more sequence than prior benchmarks and demonstrates that de novo assembly is highly accurate, achieving just one error per 100 kb.