Genome publishes our HostSeq ancestry analysis using ntRoot

We are pleased to announce the publication of our latest study, “Concordance and divergence between self-declared ancestry and genome-derived ancestry composition in 10,250 participants from the HostSeq cohort,” in Genome.

Using whole-genome sequencing data from 10,250 participants in the Canadian HostSeq initiative, we applied our ntRoot ancestry inference framework to examine the relationship between self-declared ancestry and genome-derived ancestry composition.

Rather than relying solely on discrete ancestry assignments, the study integrates continuous ancestry fraction analyses, including proportional ancestry profiles, principal component analysis, entropy-based measures of ancestry complexity, and local ancestry inference across the genome. These analyses demonstrate that while broad concordance exists for several ancestry groups, many apparent discrepancies reflect continuous admixture rather than simple categorical disagreement.

Beyond the biological findings, the study illustrates the scalability of ntRoot for population-scale whole-genome sequencing datasets and highlights the value of combining genome-derived ancestry composition with participant-reported identity to better characterize human diversity in genomic research.

Analysis scripts and additional resources associated with this work have been made available through our Zenodo record as permitted by the HostSeq data-sharing framework.