PGG.Han: the Han Chinese genome database and analysis platform

Yang Gao; Chao Zhang; Liyun Yuan; YunChao Ling; Xiaoji Wang; Chang Liu; Yuwen Pan; Xiaoxi Zhang; Xixian Ma; Yuchen Wang; Yan Lu; Kai Yuan; Wei Ye; Jiaqiang Qian; Huidan Chang; Ruifang Cao; Xiao Yang; Ling Ma; Yuanhu Ju; Long Dai; Yuanyuan Tang; Han100K Initiative; Guoqing Zhang; Shuhua Xu

doi:10.1093/nar/gkz829

PGG.Han: the Han Chinese genome database and analysis platform

Nucleic Acids Res. 2020 Jan 8;48(D1):D971-D976. doi: 10.1093/nar/gkz829.

Authors

Yang Gao^{1

2}, Chao Zhang¹, Liyun Yuan¹, YunChao Ling¹, Xiaoji Wang¹, Chang Liu¹, Yuwen Pan¹, Xiaoxi Zhang^{1

2}, Xixian Ma¹, Yuchen Wang¹, Yan Lu^{1

3}, Kai Yuan¹, Wei Ye¹, Jiaqiang Qian¹, Huidan Chang¹, Ruifang Cao¹, Xiao Yang¹, Ling Ma¹, Yuanhu Ju¹, Long Dai¹, Yuanyuan Tang¹; Han100K Initiative; Guoqing Zhang¹, Shuhua Xu^{1

2

3

4}

Affiliations

¹ Key Laboratory of Computational Biology, Bio-Med Big Data Center, CAS-MPG Partner Institute for Computational Biology, Shanghai Institute of Nutrition and Health, Shanghai Institutes for Biological Sciences, University of Chinese Academy of Sciences, Chinese Academy of Sciences, Shanghai 200031, China.
² School of Life Science and Technology, ShanghaiTech University, Shanghai 201210, China.
³ Collaborative Innovation Center of Genetics and Development, Shanghai 200438, China.
⁴ Center for Excellence in Animal Evolution and Genetics, Chinese Academy of Sciences, Kunming 650223, China.

Abstract

As the largest ethnic group in the world, the Han Chinese population is nonetheless underrepresented in global efforts to catalogue the genomic variability of natural populations. Here, we developed the PGG.Han, a population genome database to serve as the central repository for the genomic data of the Han Chinese Genome Initiative (Phase I). In its current version, the PGG.Han archives whole-genome sequences or high-density genome-wide single-nucleotide variants (SNVs) of 114 783 Han Chinese individuals (a.k.a. the Han100K), representing geographical sub-populations covering 33 of the 34 administrative divisions of China, as well as Singapore. The PGG.Han provides: (i) an interactive interface for visualization of the fine-scale genetic structure of the Han Chinese population; (ii) genome-wide allele frequencies of hierarchical sub-populations; (iii) ancestry inference for individual samples and controlling population stratification based on nested ancestry informative markers (AIMs) panels; (iv) population-structure-aware shared control data for genotype-phenotype association studies (e.g. GWASs) and (v) a Han-Chinese-specific reference panel for genotype imputation. Computational tools are implemented into the PGG.Han, and an online user-friendly interface is provided for data analysis and results visualization. The PGG.Han database is freely accessible via http://www.pgghan.org or https://www.hanchinesegenomes.org.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Asian People / genetics*
China
Databases, Genetic*
Ethnicity / genetics
Genetics, Population*
Genome, Human*
Genomics* / methods
Humans
Software
Software Design
Web Browser