Chonghao Wang

dblp:337/4239 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Enhanced Disease Susceptible Variant Identification via Short Identity by Descent Segments
abstract
Rare diseases affect millions of individuals worldwide, yet diagnostic yields for them still remain low. Among variant identification approaches, identity by descent (IBD) mapping is used to identify disease susceptible variants originating from a recent common ancestor among affected individuals, but existing IBD detection models struggle to identify these variants in short IBD segments. Here, we introduce SILO, a novel model to detect disease susceptible variants in both short and long IBD segments. SILO employs a two-stage procedure to detect IBD segments. In the first stage, SILO identifies long IBD segments based on common variants. In the second stage, SILO utilizes rare variants to detect short IBD segments using a seed-and-extend algorithm. We evaluated SILO in simulated data and real data from the 1000 Genomes Project. Our results demonstrate that SILO outperforms existing models in detecting disease susceptible variants within short IBD segments, and show comparable performance in detecting these variants within longer IBD segments. These findings highlight the potential of SILO to increase diagnostic yields for rare diseases by enhancing the identification of previously overlooked disease susceptible variants in short IBD segments. Nonetheless, we note that the detection of short IBD segments remains challenging due to limited precision, leaving room for future improvement.
Chonghao Wang, Werner Pieter Veldsman, Yufen Huang, Xiaodong Fang, Lu Zhang 0061
IEEE Trans. Comput. Biol. Bioinform.1
2025 TRAFICA: an open chromatin language model to improve transcription factor binding affinity prediction
abstract
MOTIVATION: In silico transcription factor and DNA (TF-DNA) binding affinity prediction plays a vital role in examining TF binding preferences and understanding gene regulation. The existing tools employ TF-DNA binding profiles from in vitro high-throughput technologies to predict TF-DNA binding affinity. However, TFs tend to bind to sequences in open chromatin regions in vivo, such TF binding preference is seldomly considered by these existing tools. RESULTS: In this study, we developed TRAFICA, an open chromatin language model to predict TF-DNA binding affinity by integrating sequence characteristics of open chromatin regions from ATAC-seq experiments and in vitro TF-DNA binding profiles from high-throughput technologies. We pretrained TRAFICA on over 2.8 million nucleotide sequences in open chromatin regions derived from 197 ATAC-seq experiments (115 cell lines) to learn in vivo TF binding preferences. We further fine-tuned TRAFICA using low-rank adaptation (LoRA) on PBM and HT-SELEX TF-DNA binding profiles to learn intrinsic binding preferences for specific TFs. We systematically evaluated TRAFICA and compared its predictive performance with existing prediction tools and advanced DNA language models. The experimental results demonstrated that TRAFICA significantly outperformed the others in predicting in vitro and in vivo TF-DNA binding affinity, achieving state-of-the-art performance. These findings indicate that considering the sequence characteristics from open chromatin regions could significantly improve TF-DNA binding affinity prediction. AVAILABILITY AND IMPLEMENTATION: The source code of TRAFICA and detailed tutorials are available at https://github.com/ericcombiolab/TRAFICA.
Chonghao Wang, Aiping Lyu, Lu Zhang 0061
Bioinform.2
2025 DBNX: A Machine Learning Method for Ensembling Polygenic Risk Scores and Non-Genetic Factors
abstract
Polygenic risk scoring (PRS) holds promise for improving disease prediction and medical treatments by evaluating an individual's genetic susceptibility through multiple genetic variants. However, current PRS calculation methods often excel only in specific diseases and populations, with no single approach consistently outperforming others across all contexts. Furthermore, these methods frequently overlook non-genetic factors, such as lifestyle, that also impact disease risk.We introduce an unsupervised Deep Belief Network (DBN) to aggregate PRS generated by various methods, achieving performance comparable to the Super Learner method-a supervised ensemble approach that combines predictions from multiple methods to improve outcomes. Unlike supervised methods, the DBN does not require training data and can directly ensemble the available PRS. Remarkably, on small-scale datasets, the DBN outperforms the Super Learner. Additionally, we present the DBNX model, which integrates PRS with non-genetic factors using a combination of DBN and XGBoost. DBNX produces a Composite Risk Score (CRS) that incorporates information from both PRS and non-genetic factors. In our experiments using the U.K. Biobank (UKBB) dataset across four diseases, DBNX demonstrated superior performance compared to other commonly used ensemble methods.
Xiangzhe Yuan, Chonghao Wang, Shuqin Zhu, Lu Zhang 0061
IEEE Trans. Comput. Biol. Bioinform.2
2023 A comprehensive investigation of statistical and machine learning approaches for predicting complex human diseases on genomic variants
abstract
Quantifying an individual's risk for common diseases is an important goal of precision health. The polygenic risk score (PRS), which aggregates multiple risk alleles of candidate diseases, has emerged as a standard approach for identifying high-risk individuals. Although several studies have been performed to benchmark the PRS calculation tools and assess their potential to guide future clinical applications, some issues remain to be further investigated, such as lacking (i) various simulated data with different genetic effects; (ii) evaluation of machine learning models and (iii) evaluation on multiple ancestries studies. In this study, we systematically validated and compared 13 statistical methods, 5 machine learning models and 2 ensemble models using simulated data with additive and genetic interaction models, 22 common diseases with internal training sets, 4 common diseases with external summary statistics and 3 common diseases for trans-ancestry studies in UK Biobank. The statistical methods were better in simulated data from additive models and machine learning models have edges for data that include genetic interactions. Ensemble models are generally the best choice by integrating various statistical methods. LDpred2 outperformed the other standalone tools, whereas PRS-CS, lassosum and DBSLMM showed comparable performance. We also identified that disease heritability strongly affected the predictive performance of all methods. Both the number and effect sizes of risk SNPs are important; and sample size strongly influences the performance of all methods. For the trans-ancestry studies, we found that the performance of most methods became worse when training and testing sets were from different populations.
Chonghao Wang, Werner Pieter Veldsman, Lu Zhang 0061
Briefings Bioinform.1
2023 Managing Carbon Efficiency and Carbon Equity: What Information Do We Have From Embodied Carbon Emissions?
abstract
China has committed to achieving carbon neutrality by 2060, primarily focusing on reducing carbon intensity. Understanding the regional embodied carbon emissions is critical for managing carbon efficiency and carbon equity in this process. Using the input-output and SBM-DEA models, this article first calculates China's regional embodied carbon emissions. The results reveal substantial carbon transfers between China's different regions. Therefore, designing reduction pathways solely based on production-based carbon emissions raises fairness concerns. To address this, this article employs the SBM-DEA model to calculate the regional reduction potential and marginal abatement costs using the regional embodied carbon emissions and optimize China's pathways of regional carbon emission reduction. The new pathways consider consumption-based reduction potential, marginal abatement cost, and reduction equity, and are all in line with China's carbon reduction target. These schemes are of practical significance for China to develop a more efficient and equitable regional emission reduction plan.
Chonghao Wang, Boqiang Lin
J. Glob. Inf. Manag.1
2023 Assessing the Impact of Regional Industrial Relocation in China: Based on the Information Taken From a Multi-Regional Input-Output Analysis
abstract
Through innovative application of the multi-regional input-output model (MRIO) and spatial econometric methods, this paper investigates the trends, scale, and environmental impacts of China's industrial relocation, providing new information from an input-output perspective. The findings indicate that the relocation of China's industrial sector has exhibited a distinctive trend of moving “westward” and “northward,” while the service sector has demonstrated a tendency to cluster in several developed regions. Moreover, the authors have identified that the energy efficiency in net inflow regions and other regions is affected differently by industrial relocation. Specifically, the net inflow of the industrial sector positively impacts the energy intensity of local provinces, but negatively affects neighboring provinces. Conversely, the net inflow of the service sector has the opposite effect. The research enriches the understanding of China's industrial relocation and provides targeted implications to further prove the high-quality of China's industrial relocation.
Chonghao Wang, Boqiang Lin
J. Glob. Inf. Manag.1
2022 A machine learning model for disease risk prediction by integrating genetic and non-genetic factors
abstract
Polygenic risk score (PRS) has been widely used to identify the high-risk individuals from the general population, which would be helpful for disease prevention and early treatment. Many methods have been developed to calculate PRS by weighting and aggregating the phenotype-associated risk alleles from genome-wide association studies. However, only considering genetic effects may not be sufficient for risk prediction because the disease risk is not only related to genetic factors but also non-genetic factors, e.g., diet, physical exercise et al. But it is still a challenge to integrate these genetic and non-genetic factors into a unified machine learning framework for disease risk prediction. In this paper, we proposed PRSIMD (PRS Integrating Multi-source Data), a machine learning model that applies posterior regularization to integrate genetic and non-genetic factors to improve disease risk prediction. Also, we applied Mendelian Randomization analysis to identify the causal non-genetic risk factors for the selected diseases. We applied PRSIMD to predict type 2 diabetes and coronary artery disease from UK Biobank and observed that PRSIMD was significantly better than the existing methods to calculate PRS. In addition, we observed that PRSIMD achieved the better predictive power than the composite risk score. The codes of PRSIMD are available at: https://github.con ericcombiolab/PRSIMD
Chonghao Wang, Yunpeng Cai, Ouzhou Young, Aiping Lyu, Lu Zhang 0061
BIBM2