EDBT 2026 Demo / reviewers in the wild / expert
Ping Zeng
dblp:05/5938
· DBLP profile ↗
14ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 10 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CEDR: robust consensus cancer subtyping with multi-omics data via ensemble dimensionality reductionabstractCancer is a highly heterogeneous disease underpinned by complex molecular alterations. Accurate subtyping is critical for guiding personalized treatment and improving clinical outcomes. However, multi-omics data are high-dimensional, noisy, and heterogeneous across platforms, posing major challenges for reliable subtyping. To address this, dimensionality reduction is necessary to capture underlying molecular patterns in a low-dimensional space, facilitating both computational efficiency and biological interpretation. We present Consensus subtyping method with Ensemble Dimensionality Reduction for multi-omics data integration (CEDR), a consensus subtyping framework that integrates complementary linear and nonlinear dimensionality reduction methods with robust clustering and probabilistic ensemble modeling. Different from existing dimensionality reduction techniques, our framework adopts an ensemble learning framework that integrates multiple dimensionality reduction techniques with robust clustering to achieve reliable consensus cancer subtyping. We apply Optimally Tuned Robust Improper Maximum Likelihood Estimator to the concatenated low-dimensional matrix for robust subtyping, and ensemble the result with the Mixture Model for Clustering Ensembles to identify stable subtypes. Across extensive simulations, CEDR consistently outperformed conventional dimensionality reduction-based clustering, the Cluster Of Clusters Analysis (COCA) ensemble strategy, and state-of-the-art multi-omics integration algorithms (SNF and CIMLR) in both accuracy and robustness. Application to clear cell renal cell carcinoma and lower-grade glioma revealed biologically interpretable subtypes characterized by distinctive survival outcomes, pathway activities, and immune infiltration patterns. These findings demonstrate that CEDR provides a powerful and reliable strategy for multi-omics data integration and cancer subtyping, with strong potential for broader applications in high-dimensional multimodal data analysis. Hongyan Cao, Zhaoyang Xu, Shilong Lin, Gang Du, Tong Wang 0019, Juping Wang, Ruiling Fang, Ping Zeng, Hongmei Yu, Yuehua Cui |
Briefings Bioinform. | 10 |
| 2026 | An integrative association analysis for complex diseases in underrepresented groups by leveraging the trans-ethnic genetic similarityabstractGenome-wide association studies (GWASs) have been conducted primarily in European (EUR) populations, limiting insights into underrepresented groups such as East Asian (EAS), but cross-ancestry GWASs have demonstrated high trans-ethnic genetic similarity between EUR and non-EUR populations. To enhance association analysis power in EAS populations, we propose tranScore, a novel summary-statistics-based transfer learning method that leverages trans-ethnic genetic similarity through hierarchical modeling. By considering EUR as auxiliary population, tranScore performs joint testing of genetic effects in auxiliary and target populations via well-established P-value combination procedures. Simulations demonstrate that tranScore maintains control of type I error rates and provides substantial power gains for diverse genetic architectures, showing robustness against various challenges including incomplete SNP overlap and effect heterogeneity. In the real-data application of eight diseases from the China Kadoorie Biobank (CKB), after incorporating the genetic information of the EUR population, tranScore identified significantly more genes than the traditional score test which ignored such information. Approximately 41.9% of discovered genes were replicated in the Biobank Japan cohort. Overall, tranScore represents a flexible and powerful statistical approach for association analysis of complex diseases and traits through transfer learning of shared genetic similarities between the auxiliary and target populations. Jike Qi, Hongyan Cao, Ping Zeng |
Briefings Bioinform. | 8 |
| 2025 | Multi-omics data integration for enhanced cancer subtyping via interactive multi-kernel learningabstractCancer is a highly heterogeneous disease characterized by complex molecular changes. Subtypes identified through multi-omics data hold significant promise for improving prognosis and facilitating personalized precision treatment. Recent multi-omics integration methods have mostly focused on capturing complementary information from different data types, often overlooking potential interactions between omics data. Here we develop a novel method named interactive multi-kernel learning (iMKL), which incorporates omics-omics interactions alongside heterogeneous data types under the unsupervised multi-kernel learning framework, to improve subtype identification. Using the sample-similarity kernel for each dataset, we propose a joint Hadamard product strategy to capture higher-order interactive effects from different omics data types. We applied iMKL to two renal cell carcinoma (RCC) datasets-clear renal cell carcinoma (ccRCC) and type II papillary renal cell carcinoma (type II pRCC)-both including miRNA expression, mRNA expression, and DNA methylation data. Stability analysis through random sampling of patients or features demonstrated that iMKL exhibits strong robustness and accuracy in identifying patient subtypes. The identified subtypes revealed dramatic differences in patient survival, with both ccRCC and type II pRCC classified into three distinct subtypes. The findings in the real application highlight potential biomarkers associated with adverse patient outcomes and demonstrate substantial advancement in cancer subtype identification. The iMKL method effectively identifies tumor molecular subtypes that are strongly associated with clinical features and survival rates, providing valuable insights for accurate cancer subtyping, clinical decision-making, and the realization of personalized treatment strategies. Hongyan Cao, Tong Wang 0019, Zhaoyang Xu, Gaiqin Liu, Ruiling Fang, Ping Zeng, Hongmei Yu, Yuehua Cui |
Briefings Bioinform. | 9 |
| 2025 | Polygenic prediction for underrepresented populations through transfer learning by utilizing genetic similarity shared with European populationsabstractBecause current genome-wide association studies are primarily conducted in individuals of European ancestry and information disparities exist among different populations, the polygenic score derived from Europeans thus exhibits poor transferability. Borrowing the idea of transfer learning, which enables the utilization of knowledge acquired from auxiliary samples to enhance learning capability in target samples, we propose transPGS, a novel polygenic score method, for genetic prediction in underrepresented populations by leveraging genetic similarity shared between the European and non-European populations while explaining the trans-ethnic difference in linkage disequilibrium (LD) and effect sizes. We demonstrate the usefulness and robustness of transPGS in elevated prediction accuracy via individual-level and summary-level simulations and apply it to seven continuous phenotypes and three diseases in the African, Chinese, and East Asian populations of the UK Biobank and Genetic Epidemiology Research Study on Adult Health and Aging cohorts. We further reveal that distinct LD and minor allele frequency patterns across ancestral groups are responsible for the dissatisfactory portability of PGS. Yiyang Zhu, Wenying Chen, Kexuan Zhu, Shuiping Huang, Ping Zeng |
Briefings Bioinform. | 6 |
| 2023 | Leveraging trans-ethnic genetic risk scores to improve association power for complex traits in underrepresented populationsabstractTrans-ethnic genome-wide association studies have revealed that many loci identified in European populations can be reproducible in non-European populations, indicating widespread trans-ethnic genetic similarity. However, how to leverage such shared information more efficiently in association analysis is less investigated for traits in underrepresented populations. We here propose a statistical framework, trans-ethnic genetic risk score informed gene-based association mixed model (GAMM), by hierarchically modeling single-nucleotide polymorphism effects in the target population as a function of effects of the same trait in well-studied populations. GAMM powerfully integrates genetic similarity across distinct ancestral groups to enhance power in understudied populations, as confirmed by extensive simulations. We illustrate the usefulness of GAMM via the application to 13 blood cell traits (i.e. basophil count, eosinophil count, hematocrit, hemoglobin concentration, lymphocyte count, mean corpuscular hemoglobin, mean corpuscular hemoglobin concentration, mean corpuscular volume, monocyte count, neutrophil count, platelet count, red blood cell count and total white blood cell count) in Africans of the UK Biobank (n = 3204) while utilizing genetic overlap shared in Europeans (n = 746 667) and East Asians (n = 162 255). We discovered multiple new associated genes, which had otherwise been missed by existing methods, and revealed that the trans-ethnic information indirectly contributed much to the phenotypic variance. Overall, GAMM represents a flexible and powerful statistical framework of association analysis for complex traits in underrepresented populations by integrating trans-ethnic genetic similarity across well-studied populations, and helps attenuate health inequities in current genetics research for people of minority populations. Haojie Lu, Ping Zeng |
Briefings Bioinform. | 4 |
| 2022 | UCBIS: An improved consortium blockchain information system based on UBCCSPabstractBlockchain technologies have been applied in many areas, from economics, the internet of things to the industrial internet. In order to solve the issue that the Hyperledger Fabric does not currently support Chinese Commercial Cryptographic (CCC) algorithms, we extended the Blockchain Cryptographic Service Provider (BCCSP) module in the Hyperledger Fabric by upgrading the original BCCSP module to support the CCC algorithms SM2 and SM3. Furthermore, we designed a transaction process by using UBCCSP (Upgraded BCCSP), and a new smart contract also has been presented. After that, an improved consortium blockchain information system based on UBCCSP named UCBIS (Consortium Blockchain Information System based on UBCCSP) is proposed. In the Hyperledger Fabric transaction process, the identity information and transaction data are protected by the SM2 and SM3 algorithms, moreover, SM3 is also used in the construction process of smart contracts. Our smart contracts reduce the total data amount and improve query efficiency. Finally, the information query system based on UBCCSP is implemented. After being tested and analyzed, the average time for every query is only 31.162 ms in the blockchain system, which has better performance and higher query efficiency. Yatao Yang 0001, Tianxiang Lin, Peihe Liu, Ping Zeng |
Blockchain Res. Appl. | 4 |
| 2022 | Identifying pleiotropic genes for complex phenotypes with summary statistics from a perspective of composite null hypothesis testingabstractPleiotropy has important implication on genetic connection among complex phenotypes and facilitates our understanding of disease etiology. Genome-wide association studies provide an unprecedented opportunity to detect pleiotropic associations; however, efficient pleiotropy test methods are still lacking. We here consider pleiotropy identification from a methodological perspective of high-dimensional composite null hypothesis and propose a powerful gene-based method called MAIUP. MAIUP is constructed based on the traditional intersection-union test with two sets of independent P-values as input and follows a novel idea that was originally proposed under the high-dimensional mediation analysis framework. The key improvement of MAIUP is that it takes the composite null nature of pleiotropy test into account by fitting a three-component mixture null distribution, which can ultimately generate well-calibrated P-values for effective control of family-wise error rate and false discover rate. Another attractive advantage of MAIUP is its ability to effectively address the issue of overlapping subjects commonly encountered in association studies. Simulation studies demonstrate that compared with other methods, only MAIUP can maintain correct type I error control and has higher power across a wide range of scenarios. We apply MAIUP to detect shared associated genes among 14 psychiatric disorders with summary statistics and discover many new pleiotropic genes that are otherwise not identified if failing to account for the issue of composite null hypothesis testing. Functional and enrichment analyses offer additional evidence supporting the validity of these identified pleiotropic genes associated with psychiatric disorders. Overall, MAIUP represents an efficient method for pleiotropy identification. Haojie Lu, Ping Zeng |
Briefings Bioinform. | 3 |
| 2022 | Simultaneous test and estimation of total genetic effect in eQTL integrative analysis through mixed modelsabstractIntegration of expression quantitative trait loci (eQTL) into genome-wide association studies (GWASs) is a promising manner to reveal functional roles of associated single-nucleotide polymorphisms (SNPs) in complex phenotypes and has become an active research field in post-GWAS era. However, how to efficiently incorporate eQTL mapping study into GWAS for prioritization of causal genes remains elusive. We herein proposed a novel method termed as Mixed transcriptome-wide association studies (TWAS) and mediated Variance estimation (MTV) by modeling the effects of cis-SNPs of a gene as a function of eQTL. MTV formulates the integrative method and TWAS within a unified framework via mixed models and therefore includes many prior methods/tests as special cases. We further justified MTV from another two statistical perspectives of mediation analysis and two-stage Mendelian randomization. Relative to existing methods, MTV is superior for pronounced features including the processing of direct effects of cis-SNPs on phenotypes, the powerful likelihood ratio test for assessment of joint effects of cis-SNPs and genetically regulated gene expression (GReX), two useful quantities to measure relative genetic contributions of GReX and cis-SNPs to phenotypic variance, and the computationally efferent parameter expansion expectation maximum algorithm. With extensive simulations, we identified that MTV correctly controlled the type I error in joint evaluation of the total genetic effect and proved more powerful to discover true association signals across various scenarios compared to existing methods. We finally applied MTV to 41 complex traits/diseases available from three GWASs and discovered many new associated genes that had otherwise been missed by existing methods. We also revealed that a small but substantial fraction of phenotypic variation was mediated by GReX. Overall, MTV constructs a robust and realistic modeling foundation for integrative omics analysis and has the advantage of offering more attractive biological interpretations of GWAS results. Jiahao Qiao, Yongyue Wei, Ping Zeng |
Briefings Bioinform. | 5 |
| 2022 | A comprehensive comparison of multilocus association methods with summary statistics in genome-wide association studiesabstractBACKGROUND: Multilocus analysis on a set of single nucleotide polymorphisms (SNPs) pre-assigned within a gene constitutes a valuable complement to single-marker analysis by aggregating data on complex traits in a biologically meaningful way. However, despite the existence of a wide variety of SNP-set methods, few comprehensive comparison studies have been previously performed to evaluate the effectiveness of these methods. RESULTS: We herein sought to fill this knowledge gap by conducting a comprehensive empirical comparison for 22 commonly-used summary-statistics based SNP-set methods. We showed that only seven methods could effectively control the type I error, and that these well-calibrated approaches had varying power performance under the simulation scenarios. Overall, we confirmed that the burden test was generally underpowered and score-based variance component tests (e.g., sequence kernel association test) were much powerful under the polygenic genetic architecture in both common and rare variant association analyses. We further revealed that two linkage-disequilibrium-free P value combination methods (e.g., harmonic mean P value method and aggregated Cauchy association test) behaved very well under the sparse genetic architecture in simulations and real-data applications to common and rare variant association analyses as well as in expression quantitative trait loci weighted integrative analysis. We also assessed the scalability of these approaches by recording computational time and found that all these methods can be scalable to biobank-scale data although some might be relatively slow. CONCLUSION: In conclusion, we hope that our findings can offer an important guidance on how to choose appropriate multilocus association analysis methods in post-GWAS era. All the SNP-set methods are implemented in the R package called MCA, which is freely available at https://github.com/biostatpzeng/ . Zhonghe Shao, Jiahao Qiao, Shuiping Huang, Ping Zeng |
BMC Bioinform. | 6 |
| 2021 | IUSMMT: Survival mediation analysis of gene expression with multiple DNA methylation exposures and its application to cancers of TCGAabstractEffective and powerful survival mediation models are currently lacking. To partly fill such knowledge gap, we particularly focus on the mediation analysis that includes multiple DNA methylations acting as exposures, one gene expression as the mediator and one survival time as the outcome. We proposed IUSMMT (intersection-union survival mixture-adjusted mediation test) to effectively examine the existence of mediation effect by fitting an empirical three-component mixture null distribution. With extensive simulation studies, we demonstrated the advantage of IUSMMT over existing methods. We applied IUSMMT to ten TCGA cancers and identified multiple genes that exhibited mediating effects. We further revealed that most of the identified regions, in which genes behaved as active mediators, were cancer type-specific and exhibited a full mediation from DNA methylation CpG sites to the survival risk of various types of cancers. Overall, IUSMMT represents an effective and powerful alternative for survival mediation analysis; our results also provide new insights into the functional role of DNA methylation and gene expression in cancer progression/prognosis and demonstrate potential therapeutic targets for future clinical practice. Zhonghe Shao, Shuiping Huang, Ping Zeng |
PLoS Comput. Biol. | 6 |
| 2018 | Pleiotropic mapping and annotation selection in genome-wide association studies with penalized Gaussian mixture modelsabstractMotivation: Genome-wide association studies (GWASs) have identified many genetic loci associated with complex traits. A substantial fraction of these identified loci is associated with multiple traits-a phenomena known as pleiotropy. Identification of pleiotropic associations can help characterize the genetic relationship among complex traits and can facilitate our understanding of disease etiology. Effective pleiotropic association mapping requires the development of statistical methods that can jointly model multiple traits with genome-wide single nucleic polymorphisms (SNPs) together. Results: We develop a joint modeling method, which we refer to as the integrative MApping of Pleiotropic association (iMAP). iMAP models summary statistics from GWASs, uses a multivariate Gaussian distribution to account for phenotypic correlation, simultaneously infers genome-wide SNP association pattern using mixture modeling and has the potential to reveal causal relationship between traits. Importantly, iMAP integrates a large number of SNP functional annotations to substantially improve association mapping power, and, with a sparsity-inducing penalty, is capable of selecting informative annotations from a large, potentially non-informative set. To enable scalable inference of iMAP to association studies with hundreds of thousands of individuals and millions of SNPs, we develop an efficient expectation maximization algorithm based on an approximate penalized regression algorithm. With simulations and comparisons to existing methods, we illustrate the benefits of iMAP in terms of both high association mapping power and accurate estimation of genome-wide SNP association patterns. Finally, we apply iMAP to perform a joint analysis of 48 traits from 31 GWAS consortia together with 40 tissue-specific SNP annotations generated from the Roadmap Project. Availability and implementation: iMAP is freely available at http://www.xzlab.org/software.html. Supplementary information: Supplementary data are available at Bioinformatics online. Ping Zeng, Xingjie Hao |
Bioinform. | 1 |
| 2014 | Stability SCAD: a powerful approach to detect interactions in large-scale genomic studyabstractBACKGROUND: Evidence suggests that common complex diseases may be partially due to SNP-SNP interactions, but such detection is yet to be fully established in a high-dimensional small-sample (small-n-large-p) study. A number of penalized regression techniques are gaining popularity within the statistical community, and are now being applied to detect interactions. These techniques tend to be over-fitting, and are prone to false positives. The recently developed stability least absolute shrinkage and selection operator (SLASSO) has been used to control family-wise error rate, but often at the expense of power (and thus false negative results). RESULTS: Here, we propose an alternative stability selection procedure known as stability smoothly clipped absolute deviation (SSCAD). Briefly, this method applies a smoothly clipped absolute deviation (SCAD) algorithm to multiple sub-samples, and then identifies cluster ensemble of interactions across the sub-samples. The proposed method was compared with SLASSO and two kinds of traditional penalized methods by intensive simulation. The simulation revealed higher power and lower false discovery rate (FDR) with SSCAD. An analysis using the new method on the previously published GWAS of lung cancer confirmed all significant interactions identified with SLASSO, and identified two additional interactions not reported with SLASSO analysis. CONCLUSIONS: Based on the results obtained in this study, SSCAD presents to be a powerful procedure for the detection of SNP-SNP interactions in large-scale genomic data. Jianwei Gou, Yongyue Wei, Ruyang Zhang, Yongyong Qiu, Ping Zeng, Dianke Yu, Tangchun Wu, Zhibin Hu, Dongxin Lin, Hongbing Shen, Feng Chen 0032 |
BMC Bioinform. | 7 |
| 2010 | Insecure JavaScript Detection and Analysis with Browser-Enforced Embedded RulesabstractThe JavaScript language is an interpretive programming language which is used to enhance the client-side interactivity and functionality. However, it has been much exploited by malicious parties to launch browser-based security attacks. Currently there are many security vulnerabilities assessment tools, and browsers provide sand-boxing mechanisms to protect the JavaScript code from compromising the security of the client's environment, but, unfortunately, nowadays the attacks against web applications often take advantage of the browser's own function to carry out attacks. Based on the above problems, we put forward an approach to solve the problem that is based on monitoring JavaScript code execution to detect malicious code behavior and we don't need to carry out the static analysis of JavaScript code, just compare the execution to high-level inspection rules. While visiting the website we insert the security inspection rules into the website to analyze the potential safety hazard. Ping Zeng, Jianhua Sun 0002, Hao Chen 0002 |
PDCAT | 1 |
| 2007 | A Novel Authentication Scheme Based on Trust-value Updated Model in Adhoc NetworkabstractThere are many difficulties to carry out node authentication in dynamic and self-organized adhoc network. A Trust-value Updated Model (TUM) in Layered and Grouped adhoc network Structure is adopted in this paper, and based on which, we put forward a novel authentication mechanism of leader agent and member surveillance, which can cut down the data traffic of authentication between nodes, and the calculating complexity is reduced greatly, hence the node authentication efficiency is improved, the realtime communication between the nodes is ensured. Ping Zeng |
COMPSAC (1) | 4 |