VLDB 2026 Research / reviewers in the wild / expert
Kei Hang Katie Chan
dblp:274/4441
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0001-9070-5394ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An integrative multi-omics framework for decoding microglial ecosystems in Alzheimer's diseaseabstractAbstract Background Alzheimer’s disease involves complex cellular alterations, yet current methods analyze cell states, signaling, and genetic risk in isolation, preventing systems-level understanding. Methods We developed an integrative framework combining quasi-binomial compositional analysis, scDemon [1], LIANA [2], scFates [3], and scDRS [4], applied to 12 integrated snRNA-seq datasets from human entorhinal and prefrontal cortex. Results Analysis revealed coordinated cellular alterations with inhibitory neuron depletion and microglia expansion. scDemon identified novel microglial states including a filopedia dynamics module (MYO10/PARVG). Trajectory analysis showed progression from homeostatic (P2RY12-high) to proliferative (APOE, AXL-high) and senescent (CDKN1A-high) states. LIANA implicated RTN4-LINGO1 signaling in impaired neuronal repair, while scDRS mapped disease genetic risk to microglial cells. Conclusion Our framework links cellular pathophysiology to genetic etiology, providing a blueprint for identifying therapeutic targets in neurodegenerative disease. References 1. Mathys H, Boix CA, Akay LA et al. ‘Single-cell multiregion dissection of Alzheimer’s disease.’ Nature 2024;632:858–868. 2. Dimitrov D, Schäfer PSL, Farr E et al. ‘LIANA+ provides an all-in-one framework for cell–cell communication inference.’ Nature Cell Biology 2024;26:1613–1622. 3. Faure L, Soldatov R, Kharchenko PV et al. ‘scFates: a scalable python package for advanced pseudotime and bifurcation analysis from single-cell data.’ Bioinformatics 2022;39. 4. Zhang MJ, Hou K, Dey KK et al. ‘Polygenic enrichment distinguishes disease associations of individual cells in single-cell RNA-seq data.’ Nature Genetics 2022;54:1572–1580. Chuyun Zhang, Kei Hang Katie Chan |
Briefings Bioinform. | 2 |
| 2024 | DeepGRNCS: deep learning-based framework for jointly inferring gene regulatory networks across cell subpopulationsabstractInferring gene regulatory networks (GRNs) allows us to obtain a deeper understanding of cellular function and disease pathogenesis. Recent advances in single-cell RNA sequencing (scRNA-seq) technology have improved the accuracy of GRN inference. However, many methods for inferring individual GRNs from scRNA-seq data are limited because they overlook intercellular heterogeneity and similarities between different cell subpopulations, which are often present in the data. Here, we propose a deep learning-based framework, DeepGRNCS, for jointly inferring GRNs across cell subpopulations. We follow the commonly accepted hypothesis that the expression of a target gene can be predicted based on the expression of transcription factors (TFs) due to underlying regulatory relationships. We initially processed scRNA-seq data by discretizing data scattering using the equal-width method. Then, we trained deep learning models to predict target gene expression from TFs. By individually removing each TF from the expression matrix, we used pre-trained deep model predictions to infer regulatory relationships between TFs and genes, thereby constructing the GRN. Our method outperforms existing GRN inference methods for various simulated and real scRNA-seq datasets. Finally, we applied DeepGRNCS to non-small cell lung cancer scRNA-seq data to identify key genes in each cell subpopulation and analyzed their biological relevance. In conclusion, DeepGRNCS effectively predicts cell subpopulation-specific GRNs. The source code is available at https://github.com/Nastume777/DeepGRNCS. Yahui Lei, Xingli Guo, Kei Hang Katie Chan, Lin Gao 0006 |
Briefings Bioinform. | 4 |
| 2022 | Predicting the functional effects of human non-coding variants based on stacking ensemble learningabstractPredicting the functional impact of genetic variants in non-coding regions of the human genome can aid in the elucidation of the etiology of diseases or traits. In recent years, an increasing number of methods to predict the impact of sequence variation in non-coding regions of the human genome have been developed. However, most of current studies are limited to predict specific types of non-coding variants. To address this problem, here we propose a non-coding SNVs prediction method based on stacking integration strategy. The method consists of three stacking models built using the same strategy based on different causality assumptions to predict functional, pathogenic, and cancer driver non-coding SNVs, respectively. We demonstrate that our method outperforms the other seven methods. In addition, a comparison of our proposed model with other methods for non-coding de novo mutations in autism spectrum disease reveals that our model has the highest discriminative ability, indicating that its performance is stable and superior in different scenarios. Kei Hang Katie Chan, Lin Gao 0006 |
BIBM | 3 |
| 2022 | IMRDriver: coding and non-coding cancer driver genes identification based on network propagationabstractIn cancer genomics, the identification of Cancer Driver Genes (CDGs) is a major scientific interest. CDGs can be identified by numerous methods, however the false positive rate still remains high. In addition, non-coding genes, such as miRNAs, can also operate as CDGs due to their regulatory functions in the development of cancer. In this paper, we present IMRDriver, a novel method for identifying both protein-coding and non-coding CDGs based on network propagation. The method first employs gene expression data, copy number variation data, single nucleotide variation data, and gene interaction data to construct a node-weighted gene network. Then, the network topology is combined with the reverse network propagation to rank all genes, with the top ranked genes predicted to be CDG candidates. We compared the prediction results of IMRDriver to twelve other methods and found that IMRDriver outperforms in terms of accuracy, recall, and F1 score. In addition, IMRDriver identified a number of miRNAs as non-coding CDGs, the majority of which have been verified in the scientific literature. In summary, IMRDriver is an effective approach for predicting CDGs. Source code of our paper is available at https://github.com/cczxsong/IMRDriver. Kei Hang Katie Chan, Lin Gao 0006 |
BIBM | 3 |
| 2022 | Canary: an automated tool for the conversion of MaCH imputed dosage files to PLINK filesabstractBACKGROUND: Previous studies have demonstrated the value of re-analysing publicly available genetics data with recent analytical approaches. Publicly available datasets, such as the Women's Health Initiative (WHI) offered by the database of genotypes and phenotypes (dbGaP), provide a wealthy resource for researchers to perform multiple analyses, including Genome-Wide Association Studies. Often, the genetic information of individuals in these datasets are stored in imputed dosage files output by MaCH; mldose and mlinfo files. In order for researchers to perform GWAS studies with this data, they must first be converted to a file format compatible with their tool of choice e.g., PLINK. Currently, there is no published tool which easily converts the datasets provided in MACH dosage files into PLINK-ready files. RESULTS: Herein, we present Canary a singularity-based tool which converts MaCH dosage files into PLINK-compatible files with a single line of user input at the command line. Further, we provide a detailed tutorial on preparation of phenotype files. Moreover, Canary comes with preinstalled software often used during GWAS studies, to further increase the ease-of-use of HPC systems for researchers. CONCLUSIONS: Until now, conversion of imputed data in the form of MaCH mldose and mlinfo files needed to be completed manually. Canary uses singularity container technology to allow users to automatically convert these MaCH files into PLINK compatible files. Additionally, Canary provides researchers with a platform to conduct GWAS analysis more easily as it contains essential software needed for conducting GWAS studies, such as PLINK and Bioconductor. We hope that this tool will greatly increase the ease at which researchers can perform GWAS with imputed data, particularly on HPC environments. Adam N. Bennett, Jethro Rainford, Kei Hang Katie Chan |
BMC Bioinform. | 5 |