VLDB 2026 Research / reviewers in the wild / expert
Yuk Yee Leung
dblp:87/9873 · also Yukyee Leung
· DBLP profile ↗
9ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-3047-5440ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
7 papers |
Bioinformatics and computational biology · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
functional genomics |
1.7 | 3 | 2025 | BTS: a scalable Bayesian Tissue Score for prioritizing GWAS variants and their functional contexts across >1000s of omics datasets · Bioinform. 2025 hipFG: high-throughput harmonization and integration pipeline for functional genomics data · Bioinform. 2023 SparkINFERNO: a scalable high-throughput pipeline for inferring molecular mechanisms of non-coding genetic variants · Bioinform. 2020 |
Bioinformatics and computational biology
statistical genetics |
1.3 | 2 | 2025 | BTS: a scalable Bayesian Tissue Score for prioritizing GWAS variants and their functional contexts across >1000s of omics datasets · Bioinform. 2025 SparkINFERNO: a scalable high-throughput pipeline for inferring molecular mechanisms of non-coding genetic variants · Bioinform. 2020 |
Bioinformatics and computational biology › functional genomics
functional enrichment analysis |
0.9 | 1 | 2025 | BTS: a scalable Bayesian Tissue Score for prioritizing GWAS variants and their functional contexts across >1000s of omics datasets · Bioinform. 2025 |
Bioinformatics and computational biology › statistical genetics › fine-mapping
genetic fine-mapping |
0.9 | 1 | 2025 | BTS: a scalable Bayesian Tissue Score for prioritizing GWAS variants and their functional contexts across >1000s of omics datasets · Bioinform. 2025 |
Bioinformatics and computational biology
transcriptomics |
0.7 | 2 | 2022 | Omnibus and robust deconvolution scheme for bulk RNA sequencing data integrating multiple single-cell reference sets and prior biological knowledge · Bioinform. 2022 DASHR 2.0: integrated database of human small non-coding RNA genes and mature products · Bioinform. 2019 |
Bioinformatics and computational biology › omics data analysis
cell-type deconvolution |
0.6 | 1 | 2022 | Omnibus and robust deconvolution scheme for bulk RNA sequencing data integrating multiple single-cell reference sets and prior biological knowledge · Bioinform. 2022 |
Bioinformatics and computational biology › genome annotation
non-coding RNA annotation |
0.4 | 1 | 2019 | DASHR 2.0: integrated database of human small non-coding RNA genes and mature products · Bioinform. 2019 |
Bioinformatics and computational biology › genomics
variant calling |
0.4 | 1 | 2019 | VCPA: genomic variant calling pipeline and data management tool for Alzheimer's Disease Sequencing Project · Bioinform. 2019 |
Bioinformatics and computational biology › genome annotation
genomic variant annotation |
0.3 | 1 | 2018 | Functional annotation of genomic variants in studies of late-onset Alzheimer's disease · Bioinform. 2018 |
Bioinformatics and computational biology › functional genomics
variant functional annotation |
0.3 | 1 | 2018 | Functional annotation of genomic variants in studies of late-onset Alzheimer's disease · Bioinform. 2018 |
Bioinformatics and computational biology › single-cell analysis
single-cell RNA sequencing |
0.2 | 1 | 2022 | Omnibus and robust deconvolution scheme for bulk RNA sequencing data integrating multiple single-cell reference sets and prior biological knowledge · Bioinform. 2022 |
Bioinformatics and computational biology › gene regulation › regulatory element
regulatory element annotation |
0.1 | 1 | 2020 | SparkINFERNO: a scalable high-throughput pipeline for inferring molecular mechanisms of non-coding genetic variants · Bioinform. 2020 |
Methods — techniques the papers use, named apart from their topics
bayesian model · 0.9GWAS summary statistics · 0.9indexing · 0.7data normalization · 0.7penalized regression · 0.6giggle genomic indexing · 0.4apache spark · 0.4workflow description language · 0.4unsupervised segmentation · 0.4GATK · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BTS: a scalable Bayesian Tissue Score for prioritizing GWAS variants and their functional contexts across >1000s of omics datasetsabstractMOTIVATION: statistics from genome-wide association studies (GWAS) are widely used in fine-mapping and colocalization analyses to identify causal variants and their enrichment in functional contexts, such as affected cell types and genomic features. With the expansion of functional genomic (FG) datasets, which now include hundreds of thousands of tracks across various cell and tissue types, it is critical to establish scalable algorithms integrating thousands of diverse FG annotations with GWAS results. RESULTS: We propose BTS (Bayesian Tissue Score), a novel, highly efficient algorithm uniquely designed for (i) identifying affected cell types and functional elements (context-mapping) and (ii) fine-mapping potentially causal variants in a context-specific manner using large collections of cell type-specific FG annotation tracks. BTS leverages GWAS summary statistics and annotation-specific Bayesian models to analyze genome-wide annotation tracks, including enhancers, open chromatin, and histone marks. We evaluated BTS on GWAS summary statistics for immune and cardiovascular traits, such as Inflammatory Bowel Disease (IBD), Rheumatoid Arthritis (RA), Systemic Lupus Erythematosus (SLE), and Coronary Artery Disease (CAD). Our results demonstrate that BTS is over 100× more efficient in estimating functional annotation effects and context-specific variant fine-mapping compared to existing methods. Importantly, this large-scale Bayesian approach prioritizes both known and novel annotations, cell types, genomic regions, and variants and provides valuable biological insights into the functional contexts of these diseases. AVAILABILITY AND IMPLEMENTATION: Docker image is available at https://hub.docker.com/r/wanglab/bts with preinstalled BTS R package (https://bitbucket.org/wanglab-upenn/BTS-R) and BTS GWAS summary statistics analysis pipeline (https://bitbucket.org/wanglab-upenn/bts-pipeline). Pavel P. Kuksa, Matei Ionita, Luke Carter, Jeffrey Cifello, Prabhakaran Gangadharan, Kaylyn Clark, Otto Valladares, Yuk Yee Leung, Li-San Wang |
Bioinform. | 8 |
| 2023 | hipFG: high-throughput harmonization and integration pipeline for functional genomics dataabstractSUMMARY: Preparing functional genomic (FG) data with diverse assay types and file formats for integration into analysis workflows that interpret genome-wide association and other studies is a significant and time-consuming challenge. Here we introduce hipFG (Harmonization and Integration Pipeline for Functional Genomics), an automatically customized pipeline for efficient and scalable normalization of heterogenous FG data collections into standardized, indexed, rapidly searchable analysis-ready datasets while accounting for FG datatypes (e.g. chromatin interactions, genomic intervals, quantitative trait loci). AVAILABILITY AND IMPLEMENTATION: hipFG is freely available at https://bitbucket.org/wanglab-upenn/hipFG. A Docker container is available at https://hub.docker.com/r/wanglab/hipfg. Jeffrey Cifello, Pavel P. Kuksa, Naveensri Saravanan, Otto Valladares, Li-San Wang, Yuk Yee Leung |
Bioinform. | 6 |
| 2022 | Omnibus and robust deconvolution scheme for bulk RNA sequencing data integrating multiple single-cell reference sets and prior biological knowledgeabstractMOTIVATION: Cell-type deconvolution of bulk tissue RNA sequencing (RNA-seq) data is an important step toward understanding the variations in cell-type composition among disease conditions. Owing to recent advances in single-cell RNA sequencing (scRNA-seq) and the availability of large amounts of bulk RNA-seq data in disease-relevant tissues, various deconvolution methods have been developed. However, the performance of existing methods heavily relies on the quality of information provided by external data sources, such as the selection of scRNA-seq data as a reference and prior biological information. RESULTS: We present the Integrated and Robust Deconvolution (InteRD) algorithm to infer cell-type proportions from target bulk RNA-seq data. Owing to the innovative use of penalized regression with a new evaluation criterion for deconvolution, InteRD has three primary advantages. First, it is able to effectively integrate deconvolution results from multiple scRNA-seq datasets. Second, InteRD calibrates estimates from reference-based deconvolution by taking into account extra biological information as priors. Third, the proposed algorithm is robust to inaccurate external information imposed in the deconvolution system. Extensive numerical evaluations and real-data applications demonstrate that InteRD yields more accurate and robust cell-type proportion estimates that agree well with known biology. AVAILABILITY AND IMPLEMENTATION: The proposed InteRD framework is implemented in R and the package is available at https://cran.r-project.org/web/packages/InteRD/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chixiang Chen, Yuk Yee Leung, Matei Ionita, Li-San Wang, Mingyao Li |
Bioinform. | 2 |
| 2020 | SparkINFERNO: a scalable high-throughput pipeline for inferring molecular mechanisms of non-coding genetic variantsabstractSUMMARY: We report Spark-based INFERence of the molecular mechanisms of NOn-coding genetic variants (SparkINFERNO), a scalable bioinformatics pipeline characterizing non-coding genome-wide association study (GWAS) association findings. SparkINFERNO prioritizes causal variants underlying GWAS association signals and reports relevant regulatory elements, tissue contexts and plausible target genes they affect. To achieve this, the SparkINFERNO algorithm integrates GWAS summary statistics with large-scale collection of functional genomics datasets spanning enhancer activity, transcription factor binding, expression quantitative trait loci and other functional datasets across more than 400 tissues and cell types. Scalability is achieved by an underlying API implemented using Apache Spark and Giggle-based genomic indexing. We evaluated SparkINFERNO on large GWASs and show that SparkINFERNO is more than 60 times efficient and scales with data size and amount of computational resources. AVAILABILITY AND IMPLEMENTATION: SparkINFERNO runs on clusters or a single server with Apache Spark environment, and is available at https://bitbucket.org/wanglab-upenn/SparkINFERNO or https://hub.docker.com/r/wanglab/spark-inferno. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Pavel P. Kuksa, Chien-Yueh Lee, Alexandre Amlie-Wolf, Prabhakaran Gangadharan, Elizabeth E. Mlynarski, Yi-Fan Chou, Han-Jen Lin, Heather Issen, Emily Greenfest-Allen, Otto Valladares, Yuk Yee Leung, Li-San Wang |
Bioinform. | 11 |
| 2019 | DASHR 2.0: integrated database of human small non-coding RNA genes and mature productsabstractMOTIVATION: Small non-coding RNAs (sncRNAs, <100 nts) are highly abundant RNAs that regulate diverse and often tissue-specific cellular processes by associating with transcription factor complexes or binding to mRNAs. While thousands of sncRNA genes exist in the human genome, no single resource provides searchable, unified annotation, expression and processing information for full sncRNA transcripts and mature RNA products derived from these larger RNAs. RESULTS: Our goal is to establish a complete catalog of annotation, expression, processing, conservation, tissue-specificity and other biological features for all human sncRNA genes and mature products derived from all major RNA classes. DASHR (Database of small human non-coding RNAs) v2.0 database is the first that integrates human sncRNA gene and mature products profiles obtained from multiple RNA-seq protocols. Altogether, 185 tissues/cell types and sncRNA annotations and >800 curated experiments from ENCODE and GEO/SRA across multiple RNA-seq protocols for both GRCh38/hg38 and GRCh37/hg19 assemblies are integrated in DASHR. Moreover, DASHR is the first to contain both known and novel, previously un-annotated sncRNA loci identified by unsupervised segmentation (13 times more loci with 1 678 800 total). Additionally, DASHR v2.0 adds >3 200 000 annotations for non-small RNA genes and other genomic features (long-noncoding RNAs, mRNAs, promoters, repeats). Furthermore, DASHR v2.0 introduces an enhanced user interface, interactive experiment-by-locus table view, sncRNA locus sorting and filtering by biological features. All annotation and expression information directly downloadable and accessible as UCSC genome browser tracks. AVAILABILITY AND IMPLEMENTATION: DASHR v2.0 is freely available at https://lisanwanglab.org/DASHRv2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Pavel P. Kuksa, Alexandre Amlie-Wolf, Zivadin Katanic, Otto Valladares, Li-San Wang, Yuk Yee Leung |
Bioinform. | 6 |
| 2019 | VCPA: genomic variant calling pipeline and data management tool for Alzheimer's Disease Sequencing ProjectabstractBioinformatics, https://doi.org/10.1093/bioinformatics/bty894 In the supplementary material, the phrase “The pairwise variant disconcordance rate” has been changed to “The pairwise variant concordance rate”. The updated supplementary file is now available online. Yuk Yee Leung, Otto Valladares, Yi-Fan Chou, Han-Jen Lin, Amanda Kuzma, Laura Cantwell, Liming Qu, Prabhakaran Gangadharan, William J. Salerno, Gerard D. Schellenberg, Li-San Wang |
Bioinform. | 1 |
| 2019 | VCPA: genomic variant calling pipeline and data management tool for Alzheimer's Disease Sequencing ProjectabstractSUMMARY: We report VCPA, our SNP/Indel Variant Calling Pipeline and data management tool used for the analysis of whole genome and exome sequencing (WGS/WES) for the Alzheimer's Disease Sequencing Project. VCPA consists of two independent but linkable components: pipeline and tracking database. The pipeline, implemented using the Workflow Description Language and fully optimized for the Amazon elastic compute cloud environment, includes steps from aligning raw sequence reads to variant calling using GATK. The tracking database allows users to view job running status in real time and visualize >100 quality metrics per genome. VCPA is functionally equivalent to the CCDG/TOPMed pipeline. Users can use the pipeline and the dockerized database to process large WGS/WES datasets on Amazon cloud with minimal configuration. AVAILABILITY AND IMPLEMENTATION: VCPA is released under the MIT license and is available for academic and nonprofit use for free. The pipeline source code and step-by-step instructions are available from the National Institute on Aging Genetics of Alzheimer's Disease Data Storage Site (http://www.niagads.org/VCPA). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yuk Yee Leung, Otto Valladares, Yi-Fan Chou, Han-Jen Lin, Amanda Kuzma, Laura Cantwell, Liming Qu, William J. Salerno, Gerard D. Schellenberg, Li-San Wang |
Bioinform. | 1 |
| 2018 | Functional annotation of genomic variants in studies of late-onset Alzheimer's diseaseabstractMotivation: Annotation of genomic variants is an increasingly important and complex part of the analysis of sequence-based genomic analyses. Computational predictions of variant function are routinely incorporated into gene-based analyses of rare-variants, though to date most studies use limited information for assessing variant function that is often agnostic of the disease being studied. Results: In this work, we outline an annotation process motivated by the Alzheimer's Disease Sequencing Project, illustrate the impact of including tissue-specific transcript sets and sources of gene regulatory information and assess the potential impact of changing genomic builds on the annotation process. While these factors only impact a small proportion of total variant annotations (∼5%), they influence the potential analysis of a large fraction of genes (∼25%). Availability and implementation: Individual variant annotations are available via the NIAGADS GenomicsDB, at https://www.niagads.org/genomics/ tools-and-software/databases/genomics-database. Annotations are also available for bulk download at https://www.niagads.org/datasets. Annotation processing software is available at http://www.icompbio.net/resources/software-and-downloads/. Supplementary information: Supplementary data are available at Bioinformatics online. Mariusz Butkiewicz, Elizabeth E. Blue, Yuk Yee Leung, Xueqiu Jian, Edoardo Marcora, Alan E. Renton, Amanda Kuzma, Li-San Wang, Daniel C. Koboldt, Jonathan L. Haines, William S. Bush |
Bioinform. | 3 |
| 2010 | A Multiple-Filter-Multiple-Wrapper Approach to Gene Selection and Microarray Data ClassificationabstractFilters and wrappers are two prevailing approaches for gene selection in microarray data analysis. Filters make use of statistical properties of each gene to represent its discriminating power between different classes. The computation is fast but the predictions are inaccurate. Wrappers make use of a chosen classifier to select genes by maximizing classification accuracy, but the computation burden is formidable. Filters and wrappers have been combined in previous studies to maximize the classification accuracy for a chosen classifier with respect to a filtered set of genes. The drawback of this single-filter-single-wrapper (SFSW) approach is that the classification accuracy is dependent on the choice of specific filter and wrapper. In this paper, a multiple-filter-multiple-wrapper (MFMW) approach is proposed that makes use of multiple filters and multiple wrappers to improve the accuracy and robustness of the classification, and to identify potential biomarker genes. Experiments based on six benchmark data sets show that the MFMW approach outperforms SFSW models (generated by all combinations of filters and wrappers used in the corresponding MFMW model) in all cases and for all six data sets. Some of MFMW-selected genes have been confirmed to be biomarkers or contribute to the development of particular cancers by other studies. Yuk Yee Leung, Yeung Sam Hung |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |