EDBT 2026 Demo / reviewers in the wild / expert
Xiaoquan Wen
dblp:50/1205 · also William Wen
· DBLP profile ↗
9ranked-venue papers
1as first author
3since 2021 · last 2026
0000-0001-8990-2737ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
7 papers |
Bioinformatics and computational biology · 59% Computational science and engineering · 41% | |
| Software engineering, system software, and programming languages
1 paper |
Runtime systems and virtual machines · 50% Compilers and program optimization · 50% |
Topics — the 18 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computational science and engineering
replicability analysis |
1.7 | 2 | 2026 | Diagnosing scientific replicability through probabilistic distinguishability · Bioinform. 2026 JUMP: replicability analysis of high-throughput experiments with applications to spatial transcriptomic studies · Bioinform. 2023 |
Computational science and engineering › computational reproducibility
reproducibility assessment |
1.0 | 1 | 2026 | Diagnosing scientific replicability through probabilistic distinguishability · Bioinform. 2026 |
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation |
0.8 | 1 | 2024 | PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation · ASPLOS (2) 2024 |
Bioinformatics and computational biology › biostatistics › statistical bioinformatics
statistical genomics |
0.7 | 1 | 2023 | JUMP: replicability analysis of high-throughput experiments with applications to spatial transcriptomic studies · Bioinform. 2023 |
Bioinformatics and computational biology
genomics |
0.5 | 2 | 2018 | QuASAR-MPRA: accurate allele-specific analysis for massively parallel reporter assays · Bioinform. 2018 QuASAR: quantitative allele-specific analysis of reads · Bioinform. 2015 |
Bioinformatics and computational biology › functional genomics › functional enrichment analysis
gene set enrichment analysis |
0.4 | 1 | 2020 | BAGSE: a Bayesian hierarchical model approach for gene set enrichment analysis · Bioinform. 2020 |
Bioinformatics and computational biology › functional genomics
eQTL mapping |
0.4 | 2 | 2026 | Diagnosing scientific replicability through probabilistic distinguishability · Bioinform. 2026 QuASAR: quantitative allele-specific analysis of reads · Bioinform. 2015 |
Bioinformatics and computational biology › statistical genetics
allele-specific analysis |
0.3 | 1 | 2018 | QuASAR-MPRA: accurate allele-specific analysis for massively parallel reporter assays · Bioinform. 2018 |
Bioinformatics and computational biology
gene regulation |
0.3 | 1 | 2018 | QuASAR-MPRA: accurate allele-specific analysis for massively parallel reporter assays · Bioinform. 2018 |
Bioinformatics and computational biology › gene regulation
regulatory variant analysis |
0.3 | 1 | 2018 | QuASAR-MPRA: accurate allele-specific analysis for massively parallel reporter assays · Bioinform. 2018 |
Machine learning › Deep learning architectures and training › deep learning systems
deep learning framework |
0.2 | 1 | 2024 | PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation · ASPLOS (2) 2024 |
Bioinformatics and computational biology › gene expression analysis › gene expression quantification
allele-specific expression |
0.2 | 1 | 2015 | QuASAR: quantitative allele-specific analysis of reads · Bioinform. 2015 |
Bioinformatics and computational biology › transcriptomics
spatial transcriptomics |
0.2 | 1 | 2023 | JUMP: replicability analysis of high-throughput experiments with applications to spatial transcriptomic studies · Bioinform. 2023 |
Bioinformatics and computational biology › genomics
genome-wide association study |
0.1 | 1 | 2008 | Association studies for untyped markers with TUNA · Bioinform. 2008 |
Bioinformatics and computational biology › genomics › genotyping
genotype imputation |
0.1 | 1 | 2008 | Association studies for untyped markers with TUNA · Bioinform. 2008 |
Bioinformatics and computational biology › population genetics › coalescent theory
ancestral recombination graph |
0.1 | 1 | 2005 | Association mapping and fine mapping with TreeLD · Bioinform. 2005 |
Bioinformatics and computational biology › statistical genetics › genetic association study
association mapping |
0.1 | 1 | 2005 | Association mapping and fine mapping with TreeLD · Bioinform. 2005 |
Bioinformatics and computational biology › statistical genetics
fine-mapping |
0.1 | 1 | 2005 | Association mapping and fine mapping with TreeLD · Bioinform. 2005 |
Methods — techniques the papers use, named apart from their topics
just-in-time compilation · 1.5graph compilation · 1.5bayesian p-value · 1.0bayesian model criticism · 1.0step-up procedure · 0.7false discovery rate control · 0.7expectation-maximization · 0.4empirical bayes · 0.4statistical testing · 0.3beta-binomial model · 0.3statistical learning · 0.2genotype inference · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diagnosing scientific replicability through probabilistic distinguishabilityabstractMOTIVATION: Despite the widely recognized importance of replicability in biological research, computational methods to quantify irreplicability and identify irreplicable instances remain underdeveloped. This article presents an efficient and robust computational framework to address this gap. RESULTS: To tackle the challenge of defining an acceptable level of intrinsic heterogeneity among replicable studies, we introduce a distinguishability criterion, ensuring that replicable effects, while potentially heterogeneous, can be distinguished from zero effects and maintain consistent directions with high probability. We implement a Bayesian model criticism approach, reporting a Bayesian P-value to identify potential irreplicable instances. Through numerical experiments, we demonstrate the efficacy of the proposed methods in detecting batch effects in high-throughput experiments and identifying instances of the publication bias. Finally, we apply the framework to multi-tissue eQTL data from the GTEx consortium, uncovering tissue-specific eQTLs that represent biological heterogeneity across tissues. AVAILABILITY AND IMPLEMENTATION: An R package DiscRep implementing our method is available on GitHub (https://github.com/PengWang96/DiscRep). Hongyuan Cao, Xiaoquan Wen |
Bioinform. | 3 |
| 2024 | PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationabstractThis paper introduces two extensions to the popular PyTorch machine learning framework, TorchDynamo and TorchInductor, which implement the torch.compile feature released in PyTorch 2. TorchDynamo is a Python-level just-in-time (JIT) compiler that enables graph compilation in PyTorch programs without sacrificing the flexibility of Python. It achieves this by dynamically modifying Python bytecode before execution and extracting sequences of PyTorch operations into an FX graph, which is then JIT compiled using one of many extensible backends. TorchInductor is the default compiler backend for TorchDynamo, which translates PyTorch programs into OpenAI's Triton for GPUs and C++ for CPUs. Results show that TorchDynamo is able to capture graphs more robustly than prior approaches while adding minimal overhead, and TorchInductor is able to provide a 2.27× inference and 1.41× training geometric mean speedup on an NVIDIA A100 GPU across 180+ real-world models, which outperforms six other compilers. These extensions provide a new way to apply optimizations through compilers in eager mode frameworks like PyTorch. Jason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein, Animesh Jain, Michael Voznesensky, Bin Bao, Peter Bell 0008, David Berard, Evgeni Burovski, Geeta Chauhan, Anjali Chourdia, Will Constable, Alban Desmaison, Zach DeVito, Elias Ellison, Will Feng, Jiong Gong, Michael Gschwind, Brian Hirsh, Sherlock Huang, Kshiteej Kalambarkar, Laurent Kirsch, Michael Lazos, Mario Lezcano Casado, Yanbo Liang, Jason Liang, Yinghai Lu, C. K. Luk, Bert Maher, Yunjie Pan, Christian Puhrsch, Matthias Reso, Mark Saroufim, Marcos Yukio Siraichi, Helen Suk, Shunting Zhang, Michael Suo, Phil Tillet, Xu Zhao 0004, Eikan Wang, Keren Zhou 0001, Richard Zou, Ajit Mathews, Xiaoquan Wen, Gregory Chanan, Peng Wu 0001, Soumith Chintala |
ASPLOS (2) | 46 |
| 2023 | JUMP: replicability analysis of high-throughput experiments with applications to spatial transcriptomic studiesabstractMOTIVATION: Replicability is the cornerstone of scientific research. The current statistical method for high-dimensional replicability analysis either cannot control the false discovery rate (FDR) or is too conservative. RESULTS: We propose a statistical method, JUMP, for the high-dimensional replicability analysis of two studies. The input is a high-dimensional paired sequence of p-values from two studies and the test statistic is the maximum of p-values of the pair. JUMP uses four states of the p-value pairs to indicate whether they are null or non-null. Conditional on the hidden states, JUMP computes the cumulative distribution function of the maximum of p-values for each state to conservatively approximate the probability of rejection under the composite null of replicability. JUMP estimates unknown parameters and uses a step-up procedure to control FDR. By incorporating different states of composite null, JUMP achieves a substantial power gain over existing methods while controlling the FDR. Analyzing two pairs of spatially resolved transcriptomic datasets, JUMP makes biological discoveries that otherwise cannot be obtained by using existing methods. AVAILABILITY AND IMPLEMENTATION: An R package JUMP implementing the JUMP method is available on CRAN (https://CRAN.R-project.org/package=JUMP). Pengfei Lyu, Yan Li 0061, Xiaoquan Wen, Hongyuan Cao |
Bioinform. | 3 |
| 2020 | BAGSE: a Bayesian hierarchical model approach for gene set enrichment analysisabstractMOTIVATION: Gene set enrichment analysis has been shown to be effective in identifying relevant biological pathways underlying complex diseases. Existing approaches lack the ability to quantify the enrichment levels accurately, hence preventing the enrichment information to be further utilized in both upstream and downstream analyses. A modernized and rigorous approach for gene set enrichment analysis that emphasizes both hypothesis testing and enrichment estimation is much needed. RESULTS: We propose a novel computational method, Bayesian Analysis of Gene Set Enrichment (BAGSE), for gene set enrichment analysis. BAGSE is built on a Bayesian hierarchical model and fully accounts for the uncertainty embedded in the association evidence of individual genes. We adopt an empirical Bayes inference framework to fit the proposed hierarchical model by implementing an efficient EM algorithm. Through simulation studies, we illustrate that BAGSE yields accurate enrichment quantification while achieving similar power as the state-of-the-art methods. Further simulation studies show that BAGSE can effectively utilize the enrichment information to improve the power in gene discovery. Finally, we demonstrate the application of BAGSE in analyzing real data from a differential expression experiment and a transcriptome-wide association study. Our results indicate that the proposed statistical framework is effective in aiding the discovery of potentially causal pathways and gene networks. AVAILABILITY AND IMPLEMENTATION: BAGSE is implemented using the C++ programing language and is freely available from https://github.com/xqwen/bagse/. Simulated and real data used in this paper are also available at the Github repository for reproducibility purposes. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Abhay Hukku, Corbin Quick, Francesca Luca, Roger Pique-Regi, Xiaoquan Wen |
Bioinform. | 5 |
| 2018 | QuASAR-MPRA: accurate allele-specific analysis for massively parallel reporter assaysabstractMotivation: The majority of the human genome is composed of non-coding regions containing regulatory elements such as enhancers, which are crucial for controlling gene expression. Many variants associated with complex traits are in these regions, and may disrupt gene regulatory sequences. Consequently, it is important to not only identify true enhancers but also to test if a variant within an enhancer affects gene regulation. Recently, allele-specific analysis in high-throughput reporter assays, such as massively parallel reporter assays (MPRAs), have been used to functionally validate non-coding variants. However, we are still missing high-quality and robust data analysis tools for these datasets. Results: We have further developed our method for allele-specific analysis QuASAR (quantitative allele-specific analysis of reads) to analyze allele-specific signals in barcoded read counts data from MPRA. Using this approach, we can take into account the uncertainty on the original plasmid proportions, over-dispersion, and sequencing errors. The provided allelic skew estimate and its standard error also simplifies meta-analysis of replicate experiments. Additionally, we show that a beta-binomial distribution better models the variability present in the allelic imbalance of these synthetic reporters and results in a test that is statistically well calibrated under the null. Applying this approach to the MPRA data, we found 602 SNPs with significant (false discovery rate 10%) allele-specific regulatory function in LCLs. We also show that we can combine MPRA with QuASAR estimates to validate existing experimental and computational annotations of regulatory variants. Our study shows that with appropriate data analysis tools, we can improve the power to detect allelic effects in high-throughput reporter assays. Availability and implementation: http://github.com/piquelab/QuASAR/tree/master/mpra. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available online at Bioinformatics. Cynthia A. Kalita, Gregory A. Moyerbrailean, Xiaoquan Wen, Francesca Luca, Roger Pique-Regi |
Bioinform. | 4 |
| 2015 | QuASAR: quantitative allele-specific analysis of readsabstractMOTIVATION: Expression quantitative trait loci (eQTL) studies have discovered thousands of genetic variants that regulate gene expression, enabling a better understanding of the functional role of non-coding sequences. However, eQTL studies are costly, requiring large sample sizes and genome-wide genotyping of each sample. In contrast, analysis of allele-specific expression (ASE) is becoming a popular approach to detect the effect of genetic variation on gene expression, even within a single individual. This is typically achieved by counting the number of RNA-seq reads matching each allele at heterozygous sites and testing the null hypothesis of a 1:1 allelic ratio. In principle, when genotype information is not readily available, it could be inferred from the RNA-seq reads directly. However, there are currently no existing methods that jointly infer genotypes and conduct ASE inference, while considering uncertainty in the genotype calls. RESULTS: We present QuASAR, quantitative allele-specific analysis of reads, a novel statistical learning method for jointly detecting heterozygous genotypes and inferring ASE. The proposed ASE inference step takes into consideration the uncertainty in the genotype calls, while including parameters that model base-call errors in sequencing and allelic over-dispersion. We validated our method with experimental data for which high-quality genotypes are available. Results for an additional dataset with multiple replicates at different sequencing depths demonstrate that QuASAR is a powerful tool for ASE analysis when genotypes are not available. AVAILABILITY AND IMPLEMENTATION: http://github.com/piquelab/QuASAR. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary Material is available at Bioinformatics online. Chris T. Harvey, Gregory A. Moyerbrailean, Gordon O. Davis, Xiaoquan Wen, Francesca Luca, Roger Pique-Regi |
Bioinform. | 4 |
| 2008 | Invited Keynote Talk: Set-Level Analyses for Genome-Wide Association Data
Dan L. Nicolae, Omar De la Cruz, Xiaoquan Wen, Baoguan Ke, Minsun Song |
ISBRA | 3 |
| 2008 | Association studies for untyped markers with TUNAabstractUNLABELLED: The software package TUNA (Testing UNtyped Alleles) implements a fast and efficient algorithm for testing association of genotyped and ungenotyped variants in genome-wide case-control studies. TUNA uses Linkage Disequilibrium (LD) information from existing comprehensive variation datasets such as HapMap to construct databases of frequency predictors using linear combination of haplotype frequencies of genotyped SNPs. The predictors are used to estimate untyped allele frequencies, and to perform association tests. The methods incorporated in TUNA achieve great accuracy in estimation, and the software is computationally efficient and does not demand a lot of system memory and CPU resources. AVAILABILITY: The software package is available for download from the website: http://www.stat.uchicago.edu/~wen/tuna/. Xiaoquan Wen, Dan L. Nicolae |
Bioinform. | 1 |
| 2005 | Association mapping and fine mapping with TreeLDabstractSummary: The program package TreeLD implements a unified approach to association mapping and fine mapping of complex trait loci and a novel approach to visualizing association data, based on an inferred ancestry of the sample. Fundamentally, the TreeLD approach is based on the idea that the evidence for association at a particular position is contained in the ancestral tree relating the sampled chromosomes at that position. TreeLD provides an easy-to-use interface and can be applied to case–control, TDT trio and quantitative trait data. Availability: The program TreeLD is available on the Internet from http://pritch.bsd.uchicago.edu/software.html Contact: [email protected] Sebastian Zöllner, Xiaoquan Wen, Jonathan K. Pritchard |
Bioinform. | 2 |