EDBT 2026 Demo / reviewers in the wild / expert
Qiwei Li 0001
dblp:35/8806-1
· DBLP profile ↗
13ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-1020-3050ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Advances in predicting omics profiles from imaging dataabstractWhile traditional imaging techniques, such as histopathology, are often part of clinical workflows, molecular profiling remains more difficult to conduct and is less cost-effective. Thus, the prediction of molecular 'omics' data directly from imaging has emerged as an appealing alternative. While existing reviews have mentioned image-based prediction of biomarkers within specific disease contexts, this review provides a comprehensive overview of current methods that leverage imaging to predict (i) DNA-based aberrations, (ii) bulk transcriptomic profiles, (iii) single-cell transcriptomics, and (iv) spatial transcriptomics across disease contexts and imaging modalities. To address the complexity of these predictive tasks, we find that many studies employ cutting-edge deep learning strategies for image processing, feature extraction, feature aggregation, and downstream molecular prediction. In this review, we highlight the diverse applications of both deep learning-based and modern statistical frameworks designed for image-based omics prediction. The insights gleaned from these inferred molecular data have broad clinical relevance and will continue to improve our understanding of the relationships between molecular and visual features, paving the way for new diagnostic and therapeutic applications. Alexa H. Beachum, Yuansheng Zhou, Qiwei Li 0001, Guanghua Xiao, Lin Xu 0005 |
Briefings Bioinform. | 4 |
| 2025 | MiCoDe: a web tool for performing microbiome community detection using a Bayesian weighted stochastic block modelabstractSUMMARY: The Microbiome Community Detector (MiCoDe) software is a free user-friendly web tool that is designed to cluster a network of microbial taxa into communities using a Bayesian weighted stochastic block model. MiCoDe also filters the data automatically and accounts for the challenges of microbiome high-throughput sequencing data including high-dimensionality, compositionality, zero inflation, and nonlinearity. While MiCoDe is based on a rigorous statistical unsupervised learning model, our web tool can be easily used by any investigator. Users simply upload a csv file that contains their taxonomic abundance data where rows correspond to samples and columns correspond to taxa. Then, users make a few selections regarding data transformation, network estimation, and the number of communities in order to run the online analysis. If users are unsure of what selections to make, then they can opt for the default settings as these are our recommended settings. In this paper, we discuss the motivation, methodology, implementation, and results of MiCoDe. We also discuss how MiCoDe can be adapted by the user and how it may evolve over time. Our software is a valuable tool for microbiome community detection. AVAILABILITY AND IMPLEMENTATION: MiCoDe is freely available online at https://lce.biohpc.swmed.edu/micode/, does not require installation, and is not browser-specific. Users can also work locally using our R code, which is freely available on GitHub at https://github.com/klutz920/MiCoDe. Kevin C. Lutz, Shengjie Yang, Tejasv Bedi, Michael L. Neugent, Nikita Madhavaram, Xiaowei Zhan, Nicole J. De Nisco, Qiwei Li 0001 |
Bioinform. | 9 |
| 2025 | BISON: bi-clustering of spatial omics data with feature selectionabstractMOTIVATION: The advent of next-generation sequencing-based spatially resolved transcriptomics (SRT) techniques has reshaped genomic studies by enabling high-throughput gene expression profiling while preserving spatial and morphological context. Understanding gene functions and interactions in different spatial domains is crucial, as it can enhance our comprehension of biological mechanisms, such as cancer-immune interactions and cell differentiation in various regions. It is necessary to cluster tissue regions into distinct spatial domains and identify discriminating genes (DGs) that elucidate the clustering result, referred to as spatial domain-specific DGs. Existing methods for identifying these genes typically rely on a two-stage approach, which can lead to the phenomenon known as double-dipping. RESULTS: To address the challenge, we propose a unified Bayesian latent block model that simultaneously detects a list of DGs contributing to spatial domain identification while clustering these DGs and spatial locations. The efficacy of our proposed method is validated through a series of simulation experiments, and its capability to identify DGs is demonstrated through applications to benchmark SRT datasets. AVAILABILITY AND IMPLEMENTATION: The R/C++ implementation of BISON is available at https://github.com/new-zbc/BISON. Bencong Zhu, Alberto Cassese, Marina Vannucci, Michele Guindani, Qiwei Li 0001 |
Bioinform. | 5 |
| 2024 | BayeSMART: Bayesian clustering of multi-sample spatially resolved transcriptomics dataabstractThe field of spatially resolved transcriptomics (SRT) has greatly advanced our understanding of cellular microenvironments by integrating spatial information with molecular data collected from multiple tissue sections or individuals. However, methods for multi-sample spatial clustering are lacking, and existing methods primarily rely on molecular information alone. This paper introduces BayeSMART, a Bayesian statistical method designed to identify spatial domains across multiple samples. BayeSMART leverages artificial intelligence (AI)-reconstructed single-cell level information from the paired histology images of multi-sample SRT datasets while simultaneously considering the spatial context of gene expression. The AI integration enables BayeSMART to effectively interpret the spatial domains. We conducted case studies using four datasets from various tissue types and SRT platforms, and compared BayeSMART with alternative multi-sample spatial clustering approaches and a number of state-of-the-art methods for single-sample SRT analysis, demonstrating that it surpasses existing methods in terms of clustering accuracy, interpretability, and computational efficiency. BayeSMART offers new insights into the spatial organization of cells in multi-sample SRT data. Yanghong Guo, Bencong Zhu, Ruichen Rong, Guanghua Xiao, Lin Xu 0005, Qiwei Li 0001 |
Briefings Bioinform. | 8 |
| 2021 | Spatial molecular profiling: platforms, applications and analysis toolsabstractMolecular profiling technologies, such as genome sequencing and proteomics, have transformed biomedical research, but most such technologies require tissue dissociation, which leads to loss of tissue morphology and spatial information. Recent developments in spatial molecular profiling technologies have enabled the comprehensive molecular characterization of cells while keeping their spatial and morphological contexts intact. Molecular profiling data generate deep characterizations of the genetic, transcriptional and proteomic events of cells, while tissue images capture the spatial locations, organizations and interactions of the cells together with their morphology features. These data, together with cell and tissue imaging data, provide unprecedented opportunities to study tissue heterogeneity and cell spatial organization. This review aims to provide an overview of these recent developments in spatial molecular profiling technologies and the corresponding computational methods developed for analyzing such data. Minzhe Zhang, Thomas Sheffield, Xiaowei Zhan, Qiwei Li 0001, Donghan M. Yang, Yunguan Wang, Shidan Wang, Guanghua Xiao |
Briefings Bioinform. | 4 |
| 2021 | Bayesian modeling of spatial molecular profiling data via Gaussian processabstractMOTIVATION: The location, timing and abundance of gene expression (both mRNA and proteins) within a tissue define the molecular mechanisms of cell functions. Recent technology breakthroughs in spatial molecular profiling, including imaging-based technologies and sequencing-based technologies, have enabled the comprehensive molecular characterization of single cells while preserving their spatial and morphological contexts. This new bioinformatics scenario calls for effective and robust computational methods to identify genes with spatial patterns. RESULTS: We represent a novel Bayesian hierarchical model to analyze spatial transcriptomics data, with several unique characteristics. It models the zero-inflated and over-dispersed counts by deploying a zero-inflated negative binomial model that greatly increases model stability and robustness. Besides, the Bayesian inference framework allows us to borrow strength in parameter estimation in a de novo fashion. As a result, the proposed model shows competitive performances in accuracy and robustness over existing methods in both simulation studies and two real data applications. AVAILABILITY AND IMPLEMENTATION: The related R/C++ source code is available at https://github.com/Minzhe/BOOST-GP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qiwei Li 0001, Minzhe Zhang, Guanghua Xiao |
Bioinform. | 1 |
| 2019 | GeNeCK: a web server for gene network construction and visualizationabstractBACKGROUND: Reverse engineering approaches to infer gene regulatory networks using computational methods are of great importance to annotate gene functionality and identify hub genes. Although various statistical algorithms have been proposed, development of computational tools to integrate results from different methods and user-friendly online tools is still lagging. RESULTS: We developed a web server that efficiently constructs gene networks from expression data. It allows the user to use ten different network construction methods (such as partial correlation-, likelihood-, Bayesian- and mutual information-based methods) and integrates the resulting networks from multiple methods. Hub gene information, if available, can be incorporated to enhance performance. CONCLUSIONS: GeNeCK is an efficient and easy-to-use web application for gene regulatory network construction. It can be accessed at http://lce.biohpc.swmed.edu/geneck . Minzhe Zhang, Qiwei Li 0001, Donghyeon Yu, Guanghua Xiao |
BMC Bioinform. | 2 |
| 2016 | KScons: a Bayesian approach for protein residue contact prediction using the knob-socket model of protein tertiary structureabstractMOTIVATION: By simplifying the many-bodied complexity of residue packing into patterns of simple pairwise secondary structure interactions between a single knob residue with a three-residue socket, the knob-socket construct allows a more direct incorporation of structural information into the prediction of residue contacts. By modeling the preferences between the amino acid composition of a socket and knob, we undertake an investigation of the knob-socket construct's ability to improve the prediction of residue contacts. The statistical model considers three priors and two posterior estimations to better understand how the input data affects predictions. This produces six implementations of KScons that are tested on three sets: PSICOV, CASP10 and CASP11. We compare against the current leading contact prediction methods. RESULTS: The results demonstrate the usefulness as well as the limits of knob-socket based structural modeling of protein contacts. The construct is able to extract good predictions from known structural homologs, while its performance degrades when no homologs exist. Among our six implementations, KScons MST-MP (which uses the multiple structure alignment prior and marginal posterior incorporating structural homolog information) performs the best in all three prediction sets. An analysis of recall and precision finds that KScons MST-MP improves accuracy not only by improving identification of true positives, but also by decreasing the number of false positives. Over the CASP10 and CASP11 sets, KScons MST-MP performs better than the leading methods using only evolutionary coupling data, but not quite as well as the supervised learning methods of MetaPSICOV and CoinDCA-NN that incorporate a large set of structural features. CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Qiwei Li 0001, David B. Dahl, Marina Vannucci, Hyun Joo, Jerry W. Tsai |
Bioinform. | 1 |
| 2012 | Short adjacent repeat identification based on Chemical Reaction OptimizationabstractThe analysis of short tandem repeats (STRs) in DNA sequences has become an attractive method for determining the genetic profile of an individual. Here we focus on a more general and practical issue named short adjacent repeats identification problem (SARIP), which is extended from STR by allowing short gaps between neighboring units. Presently, the best available solution to SARIP is BASARD, which uses Markov chain Monte Carlo algorithms to determine the posterior estimate. However, the computational complexity and the tendency to get stuck in a local mode lower the efficiency of BASARD and impede its wide application. In this paper, we prove that SARIP is NP-hard, and we also solve it with Chemical Reaction Optimization (CRO), a recently developed metaheuristic approach. CRO mimics the interactions of molecules in a chemical reaction and it can explore the solution space efficiently to find the optimal or near optimal solution(s). We test the CRO algorithm with both synthetic and real data, and compare its performance in mode searching with BASARD. Simulation results show that CRO enjoys dozens of times, or even a hundred times shorter computational time compared with BASARD. It is also demonstrated that CRO can obtain the global optima most of the time. Moreover, CRO is more stable in different runs, which is of great importance in practical use. Thus, CRO is by far the best method on SARIP. Albert Y. S. Lam, Victor O. K. Li, Qiwei Li 0001, Xiaodan Fan |
IEEE Congress on Evolutionary Computation | 4 |
| 2012 | Enabling Multilevel Trust in Privacy Preserving Data MiningabstractPrivacy Preserving Data Mining (PPDM) addresses the problem of developing accurate models about aggregated data without access to precise information in individual data record. A widely studied perturbation-based PPDM approach introduces random perturbation to individual values to preserve privacy before data are published. Previous solutions of this approach are limited in their tacit assumption of single-level trust on data miners. In this work, we relax this assumption and expand the scope of perturbation-based PPDM to Multilevel Trust (MLT-PPDM). In our setting, the more trusted a data miner is, the less perturbed copy of the data it can access. Under this setting, a malicious data miner may have access to differently perturbed copies of the same data through various means, and may combine these diverse copies to jointly infer additional information about the original data that the data owner does not intend to release. Preventing such diversity attacks is the key challenge of providing MLT-PPDM services. We address this challenge by properly correlating perturbation across copies at different trust levels. We prove that our solution is robust against diversity attacks with respect to our privacy goal. That is, for data miners who have access to an arbitrary collection of the perturbed copies, our solution prevent them from jointly reconstructing the original data more accurately than the best effort using any individual copy in the collection. Our solution allows a data owner to generate perturbed copies of its data for arbitrary trust levels on-demand. This feature offers data owners maximum flexibility. Minghua Chen 0001, Qiwei Li 0001, Wayne Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2011 | An MCMC algorithm for detecting short adjacent repeats shared by multiple sequencesabstractMOTIVATION: Repeats detection problems are traditionally formulated as string matching or signal processing problems. They cannot readily handle gaps between repeat units and are incapable of detecting repeat patterns shared by multiple sequences. This study detects short adjacent repeats with interunit insertions from multiple sequences. For biological sequences, such studies can shed light on molecular structure, biological function and evolution. RESULTS: The task of detecting short adjacent repeats is formulated as a statistical inference problem by using a probabilistic generative model. An Markov chain Monte Carlo algorithm is proposed to infer the parameters in a de novo fashion. Its applications on synthetic and real biological data show that the new method not only has a competitive edge over existing methods, but also can provide a way to study the structure and the evolution of repeat-containing genes. AVAILABILITY: The related C++ source code and datasets are available at http://ihome.cuhk.edu.hk/%7Eb118998/share/BASARD.zip. CONTACT: [email protected] Qiwei Li 0001, Xiaodan Fan, Tong Liang, Shuo-Yen Robert Li |
Bioinform. | 1 |
| 2010 | An automatic procedure to search highly repetitive sequences in genome as fluorescence in situ hybridization probes and its application on Brachypodium distachyonabstractFluorescence in situ hybridization (FISH) is a powerful technique that localizes specific DNA sequences on chromosomes for use in physical and genetic maps assembling, genetic counselling, species identification, etc. Highly repetitive sequences are considered to be suitable FISH probes that can avoid many potential problems of using unique sequences as FISH probes. The distinct chromosomal distributions of these highly repetitive sequences are also ideal for labelling purposes such as karyotyping. In this paper, we present an automatic computational procedure for searching highly repetitive sequences from a whole genome as FISH probes, as well as an experimental protocol to use them in FISH analysis. We successfully applied the method on the newly released genome of Brachypodium distachyon (Brachypodium) and produced satisfactory results of FISH experiment. Qiwei Li 0001, Tong Liang, Xiaodan Fan, Weichang Yu 0002, Shuo-Yen Robert Li |
BIBM | 1 |
| 2010 | An Evolutionary Monte Carlo algorithm for identifying short adjacent repeats in multiple sequencesabstractEvolutionary Monte Carlo (EMC) algorithm is an effective and powerful method to sample complicated distributions. Short adjacent repeats identification problem (SARIP), i.e., searching for the common sequence pattern in multiple DNA sequences, is considered as one of the key challenges in the field of bioinformatics. A recently proposed Markov chain Monte Carlo (MCMC) algorithm has demonstrated its effectiveness in solving SARIP. However, high computation time and inevitable local optima hinder its wide application. In this paper, we apply EMC to parallelize the MCMC algorithm to solve SARIP. Our proposed EMC scheme is implemented on a parallel platform and the simulation results show that, compared with the conventional MCMC algorithm, EMC not only improves the quality of final solution but also reduces the computation time. Qiwei Li 0001, Xiaodan Fan, Victor O. K. Li, Shuo-Yen Robert Li |
BIBM | 2 |