VLDB 2026 Research / reviewers in the wild / expert
Ruibin Xi
dblp:10/7259
· DBLP profile ↗
13ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0001-7545-7361ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 2 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 90% Computational science and engineering · 10% | |
| Artificial intelligence
2 papers |
Probabilistic and Bayesian machine learning · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Query processing and optimization · 46% Data mining · 46% Data stream processing · 7% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
1.4 | 2 | 2025 | A Generic Family of Graphical Models: Diversity, Efficiency, and Heterogeneity · ICML 2025 Estimating graphical models for count data with applications to single-cell gene network · NeurIPS 2022 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model |
0.9 | 1 | 2025 | A Generic Family of Graphical Models: Diversity, Efficiency, and Heterogeneity · ICML 2025 |
Bioinformatics and computational biology › transcriptomics › spatial transcriptomics
spatially variable gene detection |
0.9 | 1 | 2025 | Benchmarking algorithms for spatially variable gene identification in spatial transcriptomics · Bioinform. 2025 |
Bioinformatics and computational biology › transcriptomics
spatial transcriptomics |
0.9 | 1 | 2025 | Benchmarking algorithms for spatially variable gene identification in spatial transcriptomics · Bioinform. 2025 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › gaussian graphical model
precision matrix estimation |
0.6 | 1 | 2022 | Estimating graphical models for count data with applications to single-cell gene network · NeurIPS 2022 |
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference |
0.6 | 1 | 2022 | Estimating graphical models for count data with applications to single-cell gene network · NeurIPS 2022 |
Computational science and engineering
graphical models |
0.6 | 1 | 2022 | Estimating graphical models for count data with applications to single-cell gene network · NeurIPS 2022 |
Bioinformatics and computational biology › cancer genomics › copy number analysis
copy number variation detection |
0.4 | 1 | 2020 | CNV-BAC: Copy number Variation Detection in Bacterial Circular Genome · Bioinform. 2020 |
Bioinformatics and computational biology › single-cell analysis › single-cell data preprocessing
dropout imputation |
0.4 | 1 | 2020 | scRMD: imputation for single cell RNA-seq data via robust matrix decomposition · Bioinform. 2020 |
Bioinformatics and computational biology › single-cell analysis › single-cell RNA sequencing
single-cell RNA-seq analysis |
0.4 | 1 | 2020 | scRMD: imputation for single cell RNA-seq data via robust matrix decomposition · Bioinform. 2020 |
Bioinformatics and computational biology › genomics › genome sequencing
whole genome sequencing |
0.4 | 1 | 2020 | CNV-BAC: Copy number Variation Detection in Bacterial Circular Genome · Bioinform. 2020 |
Bioinformatics and computational biology › genomics › structural variation
structural variation detection |
0.3 | 1 | 2017 | SVmine improves structural variation detection by integrative mining of predictions from multiple algorithms · Bioinform. 2017 |
Query processing and optimization › OLAP
data cube |
0.1 | 1 | 2009 | Compression and Aggregation for Logistic Regression Analysis in Data Cubes · IEEE Trans. Knowl. Data Eng. 2009 |
Data mining › predictive modeling › regression
logistic regression |
0.1 | 1 | 2009 | Compression and Aggregation for Logistic Regression Analysis in Data Cubes · IEEE Trans. Knowl. Data Eng. 2009 |
Query processing and optimization
OLAP |
0.1 | 1 | 2009 | Compression and Aggregation for Logistic Regression Analysis in Data Cubes · IEEE Trans. Knowl. Data Eng. 2009 |
Data mining
pattern mining |
0.1 | 1 | 2009 | Compression and Aggregation for Logistic Regression Analysis in Data Cubes · IEEE Trans. Knowl. Data Eng. 2009 |
Data stream processing
streaming analytics |
0.0 | 1 | 2009 | Compression and Aggregation for Logistic Regression Analysis in Data Cubes · IEEE Trans. Knowl. Data Eng. 2009 |
Methods — techniques the papers use, named apart from their topics
poisson log-normal model · 1.1maximum marginal likelihood · 1.1d-trace loss · 1.1parameter estimation · 0.9marginal recoverability · 0.9robust matrix decomposition · 0.4low-rank approximation · 0.4bias normalization · 0.4sandwich alignment · 0.3local realignment · 0.3likelihood scoring · 0.3maximum likelihood estimation · 0.1first-order approximation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Generic Family of Graphical Models: Diversity, Efficiency, and HeterogeneityabstractTraditional network inference methods, such as Gaussian Graphical Models, which are built on continuity and homogeneity, face challenges when modeling discrete data and heterogeneous frameworks. Furthermore, under high-dimensionality, the parameter estimation of such models can be hindered by the notorious intractability of high-dimensional integrals. In this paper, we introduce a new and flexible device for graphical models, which accommodates diverse data types, including Gaussian, Poisson log-normal, and latent Gaussian copula models. The new device is driven by a new marginally recoverable parametric family, which can be effectively estimated without evaluating the high-dimensional integration in high-dimensional settings thanks to the marginal recoverability. We further introduce a mixture of marginally recoverable models to capture ubiquitous heterogeneous structures. We show the validity of the desirable properties of the models and the effective estimation methods, and demonstrate their advantages over the state-of-the-art network inference methods via extensive simulation studies and a gene regulatory network analysis of real single-cell RNA sequencing data. Yufei Huang 0020, Changhu Wang, Weichi Wu, Ruibin Xi |
ICML | 5 |
| 2025 | TAGET: a toolkit for analyzing full-length transcripts from single molecular sequencingabstractAbstract Single-molecule Real-time Isoform Sequencing (Iso-seq) of transcriptomes by PacBio can generate very long and accurate reads, thus providing an ideal platform for full-length transcriptome analysis.A number of computational tools have been developed for long-read sequencing data. However, integrated computational frameworks for analyzing Iso-seq data are still lacking. We present a Toolkit for Analyzing full-length GEne Transcripts (TAGET) for Iso-seq. Starting from polished high-quality transcripts (circular consensus sequences or CCSs), TAGET first aligns transcripts to the reference genome by integrating alignment results from long and short reads and further improves splice site predictions using a Convolutional Neural Network (CNN). TAGET then annotates transcripts by comparing with reference isoform databases and classifies transcripts into seven classes. Finally, TAGET estimates gene or isoform expressions and performs differential expression gene (DEG) and differential isoform usage (DIU) analysis. We evaluate the performance of TAGET using a public Iso-seq dataset and newly sequenced Iso-seq datasets from tumor patients. TAGET gives significantly more precise novel splice site prediction and enables more accurate novel isoform and gene fusion discoveries, as validated by experimental validations and comparisons with RNA-seq data. We identify and experimentally validate a differential isoform usage gene ECM1, and further show that its isoform ECM1b may be a tumor-suppressor in laryngocarcinoma. Our results demonstrate that TAGET provides a valuable computational toolkit and can be applied to many full-length transcriptome studies. Yuchao Xia, Susheng Miao, Ruibin Xi |
Briefings Bioinform. | 3 |
| 2025 | Benchmarking algorithms for spatially variable gene identification in spatial transcriptomicsabstractMOTIVATION: The rapid development of spatial transcriptomics has underscored the importance of identifying spatially variable genes. As a fundamental task in spatial transcriptomic data analysis, spatially variable gene identification has been extensively studied. However, the lack of comprehensive benchmark makes it difficult to validate the effectiveness of various algorithms scattered across a large number of studies with real-world datasets. RESULTS: In response, this article proposes a benchmark framework to evaluate algorithms for identifying spatially variable genes through the analysis of 30 synthesized and 74 real-world datasets, aiming to identify the best algorithms and their corresponding application scenarios. This framework can assist medical and life scientists in selecting suitable algorithms for their research, while also aid bioinformatics scientists in developing more powerful and efficient computational methods in spatial transcriptomic research. AVAILABILITY AND IMPLEMENTATION: The source code of this benchmarking framework is available at both Github (https://github.com/XiDsLab/svg-benchmark) and Zenodo (https://doi.org/10.5281/zenodo.15031083). In addition, all real and synthetic datasets considered in this study are also publicly available at Zenodo (https://doi.org/10.5281/zenodo.7227771). Xuanwei Chen, Qinghua Ran, Xingjie Shi, Ruibin Xi |
Bioinform. | 7 |
| 2022 | Estimating graphical models for count data with applications to single-cell gene networkabstractGraphical models such as Gaussian graphical models have been widely applied for direct interaction inference in many different areas. In many modern applications, such as single-cell RNA sequencing (scRNA-seq) studies, the observed data are counts and often contain many small counts. Traditional graphical models for continuous data are inappropriate for network inference of count data. We consider the Poisson log-normal (PLN) graphical model for count data and the precision matrix of the latent normal distribution represents the network. We propose a two-step method PLNet to estimate the precision matrix. PLNet first estimates the latent covariance matrix using the maximum marginal likelihood estimator (MMLE) and then estimates the precision matrix by minimizing the lasso-penalized D-trace loss function. We establish the convergence rate of the MMLE of the covariance matrix and further establish the convergence rate and the sign consistency of the proposed PLNet estimator of the precision matrix in the high dimensional setting. Importantly, although the PLN model is not sub-Gaussian, we show that the PLNet estimator is consistent even if the model dimension goes to infinity exponentially as the sample size increases. The performance of PLNet is evaluated and compared with available methods using simulation and gene regulatory network analysis of real scRNA-seq data. Feiyi Xiao, Huaying Fang, Ruibin Xi |
NeurIPS | 4 |
| 2020 | scRMD: imputation for single cell RNA-seq data via robust matrix decompositionabstractMOTIVATION: Single cell RNA-sequencing (scRNA-seq) technology enables whole transcriptome profiling at single cell resolution and holds great promises in many biological and medical applications. Nevertheless, scRNA-seq often fails to capture expressed genes, leading to the prominent dropout problem. These dropouts cause many problems in down-stream analysis, such as significant increase of noises, power loss in differential expression analysis and obscuring of gene-to-gene or cell-to-cell relationship. Imputation of these dropout values can be beneficial in scRNA-seq data analysis. RESULTS: In this article, we model the dropout imputation problem as robust matrix decomposition. This model has minimal assumptions and allows us to develop a computational efficient imputation method called scRMD. Extensive data analysis shows that scRMD can accurately recover the dropout values and help to improve downstream analysis such as differential expression analysis and clustering analysis. AVAILABILITY AND IMPLEMENTATION: The R package scRMD is available at https://github.com/XiDsLab/scRMD. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chong Chen 0002, Changjing Wu, Linjie Wu, Minghua Deng, Ruibin Xi |
Bioinform. | 6 |
| 2020 | CNV-BAC: Copy number Variation Detection in Bacterial Circular GenomeabstractMOTIVATION: Whole-genome sequencing (WGS) is widely used for copy number variation (CNV) detection. However, for most bacteria, their circular genome structure and high replication rate make reads more enriched near the replication origin. CNV detection based on read depth could be seriously influenced by such replication bias. RESULTS: We show that the replication bias is widespread using ∼200 bacterial WGS data. We develop CNV-BAC (CNV-Bacteria) that can properly normalize the replication bias and other known biases in bacterial WGS data and can accurately detect CNVs. Simulation and real data analysis show that CNV-BAC achieves the best performance in CNV detection compared with available algorithms. AVAILABILITY AND IMPLEMENTATION: CNV-BAC is available at https://github.com/XiDsLab/CNV-BAC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Linjie Wu, Yuchao Xia, Ruibin Xi |
Bioinform. | 4 |
| 2019 | DTMBIO 2019: The Thirteenth International Workshop on Data and Text Mining in Biomedical InformaticsabstractStarted in 2006 as a specialized workshop in the field of text mining applied to biomedical informatics, DTMBIO (ACM international workshop on Data and Text Mining in Biomedical Informatics) has been held annually in conjunction with one of the largest data management conferences, CIKM, bringing together researchers working on computer science and bioinformatics area including text mining and genomic data analysis. The purpose of DTMBIO is to foster discussions regarding the state-of-the-art applications of data and text mining on biomedical research problems. DTMBIO 2019 will help scientists navigate emerging trends and opportunities in the evolving area of informatics related techniques and problems in the context of biomedical research. Hyojung Paik, Ruibin Xi, Doheon Lee |
CIKM | 2 |
| 2017 | SVmine improves structural variation detection by integrative mining of predictions from multiple algorithmsabstractMOTIVATION: Structural variation (SV) is an important class of genomic variations in human genomes. A number of SV detection algorithms based on high-throughput sequencing data have been developed, but they have various and often limited level of sensitivity, specificity and breakpoint resolution. Furthermore, since overlaps between predictions of algorithms are low, SV detection based on multiple algorithms, an often-used strategy in real applications, has little effect in improving the performance of SV detection. RESULTS: We develop a computational tool called SVmine for further mining of SV predictions from multiple tools to improve the performance of SV detection. SVmine refines SV predictions by performing local realignment and assess quality of SV predictions based on likelihoods of the realignments. The local realignment is performed against a set of sequences constructed from the reference sequence near the candidate SV by incorporating nearby single nucleotide variations, insertions and deletions. A sandwich alignment algorithm is further used to improve the accuracy of breakpoint positions. We evaluate SVmine on a set of simulated data and real data and find that SVmine has superior sensitivity, specificity and breakpoint estimation accuracy. We also find that SVmine can significantly improve overlaps of SV predictions from other algorithms. AVAILABILITY AND IMPLEMENTATION: SVmine is available at https://github.com/xyc0813/SVmine. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yuchao Xia, Minghua Deng, Ruibin Xi |
Bioinform. | 4 |
| 2017 | Pysim-sv: a package for simulating structural variation data with GC-biasesabstractBACKGROUND: Structural variations (SVs) are wide-spread in human genomes and may have important implications in disease-related and evolutionary studies. High-throughput sequencing (HTS) has become a major platform for SV detection and simulation serves as a powerful and cost-effective approach for benchmarking SV detection algorithms. Accurate performance assessment by simulation requires the simulator capable of generating simulation data with all important features of real data, such GC biases in HTS data and various complexities in tumor data. However, no available package has systematically addressed all issues in data simulation for SV benchmarking. RESULTS: Pysim-sv is a package for simulating HTS data to evaluate performance of SV detection algorithms. Pysim-sv can introduce a wide spectrum of germline and somatic genomic variations. The package contains functionalities to simulate tumor data with aneuploidy and heterogeneous subclones, which is very useful in assessing algorithm performance in tumor studies. Furthermore, Pysim-sv can introduce GC-bias, the most important and prevalent bias in HTS data, in the simulated HTS data. CONCLUSIONS: Pysim-sv provides an unbiased toolkit for evaluating HTS-based SV detection algorithms. Yuchao Xia, Minghua Deng, Ruibin Xi |
BMC Bioinform. | 4 |
| 2016 | Evaluation of somatic copy number estimation tools for whole-exome sequencing dataabstractWhole-exome sequencing (WES) has become a standard method for detecting genetic variants in human diseases. Although the primary use of WES data has been the identification of single nucleotide variations and indels, these data also offer a possibility of detecting copy number variations (CNVs) at high resolution. However, WES data have uneven read coverage along the genome owing to the target capture step, and the development of a robust WES-based CNV tool is challenging. Here, we evaluate six WES somatic CNV detection tools: ADTEx, CONTRA, Control-FREEC, EXCAVATOR, ExomeCNV and Varscan2. Using WES data from 50 kidney chromophobe, 50 bladder urothelial carcinoma, and 50 stomach adenocarcinoma patients from The Cancer Genome Atlas, we compared the CNV calls from the six tools with a reference CNV set that was identified by both single nucleotide polymorphism array 6.0 and whole-genome sequencing data. We found that these algorithms gave highly variable results: visual inspection reveals significant differences between the WES-based segmentation profiles and the reference profile, as well as among the WES-based profiles. Using a 50% overlap criterion, 13-77% of WES CNV calls were covered by CNVs from the reference set, up to 21% of the copy gains were called as losses or vice versa, and dramatic differences in CNV sizes and CNV numbers were observed. Overall, ADTEx and EXCAVATOR had the best performance with relatively high precision and sensitivity. We suggest that the current algorithms for somatic CNV detection from WES data are limited in their performance and that more robust algorithms are needed. Jae-Yong Nam, Nayoung K. D. Kim, Sang Cheol Kim, Je-Gun Joung, Ruibin Xi, Semin Lee, Peter J. Park, Woong-Yang Park |
Briefings Bioinform. | 5 |
| 2011 | Compression and aggregation of Bayesian estimates for data intensive computing
Ruibin Xi, Yixin Chen 0001 |
Knowl. Inf. Syst. | 1 |
| 2010 | rSW-seq: Algorithm for detection of copy number alterations in deep sequencing dataabstractBACKGROUND: Recent advances in sequencing technologies have enabled generation of large-scale genome sequencing data. These data can be used to characterize a variety of genomic features, including the DNA copy number profile of a cancer genome. A robust and reliable method for screening chromosomal alterations would allow a detailed characterization of the cancer genome with unprecedented accuracy. RESULTS: We develop a method for identification of copy number alterations in a tumor genome compared to its matched control, based on application of Smith-Waterman algorithm to single-end sequencing data. In a performance test with simulated data, our algorithm shows >90% sensitivity and >90% precision in detecting a single copy number change that contains approximately 500 reads for the normal sample. With 100-bp reads, this corresponds to a ~50 kb region for 1X genome coverage of the human genome. We further refine the algorithm to develop rSW-seq, (recursive Smith-Waterman-seq) to identify alterations in a complex configuration, which are commonly observed in the human cancer genome. To validate our approach, we compare our algorithm with an existing algorithm using simulated and publicly available datasets. We also compare the sequencing-based profiles to microarray-based results. CONCLUSION: We propose rSW-seq as an efficient method for detecting copy number changes in the tumor genome. Tae-Min Kim, Lovelace J. Luquette, Ruibin Xi, Peter J. Park |
BMC Bioinform. | 3 |
| 2009 | Compression and Aggregation for Logistic Regression Analysis in Data CubesabstractLogistic regression is an important technique for analyzing and predicting data with categorical attributes. In this paper, We consider supporting online analytical processing (OLAP) of logistic regression analysis for multi-dimensional data in a data cube where it is expensive in time and space to build logistic regression models for each cell from the raw data. We propose a novel scheme to compress the data in such a way that we can reconstruct logistic regression models to answer any OLAP query without accessing the raw data. Based on a first-order approximation to the maximum likelihood estimating equations, we develop a compression scheme that compresses each base cell into a small compressed data block with essential information to support the aggregation of logistic regression models. Aggregation formulae for deriving high-level logistic regression models from lower level component cells are given. We prove that the compression is nearly lossless in the sense that the aggregated estimator deviates from the true model by an error that is bounded and approaches to zero when the data size increases. The results show that the proposed compression and aggregation scheme can make feasible OLAP of logistic regression in a data cube. Further, it supports real-time logistic regression analysis of stream data, which can only be scanned once and cannot be permanently retained. Experimental results validate our theoretical analysis and demonstrate that our method can dramatically save time and space costs with almost no degradation of the modeling accuracy. Ruibin Xi, Yixin Chen 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |