VLDB 2026 Research / reviewers in the wild / expert
Chi Zhang 0021
dblp:91/195-21
· DBLP profile ↗
23ranked-venue papers
0as first author
10since 2021 · last 2024
0000-0001-9553-0925ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 5 since 2021Artificial intelligence and machine learning · 8 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Bias-aware Boolean Matrix Factorization Using Disentangled Representation LearningabstractBoolean matrix factorization (BMF) has been widely utilized in fields such as recommendation systems, graph learning, text mining, and -omics data analysis. Traditional BMF methods decompose a binary matrix into the Boolean product of two lower-rank Boolean matrices plus homoscedastic random errors. However, real-world binary data typically involves biases arising from heterogeneous row- and column-wise signal distributions. Such biases can lead to suboptimal fitting and unexplainable predictions if not accounted for. In this study, we reconceptualize the binary data generation as the Boolean sum of three components: a binary pattern matrix, a background bias matrix influenced by heterogeneous row or column distributions, and random flipping errors. We introduce a novel Disentangled Representation Learning for Binary matrices (DRLB) method, which employs a dual auto-encoder network to reveal the true patterns. DRLB can be seamlessly integrated with existing BMF techniques to facilitate bias-aware BMF. Our experiments with both synthetic and real-world datasets show that DRLB significantly enhances the precision of traditional BMF methods while offering high scalability. Moreover, the bias matrix detected by DRLB accurately reflects the inherent biases in synthetic data, and the patterns identified in the bias-corrected real-world data exhibit enhanced interpretability. Xiao Wang 0099, Tong Zhao 0002, Yong Zang, Sha Cao, Chi Zhang 0021 |
UAI | 8 |
| 2023 | Generalized Matrix Local Low Rank Representation by Random Projection and Submatrix PropagationabstractMatrix low rank approximation is an effective method to reduce or eliminate the statistical redundancy of its components. Compared with the traditional global low rank methods such as singular value decomposition (SVD), local low rank approximation methods are more advantageous to uncover interpretable data structures when clear duality exists between the rows and columns of the matrix. Local low rank approximation is equivalent to low rank submatrix detection. Unfortunately, existing local low rank approximation methods can detect only submatrices of specific mean structure, which may miss a substantial amount of true and interesting patterns. In this work, we develop a novel matrix computational framework called RPSP (Random Probing based submatrix Propagation) that provides an effective solution for the general matrix local low rank representation problem. RPSP detects local low rank patterns that grow from small submatrices of low rank property, which are determined by a random projection approach. RPSP is supported by theories of random projection. Experiments on synthetic data demonstrate that RPSP outperforms all state-of-the-art methods, with the capacity to robustly and correctly identify the low rank matrices when the pattern has a similar mean as the background, background noise is heteroscedastic and multiple patterns present in the data. On real-world datasets, RPSP also demonstrates its effectiveness in identifying interpretable local low rank matrices. Pengdao Dang, Haiqi Zhu, Tingbo Guo, Changlin Wan, Tong Zhao 0002, Paul Salama, Sha Cao, Chi Zhang 0021 |
KDD | 9 |
| 2022 | Provable Second-Order Riemannian Gauss-Newton Method for Low-Rank Tensor Estimation ‖abstractIn this paper, we consider the estimation of a low Tucker rank tensor from a number of noisy linear measurements. We propose a Riemannian Gauss-Newton (RGN) method with fast implementations for low Tucker rank tensor estimation. Different from the generic (super)linear convergence guarantee of RGN in the literature, we prove the first quadratic convergence guarantee of RGN for low-rank tensor estimation under some mild conditions. A deterministic estimation error lower bound, which matches the upper bound, is provided that demonstrates the statistical optimality of RGN. The merit of RGN is illustrated through applications of tensor regression and tensor SVD. Yuetian Luo, Qin Ma 0003, Chi Zhang 0021, Anru Zhang |
ICASSP | 3 |
| 2022 | Bias aware probabilistic Boolean matrix factorizationabstractBoolean matrix factorization (BMF) is a combinatorial problem arising from a wide range of applications including recommendation system, collaborative filtering, and dimensionality reduction. Currently, the noise model of existing BMF methods is often assumed to be homoscedastic; however, in real world data scenarios, the deviations of observed data from their true values are almost surely diverse due to stochastic noises, making each data point not equally suitable for fitting a model. In this case, it is not ideal to treat all data points as equally distributed. Motivated by such observations, we introduce a probabilistic BMF model that recognizes the object- and feature-wise bias distribution respectively, called bias aware BMF (BABF). To the best of our knowledge, BABF is the first approach for Boolean decomposition with consideration of the feature-wise and object-wise bias in binary data. We conducted experiments on datasets with different levels of background noise, bias level, and sizes of the signal patterns, to test the effectiveness of our method in various scenarios. We demonstrated that our model outperforms the state-of-the-art factorization methods in both accuracy and efficiency in recovering the original datasets, and the inferred bias level is highly significantly correlated with true existing bias in both simulated and real world datasets. Changlin Wan, Pengdao Dang, Tong Zhao 0002, Yong Zang, Chi Zhang 0021, Sha Cao |
UAI | 5 |
| 2022 | Response to 'Letter to the Editor: on the stability and internal consistency of component-wise sparse mixture regression based clustering', Zhang et alabstractWe have recently published a clustering method, namely CSMR, based on mixture regression in the context of high-dimensional predictors [1]. The motivation of our work was to provide a scalable method to deal with high-dimensional molecular features when there exist heterogeneous relationships between the molecular features and a disease phenotype of interest. In our original work, we simulated different data environments, including sample size |$N$|, number of cluster-differentiating and cluster-specific predictors |${M}_0$|, number of clusters |$K$|, and noise level |$\sigma$|. We evaluated CSMR on the simulation datasets, based on its accuracy of clustering using Rand index, and feature selection using true positive/negative rate, as well as the consistency of predicted and observed response values using Pearson correlation. In a real-world Cancer Cell Line Encyclopedia (CCLE) dataset [2], we evaluated CSMR based on the consistency of predicted and observed response values using cross-validation. The letter by Zhang et al. raised the concerns that the performance of CSMR drops significantly when the number of cluster-differentiating predictors |${M}_0$| increases from 5 to 20, as demonstrated by evaluation metrics including the adjusted Rand index (ARI) and the clustering internal consistency (IC), both of which are advocated by Zhang et al. to be used for evaluating clustering methods. We appreciate Zhang et al.’s interests in our method and agree that their observations of the performance drop in the large |${M}_0$| case are indeed true. We would like to mention that mixture regression models, with the presence of both low- and high-dimensional features, rely on Expectation–Maximization (EM) algorithm or its variants for a solution. However, EM algorithm often converges to the maximum likelihood estimate of the mixture parameters locally [3], and using EM algorithm to solve mixture regression problem might yield poor clustering performance if the parameters are not initialized properly, and finding a good initial value is always a challenge. In other words, EM algorithm is highly sensitive to initialization, and different initial values may lead to different solutions and hence the instability. Many works have been published on establishing the convergence and stability criterion for EM algorithm [4–7]. For example, a guarantee for the EM algorithm to converge to the unique global optimum is when the likelihood is unimodal with certain regularity conditions [4]. However, the likelihood function could be multi-modal in many cases, for which only local optimum could be guaranteed; and what’s worse, sometimes the local optimum is a poor one that is far away from any global optimum of the likelihood [7]. In all, while EM algorithm has enjoyed popular use in solving mixture model, there is still space for improved theoretical understanding. Wennan Chang, Chi Zhang 0021, Sha Cao |
Briefings Bioinform. | 2 |
| 2022 | PLUS: Predicting cancer metastasis potential based on positive and unlabeled learningabstractMetastatic cancer accounts for over 90% of all cancer deaths, and evaluations of metastasis potential are vital for minimizing the metastasis-associated mortality and achieving optimal clinical decision-making. Computational assessment of metastasis potential based on large-scale transcriptomic cancer data is challenging because metastasis events are not always clinically detectable. The under-diagnosis of metastasis events results in biased classification labels, and classification tools using biased labels may lead to inaccurate estimations of metastasis potential. This issue is further complicated by the unknown metastasis prevalence at the population level, the small number of confirmed metastasis cases, and the high dimensionality of the candidate molecular features. Our proposed algorithm, called Positive and unlabeled Learning from Unbalanced cases and Sparse structures (PLUS), is the first to use a positive and unlabeled learning framework to account for the under-detection of metastasis events in building a classifier. PLUS is specifically tailored for studying metastasis that deals with the unbalanced instance allocation as well as unknown metastasis prevalence, which are not considered by other methods. PLUS achieves superior performance on synthetic datasets compared with other state-of-the-art methods. Application of PLUS to The Cancer Genome Atlas Pan-Cancer gene expression data generated metastasis potential predictions that show good agreement with the clinical follow-up data, in addition to predictive genes that have been validated by independent single-cell RNA-sequencing datasets. Junyi Zhou 0001, Wennan Chang, Changlin Wan, Xiongbin Lu, Chi Zhang 0021, Sha Cao |
PLoS Comput. Biol. | 6 |
| 2021 | Spatially and Robustly Hybrid Mixture Regression Model for Inference of Spatial DependenceabstractIn this paper, we propose a Spatial Robust Mixture Regression model to investigate the relationship between a response variable and a set of explanatory variables over the spatial domain, assuming that the relationships may exhibit complex spatially dynamic patterns that cannot be captured by constant regression coefficients. Our method integrates the robust finite mixture Gaussian regression model with spatial constraints, to simultaneously handle the spatial non-stationarity, local homogeneity, and outlier contaminations. Compared with existing spatial regression models, our proposed model assumes the existence a few distinct regression models that are estimated based on observations that exhibit similar response-predictor relationships. As such, the proposed model not only accounts for non-stationarity in the spatial trend, but also clusters observations into a few distinct and homogenous groups. This provides an advantage on interpretation with a few stationary sub-processes identified that capture the predominant relationships between response and predictor variables. Moreover, the proposed method incorporates robust procedures to handle contaminations from both regression outliers and spatial outliers. By doing so, we robustly segment the spatial domain into distinct local regions with similar regression coefficients, and sporadic locations that are purely outliers. Rigorous statistical hypothesis testing procedure has been designed to test the significance of such segmentation. Experimental results on many synthetic and real-world datasets demonstrate the robustness, accuracy, and effectiveness of our proposed method, compared with other robust finite mixture regression, spatial regression and spatial segmentation methods. Wennan Chang, Pengdao Dang, Changlin Wan, Tong Zhao 0002, Yong Zang, Chi Zhang 0021, Sha Cao |
ICDM | 9 |
| 2021 | Supervised clustering of high-dimensional data using regularized mixture modelingabstractIdentifying relationships between genetic variations and their clinical presentations has been challenged by the heterogeneous causes of a disease. It is imperative to unveil the relationship between the high-dimensional genetic manifestations and the clinical presentations, while taking into account the possible heterogeneity of the study subjects.We proposed a novel supervised clustering algorithm using penalized mixture regression model, called component-wise sparse mixture regression (CSMR), to deal with the challenges in studying the heterogeneous relationships between high-dimensional genetic features and a phenotype. The algorithm was adapted from the classification expectation maximization algorithm, which offers a novel supervised solution to the clustering problem, with substantial improvement on both the computational efficiency and biological interpretability. Experimental evaluation on simulated benchmark datasets demonstrated that the CSMR can accurately identify the subspaces on which subset of features are explanatory to the response variables, and it outperformed the baseline methods. Application of CSMR on a drug sensitivity dataset again demonstrated the superior performance of CSMR over the others, where CSMR is powerful in recapitulating the distinct subgroups hidden in the pool of cell lines with regards to their coping mechanisms to different drugs. CSMR represents a big data analysis tool with the potential to resolve the complexity of translating the clinical representations of the disease to the real causes underpinning it. We believe that it will bring new understanding to the molecular basis of a disease and could be of special relevance in the growing field of personalized medicine. Wennan Chang, Changlin Wan, Yong Zang, Chi Zhang 0021, Sha Cao |
Briefings Bioinform. | 4 |
| 2021 | SSMD: a semi-supervised approach for a robust cell type identification and deconvolution of mouse transcriptomics dataabstractDeconvolution of mouse transcriptomic data is challenged by the fact that mouse models carry various genetic and physiological perturbations, making it questionable to assume fixed cell types and cell type marker genes for different data set scenarios. We developed a Semi-Supervised Mouse data Deconvolution (SSMD) method to study the mouse tissue microenvironment. SSMD is featured by (i) a novel nonparametric method to discover data set-specific cell type signature genes; (ii) a community detection approach for fixing cell types and their marker genes; (iii) a constrained matrix decomposition method to solve cell type relative proportions that is robust to diverse experimental platforms. In summary, SSMD addressed several key challenges in the deconvolution of mouse tissue data, including: (i) varied cell types and marker genes caused by highly divergent genotypic and phenotypic conditions of mouse experiment; (ii) diverse experimental platforms of mouse transcriptomics data; (iii) small sample size and limited training data source and (iv) capable to estimate the proportion of 35 cell types in blood, inflammatory, central nervous or hematopoietic systems. In silico and experimental validation of SSMD demonstrated its high sensitivity and accuracy in identifying (sub) cell types and predicting cell proportions comparing with state-of-the-arts methods. A user-friendly R package and a web server of SSMD are released via https://github.com/xiaoyulu95/SSMD. Szu-Wei Tu, Wennan Chang, Changlin Wan, Jiashi Wang, Yong Zang, Baskar Ramdas, Reuben Kapur, Xiongbin Lu, Sha Cao, Chi Zhang 0021 |
Briefings Bioinform. | 11 |
| 2021 | IRIS-FGM: an integrative single-cell RNA-Seq interpretation system for functional gene module analysisabstractSUMMARY: Single-cell RNA-Seq (scRNA-Seq) data is useful in discovering cell heterogeneity and signature genes in specific cell populations in cancer and other complex diseases. Specifically, the investigation of condition-specific functional gene modules (FGM) can help to understand interactive gene networks and complex biological processes in different cell clusters. QUBIC2 is recognized as one of the most efficient and effective biclustering tools for condition-specific FGM identification from scRNA-Seq data. However, its limited availability to a C implementation restricted its application to only a few downstream analysis functionalities. We developed an R package named IRIS-FGM (Integrative scRNA-Seq Interpretation System for Functional Gene Module analysis) to support the investigation of FGMs and cell clustering using scRNA-Seq data. Empowered by QUBIC2, IRIS-FGM can effectively identify condition-specific FGMs, predict cell types/clusters, uncover differentially expressed genes and perform pathway enrichment analysis. It is noteworthy that IRIS-FGM can also take Seurat objects as input, facilitating easy integration with the existing analysis pipeline. AVAILABILITY AND IMPLEMENTATION: IRIS-FGM is implemented in the R environment (as of version 3.6) with the source code freely available at https://github.com/BMEngineeR/IRISFGM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yuzhou Chang, Carter Allen, Changlin Wan, Dongjun Chung, Chi Zhang 0021, Zihai Li, Qin Ma 0003 |
Bioinform. | 5 |
| 2020 | Fast and Efficient Boolean Matrix Factorization by Geometric SegmentationabstractBoolean matrix has been used to represent digital information in many fields, including bank transaction, crime records, natural language processing, protein-protein interaction, etc. Boolean matrix factorization (BMF) aims to find an approximation of a binary matrix as the Boolean product of two low rank Boolean matrices, which could generate vast amount of information for the patterns of relationships between the features and samples. Inspired by binary matrix permutation theories and geometric segmentation, we developed a fast and efficient BMF approach, called MEBF (Median Expansion for Boolean Factorization). Overall, MEBF adopted a heuristic approach to locate binary patterns presented as submatrices that are dense in 1's. At each iteration, MEBF permutates the rows and columns such that the permutated matrix is approximately Upper Triangular-Like (UTL) with so-called Simultaneous Consecutive-ones Property (SC1P). The largest submatrix dense in 1 would lie on the upper triangular area of the permutated matrix, and its location was determined based on a geometric segmentation of a triangular. We compared MEBF with other state of the art approaches on data scenarios with different density and noise levels. MEBF demonstrated superior performances in lower reconstruction error, and higher computational efficiency, as well as more accurate density patterns than popular methods such as ASSO, PANDA and Message Passing. We demonstrated the application of MEBF on both binary and non-binary data sets, and revealed its further potential in knowledge retrieving and data denoising. Changlin Wan, Wennan Chang, Tong Zhao 0002, Sha Cao, Chi Zhang 0021 |
AAAI | 6 |
| 2020 | A data denoising approach to optimize functional clustering of single cell RNA-sequencing dataabstractSingle cell RNA-sequencing (scRNA-seq) technology enables comprehensive transcriptomic profiling of thousands of cells with distinct phenotypic and physiological states in a complex tissue. Substantial efforts have been made to characterize single cells of distinct identities from scRNA-seq data, including various cell clustering techniques. While existing approaches can handle single cells in terms of different cell (sub)types at a high resolution, identification of the functional variability within the same cell type remains unsolved. In addition, there is a lack of robust method to handle the inter-subject variation that often brings severe confounding effects for the functional clustering of single cells. In this study, we developed a novel data denoising and cell clustering approach, namely CIBS, to provide biologically explainable functional classification for scRNA-seq data. CIBS is based on a systems biology model of transcriptional regulation that assumes a multi-modality distribution of the cells' activation status, and it utilizes a Boolean matrix factorization approach on the discretized expression status to robustly derive functional modules. CIBS is empowered by a novel fast Boolean Matrix Factorization method, namely PFAST, to increase the computational feasibility on large scale scRNA-seq data. Application of CIBS on two scRNA-seq datasets collected from cancer tumor micro-environment successfully identified subgroups of cancer cells with distinct expression patterns of epithelial-mesenchymal transition and extracellular matrix marker genes, which was not revealed by the existing cell clustering analysis tools. The identified cell groups were significantly associated with the clinically confirmed lymph-node invasion and metastasis events across different patients. Changlin Wan, Dongya Jia, Yue Zhao 0016, Wennan Chang, Sha Cao, Xiao Wang 0045, Chi Zhang 0021 |
BIBM | 7 |
| 2020 | Denoising Individual Bias for Fairer Binary Submatrix DetectionabstractLow rank representation of binary matrix is powerful in disentangling sparse individual-attribute associations, and has received wide applications. Existing binary matrix factorization (BMF) or co-clustering (CC) methods often assume i.i.d background noise. However, this assumption could be easily violated in real data, where heterogeneous row- or column-wise probability of binary entries results in disparate element-wise background distribution, and paralyzes the rationality of existing methods. We propose a binary data denoising framework, namely BIND, which optimizes the detection of true patterns by estimating the row- or column-wise mixture distribution of patterns and disparate background, and eliminating the binary attributes that are more likely from the background. BIND is supported by thoroughly derived mathematical property of the row- and column-wise mixture distributions. Our experiment on synthetic and real-world data demonstrated BIND effectively removes background noise and drastically increases the fairness and accuracy of state-of-the arts BMF and CC methods. Changlin Wan, Wennan Chang, Tong Zhao 0002, Sha Cao, Chi Zhang 0021 |
CIKM | 5 |
| 2020 | Geometric All-way Boolean Tensor DecompositionabstractBoolean tensor has been broadly utilized in representing high dimensional logical data collected on spatial, temporal and/or other relational domains. Boolean Tensor Decomposition (BTD) factorizes a binary tensor into the Boolean sum of multiple rank-1 tensors, which is an NP-hard problem. Existing BTD methods have been limited by their high computational cost, in applications to large scale or higher order tensors. In this work, we presented a computationally efficient BTD algorithm, namely Geometric Expansion for all-order Tensor Factorization (GETF), that sequentially identifies the rank-1 basis components for a tensor from a geometric perspective. We conducted rigorous theoretical analysis on the validity as well as algorithemic efficiency of GETF in decomposing all-order tensor. Experiments on both synthetic and real-world data demonstrated that GETF has significantly improved performance in reconstruction accuracy, extraction of latent structures and it is an order of magnitude faster than other state-of-the-art methods. Changlin Wan, Wennan Chang, Tong Zhao 0002, Sha Cao, Chi Zhang 0021 |
NeurIPS | 5 |
| 2020 | QUBIC2: a novel and robust biclustering algorithm for analyses and interpretation of large-scale RNA-Seq dataabstractMOTIVATION: The biclustering of large-scale gene expression data holds promising potential for detecting condition-specific functional gene modules (i.e. biclusters). However, existing methods do not adequately address a comprehensive detection of all significant bicluster structures and have limited power when applied to expression data generated by RNA-Sequencing (RNA-Seq), especially single-cell RNA-Seq (scRNA-Seq) data, where massive zero and low expression values are observed. RESULTS: We present a new biclustering algorithm, QUalitative BIClustering algorithm Version 2 (QUBIC2), which is empowered by: (i) a novel left-truncated mixture of Gaussian model for an accurate assessment of multimodality in zero-enriched expression data, (ii) a fast and efficient dropouts-saving expansion strategy for functional gene modules optimization using information divergency and (iii) a rigorous statistical test for the significance of all the identified biclusters in any organism, including those without substantial functional annotations. QUBIC2 demonstrated considerably improved performance in detecting biclusters compared to other five widely used algorithms on various benchmark datasets from E.coli, Human and simulated data. QUBIC2 also showcased robust and superior performance on gene expression data generated by microarray, bulk RNA-Seq and scRNA-Seq. AVAILABILITY AND IMPLEMENTATION: The source code of QUBIC2 is freely available at https://github.com/OSU-BMBL/QUBIC2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Juan Xie, Anjun Ma, Bingqiang Liu, Sha Cao, Cankun Wang, Chi Zhang 0021, Qin Ma 0003 |
Bioinform. | 8 |
| 2020 | IsoTree: A New Framework for de novo Transcriptome Assembly from RNA-seq ReadsabstractHigh-throughput sequencing of mRNA has made the deep and efficient probing of transcriptome more affordable. However, the vast amounts of short RNA-seq reads make de novo transcriptome assembly an algorithmic challenge. In this work, we present IsoTree, a novel framework for transcripts reconstruction in the absence of reference genomes. Unlike most of de novo assembly methods that build de Bruijn graph or splicing graph by connecting k- mers which are sets of overlapping substrings generated from reads, IsoTree constructs splicing graph by connecting reads directly. For each splicing graph, IsoTree applies an iterative scheme of mixed integer linear program to build a prefix tree, called isoform tree. Each path from the root node of the isoform tree to a leaf node represents a plausible transcript candidate which will be pruned based on the information of paired-end reads. Experiments showed that in most cases IsoTree performs better than other leading transcriptome assembly programs. IsoTree is available at https://github.com/Jane110111107/IsoTree. Jin Zhao 0005, Haodi Feng, Daming Zhu, Chi Zhang 0021, Ying Xu 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2020 | Corrections to "IsoTree: A New Framework for de novo Transcriptome Assembly from RNA-seq Reads"abstractPresents corrections to author information in the above named paper. Jin Zhao 0005, Haodi Feng, Daming Zhu, Chi Zhang 0021, Ying Xu 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2019 | M3S: a comprehensive model selection for multi-modal single-cell RNA sequencing dataabstractBACKGROUND: Various statistical models have been developed to model the single cell RNA-seq expression profiles, capture its multimodality, and conduct differential gene expression test. However, for expression data generated by different experimental design and platforms, there is currently lack of capability to determine the most proper statistical model. RESULTS: We developed an R package, namely Multi-Modal Model Selection (M3S), for gene-wise selection of the most proper multi-modality statistical model and downstream analysis, useful in a single-cell or large scale bulk tissue transcriptomic data. M3S is featured with (1) gene-wise selection of the most parsimonious model among 11 most commonly utilized ones, that can best fit the expression distribution of the gene, (2) parameter estimation of a selected model, and (3) differential gene expression test based on the selected model. CONCLUSION: A comprehensive evaluation suggested that M3S can accurately capture the multimodality on simulated and real single cell data. An open source package and is available through GitHub at https://github.com/zy26/M3S. Changlin Wan, Wennan Chang, Qin Ma 0003, Sha Cao, Chi Zhang 0021 |
BMC Bioinform. | 9 |
| 2019 | DTA-SiST: de novo transcriptome assembly by using simplified suffix treesabstractBACKGROUND: Alternative splicing allows the pre-mRNAs of a gene to be spliced into various mRNAs, which greatly increases the diversity of proteins. High-throughput sequencing of mRNAs has revolutionized our ability for transcripts reconstruction. However, the massive size of short reads makes de novo transcripts assembly an algorithmic challenge. RESULTS: We develop a novel radical framework, called DTA-SiST, for de novo transcriptome assembly based on suffix trees. DTA-SiST first extends contigs by reads that have the longest overlaps with the contigs' terminuses. These reads can be found in linear time of the lengths of the reads through a well-designed suffix tree structure. Then, DTA-SiST constructs splicing graphs based on contigs for each gene locus. Finally, DTA-SiST proposes two strategies to extract transcript-representing paths: a depth-first enumeration strategy and a hybrid strategy based on length and coverage. We implemented the above two strategies and compared them with the state-of-the-art de novo assemblers on both simulated and real datasets. Experimental results showed that the depth-first enumeration strategy performs always better with recall and also better with precision for smaller datasets while the hybrid strategy leads with precision for big datasets. CONCLUSIONS: DTA-SiST performs more competitive than the other compared de novo assemblers especially with precision measure, due to the read-based contig extension strategy and the elegant transcripts extraction rules. Jin Zhao 0005, Haodi Feng, Daming Zhu, Chi Zhang 0021, Ying Xu 0001 |
BMC Bioinform. | 4 |
| 2018 | DTAST: A Novel Radical Framework for de Novo Transcriptome Assembly Based on Suffix Trees
Jin Zhao 0005, Haodi Feng, Daming Zhu, Chi Zhang 0021, Ying Xu 0001 |
ICIC (1) | 4 |
| 2017 | IsoTree: De Novo Transcriptome Assembly from RNA-Seq Reads - (Extended Abstract)
Jin Zhao 0005, Haodi Feng, Daming Zhu, Chi Zhang 0021, Ying Xu 0001 |
ISBRA | 4 |
| 2017 | QUBIC: a bioconductor package for qualitative biclustering analysis of gene co-expression dataabstractMotivation: Biclustering is widely used to identify co-expressed genes under subsets of all the conditions in a large-scale transcriptomic dataset. The program, QUBIC, is recognized as one of the most efficient and effective biclustering methods for biological data interpretation. However, its availability is limited to a C implementation and to a low-throughput web interface. Results: An R implementation of QUBIC is presented here with two unique features: (i) a 82% average improved efficiency by refactoring and optimizing the source C code of QUBIC; and (ii) a set of comprehensive functions to facilitate biclustering-based biological studies, including the qualitative representation (discretization) of expression data, query-based biclustering, bicluster expanding, biclusters comparison, heatmap visualization of any identified biclusters and co-expression networks elucidation. Availability and Implementation: The package is implemented in R (as of version 3.3) and is available from Bioconductor at the URL: http://bioconductor.org/packages/QUBIC, where installation and usage instructions can be found. Contact: [email protected] Supplimentary Information: Supplementary data are available at Bioinformatics online. Juan Xie, Anne Fennell, Chi Zhang 0021, Qin Ma 0003 |
Bioinform. | 5 |
| 2016 | PUEPro: A Computational Pipeline for Prediction of Urine Excretory Proteins
Yan Wang 0028, Wei Du 0002, Yanchun Liang 0001, Xin Chen 0113, Chi Zhang 0021, Wei Pang 0001, Ying Xu 0001 |
ADMA | 5 |