EDBT 2026 Demo / reviewers in the wild / expert
Taosheng Xu
dblp:209/7393
· DBLP profile ↗
17ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0001-8283-7274ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fourier-enhanced semi-supervised proxy learning for ultra-fine-grained novel class discovery
Qiupu Chen, Hongkui Jiang, Lin Jiao, Taosheng Xu, Rujing Wang |
Pattern Recognit. | 5 |
| 2025 | DBiSeNet: Dual bilateral segmentation network for real-time semantic segmentationabstractBilateral networks have shown effectiveness and efficiency for real-time semantic segmentation. However, the single bilateral architecture exhibits limitations in capturing multi-scale feature representations and addressing misalignment issues during spatial and contextual feature fusion, thereby constraining segmentation accuracy. To address these challenges, we propose a novel dual bilateral segmentation network (DBiSeNet) that incorporates an additional bilateral branch into the original architecture. The additional (high-scale) bilateral operating at high resolution to preserve fine-grained details and responsible for thin object prediction, while the original (low-scale) bilateral maintains an enlarged receptive field to capture global context for large object segmentation. Furthermore, we introduce an aligned and refined feature fusion module to mitigate feature misalignment within each bilateral branch. To optimize the final prediction, we design a dual prediction fusion module that utilizes the low-scale segmentation results as a baseline and adaptively incorporates complementary information from high-scale predictions. Extensive experiments on the Cityscapes and CamVid datasets validate the effectiveness of DBiSeNet in achieving an optimal balance between accuracy and inference speed. In particular, on a single RTX3090 GPU, DBiSeNet2 yields 75.6% mIoU at 225.9 FPS on Cityscapes test set and 75.7% mIoU at 203.4 FPS on CamVid test set. Taosheng Xu |
Comput. Vis. Image Underst. | 4 |
| 2025 | Weighted Sparse Partial Least Squares With Joint Sample and Feature Selection for Integrating Multi-Omics DataabstractSparse Partial Least Squares (sPLS) is a common dimensionality reduction technique for data fusion, which projects data samples from two views by seeking linear combinations with a small number of variables with the maximum variance. However, sPLS extracts the combinations between two data sets with all data samples so that it cannot detect latent subsets of samples. To extend the application of sPLS by identifying a specific subset of samples and remove outliers, we propose an $\ell _\infty /\ell _{0}$-norm constrained weighted sparse PLS ($\ell _\infty /\ell _{0}$-wsPLS) method for joint sample and feature selection, where the $\ell _\infty /\ell _{0}$-norm constrains are used to select a subset of samples. We prove that the $\ell _\infty /\ell _{0}$-norm constrains have the Kurdyka-Łojasiewicz property so that a globally convergent algorithm is developed to solve it. Moreover, multi-view data with a same set of samples can be available in various real problems. To this end, we extend the $\ell _\infty /\ell _{0}$-wsPLS model and propose two multi-view wsPLS models for multi-view data fusion. We develop an efficient iterative algorithm for each multi-view wsPLS model and show its convergence property. As well as numerical and biomedical data experiments demonstrate the efficiency of the proposed methods. Wenwen Min, Taosheng Xu, Chris Ding |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2024 | Masked Conditional Diffusion Model with GNN for Spatial Transcriptomics Data ImputationabstractSpatially resolved transcriptomics represents a significant advancement in single-cell analysis by offering both gene expression data and their corresponding physical locations. However, this high degree of spatial resolution entails a drawback, as the resulting spatial transcriptomic data at the cellular level is notably plagued by a high incidence of missing values. Furthermore, most existing imputation methods either overlook the spatial information between spots or compromise the overall gene expression data distribution. To address these challenges, our primary focus is on effectively utilizing the spatial location information within spatial transcriptomic data to impute missing values, while preserving the overall data distribution. We introduce stMCDI, a masked conditional diffusion model for spatial transcriptomics data imputation, which employs a denoising network trained using randomly masked data portions as guidance, with the unmasked data serving as conditions. Additionally, it utilizes a GNN encoder to integrate the spatial position information, thereby enhancing model performance. Compared with baseline methods, our model achieves state-of-the-art performance in all evaluation metrics on six real-world datasets. The results obtained from spatial transcriptomics datasets elucidate the performance of our methods relative to existing approaches. Our code can be accessed at https://github.com/wenwenmin/stMCDI. Wenwen Min, Shunfang Wang, Changmiao Wang, Taosheng Xu |
BIBM | 5 |
| 2024 | scASDC: Attention Enhanced Structural Deep Clustering for Single-cell RNA-seq DataabstractSingle-cell RNA sequencing (scRNA-seq) data analysis is pivotal for understanding cellular heterogeneity. However, the high sparsity and complex noise patterns inherent in scRNA-seq data present significant challenges for traditional clustering methods. To address these issues, we propose a deep clustering method, Attention-Enhanced Structural Deep Embedding Graph Clustering (scASDC), which integrates multiple advanced modules to improve clustering accuracy and robustness. Our approach employs a multi-layer graph convolutional network (GCN) to capture high-order structural relationships between cells, termed as the graph autoencoder module. We introduce a ZINB-based autoencoder module that extracts content information from the data and learns latent representations of gene expression. These modules are further integrated through an attention fusion mechanism, ensuring effective combination of gene expression and structural information at each layer of the GCN. Additionally, a self-supervised learning module is incorporated to enhance the robustness of the learned embeddings. Extensive experiments demonstrate that scASDC outperforms existing state-of-the-art methods, providing a robust and effective solution for single-cell clustering tasks. All code and public datasets used in this paper are available at https://github.com/wenwenmin/scASDC. Wenwen Min, Taosheng Xu, Guangsheng Wu, Shunfang Wang |
BIBM | 4 |
| 2023 | Multimodal attention-based variational autoencoder for clinical risk predictionabstractPrediction of survival risk in cancer patients is crucial for understanding the underlying mechanisms of canceration in different stages. Previous studies mainly relied on single-modal omics data due to technological constraints. However, with the increasing availability of cancer omics data, researchers have focused on the use of multi-omics and multimodal data for survival analysis. The application of deep learning methods has become an option for the prediction of clinical risk. Recent advances in the attention mechanism and the variational autoencoder (VAE) have made them promising for analyzing cancer omics data. However, VAE has limitations in disregarding the importance of different features between modalities, and the introduction of an attention mechanism could address this limitation. In this study, we propose a Multimodal Attention-based VAE (MAVAE) deep learning framework using cross-modal multihead attention to integrate cancer multi-omics data for clinical risk prediction. We evaluated our approach on eight TCGA datasets. We find that (1) MAVAE outperforms traditional machine learning and recent deep learning methods; (2) Multi-modal data yields better classification performance than single-modal data; (3) The multi-head attention mechanism improves the decision-making process; (4) Clinical and genetic data are the most important modal data. Our implementation of MAVAE is available at https://github.com/wenwenmin/MAVAE. Taosheng Xu, Jun Wan 0005, Wenwen Min |
BIBM | 2 |
| 2023 | Structured Sparse Non-Negative Matrix Factorization With $\ell _{2,0}$ℓ2,0-NormabstractNon-negative matrix factorization (NMF) is a powerful tool for dimensionality reduction and clustering. However, the interpretation of the clustering result from NMF is difficult, especially for the high-dimensional biological data without effective feature selection. To address this problem, we introduce a row-sparse NMF with$\ell _{2,0}$-norm constraint (NMF$\_\ell _{20}$), where the basis matrix$\bm {W}$is constrained by using the$\ell _{2,0}$-norm constraint such that$\bm {W}$has a row-sparsity pattern with feature selection. However, it is a challenge to solve the model, because the$\ell _{2,0}$-norm constraint is a non-convex and non-smooth function. Fortunately, we prove that the$\ell _{2,0}$-norm constraint satisfies the Kurdyka-Łojasiewicz property. Based on this finding, we present a proximal alternating linearized minimization algorithm and its monotone accelerated version to solve the NMF$\_\ell _{20}$model. In addition, we further present a orthogonal NMF with$\ell _{2,0}$-norm constraint (ONMF$\_\ell _{20}$) to enhance the clustering performance by using a non-negative orthogonal constraint. The ONMF$\_\ell _{20}$model is solved by transforming into a series of constrained and penalized matrix factorization problems. The convergence and guarantees for these proposed algorithms are proved and the computational complexity is well evaluated. The results on numerical and scRNA-seq datasets demonstrate the efficiency of our methods in comparison with existing methods. Wenwen Min, Taosheng Xu, Tsung-Hui Chang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Uncovering the roles of microRNAs/lncRNAs in characterising breast cancer subtypes and prognosisabstractBACKGROUND: Accurate prognosis and identification of cancer subtypes at molecular level are important steps towards effective and personalised treatments of breast cancer. To this end, many computational methods have been developed to use gene (mRNA) expression data for breast cancer subtyping and prognosis. Meanwhile, microRNAs (miRNAs) and long non-coding RNAs (lncRNAs) have been extensively studied in the last 2 decades and their associations with breast cancer subtypes and prognosis have been evidenced. However, it is not clear whether using miRNA and/or lncRNA expression data helps improve the performance of gene expression based subtyping and prognosis methods, and this raises challenges as to how and when to use these data and methods in practice. RESULTS: In this paper, we conduct a comparative study of 35 methods, including 12 breast cancer subtyping methods and 23 breast cancer prognosis methods, on a collection of 19 independent breast cancer datasets. We aim to uncover the roles of miRNAs and lncRNAs in breast cancer subtyping and prognosis from the systematic comparison. In addition, we created an R package, CancerSubtypesPrognosis, including all the 35 methods to facilitate the reproducibility of the methods and streamline the evaluation. CONCLUSIONS: The experimental results show that integrating miRNA expression data helps improve the performance of the mRNA-based cancer subtyping methods. However, miRNA signatures are not as good as mRNA signatures for breast cancer prognosis. In general, lncRNA expression data does not help improve the mRNA-based methods in both cancer subtyping and cancer prognosis. These results suggest that the prognostic roles of miRNA/lncRNA signatures in the improvement of breast cancer prognosis needs to be further verified. Buu Minh Thanh Truong, Taosheng Xu, Lin Liu 0003, Jiuyong Li, Thuc Duy Le |
BMC Bioinform. | 3 |
| 2021 | Exploring cell-specific miRNA regulation with single-cell miRNA-mRNA co-sequencing dataabstractBACKGROUND: Existing computational methods for studying miRNA regulation are mostly based on bulk miRNA and mRNA expression data. However, bulk data only allows the analysis of miRNA regulation regarding a group of cells, rather than the miRNA regulation unique to individual cells. Recent advance in single-cell miRNA-mRNA co-sequencing technology has opened a way for investigating miRNA regulation at single-cell level. However, as currently single-cell miRNA-mRNA co-sequencing data is just emerging and only available at small-scale, there is a strong need of novel methods to exploit existing single-cell data for the study of cell-specific miRNA regulation. RESULTS: In this work, we propose a new method, CSmiR (Cell-Specific miRNA regulation) to combine single-cell miRNA-mRNA co-sequencing data and putative miRNA-mRNA binding information to identify miRNA regulatory networks at the resolution of individual cells. We apply CSmiR to the miRNA-mRNA co-sequencing data in 19 K562 single-cells to identify cell-specific miRNA-mRNA regulatory networks for understanding miRNA regulation in each K562 single-cell. By analyzing the obtained cell-specific miRNA-mRNA regulatory networks, we observe that the miRNA regulation in each K562 single-cell is unique. Moreover, we conduct detailed analysis on the cell-specific miRNA regulation associated with the miR-17/92 family as a case study. The comparison results indicate that CSmiR is effective in predicting cell-specific miRNA targets. Finally, through exploring cell-cell similarity matrix characterized by cell-specific miRNA regulation, CSmiR provides a novel strategy for clustering single-cells and helps to understand cell-cell crosstalk. CONCLUSIONS: To the best of our knowledge, CSmiR is the first method to explore miRNA regulation at a single-cell resolution level, and we believe that it can be a useful method to enhance the understanding of cell-specific miRNA regulation. Junpeng Zhang 0001, Lin Liu 0003, Taosheng Xu, Chunwen Zhao, Sijing Li, Jiuyong Li, Nini Rao, Thuc Duy Le |
BMC Bioinform. | 3 |
| 2020 | Correction to: Identifying miRNA synergism using multiple-intervention causal inferenceabstractAfter publication of this supplement article [1], it was brought to our attention that the Fig. 3 was incorrect. The correct Fig. 3 is as below. Junpeng Zhang 0001, Vu Viet Hoang Pham, Lin Liu 0003, Taosheng Xu, Buu Minh Thanh Truong, Jiuyong Li, Nini Rao, Thuc Duy Le |
BMC Bioinform. | 4 |
| 2020 | A novel single-cell based method for breast cancer prognosisabstractBreast cancer prognosis is challenging due to the heterogeneity of the disease. Various computational methods using bulk RNA-seq data have been proposed for breast cancer prognosis. However, these methods suffer from limited performances or ambiguous biological relevance, as a result of the neglect of intra-tumor heterogeneity. Recently, single cell RNA-sequencing (scRNA-seq) has emerged for studying tumor heterogeneity at cellular levels. In this paper, we propose a novel method, scPrognosis, to improve breast cancer prognosis with scRNA-seq data. scPrognosis uses the scRNA-seq data of the biological process Epithelial-to-Mesenchymal Transition (EMT). It firstly infers the EMT pseudotime and a dynamic gene co-expression network, then uses an integrative model to select genes important in EMT based on their expression variation and differentiation in different stages of EMT, and their roles in the dynamic gene co-expression network. To validate and apply the selected signatures to breast cancer prognosis, we use them as the features to build a prediction model with bulk RNA-seq data. The experimental results show that scPrognosis outperforms other benchmark breast cancer prognosis methods that use bulk RNA-seq data. Moreover, the dynamic changes in the expression of the selected signature genes in EMT may provide clues to the link between EMT and clinical outcomes of breast cancer. scPrognosis will also be useful when applied to scRNA-seq datasets of different biological processes other than EMT. Lin Liu 0003, Gregory J. Goodall, Andreas W. Schreiber, Taosheng Xu, Jiuyong Li, Thuc Duy Le |
PLoS Comput. Biol. | 5 |
| 2020 | LMSM: A modular approach for identifying lncRNA related miRNA sponge modules in breast cancerabstractUntil now, existing methods for identifying lncRNA related miRNA sponge modules mainly rely on lncRNA related miRNA sponge interaction networks, which may not provide a full picture of miRNA sponging activities in biological conditions. Hence there is a strong need of new computational methods to identify lncRNA related miRNA sponge modules. In this work, we propose a framework, LMSM, to identify LncRNA related MiRNA Sponge Modules from heterogeneous data. To understand the miRNA sponging activities in biological conditions, LMSM uses gene expression data to evaluate the influence of the shared miRNAs on the clustered sponge lncRNAs and mRNAs. We have applied LMSM to the human breast cancer (BRCA) dataset from The Cancer Genome Atlas (TCGA). As a result, we have found that the majority of LMSM modules are significantly implicated in BRCA and most of them are BRCA subtype-specific. Most of the mediating miRNAs act as crosslinks across different LMSM modules, and all of LMSM modules are statistically significant. Multi-label classification analysis shows that the performance of LMSM modules is significantly higher than baseline's performance, indicating the biological meanings of LMSM modules in classifying BRCA subtypes. The consistent results suggest that LMSM is robust in identifying lncRNA related miRNA sponge modules. Moreover, LMSM can be used to predict miRNA targets. Finally, LMSM outperforms a graph clustering-based strategy in identifying BRCA-related modules. Altogether, our study shows that LMSM is a promising method to investigate modular regulatory mechanism of sponge lncRNAs from heterogeneous data. Junpeng Zhang 0001, Taosheng Xu, Lin Liu 0003, Chunwen Zhao, Sijing Li, Jiuyong Li, Nini Rao, Thuc Duy Le |
PLoS Comput. Biol. | 2 |
| 2019 | Identifying miRNA-mRNA regulatory relationships in breast cancer with invariant causal predictionabstractBACKGROUND: microRNAs (miRNAs) regulate gene expression at the post-transcriptional level and they play an important role in various biological processes in the human body. Therefore, identifying their regulation mechanisms is essential for the diagnostics and therapeutics for a wide range of diseases. There have been a large number of researches which use gene expression profiles to resolve this problem. However, the current methods have their own limitations. Some of them only identify the correlation of miRNA and mRNA expression levels instead of the causal or regulatory relationships while others infer the causality but with a high computational complexity. To overcome these issues, in this study, we propose a method to identify miRNA-mRNA regulatory relationships in breast cancer using the invariant causal prediction. The key idea of invariant causal prediction is that the cause miRNAs of their target mRNAs are the ones which have persistent causal relationships with the target mRNAs across different environments. RESULTS: In this research, we aim to find miRNA targets which are consistent across different breast cancer subtypes. Thus, first of all, we apply the Pam50 method to categorize BRCA samples into different "environment" groups based on different cancer subtypes. Then we use the invariant causal prediction method to find miRNA-mRNA regulatory relationships across subtypes. We validate the results with the miRNA-transfected experimental data and the results show that our method outperforms the state-of-the-art methods. In addition, we also integrate this new method with the Pearson correlation analysis method and Lasso in an ensemble method to take the advantages of these methods. We then validate the results of the ensemble method with the experimentally confirmed data and the ensemble method shows the best performance, even comparing to the proposed causal method. CONCLUSIONS: This research found miRNA targets which are consistent across different breast cancer subtypes. Further functional enrichment analysis shows that miRNAs involved in the regulatory relationships predicated by the proposed methods tend to synergistically regulate target genes, indicating the usefulness of these methods, and the identified miRNA targets could be used in the design of wet-lab experiments to discover the causes of breast cancer. Vu Viet Hoang Pham, Junpeng Zhang 0001, Lin Liu 0003, Buu Minh Thanh Truong, Taosheng Xu, Trung T. Nguyen, Jiuyong Li, Thuc Duy Le |
BMC Bioinform. | 5 |
| 2019 | miRspongeR: an R/Bioconductor package for the identification and analysis of miRNA sponge interaction networks and modulesabstractBACKGROUND: A microRNA (miRNA) sponge is an RNA molecule with multiple tandem miRNA response elements that can sequester miRNAs from their target mRNAs. Despite growing appreciation of the importance of miRNA sponges, our knowledge of their complex functions remains limited. Moreover, there is still a lack of miRNA sponge research tools that help researchers to quickly compare their proposed methods with other methods, apply existing methods to new datasets, or select appropriate methods for assisting in subsequent experimental design. RESULTS: To fill the gap, we present an R/Bioconductor package, miRspongeR, for simplifying the procedure of identifying and analyzing miRNA sponge interaction networks and modules. It provides seven popular methods and an integrative method to identify miRNA sponge interactions. Moreover, it supports the validation of miRNA sponge interactions and the identification of miRNA sponge modules, as well as functional enrichment and survival analysis of miRNA sponge modules. CONCLUSIONS: This package enables researchers to quickly evaluate their new methods, apply existing methods to new datasets, and consequently speed up miRNA sponge research. Junpeng Zhang 0001, Lin Liu 0003, Taosheng Xu, Chunwen Zhao, Jiuyong Li, Thuc Duy Le |
BMC Bioinform. | 3 |
| 2019 | Identifying miRNA synergism using multiple-intervention causal inferenceabstractBACKGROUND: Studying multiple microRNAs (miRNAs) synergism in gene regulation could help to understand the regulatory mechanisms of complicated human diseases caused by miRNAs. Several existing methods have been presented to infer miRNA synergism. Most of the current methods assume that miRNAs with shared targets at the sequence level are working synergistically. However, it is unclear if miRNAs with shared targets are working in concert to regulate the targets or they individually regulate the targets at different time points or different biological processes. A standard method to test the synergistic activities is to knock-down multiple miRNAs at the same time and measure the changes in the target genes. However, this approach may not be practical as we would have too many sets of miRNAs to test. RESULTS: n this paper, we present a novel framework called miRsyn for inferring miRNA synergism by using a causal inference method that mimics the multiple-intervention experiments, e.g. knocking-down multiple miRNAs, with observational data. Our results show that several miRNA-miRNA pairs that have shared targets at the sequence level are not working synergistically at the expression level. Moreover, the identified miRNA synergistic network is small-world and biologically meaningful, and a number of miRNA synergistic modules are significantly enriched in breast cancer. Our further analyses also reveal that most of synergistic miRNA-miRNA pairs show the same expression patterns. The comparison results indicate that the proposed multiple-intervention causal inference method performs better than the single-intervention causal inference method in identifying miRNA synergistic network. CONCLUSIONS: Taken together, the results imply that miRsyn is a promising framework for identifying miRNA synergism, and it could enhance the understanding of miRNA synergism in breast cancer. Junpeng Zhang 0001, Vu Viet Hoang Pham, Lin Liu 0003, Taosheng Xu, Buu Minh Thanh Truong, Jiuyong Li, Nini Rao, Thuc Duy Le |
BMC Bioinform. | 4 |
| 2018 | miRBaseConverter: an R/Bioconductor package for converting and retrieving miRNA name, accession, sequence and family information in different versions of miRBaseabstractBACKGROUND: miRBase is the primary repository for published miRNA sequence and annotation data, and serves as the "go-to" place for miRNA research. However, the definition and annotation of miRNAs have been changed significantly across different versions of miRBase. The changes cause inconsistency in miRNA related data between different databases and articles published at different times. Several tools have been developed for different purposes of querying and converting the information of miRNAs between different miRBase versions, but none of them individually can provide the comprehensive information about miRNAs in miRBase and users will need to use a number of different tools in their analyses. RESULTS: We introduce miRBaseConverter, an R package integrating the latest miRBase version 22 available in Bioconductor to provide a suite of functions for converting and retrieving miRNA name (ID), accession, sequence, species, version and family information in different versions of miRBase. The package is implemented in R and available under the GPL-2 license from the Bioconductor website ( http://bioconductor.org/packages/miRBaseConverter/ ). A Shiny-based GUI suitable for non-R users is also available as a standalone application from the package and also as a web application at http://nugget.unisa.edu.au:3838/miRBaseConverter . miRBaseConverter has a built-in database for querying miRNA information in all species and for both pre-mature and mature miRNAs defined by miRBase. In addition, it is the first tool for batch querying the miRNA family information. The package aims to provide a comprehensive and easy-to-use tool for miRNA research community where researchers often utilize published miRNA data from different sources. CONCLUSIONS: The Bioconductor package miRBaseConverter and the Shiny-based web application are presented to provide a suite of functions for converting and retrieving miRNA name, accession, sequence, species, version and family information in different versions of miRBase. The package will serve a wide range of applications in miRNA research and could provide a full view of the miRNAs of interest. Taosheng Xu, Lin Liu 0003, Junpeng Zhang 0001, Weijia Zhang 0001, Jie Gui, Kui Yu, Jiuyong Li, Thuc Duy Le |
BMC Bioinform. | 1 |
| 2017 | CancerSubtypes: an R/Bioconductor package for molecular cancer subtype identification, validation and visualizationabstractSUMMARY: Identifying molecular cancer subtypes from multi-omics data is an important step in the personalized medicine. We introduce CancerSubtypes, an R package for identifying cancer subtypes using multi-omics data, including gene expression, miRNA expression and DNA methylation data. CancerSubtypes integrates four main computational methods which are highly cited for cancer subtype identification and provides a standardized framework for data pre-processing, feature selection, and result follow-up analyses, including results computing, biology validation and visualization. The input and output of each step in the framework are packaged in the same data format, making it convenience to compare different methods. The package is useful for inferring cancer subtypes from an input genomic dataset, comparing the predictions from different well-known methods and testing new subtype discovery methods, as shown with different application scenarios in the Supplementary Material. AVAILABILITY AND IMPLEMENTATION: The package is implemented in R and available under GPL-2 license from the Bioconductor website (http://bioconductor.org/packages/CancerSubtypes/). CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Taosheng Xu, Thuc Duy Le, Lin Liu 0003, Rujing Wang, Bing-Yu Sun, Antonio Colaprico, Gianluca Bontempi, Jiuyong Li |
Bioinform. | 1 |