EDBT 2026 Demo / reviewers in the wild / expert
Bo Yang 0041
dblp:46/999-41
· DBLP profile ↗
22ranked-venue papers
9as first author
13since 2021 · last 2026
0000-0001-6200-8727ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep embedding with consensus self-expressiveness for subspace clustering
Meng Li 0092, Bo Yang 0041 |
Neurocomputing | 2 |
| 2026 | AugGCL: Multimodal graph learning for spatial transcriptomics analysis with enhanced gene and morphological dataabstractSpatial transcriptomics enables the measurement of gene expression in intact tissues. Despite this, reconstructing anatomically accurate spatial domains remains challenging, primarily due to expression sparsity, complex tissue architecture that is characterized by sharp boundaries and long-range continuity, and weak spatial signals. Traditional pipelines typically rely on expression-driven clustering and spatial smoothing, which underperform at boundaries and in sparse regions while neglecting morphological information. To address these challenges, AugGCL is proposed, an augmented graph-convolutional learning framework that enhances spatial structure decoding and gene expression reconstruction through targeted augmentation of both gene and image data. A key component of AugGCL is neighborhood information aggregation mechanism, which integrates expression similarity and spatial proximity to construct a weighted graph and an enhanced expression matrix, addressing sparsity without sacrificing boundary clarity. Additionally, a two stream weighted graph convolutional network jointly models refined gene features and image-derived morphological information, with image-aware auxiliary reconstructions enhancing weak spatial signals and sharpening boundaries. On datasets from the human dorsolateral prefrontal cortex, breast cancer, and mouse embryo, AugGCL outperforms baseline methods across multiple metrics, showing robustness and generalization across a range of datasets. Downstream analysis validated the reliability of the method, confirming its effectiveness in cell annotation, functional enrichment, and mechanistic studies. AugGCL generates clearer spatial domains and significantly advances the application of spatial transcriptomics in tissue structure and disease research. Tengfei Ji, Bo Yang 0041, Huazhe Yang, Yizhuo Liu |
PLoS Comput. Biol. | 2 |
| 2025 | Towards Lightweight Time Series Forecasting: A Patch-Wise Transformer with Weak Data EnrichingabstractPatch-wise Transformer based time series forecasting achieves superior accuracy. However, this superiority relies heavily on intricate model design with massive parameters, rendering both training and inference expensive, thus preventing their deployments on edge devices with limited resources and low latency requirements. In addition, existing methods often work in an autoregressive manner, which take into account only historical values, but ignore valuable, easy-to-obtain context information, such as weather forecasts, date and time of day. To contend with the two limitations, we propose LiPFormer, a novel Lightweight Patch-wise Transformer with weak data enriching. First, to simplify the Transformer backbone, LiPFormer employs a novel lightweight cross-patch attention and a linear transformationbased attention to eliminate Layer Normalization and Feed Forward Network, two heavy components in existing Transformers. Second, we propose a lightweight, weak data enriching module to provide additional, valuable weak supervision to the training. It enhances forecasting accuracy without significantly increasing model complexity as it does not involve expensive, human-labeling but using easily accessible context information. This facilitates the weak data enriching to plug-and-play on existing models. Extensive experiments on nine benchmark time series datasets demonstrate that LiPFormer outperforms state-of-the-art methods in accuracy, while significantly reducing parameter scale, training duration, and GPU memory usage. Deployment on an edge device reveals that LiPFormer takes only 1/3 inference time compared to classic Transformers. In addition, we demonstrate that the weak data enriching can integrate seamlessly into various Transformer based models to enhance their accuracy, suggesting its generality. Meng Wang 0015, Jintao Yang, Bin Yang 0002, Hui Li 0005, Tongxin Gong, Bo Yang 0041, Jiangtao Cui |
ICDE | 6 |
| 2025 | MC2LS: Towards Efficient Collective Location Selection in Competition: (Extended Abstract)abstractCollective Location Selection (CLS) aims to identify$k$optimal sites for facility establishment to collectively maximize user attraction. Traditional CLS approaches often overlook user mobility and inter-facility competition, critical factors in real-world scenarios. This paper introduces MC2LS, the first effort on CLS that addresses these gaps by considering user mobility and peer competition. Solving MC2LS is nontrivial due to its NP-hardness. To overcome the challenge of pruning multi-point users with highly overlapping minimum boundary rectangles (MBRs), we develop a position count threshold and two square-based pruning rules. We propose IQuad-tree, a user-MBR-free index, to benefit the hierarchical and batch-wise properties of the pruning rules. We present an$(1-\frac{1}{e})$-approximate greedy solution to MC2LS, and empirical studies demonstrate the superiority of our proposed solution over the state-of-the-art techniques. Meng Wang 0015, Mengfei Zhao, Hui Li 0005, Jiangtao Cui, Bo Yang 0041, Tao Xue 0001 |
ICDE | 5 |
| 2025 | Multi-layer matrix factorization for cancer subtyping using full and partial multi-omics datasetabstractCancer, with its inherent heterogeneity, is commonly categorized into distinct subtypes based on unique traits, cellular origins, and molecular markers specific to each type. However, current studies primarily rely on complete multi-omics datasets for predicting cancer subtypes, often overlooking predictive performance in cases where some omics data may be missing and neglecting implicit relationships across multiple layers of omics data integration. This paper introduces Multi-Layer Matrix Factorization (MLMF), a novel approach for cancer subtyping that employs multi-omics data clustering. MLMF initially processes multi-omics feature matrices by performing multi-layer linear or nonlinear factorization, decomposing the original data into latent feature representations unique to each omics type. These latent representations are subsequently fused into a consensus form, on which spectral clustering is performed to determine subtypes. Additionally, MLMF incorporates a class indicator matrix to handle missing omics data, creating a unified framework that can manage both complete and incomplete multi-omics data. Extensive experiments conducted on 12 multi-omics cancer datasets, both complete and with missing values, demonstrate that MLMF achieves results that are comparable to or surpass the performance of several state-of-the-art approaches. MLMF is open source and available at (https://github.com/renyingxuan/MLMF.git). Yingxuan Ren, Fengtao Ren, Bo Yang 0041 |
Briefings Bioinform. | 3 |
| 2025 | Multi-view multi-level contrastive graph convolutional network for cancer subtyping on multi-omics dataabstractCancer is a highly diverse group of diseases, and each type of cancer can be further divided into various subtypes according to specific characteristics, cellular origins, and molecular markers. Subtyping helps in tailoring treatment and prognosis accuracy. However, the existing studies are more concerned with integrating different omics data to discover potential connections, but ignoring the relationships between consensus information and individual information within each omics level during the integration process. To this end, we propose a novel fusion-free method called multi-view multi-level contrastive graph convolutional network (M$^{2}$CGCN) for cancer subtyping. M$^{2}$CGCN learns multi-level features, i.e. high-level and low-level features, respectively. The low-level features from each view capture the intrinsic information in each omics by reconstruction of node attribute and graph structures. The high-level features achieve cancer subtyping via contrastive learning. Comprehensive experiments were performed on 34 multi-omics cancer datasets. The findings indicate that M$^{2}$CGCN achieves results comparable to or surpassing many state-of-the-art methods. Bo Yang 0041, Chenxi Cui, Feiyue Gao |
Briefings Bioinform. | 1 |
| 2025 | PGC-CSS: A parallel graph clustering framework with collaborative self-supervision
Meng Li 0092, Jun Wu 0024, Bo Yang 0041 |
Knowl. Based Syst. | 3 |
| 2025 | MC$^{2}$2LS: Towards Efficient Collective Location Selection in CompetitionabstractCollective Location Selection (CLS) has received significant research attention in the spatial database community due to its wide range of applications. The CLS problem selects a group ofkpreferred locations among candidate sites to establish facilities, aimed at collectively attracting the maximum number of users. Existing studies commonly assume every user is located in a fixed position, without considering the competition between peer facilities. Unfortunately, in real markets, users are mobile and choose to patronize from a host of competitors, making traditional techniques unavailable. To this end, this paper presents the first effort on a CLS problem in competition scenarios, calledmc$^{2}$2ls, taking into account the mobility factor. Solvingmc$^{2}$2lsis a non-trivial task due to its NP-hardness. To overcome the challenge of pruning multi-point users with highly overlapped minimum boundary rectangles (MBRs), we exploit a position count threshold and design two square-based pruning rules. We introduce IQuad-tree, a user-MBR-free index, to benefit the hierarchical and batch-wise properties of the pruning rules. We propose an$(1-\frac{1}{e})$-approximate greedy solution tomc$^{2}$2lsand incorporate a candidate-pruning strategy to further accelerate the computation for handling skewed datasets. Extensive experiments are conducted on real datasets, demonstrating the superiority of our proposed pruning rules and solution compared to the state-of-the-art techniques. Meng Wang 0015, Mengfei Zhao, Hui Li 0005, Jiangtao Cui, Bo Yang 0041, Tao Xue 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | MRGCN: cancer subtyping with multi-reconstruction graph convolutional network using full and partial multi-omics datasetabstractMOTIVATION: Cancer is a molecular complex and heterogeneous disease. Each type of cancer is usually composed of several subtypes with different treatment responses and clinical outcomes. Therefore, subtyping is a crucial step in cancer diagnosis and therapy. The rapid advances in high-throughput sequencing technologies provide an increasing amount of multi-omics data, which benefits our understanding of cancer genetic architecture, and yet poses new challenges in multi-omics data integration. RESULTS: We propose a graph convolutional network model, called MRGCN for multi-omics data integrative representation. MRGCN simultaneously encodes and reconstructs multiple omics expression and similarity relationships into a shared latent embedding space. In addition, MRGCN adopts an indicator matrix to denote the situation of missing values in partial omics, so that the full and partial multi-omics processing procedures are combined in a unified framework. Experimental results on 11 multi-omics datasets show that cancer subtypes obtained by MRGCN with superior enriched clinical parameters and log-rank test P-values in survival analysis over many typical integrative methods. AVAILABILITY AND IMPLEMENTATION: https://github.com/Polytech-bioinf/MRGCN.git https://figshare.com/articles/software/MRGCN/23058503. Bo Yang 0041, Meng Wang 0001, Xueping Su |
Bioinform. | 1 |
| 2022 | Deep structure integrative representation of multi-omics data for cancer subtypingabstractMOTIVATION: Cancer is a heterogeneous group of diseases. Cancer subtyping is a crucial and critical step to diagnosis, prognosis and treatment. Since high-throughput sequencing technologies provide an unprecedented opportunity to rapidly collect multi-omics data for the same individuals, an urgent need in current is how to effectively represent and integrate these multi-omics data to achieve clinically meaningful cancer subtyping. RESULTS: We propose a novel deep learning model, called Deep Structure Integrative Representation (DSIR), for cancer subtypes dentification by integrating representation and clustering multi-omics data. DSIR simultaneously captures the global structures in sparse subspace and local structures in manifold subspace from multi-omics data and constructs a consensus similarity matrix by utilizing deep neural networks. Extensive tests are performed in 12 different cancers on three levels of omics data from The Cancer Genome Atlas. The results demonstrate that DSIR obtains more significant performances than the state-of-the-art integrative methods. AVAILABILITY AND IMPLEMENTATION: https://github.com/Polytech-bioinf/Deep-structure-integrative-representation.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bo Yang 0041, Xueping Su |
Bioinform. | 1 |
| 2022 | Tensor-based multi-view clustering with consistency exploration and diversity regularization
Wenyu Hao, Shanmin Pang, Bo Yang 0041, Jianru Xue |
Knowl. Based Syst. | 3 |
| 2021 | Deep Subspace Mutual Learning for cancer subtypes predictionabstractMOTIVATION: Precise prediction of cancer subtypes is of significant importance in cancer diagnosis and treatment. Disease etiology is complicated existing at different omics levels; hence integrative analysis provides a very effective way to improve our understanding of cancer. RESULTS: We propose a novel computational framework, named Deep Subspace Mutual Learning (DSML). DSML has the capability to simultaneously learn the subspace structures in each available omics data and in overall multi-omics data by adopting deep neural networks, which thereby facilitates the subtype's prediction via clustering on multi-level, single-level and partial-level omics data. Extensive experiments are performed in five different cancers on three levels of omics data from The Cancer Genome Atlas. The experimental analysis demonstrates that DSML delivers comparable or even better results than many state-of-the-art integrative methods. AVAILABILITY AND IMPLEMENTATION: An implementation and documentation of the DSML is publicly available at https://github.com/polytechnicXTT/Deep-Subspace-Mutual-Learning.git. Bo Yang 0041, Tingting Xin, Shanmin Pang, Meng Wang 0001 |
Bioinform. | 1 |
| 2021 | Integrating Multi-Omic Data With Deep Subspace Fusion Clustering for Cancer Subtype PredictionabstractOne type of cancer usually consists of several subtypes with distinct clinical implications, thus the cancer subtype prediction is an important task in disease diagnosis and therapy. Utilizing one type of data from molecular layers in biological system to predict is difficult to bridge the cancer genome to cancer phenotypes, since the genome is neither simple nor independent but rather complicated and dysregulated from multiple molecular mechanisms. Similarity Network Fusion (SNF) has been recently proposed to integrate diverse omics data for improving the understanding of tumorigenesis. SNF adopts Euclidean distance to measure the similarity between patients, which shows some limitations. In this article, we introduce a novel prediction technique as an extension of SNF, namely Deep Subspace Fusion Clustering (DSFC). DSFC utilizes auto-encoder and data self-expressiveness approaches to guide a deep subspace model, which can achieve effective expression of discriminative similarity between patients. As a result, the dissimilarity between inter-cluster is delivered and enhanced compactness of intra-cluster is achieved at the same time. The validity of DSFC is examined by extensive simulations over six different cancer through three levels omics data. The survival analysis demonstrates that DSFC delivers comparable or even better results than many state-of-the-art integrative methods. Bo Yang 0041, Shanmin Pang, Xuequn Shang 0001, Xueqing Zhao, Minghui Han |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | Deep Denoising Subspace Single-Cell Clustering
Bo Yang 0041 |
ICONIP (5) | 2 |
| 2020 | Deep Denoising Sparse CodingabstractSingle-cell Ribonucleic Acid sequencing (scRNA-seq) has great potential to discover cell types, identify cell states, trace development lineages, and reconstruct the spatial organization of cells. Clustering transcriptomes profiled by scRNA-seq has been routinely conducted to reveal cell heterogeneity and diversity. In fact, scRNA-seq data contain an abundance of dropout events that lead to zero expression measurements. These dropout events may be the result of technical sampling effects or real biology arising from stochastic transcriptional activity. Therefore clustering analysis of scRNA-seq data remains a statistical and computational challenge. Here, we have developed Deep Denoising Sparse Coding (DDSC), a deep clustering method combine autoencoder and sparse coding approach. Based on six real datasets from five representative single-cell sequencing platforms, DDSC outperformed some state-of-the-art methods under various clustering performance metrics and exhibited improved scalability. Its accuracy and efficiency make DDSC a promising algorithm for clustering large-scale scRNA-seq data. Bo Yang 0041 |
ICTAI | 2 |
| 2020 | Spatial-Content Image Search in Complex ScenesabstractAlthough the topic of image search has been heavily studied in the last two decades, many works have focused on either instance-level retrieval or semantic-level retrieval. In this work, we develop a novel visually similar spatial-semantic method, namely spatial-content image search, to search images that not only share the same spatial-semantics but also enjoy visual consistency as the query image in complex scenes. We achieve the goal by capturing spatial-semantic concepts as well as the visual representation of each concept contained in an image. Specifically, we first generate a set of bounding boxes and their category labels representing spatial-semantic constraints with YOLOV3, and then obtain visual content of each bounding box with deep features extracted from a convolutional neural network. After that, we customize a similarity computation method that evaluates the relevance between dataset images and input queries according to the developed image representations. Experimental results on two large-scale benchmark retrieval datasets with images consisting of multiple objects demonstrate that our method provides an effective way to query image databases. Our code is available at https://github.com/MaJinWakeUp/spatial-content. Shanmin Pang, Bo Yang 0041, Jihua Zhu, Yaochen Li |
WACV | 3 |
| 2020 | Structured feature for multi-label learning
Bo Yang 0041, Tingting Xin, Minghui Han, Xueqing Zhao, Jinguang Chen |
Neurocomputing | 1 |
| 2018 | Deep Subspace Similarity Fusion for the Prediction of Cancer Subtypes
Bo Yang 0041, Shuhui Liu, Shanmin Pang, Chenpai Pang, Xuequn Shang 0001 |
BIBM | 1 |
| 2017 | Graph regularized nonnegative sparse coding using incoherent dictionary for approximate nearest neighbor search
Ming Xiang, Bo Yang 0041 |
Pattern Recognit. | 3 |
| 2017 | Low-rank preserving embedding
Ming Xiang, Bo Yang 0041 |
Pattern Recognit. | 3 |
| 2017 | Isometric hashing for image retrieval
Bo Yang 0041, Xuequn Shang 0001, Shanmin Pang |
Signal Process. Image Commun. | 1 |
| 2016 | Multi-manifold Discriminant Isomap for visualization and classification
Bo Yang 0041, Ming Xiang |
Pattern Recognit. | 1 |