VLDB 2026 Research / reviewers in the wild / expert
Yunhe Wang 0002
dblp:63/8217-2
· DBLP profile ↗
22ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0002-0013-4530ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A niching archive-assisted evolutionary algorithm for multimodal feature selection in high-dimensional data classification
Yunhe Wang 0002, Zhengyu Du, Zeming Zhou, Xubin Wang 0001, Shengxiang Yang |
Knowl. Based Syst. | 1 |
| 2026 | Hierarchical Multi-View Graph Diffusion Weighted Model for Cancer Subtype IdentificationabstractAccurate cancer subtype identification is crucial for personalized medicine, as it enables precise diagnosis based on molecular characteristics. With the advent of large-scale multi-omics data from various resources, researchers now have unprecedented opportunities to explore cancer subtypes comprehensively. However, the inherent complexity, high dimensionality, and heterogeneity of these datasets present significant statistical and computational challenges, often leading to suboptimal clustering performance when inter-omics heterogeneity is overlooked. To address these challenges, we propose a novel method called the Hierarchical Multi-view Graph Diffusion Weighted (HMGDW) model for cancer subtype identification. Our approach begins with the generation of multiple base clusterings through random feature sampling, effectively mitigating the impact of high dimensionality. These base clusterings are subsequently integrated via a late integration strategy to yield the consensus clustering result. Then, we introduce a graph diffusion weighted mechanism that prioritizes views with the most significant contributions to the unified graph representation. Lastly, we conducted extensive experiments on both generic multi-view datasets and multi-omics cancer multi-omics datasets. The experimental results demonstrate that HMGDW consistently outperforms several state-of-the-art methods, achieving robust and accurate clustering. Additionally, a case study on the acute myeloid leukemia (AML) dataset validates the practical efficacy of our model in identifying clinically relevant subtypes. Yunhe Wang 0002, Zhengyu Du, Yanchi Su, Xiangtao Li |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Gradient-based federated Bayesian optimization
Junhua Gu, Qiqi Liu, Yunhe Wang 0002, Yaochu Jin |
Knowl. Based Syst. | 5 |
| 2025 | Evolving Dual-Directional Multiobjective Feature Selection for High-Dimensional Gene Expression DataabstractHigh-dimensional gene expression data has gained considerable attention in diverse medical fields such as disease diagnosis, with the challenges of the dimensionality curse and exponentially growing computation. To analyze the data, feature selection is an essential step by reducing the dimensionality. However, most feature selection algorithms for high-dimensional gene expression data still suffer from low classification and poor generalization ability. An evolutionary algorithm is an effective paradigm for enhancing global search capability in feature selection. Inspired by the evolutionary algorithm Competitive Swarm Optimization, we propose a Multiobjective Dual-directional Competitive Swarm Optimization (MODCSO) method for feature selection from high-dimensional gene expression data. First, we design a competitive swarm optimization algorithm framework based on multi-objective optimization to evolve three objective functions simultaneously. Then, we introduce a dual-directional learning strategy that trains particles within the loser group using two distinct learning strategies. To assess the effectiveness and efficiency of the suggested algorithm, we evaluate MODCSO through extensive experiments on twenty high-dimensional gene expression datasets and three real-world biological datasets. Compared to various leading feature selection algorithms, our proposed algorithm MODCSO exhibits superior competitiveness for the high-dimensional feature selection task. Moreover, we provide other extensive analyses to demonstrate further the robustness and biological interpretability of MODCSO in handling high-dimensional gene expression data. Yunhe Wang 0002, Zhengyu Du, Wenyuan Xiao, Hongpu Liu, Liang Yang 0002 |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Path Planning for Unmanned Aerial Vehicle Using Enhanced Dynamic Group Based Collaborative Optimization Algorithm
Wenyuan Xiao, Dong Wang 0074, Hongpu Liu, Yunhe Wang 0002 |
ICIC (1) | 5 |
| 2024 | PredGCN: a Pruning-enabled Gene-Cell Net for automatic cell annotation of single cell transcriptome dataabstractMOTIVATION: The annotation of cell types from single-cell transcriptomics is essential for understanding the biological identity and functionality of cellular populations. Although manual annotation remains the gold standard, the advent of automatic pipelines has become crucial for scalable, unbiased, and cost-effective annotations. Nonetheless, the effectiveness of these automatic methods, particularly those employing deep learning, significantly depends on the architecture of the classifier and the quality and diversity of the training datasets. RESULTS: To address these limitations, we present a Pruning-enabled Gene-Cell Net (PredGCN) incorporating a Coupled Gene-Cell Net (CGCN) to enable representation learning and information storage. PredGCN integrates a Gene Splicing Net (GSN) and a Cell Stratification Net (CSN), employing a pruning operation (PrO) to dynamically tackle the complexity of heterogeneous cell identification. Among them, GSN leverages multiple statistical and hypothesis-driven feature extraction methods to selectively assemble genes with specificity for scRNA-seq data while CSN unifies elements based on diverse region demarcation principles, exploiting the representations from GSN and precise identification from different regional homogeneity perspectives. Furthermore, we develop a multi-objective Pareto pruning operation (Pareto PrO) to expand the dynamic capabilities of CGCN, optimizing the sub-network structure for accurate cell type annotation. Multiple comparison experiments on real scRNA-seq datasets from various species have demonstrated that PredGCN surpasses existing state-of-the-art methods, including its scalability to cross-species datasets. Moreover, PredGCN can uncover unknown cell types and provide functional genomic analysis by quantifying the influence of genes on cell clusters, bringing new insights into cell type identification and characterizing scRNA-seq data from different perspectives. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/IrisQi7/PredGCN and test data is available at https://figshare.com/articles/dataset/PredGCN/25251163. Yunhe Wang 0002, Yujian Huang, Xiangtao Li |
Bioinform. | 2 |
| 2024 | Evolving pathway activation from cancer gene expression data using nature-inspired ensemble optimization
Xubin Wang 0001, Yunhe Wang 0002, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
Expert Syst. Appl. | 2 |
| 2024 | Exhaustive Exploitation of Nature-Inspired Computation for Cancer Screening in an Ensemble MannerabstractAccurate screening of cancer types is crucial for effective cancer detection and precise treatment selection. However, the association between gene expression profiles and tumors is often limited to a small number of biomarker genes. While computational methods using nature-inspired algorithms have shown promise in selecting predictive genes, existing techniques are limited by inefficient search and poor generalization across diverse datasets. This study presents a framework termed Evolutionary Optimized Diverse Ensemble Learning (EODE) to improve ensemble learning for cancer classification from gene expression data. The EODE methodology combines an intelligent grey wolf optimization algorithm for selective feature space reduction, guided random injection modeling for ensemble diversity enhancement, and subset model optimization for synergistic classifier combinations. Extensive experiments were conducted across 35 gene expression benchmark datasets encompassing varied cancer types. Results demonstrated that EODE obtained significantly improved screening accuracy over individual and conventionally aggregated models. The integrated optimization of advanced feature selection, directed specialized modeling, and cooperative classifier ensembles helps address key challenges in current nature-inspired approaches. This provides an effective framework for robust and generalized ensemble learning with gene expression biomarkers. Xubin Wang 0001, Yunhe Wang 0002, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | scBGEDA: deep single-cell clustering analysis via a dual denoising autoencoder with bipartite graph ensemble clusteringabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) is an increasingly popular technique for transcriptomic analysis of gene expression at the single-cell level. Cell-type clustering is the first crucial task in the analysis of scRNA-seq data that facilitates accurate identification of cell types and the study of the characteristics of their transcripts. Recently, several computational models based on a deep autoencoder and the ensemble clustering have been developed to analyze scRNA-seq data. However, current deep autoencoders are not sufficient to learn the latent representations of scRNA-seq data, and obtaining consensus partitions from these feature representations remains under-explored. RESULTS: To address this challenge, we propose a single-cell deep clustering model via a dual denoising autoencoder with bipartite graph ensemble clustering called scBGEDA, to identify specific cell populations in single-cell transcriptome profiles. First, a single-cell dual denoising autoencoder network is proposed to project the data into a compressed low-dimensional space and that can learn feature representation via explicit modeling of synergistic optimization of the zero-inflated negative binomial reconstruction loss and denoising reconstruction loss. Then, a bipartite graph ensemble clustering algorithm is designed to exploit the relationships between cells and the learned latent embedded space by means of a graph-based consensus function. Multiple comparison experiments were conducted on 20 scRNA-seq datasets from different sequencing platforms using a variety of clustering metrics. The experimental results indicated that scBGEDA outperforms other state-of-the-art methods on these datasets, and also demonstrated its scalability to large-scale scRNA-seq datasets. Moreover, scBGEDA was able to identify cell-type specific marker genes and provide functional genomic analysis by quantifying the influence of genes on cell clusters, bringing new insights into identifying cell types and characterizing the scRNA-seq data from different perspectives. AVAILABILITY AND IMPLEMENTATION: The source code of scBGEDA is available at https://github.com/wangyh082/scBGEDA. The software and the supporting data can be downloaded from https://figshare.com/articles/software/scBGEDA/19657911. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yunhe Wang 0002, Zhuohan Yu, Shaochuan Li, Chuang Bian, Ka-Chun Wong, Xiangtao Li |
Bioinform. | 1 |
| 2023 | Convex combination multiple populations competitive swarm optimization for moving target search using UAVs
Tianxi Ma, Yunhe Wang 0002, Xiangtao Li |
Inf. Sci. | 2 |
| 2022 | ZINB-Based Graph Embedding Autoencoder for Single-Cell RNA-Seq InterpretationsabstractSingle-cell RNA sequencing (scRNA-seq) provides high-throughput information about the genome-wide gene expression levels at the single-cell resolution, bringing a precise understanding on the transcriptome of individual cells. Unfortunately, the rapidly growing scRNA-seq data and the prevalence of dropout events pose substantial challenges for cell type annotation. Here, we propose a single-cell model-based deep graph embedding clustering (scTAG) method, which simultaneously learns cell–cell topology representations and identifies cell clusters based on deep graph convolutional network. scTAG integrates the zero-inflated negative binomial (ZINB) model into a topology adaptive graph convolutional autoencoder to learn the low-dimensional latent representation and adopts Kullback–Leibler (KL) divergence for the clustering tasks. By simultaneously optimizing the clustering loss, ZINB loss, and the cell graph reconstruction loss, scTAG jointly optimizes cluster label assignment and feature learning with the topological structures preserved in an end-to-end manner. Extensive experiments on 16 single-cell RNA-seq datasets from diverse yet representative single-cell sequencing platforms demonstrate the superiority of scTAG over various state-of-the-art clustering methods. Zhuohan Yu, Yifu Lu, Yunhe Wang 0002, Fan Tang, Ka-Chun Wong, Xiangtao Li |
AAAI | 3 |
| 2022 | Exploring high-throughput biomolecular data with multiobjective robust continuous clustering
Yunhe Wang 0002, Ka-Chun Wong, Xiangtao Li |
Inf. Sci. | 1 |
| 2022 | A self-adaptive weighted differential evolution approach for large-scale feature selection
Xubin Wang 0001, Yunhe Wang 0002, Ka-Chun Wong, Xiangtao Li |
Knowl. Based Syst. | 2 |
| 2022 | UCAV Path Planning for Avoiding Obstacles using Cooperative Co-evolution Spider Monkey Optimization
Yunhe Wang 0002, Xiangtao Li |
Knowl. Based Syst. | 2 |
| 2022 | Evolutionary Multiobjective Clustering Algorithms With Ensemble for Patient StratificationabstractPatient stratification has been studied widely to tackle subtype diagnosis problems for effective treatment. Due to the dimensionality curse and poor interpretability of data, there is always a long-lasting challenge in constructing a stratification model with high diagnostic ability and good generalization. To address these problems, this article proposes two novel evolutionary multiobjective clustering algorithms with ensemble (NSGA-II-ECFE and MOEA/D-ECFE) with four cluster validity indices used as the objective functions. First, an effective ensemble construction method is developed to enrich the ensemble diversity. After that, an ensemble clustering fitness evaluation (ECFE) method is proposed to evaluate the ensembles by measuring the consensus clustering under those four objective functions. To generate the consensus clustering, ECFE exploits the hybrid co-association matrix from the ensembles and then dynamically selects the suitable clustering algorithm on that matrix. Multiple experiments have been conducted to demonstrate the effectiveness of the proposed algorithm in comparison with seven clustering algorithms, twelve ensemble clustering approaches, and two multiobjective clustering algorithms on 55 synthetic datasets and 35 real patient stratification datasets. The experimental results demonstrate the competitive edges of the proposed algorithms over those compared methods. Furthermore, the proposed algorithm is applied to extend its advantages by identifying cancer subtypes from five cancer-related single-cell RNA-seq datasets. Yunhe Wang 0002, Xiangtao Li, Ka-Chun Wong, Yi Chang 0001, Shengxiang Yang |
IEEE Trans. Cybern. | 1 |
| 2022 | Multiobjective Deep Clustering and its Applications in Single-cell RNA-seq DataabstractSingle-cell RNA sequencing is a transformative technology that enables us to study the heterogeneity of the tissue at the cellular level. Clustering is used as the key computational approach to group cells under the transcriptome profiles from single-cell RNA-seq data. However, accurate identification of distinct cell types is facing the challenge of high dimensionality, and it could cause uninformative clusters when clustering is directly applied on the original transcriptome. To address such challenge, an evolutionary multiobjective deep clustering (EMDC) algorithm is proposed to identify single-cell RNA-seq data in this study. First, EMDC removes redundant and irrelevant genes by applying the differential gene expression analysis to identify differentially expressed genes across biological conditions. After that, a deep autoencoder is proposed to project the high-dimensional data into different low-dimensional nonlinear embedding subspaces under different bottleneck layers. Then, the basic clustering algorithm is applied in those nonlinear embedding subspaces to generate some basic clustering results to produce the cluster ensemble. To lessen the unnecessary cost produced by those clusterings in the ensemble, the multiobjective evolutionary optimization is designed to prune the basic clustering results in the ensemble, unleashing its cell type discovery performance under three objective functions. Multiple experiments have been conducted on 30 synthetic single-cell RNA-seq datasets and six real single-cell RNA-seq datasets, which reveal that EMDC outperforms eight other clustering methods and three multiobjective optimization algorithms in cell type identification. In addition, we have conducted extensive comparisons to effectively demonstrate the impact of each component in our proposed EMDC. Yunhe Wang 0002, Chuang Bian, Ka-Chun Wong, Xiangtao Li, Shengxiang Yang |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2021 | Identification of pan-cancer Ras pathway activation with deep learningabstractThe identification of hidden responders is often an essential challenge in precision oncology. A recent attempt based on machine learning has been proposed for classifying aberrant pathway activity from multiomic cancer data. However, we note several critical limitations there, such as high-dimensionality, data sparsity and model performance. Given the central importance and broad impact of precision oncology, we propose nature-inspired deep Ras activation pan-cancer (NatDRAP), a deep neural network (DNN) model, to address those restrictions for the identification of hidden responders. In this study, we develop the nature-inspired deep learning model that integrates bulk RNA sequencing, copy number and mutation data from PanCanAltas to detect pan-cancer Ras pathway activation. In NatDRAP, we propose to synergize the nature-inspired artificial bee colony algorithm with different gradient-based optimizers in one framework for optimizing DNNs in a collaborative manner. Multiple experiments were conducted on 33 different cancer types across PanCanAtlas. The experimental results demonstrate that the proposed NatDRAP can provide superior performance over other benchmark methods with strong robustness towards diagnosing RAS aberrant pathway activity across different cancer types. In addition, gene ontology enrichment and pathological analysis are conducted to reveal novel insights into the RAS aberrant pathway activity identification and characterization. NatDRAP is written in Python and available at https://github.com/lixt314/NatDRAP1. Xiangtao Li, Shaochuan Li, Yunhe Wang 0002, Shixiong Zhang 0002, Ka-Chun Wong |
Briefings Bioinform. | 3 |
| 2021 | Identification of haploinsufficient genes from epigenomic data using deep forestabstractHaploinsufficiency, wherein a single allele is not enough to maintain normal functions, can lead to many diseases including cancers and neurodevelopmental disorders. Recently, computational methods for identifying haploinsufficiency have been developed. However, most of those computational methods suffer from study bias, experimental noise and instability, resulting in unsatisfactory identification of haploinsufficient genes. To address those challenges, we propose a deep forest model, called HaForest, to identify haploinsufficient genes. The multiscale scanning is proposed to extract local contextual representations from input features under Linear Discriminant Analysis. After that, the cascade forest structure is applied to obtain the concatenated features directly by integrating decision-tree-based forests. Meanwhile, to exploit the complex dependency structure among haploinsufficient genes, the LightGBM library is embedded into HaForest to reveal the highly expressive features. To validate the effectiveness of our method, we compared it to several computational methods and four deep learning algorithms on five epigenomic data sets. The results reveal that HaForest achieves superior performance over the other algorithms, demonstrating its unique and complementary performance in identifying haploinsufficient genes. The standalone tool is available at https://github.com/yangyn533/HaForest. Shaochuan Li, Yunhe Wang 0002, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
Briefings Bioinform. | 3 |
| 2021 | Evolving Multiobjective Cancer Subtype Diagnosis From Cancer Gene Expression DataabstractDetection and diagnosis of cancer are especially essential for early prevention and effective treatments. Many studies have been proposed to tackle the subtype diagnosis problems with those data, which often suffer from low diagnostic ability and bad generalization. This article studies a multiobjective PSO-based hybrid algorithm (MOPSOHA) to optimize four objectives including the number of features, the accuracy, and two entropy-based measures: the relevance and the redundancy simultaneously, diagnosing the cancer data with high classification power and robustness. First, we propose a novel binary encoding strategy to choose informative gene subsets to optimize those objective functions. Second, a mutation operator is designed to enhance the exploration capability of the swarm. Finally, a local search method based on the "best/1" mutation operator of differential evolutionary algorithm (DE) is employed to exploit the neighborhood area with sparse high-quality solutions since the base vector always approaches to some good promising areas. In order to demonstrate the effectiveness of MOPSOHA, it is tested on 41 cancer datasets including thirty-five cancer gene expression datasets and six independent disease datasets. Compared MOPSOHA with other state-of-the-art algorithms, the performance of MOPSOHA is superior to other algorithms in most of the benchmark datasets. Yunhe Wang 0002, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | Nature-inspired multiobjective patient stratification from cancer gene expression data
Yunhe Wang 0002, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
Inf. Sci. | 1 |
| 2020 | Cancer molecular subtype classification from hypervolume-based discrete evolutionary optimization
Yunhe Wang 0002, Shaochuan Li, Zhiqiang Ma 0003, Xiangtao Li |
Neural Comput. Appl. | 1 |
| 2016 | BAMOKNN: A novel computational method for predicting the apoptosis protein locationsabstractIn this paper, we propose a novel hybrid binary animal migration optimization (BAMO) with k-nearest neighbor approach (KNN) to predict apoptosis protein sequences using statistical factors and dipeptide composition. Binary animal migration optimization is used for selecting a near-optimal subset of informative features that is most relevant for the classification. K-nearest neighbor approach is used as the classifier with the jackknife cross-validation. Finally, BAMOKNN is tested on a dataset including 317 proteins. Our method achieves the accuracy of 92.43%. Then, our model also tests on a testing dataset including 98 apoptosis proteins and obtains the accuracy of 94.90%. High prediction accuracy and successful prediction of apoptosis proteins suggest that BAMOKNN can be a useful approach to identify apoptosis protein locations. Xiangtao Li, Shijing Ma, Yunhe Wang 0002 |
BIBM | 3 |