Daoyuan Wang

dblp:210/4831 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Geometry-Aware Variational Information Maximization for Deep Incomplete Multi-view Clustering
abstract
Incomplete multi-view clustering (IMVC) aims to group data into meaningful clusters when each sample is only partially observed across multiple views. Most existing methods either rely on imputation strategies that may introduce noise and distort the underlying data distribution, or adopt cross-view alignment techniques that focus on pairwise relationships, often resulting in suboptimal representations and unstable clustering performance. In this paper, we propose Geometry-Aware Variational Information Maximization for Deep Incomplete Multi-view Clustering (GAVIM), a novel imputation-free variational framework that enables robust and coherent incomplete multi-view clustering. Specifically, GAVIM leverages mutual information maximization to preserve the high mutual information between the available multi-view data and the shared embedding. Moreover, we explicitly retain local geometric consistency within each view-specific latent space under the guidance of an adaptive global supervision signal. Lastly, GAVIM aligns all views simultaneously using a Gramian representation alignment measure, ensuring coherent structure across modalities and promoting unified, semantically meaningful representations. Extensive experiments on five benchmark IMVC datasets with varying levels of view incompleteness demonstrate that GAVIM consistently outperforms state-of-the-art methods in clustering accuracy and representation quality.
Wenlan Chen, Daoyuan Wang, Fei Guo 0001, Cheng Liang 0001
AAAI3
2026 Time-Surface Self-Attention: Restoring Temporal Connectivity in Spiking Transformers
abstract
The recent integration of Spiking Neural Networks (SNNs) with self-attention mechanisms has yielded promising results. However, existing Spiking Transformers often overlook the intrinsic temporal continuity of spike trains, typically treating the time dimension merely as a batch extension and computing attention based solely on instantaneous spikes. This frame-by-frame processing paradigm fails to exploit the historical context accumulated in neuronal membrane potentials, resulting in attention maps that fluctuate excessively. To address these challenges, we propose a novel Time-Surface Self-Attention (TSSA) mechanism that restores temporal connectivity within self-attention for spiking neural networks. Instead of computing attention on instantaneous discrete spikes, TSSA projects the binary Query and Key spike streams into a continuous temporal manifold via a memory-decayer module, effectively measuring temporal synchronization. Crucially, the Value branch remains in the sparse spike domain to ensure energy-efficient accumulation. We further design an adaptive multi-scale decay strategy, assigning distinct time constants to different attention heads to enforce a structural prior of diverse temporal receptive fields. Extensive experiments on both static datasets and neuromorphic benchmarks demonstrate that our approach achieves state-of-the-art accuracy and superior robustness against noise.
Tiantian Xiao, Hongbin Lv, Daoyuan Wang
ICMR4
2026 High-order correlation and consistency-aware multi-view clustering via anchor graph learning
Cheng Liang 0001, Wenchao Zang, Daoyuan Wang, Fei Guo 0001
Neural Networks3
2026 Incomplete Multi-View Clustering via Robust Representation Learning and Tensor-Based Co-Regularization
abstract
As incompleteness is common in real-world data, incomplete multi-view clustering is of great significance in the unsupervised learning field because it allows the partitioning of multi-view data with missing information into distinct groups. In this paper, we propose a novel generalized framework for incomplete multi-view clustering based on robust representation learning and tensor-based co-regularization (RRLTCR). Specifically, a robust principal component analysis is first used to learn a robust representation for each view. To explore high-order relationships among views, the view-specific spectral embeddings are stacked into a third-order tensor with a Schattenp-norm constraint. By spreading the complementary information of the high-quality available data from each view on a global scale, our model is able to alleviate the adverse effects of data noise and uncover the underlying common cluster structure. An effective iterative optimization strategy is developed to efficiently solve our model. According to the experimental results on seven datasets, our proposed framework has the potential to improve the clustering performance for a variety of incomplete multi-view clustering problems. Our research work brings a generalized framework for incomplete multi-view clustering, which can also assist in exploring the large cohort of existing incomplete multimodality datasets for other downstream tasks.
Cheng Liang 0001, Daoyuan Wang, Fei Guo 0001, Shichao Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Prototype-Calibrated Multimodal Fusion with Relational Consistency for Spatial Domain Identification in Spatial Transcriptomics
abstract
Spatial transcriptomics technologies produce rich multimodal data, including gene expression profiles, spatial coordinates, and histological images, offering unprecedented opportunities to characterize tissue organization. However, effectively integrating these heterogeneous modalities remains challenging due to the complex molecular and morphological interplay. To address this, we propose PCMF-RC, a novel Prototype-Calibrated Multimodal Fusion framework with Relational Consistency, designed for accurate spatial domain identification. Our approach first partitions high-resolution histopathological images into patches centered on spatial transcriptomics spots and extracts spatially aware visual features using a pre-trained visual state space model. We then obtain modality-specific latent embeddings where we separately encode gene expression and image features via graph convolutional networks. To improve semantic consistency, a cross-view prototype matching module that aligns cluster prototypes across modalities is introduced to mitigate prototype shift. We further design a relational consistency contrastive learning module to effectively enforce alignment of spatial relational structures, reducing distributional discrepancies and enhancing robustness. Extensive experiments on multiple spatial transcriptomics datasets demonstrate that PCMF-RC outperforms existing methods in spatial domain delineation, offering a robust and generalizable solution for multimodal spatial omics integration.
Daoyuan Wang, Guanghui Li 0003, Cheng Liang 0001
BIBM2
2025 Image-Enhanced Hybrid Encoding with Reinforced Contrastive Learning for Spatial Domain Identification in Spatial Transcriptomics
abstract
Spatial transcriptomics integrates spatial, gene expression, and multichannel immunohistochemistry image data, enabling advanced insights into cellular organization. However, existing methods often struggle to effectively fuse these multimodal data, limiting their potential for accurate spatial domain identification. Here, we propose IE-HERCL (Image-Enhanced Hybrid Encoding with Reinforced Contrastive Learning), a novel framework designed to address this challenge. Specifically, IE-HERCL employs hybrid encoding to capture both the non-spatial features and spatial dependencies for both gene and image modalities via autoencoders and GraphSAGE, respectively. These features are then fused using cross-view attention mechanisms to generate the unified informative embedding. To enhance the representation learning capability, we introduce a reinforced contrastive learning strategy to mitigate the influences of false negative samples, where we detect potential positive counterparts with high-order random walks. In addition, the cluster alignment is dynamically refined through optimal transport, which ensures that the fused consensus representation is coherent and robust, enabling accurate spatial domain identification. Our approach achieves state-of-the-art performance on five image-enhanced spatial transcriptomics datasets, demonstrating its robustness and effectiveness in multimodal integration and spatial domain identification. IE-HERCL offers a powerful and innovative solution for advancing spatial transcriptomics analysis. The code is released on https://github.com/wdyi701/IE-HERCL.
Daoyuan Wang, Wenlan Chen, Cheng Liang 0001, Fei Guo 0001
IJCAI1
2025 Disentangled Cross-Modal Representation Learning with Enhanced Mutual Supervision
abstract
Cross-modal representation learning aims to extract semantically aligned representations from heterogeneous modalities such as images and text. Existing multimodal VAE-based models often suffer from limited capability to align heterogeneous modalities or lack sufficient structural constraints to clearly separate the modality-specific and shared factors. In this work, we propose a novel framework, termed **D**isentangled **C**ross-**M**odal Representation Learning with **E**nhanced **M**utual Supervision (DCMEM). Specifically, our model disentangles the common and distinct information across modalities and regularizes the shared representation learned from each modality in a mutually supervised manner. Moreover, we incorporate the information bottleneck principle into our model to ensure that the shared and modality-specific factors encode exclusive yet complementary information. Notably, our model is designed to be trainable on both complete and partial multimodal datasets with a valid Evidence Lower Bound. Extensive experimental results demonstrate significant improvements of our model over existing methods on various tasks including cross-modal generation, clustering, and classification.
Wenlan Chen, Daoyuan Wang, Fei Guo 0001, Cheng Liang 0001
NeurIPS3
2025 Unsupervised multi-view feature selection based on weighted low-rank tensor learning and its application in multi-omics datasets
Daoyuan Wang, Lianzhi Wang, Wenlan Chen, Hong Wang 0015, Cheng Liang 0001
Eng. Appl. Artif. Intell.1
2025 Robust high-order graph learning for incomplete multi-view clustering
Daoyuan Wang, Fujian Ren, Yuntang Zhuang, Cheng Liang 0001
Expert Syst. Appl.1
2024 Sequence-level Semantic Representation Fusion for Recommender Systems
Lanling Xu, Zhen Tian 0001, Bingqian Li, Junjie Zhang 0009, Daoyuan Wang, Jinpeng Wang 0001, Wayne Xin Zhao
CIKM5
2024 Robust Tensor Subspace Learning for Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering has represented a significant role in grouping real images. In this study, a novel robust tensor subspace learning (RTSL) is proposed for incomplete multi-view clustering. Specifically, the missing samples within views are first recovered by matrix factorization. The recovered information is utilized for latent representations learning. And then, the obtained latent representations are organized from all views into a third-order tensor and the intrinsic sample relations are captured with tensor linear representation. Moreover, a low-rank sample coefficient tensor is sought to capture high-order connections among views by imposing the tensor nuclear norm. Compared with traditional learning paradigms in the vector space, the sample relations within each view as well as across views could be preserved with the aid of robust tensor subspace learning. As a result, our model can simultaneously handle the missing samples and exploit the intrinsic correlations, leading to enhanced representation capability and better quality of the recovered data. We design an efficient iterative optimization strategy to solve the proposed method. Experimental results on eight datasets show that our model outperforms other competing approaches.
Cheng Liang 0001, Daoyuan Wang, Huaxiang Zhang 0001, Shichao Zhang 0001, Fei Guo 0001
IEEE Trans. Knowl. Data Eng.2
2023 Reinforcement Learning Model for Managing Noninvasive Ventilation Switching Policy
abstract
Noninvasive ventilation (NIV) has been recognized as a first-line treatment for respiratory failure in patients with chronic obstructive pulmonary disease (COPD) and hypercapnia respiratory failure, which can reduce mortality and burden of intubation. However, during the long-term NIV process, failure to respond to NIV may cause overtreatment or delayed intubation, which is associated with increased mortality or costs. Optimal strategies for switching regime in the course of NIV treatment remain to be explored.For the goal of reducing 28-day mortality of the patients undergoing NIV, Double Dueling Deep Q Network (D3QN) of offline-reinforcement learning algorithm was adopted to develop an optimal regime model for making treatment decisions of discontinuing ventilation, continuing NIV, or intubation. The model was trained and tested using the data from Multi-Parameter Intelligent Monitoring in Intensive Care III (MIMIC-III) and evaluated by the practical strategies. Furthermore, the applicability of the model in majority disease subgroups (Catalogued by International Classification of Diseases, ICD) was investigated. Compared with physician's strategies, the proposed model achieved a higher expected return score (4.25 vs. 2.68) and its recommended treatments reduced the expected mortality from 27.82% to 25.44% in all NIV cases. In particular, for these patients finally received intubation in practice, if the model also supported the regime, it would warn of switching to intubation 13.36 hours earlier than clinicians (8.64 vs. 22 hours after the NIV treatment), granting a 21.7% reduction in estimated mortality. In addition, the model was applicable across various disease groups with distinguished achievement in dealing with respiratory disorders. The proposed model is promising to dynamically provide personalized optimal NIV switching regime for patients undergoing NIV with the potential of improving treatment outcomes.
Daoyuan Wang, Molei Yan, Yanfei Shen, Luping Fang, Guolong Cai, Gangmin Ning
IEEE J. Biomed. Health Informatics2