EDBT 2026 Demo / reviewers in the wild / expert
Yiming Wang 0007
dblp:71/3182-7
· DBLP profile ↗
28ranked-venue papers
9as first author
28since 2021 · last 2026
0000-0002-8765-7640ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Local geometry-enhanced anchor learning for multi-view clustering
Zisen Kong, Zhiqiang Fu, Dongxia Chang, Yiming Wang 0007, Pengyuan Li 0013, Yao Zhao 0001 |
Neurocomputing | 4 |
| 2026 | Anchor-based disentanglement framework for incremental multi-view clustering
Pengyuan Li 0013, Zisen Kong, Dongxia Chang, Yiming Wang 0007, Yao Zhao 0001 |
Neural Networks | 5 |
| 2026 | Tensorial Multi-View Clustering via Alternative Rank Minimization and Inter-View AlignmentabstractTensor-based multi-view clustering is a popular approach. It can enhance representation learning by exploring higher-order correlations among views. However, two key issues remain unsolved. First, minimizing the tensor rank is a complex multi-objective optimization problem, so finding a suitable optimization strategy is an open problem. Moreover, most tensor methods require two phases to obtain the consensus matrix, which usually leads to suboptimal performance. To address these issues, we propose a Tensorial Multi-view Clustering via Alternative Rank Minimization and Inter-view Alignment (ARIA), in which multiple low-rank matrices and the consistent matrix are jointly optimized in a unified framework. Specifically, we stack the representations obtained from different views into a higher-order tensor. Then, a non-convex alternative rank-minimizing regularization is introduced to achieve a tighter approximation of the rank function. Besides, we impose intra-view alignment constraints to establish a connection between inter-view and intra-view. Unlike the previous method, it is a one-step strategy to obtain the consensus representation. Notably, our approach requires only linear complexity, and thus it can be successfully applied in large-scale clustering tasks. Extensive experiments validate the effectiveness and scalability of the proposed method. The code for ARIA is publicly available athttps://github.com/zskong/ARIA. Zisen Kong, Dongxia Chang, Yiming Wang 0007, Pengyuan Li 0013, Yao Zhao 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2026 | Disentangled Contrastive Multi-View Clustering via Semantic Relevance Invariance
Pengyuan Li 0013, Dongxia Chang, Yiming Wang 0007, Zisen Kong, Linhua Kong, Yao Zhao 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2026 | Deep Multi-View Clustering With Intra-View Similarity and Cross-View Correlation LearningabstractDeep multi-view clustering (MVC) has gained widespread attention as it can effectively mine consistent information from multiple views and improve clustering performance. However, view bias often exists between views (i.e., the quality differences between views). Treating all views equally inevitably destroys structural information when simply concatenating or summing the embedded representation of multiple views. To alleviate this issue, we propose a deep multi-view clustering with intra-view similarity and cross-view correlation learning (MISCC), facilitating the intra-view discriminability and inter-view complementarity. Specifically, we utilize the intra-view inherent structure information to dynamically identify semantically similar samples within each view. By aggregating their embedding representations, fine-grained structures are enhanced to boost intra-cluster compactness and inter-cluster separation. Then, we construct a cross-view correlation learning module to align semantically related views while preserving the distinctive features of irrelevant views. Based on them, a centralized clustering alignment strategy is proposed to align the similarity distribution and clustering structure between each view and the unified view, balancing the diverse information among multiple views. By jointly training these modules, the unified representation is optimized to capture more discriminative information from multiple views. Extensive experiments conducted on eleven multi-view datasets demonstrate that MISCC outperforms the state-of-the-art clustering methods. Pengyuan Li 0013, Dongxia Chang, Yiming Wang 0007, Man Liu 0003, Zisen Kong, Linhua Kong, Yao Zhao 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | MCLiD: Multi-Target and Container-Independent Liquid Sensing via mmWave and Camera FusionabstractLiquid sensing is critical for food safety and public security. Although mmWave-based approaches enable non-invasive and high-accuracy sensing, they are typically limited to single-target and fixed-container scenarios, restricting their applicability in real-world scenarios. In this paper, we present MCLiD, a multi-modal liquid sensing framework that fuses mmWave radar and camera data to achieve simultaneous multi-target and container-independent liquid identification. The basic idea is to leverage camera-captured object positions and container information to guide mmWave data processing, generating robust and discriminative liquid-specific representations for identification. MCLiD addresses a series of practical challenges and integrates three specialized modules for image-mmWave signature construction, liquid-specific feature extraction, and identification. Experimental results show that MCLiD achieves an average accuracy of 96.46 % across all combinations of 10 liquids and 7 container types. In multi-target scenarios, it maintains 96.4% accuracy for two concurrent liquids and 94.02% for five. These results indicate that MCLiD could enable rapid, non-invasive liquid detection for food safety and high-throughput public security applications. Jiawen Gai, Cheng Peng 0019, Zhekai Xu, Kaiyan Cui, Yiming Wang 0007, Fu Xiao 0001 |
ICPADS | 5 |
| 2025 | AEMVC: Mitigate Imbalanced Embedding Space in Multi-view ClusteringabstractMulti-view clustering (MVC) has gained extensive attention for its capacity to handle heterogeneous data. However, current autoencoder-based MVC methods suffer from a limitation: embedding space exhibits severe imbalances in the efficacy of feature direction, creating a long-tailed singular value distribution where few directions dominate. To mitigate this, we introduce a novel Activate-Then-Eliminate Strategy for Multi-View Clustering (AEMVC), inspired by the observation that balanced feature directions can facilitate enhancing discrimination of learned representations. AEMVC dynamically adjusts the contributions of different feature directions through two keys: a Feature Activation Module that narrows singular value discrepancies to prevent dominant directions from controlling clustering decisions, and an Inter-view Mutual Supervision strategy that filters redundant information by adaptively determining view-specific thresholds based on cross-view consistency. By activating more feature directions and eliminating each view's adverse factors, AEMVC achieves more balanced and discriminative embedding representations. Extensive experiments on seven multi-view benchmarks validate AEMVC's effectiveness, demonstrating substantial improvements over state-of-the-art methods. Pengyuan Li 0013, Man Liu 0003, Dongxia Chang, Yiming Wang 0007, Zisen Kong, Yao Zhao 0001 |
ACM Multimedia | 4 |
| 2025 | Learning from Disjoint Views: A Contrastive Prototype Matching Network for Fully Incomplete Multi-View ClusteringabstractMulti-view clustering aims to enhance clustering performance by leveraging information from diverse sources. However, its practical application is often hindered by a barrier: the lack of correspondences across views. This paper focuses on the understudied problem of fully incomplete multi-view clustering (FIMC), a scenario where existing methods fail due to their reliance on partial alignment. To address this problem, we introduce the Contrastive Prototype Matching Network (CPMN), a novel framework that establishes a new paradigm for cross-view alignment based on matching high-level categorical structures. Instead of aligning individual instances, CPMN performs a more robust cluster prototype alignment. CPMN first employs a correspondence-free graph contrastive learning approach, leveraging mutual $k$-nearest neighbors (MNN) to uncover intrinsic data structures and establish initial prototypes from entirely unpaired views. Building on the prototypes, we introduce a cross-view prototype graph matching stage to resolve category misalignment and forge a unified clustering structure. Finally, guided by this alignment, we devise a prototype-aware contrastive learning mechanism to promote semantic consistency, replacing the reliance on the initial MNN-based structural similarity. Extensive experiments on benchmark datasets demonstrate that our method significantly outperforms various baselines and ablation variants, validating its effectiveness. Yiming Wang 0007, Qun Li 0002, Dongxia Chang, Jie Wen 0001, Hua Dai 0003, Fu Xiao 0001, Yao Zhao 0001 |
NeurIPS | 1 |
| 2025 | PS-CoT-Adapter: adapting plan-and-solve chain-of-thought for ScienceQA
Qun Li 0002, Fu Xiao 0001, Yiming Wang 0007, Xinping Gao, Bir Bhanu |
Sci. China Inf. Sci. | 4 |
| 2025 | DCMVC: Dual contrastive multi-view clustering
Pengyuan Li 0013, Dongxia Chang, Zisen Kong, Yiming Wang 0007, Yao Zhao 0001 |
Neurocomputing | 4 |
| 2025 | Dual-space Co-training for Large-scale Multi-view Clustering
Zisen Kong, Zhiqiang Fu, Dongxia Chang, Yiming Wang 0007, Yao Zhao 0001 |
Pattern Recognit. | 4 |
| 2025 | Reordered $k$-Means: A New Baseline for View-Unaligned Multi-View ClusteringabstractMost current multi-view clustering methods necessitate that a sample's features be view-aligned or at least partially aligned across different views. Regrettably, real-world applications often fail to meet this requirement due to spatial, temporal, or spatiotemporal mismatches, resulting in the view-unaligned issue. To tackle this issue, we conceptualize the view-unaligned problem and demonstrate that it can be transformed into a view-aligned problem through reordering. Building on this concept, we introduce an innovative reorder matrix that realigns view-unaligned features. Utilizing these realigned features, we develop a sophisticated and efficient approach called Reordered$k$-means (RKM), which merges NMF with$k$-means. Unlike traditional$k$-means, our method converts the binary challenge into an$\ell _{0}$problem, confirming the merit of this advancement. Furthermore, RKM's efficacy is affirmed on benchmarks, indicating substantial enhancements in handling the view-unaligned issue and maintaining competitive results with view-aligned problems. Zhiqiang Fu, Yao Zhao 0001, Dongxia Chang, Yiming Wang 0007, Jie Wen 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | A Category-Driven Contrastive Recovery Network for Double Incomplete Multi-View Multi-Label ClassificationabstractIn the field of multi-view multi-label learning, the challenges of incomplete views and missing labels are prevalent due to the complexity of manual labeling and data acquisition errors. These challenges significantly reduce the quality of latent representations and hinder prediction by multi-label classification. To address this issue, we propose a novel Category-driven Semi-supervised Contrastive Recovery (CSCR) framework in this study. Our framework aims to fully integrate existing label information into incomplete representation learning and classification. Specifically, to address the limitations posed by incomplete views and labels, we construct a label coincidence matrix based on existing labels, which serves as a similarity matrix in subsequent semi-supervised contrastive learning and multi-view classification. By leveraging this matrix, we design a semi-supervised multi-view contrastive learning module, which constructs sample pairs on the basis of inter-view correspondences and label similarity. It learns discriminative latent representations without the need for data augmentation. A weighted multi-label classification module is subsequently employed to integrate the predictions from each view to obtain the final classification result. Experimental evaluations on five challenging datasets demonstrate the superiority of our model over existing state-of-the-art methods. Yiming Wang 0007, Qun Li 0002, Dongxia Chang, Jie Wen 0001, Fu Xiao 0001, Yao Zhao 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Temporal-Enhanced Radar and Camera Fusion for Object DetectionabstractRecently, object detection methods based on multi-modal fusion have gained widespread adoption in autonomous driving, proving to be valuable for detecting objects in dynamic environments. Among them, millimeter wave (mmWave) radar is commonly utilized as an effective complement to cameras, as it is almost unaffected by harsh weather conditions. However, current approaches that fuse mmWave radar and camera often overlook the correlation between the two modalities, failing to fully exploit their complementary features. To address this, we propose a temporal-enhanced radar and camera fusion network to explore the correlation between these two modalities and learn a comprehensive representation for object detection. In our model, a temporal fusion model is introduced to fuse mmWave radar features from different moments, thus mitigating the problem of mmWave radar point-object mismatch due to object movement. Moreover, a new correlation-based fusion strategy using the dedicated mask cross-attention is proposed to fuse mmWave radar and vision features more effectively. Finally, we design a gate feature pyramid network that selects shallow texture information based on deep semantic information to obtain more representative features. The experimental results on the nuScenes benchmark demonstrate the effectiveness of our proposed method. Linhua Kong, Yiming Wang 0007, Dongxia Chang, Yao Zhao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Partially View-Aligned Representation Learning via Cross-View Graph Contrastive NetworkabstractMulti-view representation learning, aimed at uncovering the inherent structure within multi-view data, has developed rapidly in recent years. In practice, due to temporal and spatial desynchronization, it is common that only part of the data is aligned between views, which leads to thePartial View Alignment(PVA) problem. To address the challenge of representation learning on partially view-aligned multi-view data, we propose a new cross-view graph contrastive learning network, which integrates multi-view information to align data and learn latent representations. First, view-specific autoencoders are used to construct an end-to-end multi-view representation learning framework for learning specific view representations. Furthermore, to achieve cluster-level alignment, we introduce a cross-view graph contrastive learning module to guide the learning of discriminative representations. Compared to the existing methods, the proposed cluster-level alignment method successfully extends the view alignment to more than two views. Meanwhile, the results of clustering and classification experiments on several popular multi-view datasets can also illustrate the effectiveness and superiority of the proposed method. Yiming Wang 0007, Dongxia Chang, Zhiqiang Fu, Jie Wen 0001, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Seeing All From a Few: Nodes Selection Using Graph Pooling for Graph ClusteringabstractRecently, there has been considerable research interest in graph clustering aimed at data partition using graph information. However, one limitation of most graph-based methods is that they assume that the graph structure to operate is reliable. However, there are inevitably some edges in the graph that are not conducive to graph clustering, which we call spurious edges. This brief is the first attempt to employ the graph pooling technique for node clustering to the best of our knowledge. In this brief, we propose a novel dual graph embedding network (DGEN), which is designed as a two-step graph encoder connected by a graph pooling layer to learn the graph embedding. In DGEN, we assume that if a node and its nearest neighboring node are close to the same clustering center, this node is informative, and this edge can be considered as a cluster-friendly edge. Based on this assumption, the neighbor cluster pooling (NCPool) is devised to select the most informative subset of nodes and the corresponding edges based on the distance of nodes and their nearest neighbors to the cluster centers. This can effectively alleviate the impact of the spurious edges on the clustering. Finally, to obtain the clustering assignment of all nodes, a classifier is trained using the clustering results of the selected nodes. Experiments on five benchmark graph datasets demonstrate the superiority of the proposed method over state-of-the-art algorithms. Yiming Wang 0007, Dongxia Chang, Zhiqiang Fu, Yao Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Projection-preserving block-diagonal low-rank representation for subspace clustering
Zisen Kong, Dongxia Chang, Zhiqiang Fu, Jiapeng Wang 0002, Yiming Wang 0007, Yao Zhao 0001 |
Neurocomputing | 5 |
| 2023 | Learning a bi-directional discriminative representation for deep clustering
Yiming Wang 0007, Dongxia Chang, Zhiqiang Fu, Yao Zhao 0001 |
Pattern Recognit. | 1 |
| 2023 | Incomplete Multiview Clustering via Cross-View Relation TransferabstractIn this paper, we consider the problem of multi-view clustering on incomplete views. Compared with complete multi-view clustering, the view-missing problem increases the difficulty of learning common representations from different views. To address the challenge, we propose a novel incomplete multi-view clustering framework, which incorporates cross-view relation transfer and multi-view fusion learning. Specifically, based on the consistency existing in multi-view data, we devise a cross-view relation transfer-based completion module, which transfers known similar inter-instance relationships to the missing view and infers the missing data via graph networks based on the transferred relationship graph. Then the view-specific encoders are designed to extract the recovered multi-view data, and an attention-based fusion layer is introduced to obtain the common representation. Moreover, to reduce the impact of the error caused by the inconsistency between views and obtain a better clustering structure, a joint clustering layer is introduced to optimize recovery and clustering simultaneously. Extensive experiments conducted on several real datasets demonstrate the effectiveness of the proposed method. Yiming Wang 0007, Dongxia Chang, Zhiqiang Fu, Jie Wen 0001, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Latent Low-Rank Representation With Weighted Distance Penalty for ClusteringabstractLatent low-rank representation (LatLRR) is a critical self-representation technique that improves low-rank representation (LRR) by using observed and unobserved samples. It can simultaneously learn the low-dimensional structure embedded in the data space and capture the salient features. However, LatLRR ignores the local geometry structure and can be affected by the noise and redundancy in the original data space. To solve the above problems, we propose a latent LRR with weighted distance penalty (LLRRWD) for clustering in this article. First, a weighted distance is proposed to enhance the original Euclidean distance by enlarging the distance among the unconnected samples, which can enhance the discriminitation of the distance among the samples. By leveraging on the weighted distance, a weighted distance penalty is introduced to the LatLRR model to enable the method to preserve both the local geometric information and global information, improving discrimination of the learned affinity matrix. Moreover, a weight matrix is imposed on the sparse error norm to reduce the effect of noise and redundancy. Experimental results based on several benchmark databases show the effectiveness of our method in clustering. Zhiqiang Fu, Yao Zhao 0001, Dongxia Chang, Yiming Wang 0007, Jie Wen 0001 |
IEEE Trans. Cybern. | 4 |
| 2023 | Graph Contrastive Partial Multi-View ClusteringabstractWith the diversity of information acquisition, data is stored and transmitted in an increasing number of modalities. Nevertheless, it is not unusual for parts of the data to be lost in some views due to unavoidable acquisition, transmission or storage errors. In this paper, we propose an augmentation-free graph contrastive learning framework to solve the problem of partial multi-view clustering. Notably, we suppose that the representations of similar samples (i.e., belonging to the same cluster) should be similar. This is distinct from the general unsupervised contrastive learning that assumes an image and its augmentations share a similar representation. Specifically, relation graphs are constructed using the nearest neighbors to identify existing similar samples, then the constructed inter-instance relation graphs are transferred to the missing views to build graphs on the corresponding missing data. Subsequently, two main components, within-view graph contrastive learning and cross-view graph consistency learning, are devised to maximize the mutual information of different views within a cluster. The proposed approach elevates instance-level contrastive learning and missing data inference to the cluster-level, effectively mitigating the impact of individual missing data on clustering. Experiments on several challenging datasets demonstrate the superiority of our proposed methods. Yiming Wang 0007, Dongxia Chang, Zhiqiang Fu, Jie Wen 0001, Yao Zhao 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Consistent Multiple Graph Embedding for Multi-View ClusteringabstractGraph-based multi-view clustering aiming to obtain a partition of data across multiple views, has received considerable attention in recent years. Although great efforts have been made for graph-based multi-view clustering, it is still challenging to fuse characteristics from various views to learn a common representation for clustering. In this paper, we propose a novel Consistent Multiple Graph Embedding Clustering framework (CMGEC). Specifically, a multiple graph auto-encoder (M-GAE) is designed to flexibly encode the complementary information of multi-view data using a multi-graph attention fusion encoder. To guide the learned common representation maintaining the similarity of the neighboring characteristics in each view, a Multi-view Mutual Information Maximization module (MMIM) is introduced. Furthermore, a graph fusion network (GFN) is devised to explore the relationship among graphs from different views and provide a common consensus graph needed in M-GAE. By jointly training these models, the common representation can be obtained, which encodes more complementary information from multiple views and depicts data more comprehensively. Experiments on three types of multi-view datasets demonstrate CMGEC outperforms the state-of-the-art clustering methods. Yiming Wang 0007, Dongxia Chang, Zhiqiang Fu, Yao Zhao 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | One-step Low-Rank Representation for ClusteringabstractExisting low-rank representation-based methods adopt a two-step framework, which must employ an extra clustering method to gain labels after representation learning. In this paper, a novel one-step representation-based method, i.e., One-step Low-Rank Representation (OLRR), is proposed to capture multi-subspace structures for clustering. OLRR integrates the low-rank representation model and clustering into a unified framework. Thus it can jointly learn the low-rank subspace structure embedded in the database and gain the clustering results. In particular, by approximating the representation matrix with two same clustering indicator matrices, OLRR can directly show the probability of samples belonging to each cluster. Further, a probability penalty is introduced to ensure that the samples with smaller distances are more inclined to be in the same cluster, thus enhancing the discrimination of the clustering indicator matrix and resulting in a more favorable clustering performance. Moreover, to enhance the robustness against noise, OLRR uses the probability to guide denoising and then performs representation learning and clustering in a recovered clean space. Extensive experiments well demonstrate the robustness and effectiveness of OLRR. Our code is publicly available at: https://github.com/fuzhiqiang1230/OLRR. Zhiqiang Fu, Yao Zhao 0001, Dongxia Chang, Yiming Wang 0007, Jie Wen 0001, Xingxing Zhang 0001, Guodong Guo |
ACM Multimedia | 4 |
| 2022 | Contour enhanced image super-resolution
Linhua Kong, Yiming Wang 0007, Dongxia Chang, Yao Zhao 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | Auto-weighted low-rank representation for clustering
Zhiqiang Fu, Yao Zhao 0001, Dongxia Chang, Xingxing Zhang 0001, Yiming Wang 0007 |
Knowl. Based Syst. | 5 |
| 2021 | Double Low-Rank Representation With Projection Distance Penalty for ClusteringabstractThis paper presents a novel, simple yet robust self-representation method, i.e., Double Low-Rank Representation with Projection Distance penalty (DLRRPD) for clustering. With the learned optimal projected representations, DLRRPD is capable of obtaining an effective similarity graph to capture the multi-subspace structure. Besides the global low-rank constraint, the local geometrical structure is additionally exploited via a projection distance penalty in our DLRRPD, thus facilitating a more favorable graph. Moreover, to improve the robustness of DLRRPD to noises, we introduce a Laplacian rank constraint, which can further encourage the learned graph to be more discriminative for clustering tasks. Meanwhile, Frobenius norm (instead of the popularly used nuclear norm) is employed to enforce the graph to be more block-diagonal with lower complexity. Extensive experiments have been conducted on synthetic, real, and noisy data to show that the proposed method outperforms currently available alternatives by a margin of 1.0%~10.1%. Zhiqiang Fu, Yao Zhao 0001, Dongxia Chang, Xingxing Zhang 0001, Yiming Wang 0007 |
CVPR | 5 |
| 2021 | A new blind image denoising method based on asymmetric generative adversarial networkabstractAbstract Image denoising is a classical topic in computer vision. In recent years, with the development of deep learning, image denoising methods based on discriminative learning have received more attention. In this paper, a new blind image denoising method based on the asymmetric generative adversarial network (ID‐AGAN) is proposed. In the new method, the adversarial learning is used to optimise the high‐dimensional image information denoising, so as to balance the noise removal and detail retention. In order to overcome the unstability of the GAN training and improve the discriminative ability of the discriminating model, an image downsampling layer is added between the generating model and the discriminating model. Moreover, a multi‐scale feature downsampling layer is utilised to extract the feature of the entire image and reducing the effect of noise on training images. Extensive experiments are conducted to verify the performance of the ID‐AGAN algorithm. The results demonstrate that authors' method has high performance and flexibility. Yiming Wang 0007, Dongxia Chang, Yao Zhao 0001 |
IET Image Process. | 1 |
| 2021 | A hierarchical weighted low-rank representation for image clustering and classification
Zhiqiang Fu, Yao Zhao 0001, Dongxia Chang, Yiming Wang 0007 |
Pattern Recognit. | 4 |