EDBT 2026 Demo / reviewers in the wild / expert
Xiaoli Wang 0003
dblp:31/6192-3
· DBLP profile ↗
13ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0001-9336-1013ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multimodal Graph Representation Learning with Dynamic Information PathwaysabstractMultimodal graphs, where nodes contain heterogeneous features such as images and text, are increasingly common in real-world applications. Effectively learning on such graphs requires both adaptive intra-modal message passing and efficient inter-modal aggregation. However, most existing approaches to multimodal graph learning are typically extended from conventional graph neural networks and rely on static structures or dense attention, which limit flexibility and expressive node embedding learning. In this paper, we propose a novel multimodal graph representation learning framework with Dynamic information Pathways (DiP). By introducing modality-specific pseudo nodes, DiP enables dynamic message routing within each modality via proximity-guided pseudo-node interactions and captures inter-modality dependence through efficient information pathways in a shared state space. This design achieves adaptive, expressive, and sparse message propagation across modalities with linear complexity. We conduct the link prediction and node classification tasks to evaluate performance and carry out full experimental analyses. Extensive experiments across multiple benchmarks demonstrate that DiP consistently outperforms baselines. Xiaobin Hong 0002, Mingkai Lin, Xiaoli Wang 0003, Chaoqun Wang 0012 |
AAAI | 3 |
| 2026 | Leveraging explicit priors for guided learning under data scarcity
Xiaoliang Zhou, Yongli Wang 0002, Anqi Huang 0001, Xiaoli Wang 0003 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Global-Semantic Alignment Distillation for Partial Multi-view ClassificationabstractPartial multi-view classification (PMvC) poses a significant challenge due to the incomplete nature of multi-view data, which complicates effective information fusion and accurate classification. Existing PMvC methods typically rely on heuristic evaluations of view informativeness to achieve global alignment for downstream classification tasks. However, these approaches suffer from two critical issues: information redundancy and semantic misalignment. The complexity of missing data not only leads to over-reliance on redundant or less informative views but also exacerbates semantic misalignment across views, making it difficult for existing methods to effectively capture and discriminate the class-related features. To address these issues, this work proposes a novel GLobal-semantic Alignment Distillation (GLAD) model for partial multi-view classification without requiring imputation. Our approach incorporates a self-distillation mechanism that enables the model to extract informative features and achieve global semantic alignment across views. The key insight of GLAD is leveraging labels as semantic anchors to guide the alignment of partial multi-view features. By integrating labels with extracted features via a cross-attention mechanism, we generate ideal embeddings that consistently capture global semantics across views. These embeddings then serve as intermediate supervision for distilling the student model, ensuring robust semantic alignment even with missing views. We further introduce a margin-aware weighting strategy to enhance the model's discriminative ability. Extensive experimental results validate the effectiveness and superiority of the proposed method, showcasing significant improvements in classification performance over existing techniques. Xiaoli Wang 0003, Anqi Huang 0001, Yongli Wang 0002, Guanzhou Ke, Xiaobin Hong 0002, Jun Liu 0036 |
AAAI | 1 |
| 2025 | Knowledge Bridger: Towards Training-Free Missing Modality CompletionabstractPrevious successful approaches to missing modality completion rely on carefully designed fusion techniques and extensive pre-training on complete data, which can limit their generalizability in out-of-domain (OOD) scenarios. In this study, we pose a new challenge: can we develop a missing modality completion model that is both resource-efficient and robust to OOD generalization? To address this, we present a training-free framework for missing modality completion that leverages large multimodal model (LMM). Our approach, termed the "Knowledge Bridger", is modality-agnostic and integrates generation and ranking of missing modalities. By defining domain-specific priors, our method automatically extracts structured information from available modalities to construct knowledge graphs. These extracted graphs connect the missing modality generation and ranking modules through the LMM, resulting in high-quality imputations of missing modalities. Experimental results across both general and medical domains show that our approach consistently outperforms competing methods, including in OOD generalization. Additionally, our knowledge-driven generation and ranking techniques demonstrate superiority over variants that directly employ LMMs for generation and ranking, offering insights that may be valuable for applications in other domains. Guanzhou Ke, Shengfeng He, Xiaoli Wang 0003, Bo Wang 0057, Guoqing Chao, Yuanyang Zhang, Hexing Su |
CVPR | 3 |
| 2025 | Graph-in-graph discriminative feature enhancement network for fine-grained visual classification
Yupeng Wang 0004, Can Xu 0006, Yongli Wang 0002, Xiaoli Wang 0003, Weiping Ding 0001 |
Appl. Intell. | 4 |
| 2025 | An innovative multi-view collaborative optimization framework for Weighted Naive Bayes
Xiaoliang Zhou, Yongli Wang 0002, Anqi Huang 0001, Xiaoli Wang 0003 |
Knowl. Based Syst. | 5 |
| 2024 | Rethinking Multi-View Representation Learning via Distilled DisentanglingabstractMulti-view representation learning aims to derive robust representations that are both view-consistent and view-specific from diverse data sources. This paper presents an in-depth analysis of existing approaches in this domain, highlighting a commonly overlooked aspect: the redundancy between view-consistent and view-specific representations. To this end, we propose an innovative framework for multi-view representation learning, which incorporates a technique we term ‘distilled disentangling’. Our method introduces the concept of masked cross-view prediction, enabling the extraction of compact, high-quality view-consistent representations from various sources without incurring extra computational overhead. Additionally, we develop a distilled disentangling module that efficiently filters out consistency-related information from multi-view representations, resulting in purer view-specific representations. This approach significantly reduces redundancy between view-consistent and view-specific representations, enhancing the overall efficiency of the learning process. Our empirical evaluations reveal that higher mask ratios substantially improve the quality of view-consistent representations. Moreover, we find that reducing the dimensionality of view-consistent representations relative to that of view-specific representations further refines the quality of the combined representations. Our code is accessible at: https://github.com/Guanzhou-Ke/MRDD. Guanzhou Ke, Bo Wang 0057, Xiaoli Wang 0003, Shengfeng He |
CVPR | 3 |
| 2024 | DVF:Multi-agent Q-learning with difference value factorization
Anqi Huang 0001, Yongli Wang 0002, Jianghui Sang, Xiaoli Wang 0003, Yupeng Wang 0004 |
Knowl. Based Syst. | 4 |
| 2024 | A Clustering-Guided Contrastive Fusion for Multi-View Representation LearningabstractMulti-view representation learning aims to extract comprehensive information from multiple sources. It has achieved significant success in applications such as video understanding and 3D rendering. However, how to improve the robustness and generalization of multi-view representations from unsupervised and incomplete scenarios remains an open question in this field. In this study, we discovered a positive correlation between the semantic distance of multi-view representations and the tolerance for data corruption. Moreover, we found that the information ratio of consistency and complementarity significantly impacts the performance of discriminative and generative tasks related to multi-view representations. Based on these observations, we propose an end-to-end CLustering-guided cOntrastiVE fusioN (CLOVEN) method, which enhances the robustness and generalization of multi-view representations simultaneously. To balance consistency and complementarity, we design an asymmetric contrastive fusion module. The module first combines all view-specific representations into a comprehensive representation through a scaling fusion layer. Then, the information of the comprehensive representation and view-specific representations is aligned via contrastive learning loss function, resulting in a view-common representation that includes both consistent and complementary information. We prevent the module from learning suboptimal solutions by not allowing information alignment between view-specific representations. We design a clustering-guided module that encourages the aggregation of semantically similar views. This action reduces the semantic distance of the view-common representation. We quantitatively and qualitatively evaluate CLOVEN on five datasets, demonstrating its superiority over 13 other competitive multi-view learning methods in terms of clustering and classification performance. In the data-corrupted scenario, our proposed method resists noise interference better than competitors. Additionally, the visualization demonstrates that CLOVEN succeeds in preserving the intrinsic structure of view-specific representations and improves the compactness of view-common representations. Our code can be found athttps://github.com/guanzhou-ke/cloven. Guanzhou Ke, Guoqing Chao, Xiaoli Wang 0003, Chenyang Xu 0007, Yongqi Zhu, Yang Yu 0058 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Trusted Semi-Supervised Multi-View Classification With Contrastive LearningabstractSemi-supervised multi-view learning is a remarkable but challenging task. Existing semi-supervised multi-view classification (SMVC) approaches mainly focus on performance improvement while ignoring decision reliability, which limits their deployment in safety-critical applications. Although several trusted multi-view classification methods are proposed recently, they rely on manual annotations. Therefore, this work emphasizes trusted multi-view classification learning under semi-supervised conditions. Different from existing SMVC methods, this work jointly models class probabilities and uncertainties based on evidential deep learning to formulate view-specific opinions. Moreover, unlike previous works that explore cross-view consistency in a single schema, this work proposes a multi-level consistency constraint. Specifically, we explore instance-level consistency on the view-specific representation space and category-level consistency on opinions from multiple views. Our proposed trusted graph-based contrastive loss nicely establishes the relationship between joint opinions and view-specific representations, which enables view-specific representations to enjoy a good manifold to improve classification performance. Overall, the proposed approach provides reliable and superior semi-supervised multiview classification decisions. Extensive experiments demonstrate the effectiveness, reliability and robustness of the proposed model. Xiaoli Wang 0003, Yongli Wang 0002, Yupeng Wang 0004, Anqi Huang 0001, Jun Liu 0036 |
IEEE Trans. Multim. | 1 |
| 2024 | Computer vision-driven forest wildfire and smoke recognition via IoT drone cameras
Yupeng Wang 0004, Yongli Wang 0002, Can Xu 0006, Xiaoli Wang 0003 |
Wirel. Networks | 4 |
| 2023 | Disentangling Multi-view Representations Beyond Inductive BiasabstractMulti-view (or -modality) representation learning aims to understand the relationships between different view representations. Existing methods disentangle multi-view representations into consistent and view-specific representations by introducing strong inductive biases, which can limit their generalization ability. In this paper, we propose a novel multi-view representation disentangling method that aims to go beyond inductive biases, ensuring both interpretability and generalizability of the resulting representations. Our method is based on the observation that discovering multi-view consistency in advance can determine the disentangling information boundary, leading to a decoupled learning objective. We also found that the consistency can be easily extracted by maximizing the transformation invariance and clustering consistency between views. These observations drive us to propose a two-stage framework. In the first stage, we obtain multi-view consistency by training a consistent encoder to produce semantically-consistent representations across views as well as their corresponding pseudo-labels. In the second stage, we disentangle specificity from comprehensive representations by minimizing the upper bound of mutual information between consistent and comprehensive representations. Finally, we reconstruct the original data by concatenating pseudo-labels and view-specific representations. Our experiments on four multi-view datasets demonstrate that our proposed method outperforms 12 comparison methods in terms of clustering and classification performance. The visualization results also show that the extracted consistency and specificity are compact and interpretable. Our code can be found at https://github.com/Guanzhou-Ke/DMRIB. Guanzhou Ke, Yang Yu 0058, Guoqing Chao, Xiaoli Wang 0003, Chenyang Xu 0007, Shengfeng He |
ACM Multimedia | 4 |
| 2022 | MMatch: Semi-Supervised Discriminative Representation Learning for Multi-View ClassificationabstractSemi-supervised multi-view learning has been an important research topic due to its capability to exploit complementary information from unlabeled multi-view data. This work proposes MMatch, a new semi-supervised discriminative representation learning method for multi-view classification. Unlike existing multi-view representation learning methods that seldom consider the negative impact caused by particular views with unclear classification structures (weak discriminative views). MMatch jointly learns view-specific representations and class probabilities of training data. The representations concatenated to integrate multiple views’ information to form a global representation. Moreover, MMatch performs the smoothness constraint on the class probabilities of the global representation to improve pseudo labels, whereas the pseudo labels regularize the structure of view-specific representations. A discriminative global representation is mined with the training process, and the negative impact of weak discriminative views is overcome. Besides, MMatch learns consistent classification while preserving diverse information from multiple views. Experiments on several multi-view datasets demonstrate the effectiveness of MMatch. Xiaoli Wang 0003, Liyong Fu, Yudong Zhang 0001, Yongli Wang 0002, Zechao Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |