EDBT 2026 Demo / reviewers in the wild / expert
Ruiting Dai
dblp:341/2778
· DBLP profile ↗
10ranked-venue papers
8as first author
10since 2021 · last 2026
0000-0002-8944-6759ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HiPMD: Hierarchical Prompting meets Multi-Layer Distillation for Incomplete Multimodal Learning
Lisi Mo, Aoqiang Jian, Ruofan Zhou, Letian Xu, Ruiting Dai |
ICC | 5 |
| 2026 | Anchor Drift No More: Hierarchical Consistency-Guided Prompt Distillation for Incomplete Multimodal LearningabstractWeb-scale content is rich in modalities yet frequently incomplete due to device limits, transmission errors, or privacy controls, making learning with missing modalities a core challenge. Prior reconstruction and alignment strategies often fail to preserve a stable class geometry when inputs are partial, leading to anchor drift -- a shift of class prototypes between complete and incomplete views that distorts the shared representation space and degrades generalization. We introduce HiCoD (Hierarchical Consistency-Guided Pro mpt Distillation), which learns a robust, class-anchored semantic space. HiCoD combines: (1) a modality-aware semantic graph that restores cross-modal structure under partial observations; (2) dual-level anchoring that unifies large-language-model–derived global category prototypes with top-K local exemplars to balance cross-modal coherence and modality-specific detail; and (3) multi-level distillation that aligns unimodal features, fused embeddings, and prompt-completed signals within a single anchor space. Across CMU-MOSI, CMU-MOSEI, and additional benchmarks, HiCoD sets a new state of the art under both fixed-pattern and random missingness, improving Acc-2 by up to 6.4 points over MPLMM and remaining robust when key modalities are absent. Ruiting Dai, Zesen Cai, Lisi Mo, Guiduo Duan, Keren Shi, Tao He 0007 |
WWW | 1 |
| 2026 | Multimodal learning with missing modalities: Can progressive diffusion achieve distribution-consistent learning from incomplete multimodal data?
Ruiting Dai, Aoqiang Jian, Ruoxuan Zhang, Yuang Jia, Lisi Mo, Siyu Zhan, Tao He 0007 |
Knowl. Based Syst. | 1 |
| 2025 | Unbiased Missing-Modality Multimodal Learning
Ruiting Dai, Yandong Yan, Lisi Mo, Ke Qin, Tao He 0007 |
ICCV | 1 |
| 2025 | RobustPT: Dynamic Disentanglement Prompt Tuning in Vision-Language Models with Missing ModalitiesabstractRecently, prompt tuning has garnered considerable attention due to its success across various Vision-Language (VL) tasks. However, unimodal prompts, coupled prompts, and joint prompts in these models often lead to suboptimal performance due to differences in information density and complexity between modalities. Particularly, in scenarios with missing modalities, these prompt-based approaches tend to exacerbate 'Channel Bias'-a phonomenon where models overly rely on specific feature (such as unmissing-modal feature) channels from the base tasks, thereby undermining the model's ability to capture crucial shared knowledge applicable to new tasks and affecting its generalizability. To address this challenge, we propose RobustPT, a dynamic disentanglement prompt tuning model designed to enhance the robustness of VL models under modality missing conditions. RobustPT utilizes a multi-channel prompting mechanism to dynamically disentangle and align prompts. Specifically, RobustPT is divided into single-channel tuning and alignment-channel tuning, where prompts for each modality run independently in sequence to delve deeply into their intrinsic characteristics, followed by an integration through a non-strong coupling strategy to effectively balance information contributions and enhance overall performance. Extensive experiments demonstrate that our RobustPT achieve significant improvements over the current state-of-the-art across all benchmark datasets. Our codes are available at https://github.com/Trae1ounG/RobustPT. Ruiting Dai, Yuqiao Tan, Lisi Mo, Tao He 0007, Ke Qin, Shuang Liang 0002 |
ICMR | 1 |
| 2025 | γ-CRD: Gamma-Cooperative Retrieval Diffusion Model for Robust Incomplete Multimodal LearningabstractMultimodal learning in open environments faces significant challenges due to modality incompleteness and noise interference. Current prompt engineering emphasizes modality absence over task-instance contextualization that hinders cross-modal knowledge transfer, whereas conditional generation approaches for missing modality recovery exhibit an excessive reliance on the quality of available modalities. To address these issues, we propose a Gamma-Cooperative Retrieval Diffusion model (γ-CRD), inspired by the human brain's multi-source contextual completion mechanism, which leverages a retrieval-augmented prompt generation framework and normal-inverse Gamma noise modeling to enhance robustness in incomplete multimodal learning. Specifically, it consists of three modules: (1) Retrieval-Augmented Contextualization: Construct a multimodal memory bank and retrieve relevant instances via similarity calculation under a gating mechanism to augment the contextualization of missing modalities. (2) Prompt-Driven Diffusion Generation Module: Builds prompts based on retrieval results and incorporates them into a denoising diffusion probabilistic model through an attention mechanism to enhance contextualized knowledge transfer and generate missing modalities. (3) Inverse-Gamma Noise Optimization Module: Model a mixed normal-inverse gamma distribution, which is dynamically aware of noise and enables uncertainty estimation in multimodal fusion, ensuring robust and reliable multimodal regression. Extensive experiments on three real-world datasets demonstrate that γ-CRD consistently outperforms state-of-the-art baselines, especially achieving a 5.7% improvement in accuracy compared to the leading model e.g., IMDer [1]. Ruiting Dai, Wenwei Zhu, Haoran Meng, Zhengdao Yuan, Yandong Yan, Lisi Mo |
ICMR | 1 |
| 2025 | Unbiased multimodal intent recognition with auxiliary rationale generation
Ruiting Dai, Guiduo Duan, Ke Qin, Tao He 0007 |
Neurocomputing | 2 |
| 2025 | A unified cross-source context enhancement model for multi-source fake news detection
Ruiting Dai, Haoran Meng, Zhengdao Yuan, Lisi Mo, Wenwei Zhu, Tao He 0007 |
Knowl. Based Syst. | 1 |
| 2024 | G-SAP: Graph-based Structure-Aware Prompt Learning over Heterogeneous Knowledge for Commonsense ReasoningabstractCommonsense question answering has demonstrated considerable potential across various applications like assistants and social robots.Although fully fine-tuned Pre-trained Language Model(PLM) has achieved remarkable performance in commonsense reasoning, their tendency to excessively prioritize textual information hampers the precise transfer of structural knowledge and undermines interpretability.Some studies have explored combining Language Models (LM) with Knowledge Graphs (KGs) by coarsely fusing the two modalities to perform Graph Neural Network (GNN)-based reasoning that lacks a profound interaction between heterogeneous modalities.In this paper, we propose a novel Graph-based Structure-Aware Prompt Learning Model for commonsense reasoning, named G-SAP, aiming to maintain a balance between heterogeneous knowledge and enhance the cross-modal interaction within the LM+GNNs model.In particular, an evidence graph is constructed by integrating multiple knowledge sources, i.e.ConceptNet, Wikipedia, and Cambridge Dictionary to boost the performance.Afterward, a structure-aware frozen PLM is employed to fully incorporate the structured and textual information from the evidence graph, where the generation of prompts is driven Ruiting Dai, Yuqiao Tan, Lisi Mo, Shuang Liang 0002, Guohao Huo, Yao Cheng 0013 |
ICMR | 1 |
| 2023 | Multi-modal Representation Learning for Social Post Location InferenceabstractInferring geographic locations via social posts is essential for many practical location-based applications such as product marketing, point-of-interest recommendation, and infector tracking for COVID-19. Unlike image-based location retrieval or social-post text embedding-based location inference, the combined effect of multi-modal information (i.e., post images, text, and hashtags) for social post positioning receives less attention. In this work, we collect real datasets of social posts with images, texts, and hashtags from Instagram and propose a novel Multi-modal Representation Learning Framework (MRLF) capable of fusing different modalities of social posts for location inference. MRLF integrates a multi-head attention mechanism to enhance location-salient information extraction while significantly improving location inference compared with single domain-based methods. To overcome the noisy user-generated textual content, we introduce a novel attention-based character-aware module that considers the relative dependencies between characters of social post texts and hashtags for flexible multi-model information fusion. The experimental results show that MRLF can make accurate location predictions and open a new door to understanding the multi-modal data of social posts for online inference tasks. Ruiting Dai, Xucheng Luo, Lisi Mo, Wanlun Ma, Fan Zhou 0002 |
ICC | 1 |