Lisi Mo

dblp:349/4411 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
10since 2021 · last 2026
0009-0008-4742-4456ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HiPMD: Hierarchical Prompting meets Multi-Layer Distillation for Incomplete Multimodal Learning
Lisi Mo, Aoqiang Jian, Ruofan Zhou, Letian Xu, Ruiting Dai
ICC1
2026 Anchor Drift No More: Hierarchical Consistency-Guided Prompt Distillation for Incomplete Multimodal Learning
abstract
Web-scale content is rich in modalities yet frequently incomplete due to device limits, transmission errors, or privacy controls, making learning with missing modalities a core challenge. Prior reconstruction and alignment strategies often fail to preserve a stable class geometry when inputs are partial, leading to anchor drift -- a shift of class prototypes between complete and incomplete views that distorts the shared representation space and degrades generalization. We introduce HiCoD (Hierarchical Consistency-Guided Pro mpt Distillation), which learns a robust, class-anchored semantic space. HiCoD combines: (1) a modality-aware semantic graph that restores cross-modal structure under partial observations; (2) dual-level anchoring that unifies large-language-model–derived global category prototypes with top-K local exemplars to balance cross-modal coherence and modality-specific detail; and (3) multi-level distillation that aligns unimodal features, fused embeddings, and prompt-completed signals within a single anchor space. Across CMU-MOSI, CMU-MOSEI, and additional benchmarks, HiCoD sets a new state of the art under both fixed-pattern and random missingness, improving Acc-2 by up to 6.4 points over MPLMM and remaining robust when key modalities are absent.
Ruiting Dai, Zesen Cai, Lisi Mo, Guiduo Duan, Keren Shi, Tao He 0007
WWW3
2026 Multimodal learning with missing modalities: Can progressive diffusion achieve distribution-consistent learning from incomplete multimodal data?
Ruiting Dai, Aoqiang Jian, Ruoxuan Zhang, Yuang Jia, Lisi Mo, Siyu Zhan, Tao He 0007
Knowl. Based Syst.5
2026 TopoKG: Infer Internet AS-Level Topology From Global Perspective
abstract
Internet Autonomous System (AS) level topology includes AS topology structure and AS business relationships, describes the essence of Internet inter-domain routing, and is the basis for Internet operation and management research. Although the latest topology inference methods have made significant progress, those relying solely on local information struggle to eliminate inference errors caused by observation bias and data noise due to their lack of a global perspective. In contrast, we not only leverage local AS link features but also re-examine the hierarchical structure of Internet AS-level topology, proposing a novel inference method called topoKG. TopoKG introduces a knowledge graph to represent the relationships between different elements on a global scale and the business routing strategies of ASes at various tiers, which effectively reduces inference errors resulting from observation bias and data noise by incorporating a global perspective. First, we construct an Internet AS-level topology knowledge graph to represent relevant data, enabling us to better leverage the global perspective and uncover the complex relationships among multiple elements. Next, we employ knowledge graph meta paths to measure the similarity of AS business routing strategies and introduce this global perspective constraint to infer the AS business relationships and hierarchical structure iteratively. Additionally, we embed the entire knowledge graph upon completing the iteration and conduct knowledge inference to derive AS business relationships. This approach captures global features and more intricate relational patterns within the knowledge graph, further enhancing the accuracy of AS-level topology inference. Compared to the state-of-the-art methods, our approach achieves more accurate AS-level topology inference, reducing the average inference error across various AS link types by up to 1.2 to 4.4 times.
Lisi Mo, Gaolei Fei, Yunpeng Zhou, Ming Xian, Xuemeng Zhai, Guangmin Hu
IEEE Trans. Netw. Serv. Manag.2
2025 Unbiased Missing-Modality Multimodal Learning
Ruiting Dai, Yandong Yan, Lisi Mo, Ke Qin, Tao He 0007
ICCV4
2025 RobustPT: Dynamic Disentanglement Prompt Tuning in Vision-Language Models with Missing Modalities
abstract
Recently, prompt tuning has garnered considerable attention due to its success across various Vision-Language (VL) tasks. However, unimodal prompts, coupled prompts, and joint prompts in these models often lead to suboptimal performance due to differences in information density and complexity between modalities. Particularly, in scenarios with missing modalities, these prompt-based approaches tend to exacerbate 'Channel Bias'-a phonomenon where models overly rely on specific feature (such as unmissing-modal feature) channels from the base tasks, thereby undermining the model's ability to capture crucial shared knowledge applicable to new tasks and affecting its generalizability. To address this challenge, we propose RobustPT, a dynamic disentanglement prompt tuning model designed to enhance the robustness of VL models under modality missing conditions. RobustPT utilizes a multi-channel prompting mechanism to dynamically disentangle and align prompts. Specifically, RobustPT is divided into single-channel tuning and alignment-channel tuning, where prompts for each modality run independently in sequence to delve deeply into their intrinsic characteristics, followed by an integration through a non-strong coupling strategy to effectively balance information contributions and enhance overall performance. Extensive experiments demonstrate that our RobustPT achieve significant improvements over the current state-of-the-art across all benchmark datasets. Our codes are available at https://github.com/Trae1ounG/RobustPT.
Ruiting Dai, Yuqiao Tan, Lisi Mo, Tao He 0007, Ke Qin, Shuang Liang 0002
ICMR3
2025 γ-CRD: Gamma-Cooperative Retrieval Diffusion Model for Robust Incomplete Multimodal Learning
abstract
Multimodal learning in open environments faces significant challenges due to modality incompleteness and noise interference. Current prompt engineering emphasizes modality absence over task-instance contextualization that hinders cross-modal knowledge transfer, whereas conditional generation approaches for missing modality recovery exhibit an excessive reliance on the quality of available modalities. To address these issues, we propose a Gamma-Cooperative Retrieval Diffusion model (γ-CRD), inspired by the human brain's multi-source contextual completion mechanism, which leverages a retrieval-augmented prompt generation framework and normal-inverse Gamma noise modeling to enhance robustness in incomplete multimodal learning. Specifically, it consists of three modules: (1) Retrieval-Augmented Contextualization: Construct a multimodal memory bank and retrieve relevant instances via similarity calculation under a gating mechanism to augment the contextualization of missing modalities. (2) Prompt-Driven Diffusion Generation Module: Builds prompts based on retrieval results and incorporates them into a denoising diffusion probabilistic model through an attention mechanism to enhance contextualized knowledge transfer and generate missing modalities. (3) Inverse-Gamma Noise Optimization Module: Model a mixed normal-inverse gamma distribution, which is dynamically aware of noise and enables uncertainty estimation in multimodal fusion, ensuring robust and reliable multimodal regression. Extensive experiments on three real-world datasets demonstrate that γ-CRD consistently outperforms state-of-the-art baselines, especially achieving a 5.7% improvement in accuracy compared to the leading model e.g., IMDer [1].
Ruiting Dai, Wenwei Zhu, Haoran Meng, Zhengdao Yuan, Yandong Yan, Lisi Mo
ICMR7
2025 A unified cross-source context enhancement model for multi-source fake news detection
Ruiting Dai, Haoran Meng, Zhengdao Yuan, Lisi Mo, Wenwei Zhu, Tao He 0007
Knowl. Based Syst.4
2024 G-SAP: Graph-based Structure-Aware Prompt Learning over Heterogeneous Knowledge for Commonsense Reasoning
abstract
Commonsense question answering has demonstrated considerable potential across various applications like assistants and social robots.Although fully fine-tuned Pre-trained Language Model(PLM) has achieved remarkable performance in commonsense reasoning, their tendency to excessively prioritize textual information hampers the precise transfer of structural knowledge and undermines interpretability.Some studies have explored combining Language Models (LM) with Knowledge Graphs (KGs) by coarsely fusing the two modalities to perform Graph Neural Network (GNN)-based reasoning that lacks a profound interaction between heterogeneous modalities.In this paper, we propose a novel Graph-based Structure-Aware Prompt Learning Model for commonsense reasoning, named G-SAP, aiming to maintain a balance between heterogeneous knowledge and enhance the cross-modal interaction within the LM+GNNs model.In particular, an evidence graph is constructed by integrating multiple knowledge sources, i.e.ConceptNet, Wikipedia, and Cambridge Dictionary to boost the performance.Afterward, a structure-aware frozen PLM is employed to fully incorporate the structured and textual information from the evidence graph, where the generation of prompts is driven
Ruiting Dai, Yuqiao Tan, Lisi Mo, Shuang Liang 0002, Guohao Huo, Yao Cheng 0013
ICMR3
2023 Multi-modal Representation Learning for Social Post Location Inference
abstract
Inferring geographic locations via social posts is essential for many practical location-based applications such as product marketing, point-of-interest recommendation, and infector tracking for COVID-19. Unlike image-based location retrieval or social-post text embedding-based location inference, the combined effect of multi-modal information (i.e., post images, text, and hashtags) for social post positioning receives less attention. In this work, we collect real datasets of social posts with images, texts, and hashtags from Instagram and propose a novel Multi-modal Representation Learning Framework (MRLF) capable of fusing different modalities of social posts for location inference. MRLF integrates a multi-head attention mechanism to enhance location-salient information extraction while significantly improving location inference compared with single domain-based methods. To overcome the noisy user-generated textual content, we introduce a novel attention-based character-aware module that considers the relative dependencies between characters of social post texts and hashtags for flexible multi-model information fusion. The experimental results show that MRLF can make accurate location predictions and open a new door to understanding the multi-modal data of social posts for online inference tasks.
Ruiting Dai, Xucheng Luo, Lisi Mo, Wanlun Ma, Fan Zhou 0002
ICC4