Meirong Ding

dblp:117/1552 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Enhancing Incomplete Multimodal Learning via Modal Complementary Recovering
abstract
Multimodal learning presents significant challenges arising from the unpredictable absence of modalities during both training and testing phases. Existing recovery methods struggle to leverage the available data, which can introduce additional noise during the recovery process and degrade performance. To mitigate these issues, we introduce a novel Modal Complementary Recovering (MCR) paradigm that strategically integrates both Complementary Graph-based Recovery (CGR) and Topological Low-Rank Adaptation (ToRA) mechanisms to enhance the effectiveness and reliability of incomplete multimodal learning. For effective exploitation of the complementarity among different modalities, the main objective of CGR is to employ bidirectional mapping flows trained on a small subset of complete data to learn complementary graphs across all modalities. By constructing entity-relationship diagram specific to dataset, ToRA is designed to enhance the fine-tuning process by incorporating topological prompt. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our MCR paradigm in comparison to state-of-the-art baselines.
Meirong Ding, Chuang Zou, Wenxiu Cai, Bingzhi Chen
ICASSP1
2025 Local Feature Alignment Prompt-Tuning for Few-shot Multimodal Aspect Sentiment Analysis
abstract
Multi-modal aspect-oriented sentiment classification (MASC) is a fine-grain task, which aims to detect the sentiment polarity of specific aspect. However, conventional studies suffer from two issues. It is difficult to collect the annotated multi-modal data in fine-grained domains. Meanwhile, the local information corresponding to aspect words in the image has not been mined, and redundant information can affect the accuracy of fine-grained sentiment analysis. To alleviate the above two issues, we propose a Prompt-tuning method based on Alignment between Aspect and Local Images (PAALI). Our approach introduces a novel multi-modal prompt template to bridge the gap between text and visual data modalities. Furthermore, we employ a strategy of randomly masking image patches to align them with aspect word features, capturing deeper semantic information. Extensive experiments on multiple benchmark datasets in few-shot settings consistently demonstrate the superiority and robustness of our PAALI method over state-of-the-art competitors.
Meirong Ding, Chuang Zou
ICASSP1
2025 Dual-Path Consistency Unsupervised Domain Adaptation for Nighttime Semantic Segmentation
abstract
Nighttime semantic segmentation is an indispensable component in practical applications, such as automated vehicles. However, it is often hindered by the lack of annotations due to interference caused by inadequate lighting or exposure. To overcome these difficulties, we propose a Dual-Path Consistency (DPC) unsupervised domain adaptation (UDA) approach. One path is Image Darkening Path (IDP), in which feature representations of original images and darkened images extracted from the feature encoder are leveraged to maintain cross-domain style consistency. Another path is Image Masking Path (IMP), in which the masked images are reconstructed under the guidance of pseudo-labels, aiming to maintain content consistency in an entirely identical scenario. Extensive experiments on four bench-marks demonstrate the superior performance of the proposed DPC for nighttime semantic segmentation.
Yuwu Lu, Jicong Lang, Meirong Ding
ICASSP3
2024 Medical Vision-Language Representation Learning with Cross-Modal Multi-Teacher Contrastive Distillation
abstract
Medical vision-language representation learning has garnered considerable attention owing to its applicability to extracting generic representations from the image and text modality. However, it still remains challenging to acquire a more comprehensive understanding of intra- and inter-modal semantic knowledge. In this paper, we propose a Cross-Modal Multi-Teacher Contrastive Distillation (CMCD) architecture, which aims to comprehensively learn medical vision-language representation in a unified multi-teacher framework. Specifically, a cross-modal knowledge distillation (CKD) module is designed to refine reconstructed semantics under an additional supervision signal generated by momentum teachers from the other modality, achieving more robust semantic interaction across modalities. To better alleviate the heterogeneity and semantic gaps, the multi-level contrastive learning (MCL) module is conceived to align features of both intra- and inter-modal via contrastive learning from multi-level perspectives. Extensive experiments on two medical downstream tasks, i.e., Med-VQA and Med-ITC, demonstrate that our CMCD consistently outperforms the state-of-the-art methods.
Bingzhi Chen, Yishu Liu 0001, Jiahui Pan 0003, Meirong Ding
ICASSP6
2024 Robust Visual Question Answering With Contrastive-Adversarial Consistency Constraints
abstract
Visual cues and question semantics contribute to final answer predictions from distinct perspectives. However, inherent language bias confounds the relationship between visual and question cues, leading to a misguided preference for question semantics. Different from the existing studies that focus on inter-class discrimination, this paper proposes a robust visual question answering framework with contrastive-adversarial consistency constraints (CACC) at both inter- and intra-instance levels. From a fine-grained instance-level perspective, our approach initially introduces an effective inter-instance contrastive constraint to perform adaptive bias rectification. To enhance intra-instance invariance and reduce information redundancy, we refine the concept of semantic structure relationships by constructing intra-instance adversarial constraints using the Hilbert-Schmidt Independence Criterion (HSIC) independence criterion. Benefitting from both inter- and intra-instance perspectives, our method can effectively alleviate these language biases, enhancing the overall robustness of the representation. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our CACC over state-of-the-art baselines.
Meirong Ding, Yishu Liu 0001, Guangming Lu 0002, Bingzhi Chen
ICME2
2023 RFM: response-aware feedback mechanism for background based conversation
Jiatao Chen, Zhibin Du, Huimin Deng, Mayi Xu, Zibang Gan, Meirong Ding
Appl. Intell.7
2022 Ultra-short-Term Load Forecasting Model Based on VMD and TGCN-GRU
Meirong Ding, Gaoyan Cai, Wensheng Gan
IEA/AIE1