Cam-Van Thi Nguyen

dblp:191/0595 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2026
0009-0001-9675-2105ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Counterfactual Understanding via Retrieval-Aware Multimodal Modeling for Time-to-Event Survival Prediction
Ha-Anh Hoang Nguyen, Tri-Duc Phan Le, Duc-Hoang Pham, Huy-Son Nguyen, Cam-Van Thi Nguyen, Duc-Trong Le, Hoang-Quynh Le
ECIR (3)5
2026 From Top-1 to Top-K: A Reproducibility Study and Benchmarking of Counterfactual Explanations for Recommender Systems
abstract
Counterfactual explanations (CEs) provide an intuitive way to understand recommender systems by identifying minimal modifications to user-item interactions that alter recommendation outcomes. Existing CE methods for recommender systems, however, have been evaluated under heterogeneous protocols, using different datasets, recommenders, metrics, and even explanation formats, which hampers reproducibility and fair comparison. Our paper systematically reproduces, re-implement, and re-evaluate eleven state-of-the-art CE methods for recommender systems, covering both native explainers (e.g., LIME-RS, SHAP, PRINCE, ACCENT, LXR, GREASE) and specific graph-based explainers originally proposed for GNNs. Here, a unified benchmarking framework is proposed to assess explainers along three dimensions: explanation format (implicit vs. explicit), evaluation level (item-level vs. list-level), and perturbation scope (user interaction vectors vs. user-item interaction graphs). Our evaluation protocol includes effectiveness, sparsity, and computational complexity metrics, and extends existing item-level assessments to top-K list-level explanations. Through extensive experiments on three real-world datasets and six representative recommender models, we analyze how well previously reported strengths of CE methods generalize across diverse setups. We observe that the trade-off between effectiveness and sparsity depends strongly on the specific method and evaluation setting, particularly under the explicit format; in addition, explainer performance remains largely consistent across item level and list level evaluations, and several graph-based explainers exhibit notable scalability limitations on large recommender graphs. Our results refine and challenge earlier conclusions about the robustness and practicality of CE generation methods in recommender systems: https://github.com/L2R-UET/CFExpRec.
Khac-Manh Thai, Duc-Hoang Pham, Huy-Son Nguyen, Cam-Van Thi Nguyen, Masoud Mansoury, Duc-Trong Le, Hoang-Quynh Le
SIGIR6
2026 Divide and Refine: Enhancing Multimodal Representation and Explainability for Emotion Recognition in Conversation
abstract
Multimodal emotion recognition in conversation (MERC) requires representations that effectively integrate signals from multiple modalities. These signals include modality-specific cues, information shared across modalities, and interactions that emerge only when modalities are combined. In information-theoretic terms, these correspond to unique, redundant, and synergistic contributions. An ideal representation should leverage all three, yet achieving such balance remains challenging. Recent advances in contrastive learning and augmentation-based methods have made progress, but they often overlook the role of data preparation in preserving these components. In particular, applying augmentations directly to raw inputs or fused embeddings can blur the boundaries between modality-unique and cross-modal signals. To address this challenge, we propose a two-phase framework Divide and Refine (DnR). In the Divide phase, each modality is explicitly decomposed into uniqueness, pairwise redundancy, and synergy. In the Refine phase, tailored objectives enhance the informativeness of these components while maintaining their distinct roles. The refined representations are plug-and-play compatible with diverse multimodal pipelines. Extensive experiments on IEMOCAP and MELD demonstrate consistent improvements across multiple MERC backbones. These results highlight the effectiveness of explicitly dividing, refining, and recombining multimodal representations as a principled strategy for advancing emotion recognition. Our implementation is available at https://github.com/mattam301/DnR-WACV2026
Anh-Tuan Mai, Cam-Van Thi Nguyen, Duc-Trong Le
WACV2
2026 Leveraging self-paced curriculum learning for enhanced modality balance in multimodal conversational emotion recognition
Phuong-Anh Nguyen, The-Son Le, Duc-Trong Le, Cam-Van Thi Nguyen
Neural Comput. Appl.4
2025 Multi-modal Adaptive Mixture of Experts for Cold-start Recommendation
abstract
Recommendation systems have faced significant challenges in cold-start scenarios, where new items with a limited history of interaction need to be effectively recommended to users. Though multimodal data (e.g., images, text, audio, etc.) offer rich information to address this issue, existing approaches often employ simplistic integration methods such as concatenation, average pooling, or fixed weighting schemes, which fail to capture the complex relationships between modalities. Our study proposes a novel Mixture of Experts framework for multimodal cold-start recommendation (MAMEX), which dynamically leverages latent representation from different modalities. MAMEX utilizes modality-specific expert networks and introduces a learnable gating mechanism that adaptively weights the contribution of each modality based on its content characteristics. This approach enables MAMEX to emphasize the most informative modalities for each item while maintaining robustness when certain modalities are less relevant or missing. Extensive experiments on benchmark datasets show that MAMEX outperforms state-of-the-art models with superior accuracy and adaptability.
Van-Khang Nguyen 0005, Duc-Hoang Pham, Huy-Son Nguyen, Cam-Van Thi Nguyen, Hoang-Quynh Le, Duc-Trong Le
CIKM4
2025 BRIDGE: Bundle Recommendation via Instruction-Driven Generation
abstract
Bundle recommendation aims to suggest a set of interconnected items to users. However, diverse interaction types and sparse interaction matrices often pose challenges for previous approaches in accurately predicting user-bundle adoptions. Inspired by the distant supervision strategy and generative paradigm, we propose BRIDGE, a novel framework for bundle recommendation. It consists of two main components, namely the item-sensitive instruction generation and the pseudo bundle generation modules. Inspired by the distant supervision approach, the former is to generate more auxiliary information, e.g., sampled item-sensitive instruction, for training without using external data. This information is subsequently aggregated with collaborative signals from user historical interactions to create pseudo ‘ideal’ bundles. This capability allows BRIDGE to explore all aspects of bundles, rather than being limited to existing real-world bundles. It effectively bridging the gap between user imagination and predefined bundles, hence improving the bundle recommendation performance. Experimental results and analyses validate the superiority of BRIDGE over state-of-the-art methods across four benchmark datasets. Our implementation is available at https://github.com/Rec4Fun/BRIDGE.
Tuan-Nghia Bui, Huy-Son Nguyen, Cam-Van Thi Nguyen, Hoang-Quynh Le, Duc-Trong Le
ECAI3
2025 Mi-CGA: Cross-modal Graph Attention Network for robust emotion recognition in the presence of incomplete modalities
Cam-Van Thi Nguyen, Hai-Dang Kieu, Quang-Thuy Ha, Xuan-Hieu Phan, Duc-Trong Le
Neurocomputing1
2024 Curriculum Learning Meets Directed Acyclic Graph for Multimodal Emotion Recognition
abstract
Emotion recognition in conversation (ERC) is a crucial task in natural language processing and affective computing. This paper proposes MultiDAG+CL, a novel approach for Multimodal Emotion Recognition in Conversation (ERC) that employs Directed Acyclic Graph (DAG) to integrate textual, acoustic, and visual features within a unified framework. The model is enhanced by Curriculum Learning (CL) to address challenges related to emotional shifts and data imbalance. Curriculum learning facilitates the learning process by gradually presenting training samples in a meaningful order, thereby improving the model’s performance in handling emotional variations and data imbalance. Experimental results on the IEMOCAP and MELD datasets demonstrate that the MultiDAG+CL models outperform baseline models. We release the code for and experiments: https://github.com/vanntc711/MultiDAG-CL.
Cam-Van Thi Nguyen, Cao-Bach Nguyen, Duc-Trong Le, Quang-Thuy Ha
LREC/COLING1
2024 Ada2I: Enhancing Modality Balance for Multimodal Conversational Emotion Recognition
abstract
Multimodal Emotion Recognition in Conversations (ERC) is a typical multimodal learning task in exploiting various data modalities concurrently. Prior studies on effective multimodal ERC encounter challenges in addressing modality imbalances and optimizing learning across modalities. Dealing with these problems, we present a novel framework named Ada2I, which consists of two inseparable modules namely Adaptive Feature Weighting (AFW) and Adaptive Modality Weighting (AMW) for feature-level and modality-level balancing respectively via leveraging both Inter- and Intra-modal interactions. Additionally, we introduce a refined disparity ratio as part of our training optimization strategy, a simple yet effective measure to assess the overall discrepancy of the model's learning process when handling multiple modalities simultaneously. Experimental results validate the effectiveness of Ada2I with state-of-the-art performance compared to baselines on three benchmark datasets, particularly in addressing modality imbalances.
Cam-Van Thi Nguyen, The-Son Le, Anh-Tuan Mai, Duc-Trong Le
ACM Multimedia1
2024 A Dual-Module Denoising Approach with Curriculum Learning for Enhancing Multimodal Aspect-Based Sentiment Analysis
Nguyen Van Doan, Dat Tran Nguyen, Cam-Van Thi Nguyen
PACLIC3
2024 TECO: Improving Multimodal Intent Recognition with Text Enhancement through Commonsense Knowledge Extraction
Quynh-Mai Thi Nguyen, Lan-Nhi Thi Nguyen, Cam-Van Thi Nguyen
PACLIC3
2023 Vietnamese Multidocument Summarization Using Subgraph Selection-Based Approach with Graph-Informed Self-attention Mechanism
Tam Doan Thanh, Cam-Van Thi Nguyen, Huu-Thin Nguyen, Mai-Vu Tran, Quang-Thuy Ha
ACIIDS (2)2
2023 Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction
abstract
Emotion recognition is a crucial task for human conversation understanding.It becomes more challenging with the notion of multimodal data, e.g., language, voice, and facial expressions.As a typical solution, the globaland the local context information are exploited to predict the emotional label for every single sentence, i.e., utterance, in the dialogue.Specifically, the global representation could be captured via modeling of cross-modal interactions at the conversation level.The local one is often inferred using the temporal information of speakers or emotional shifts, which neglects vital factors at the utterance level.Additionally, most existing approaches take fused features of multiple modalities in an unified input without leveraging modalityspecific representations.Motivating from these problems, we propose the Relational Temporal Graph Neural Network with Auxiliary Cross-Modality Interaction (CORECT), an novel neural network framework that effectively captures conversation-level cross-modality interactions and utterance-level temporal dependencies with the modality-specific manner for conversation understanding.Extensive experiments demonstrate the effectiveness of CORECT via its stateof-the-art results on the IEMOCAP and CMU-MOSEI datasets for the multimodal ERC task.
Cam-Van Thi Nguyen, Anh-Tuan Mai, The-Son Le, Hai-Dang Kieu, Duc-Trong Le
EMNLP1