EDBT 2026 Demo / reviewers in the wild / expert
Duc-Trong Le
dblp:155/5277
· DBLP profile ↗
30ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0003-4621-8956ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 4 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Counterfactual Understanding via Retrieval-Aware Multimodal Modeling for Time-to-Event Survival Prediction
Ha-Anh Hoang Nguyen, Tri-Duc Phan Le, Duc-Hoang Pham, Huy-Son Nguyen, Cam-Van Thi Nguyen, Duc-Trong Le, Hoang-Quynh Le |
ECIR (3) | 6 |
| 2026 | FedCKD: A Knowledge Distillation Approach to Cross-Client Learning in Federated Learning with Label-Exclusive Datasets
Minh-Chau Le, Hoang-Quynh Le, Duc-Trong Le, Tram Truong Huu |
PAKDD (2) | 3 |
| 2026 | From Top-1 to Top-K: A Reproducibility Study and Benchmarking of Counterfactual Explanations for Recommender SystemsabstractCounterfactual explanations (CEs) provide an intuitive way to understand recommender systems by identifying minimal modifications to user-item interactions that alter recommendation outcomes. Existing CE methods for recommender systems, however, have been evaluated under heterogeneous protocols, using different datasets, recommenders, metrics, and even explanation formats, which hampers reproducibility and fair comparison. Our paper systematically reproduces, re-implement, and re-evaluate eleven state-of-the-art CE methods for recommender systems, covering both native explainers (e.g., LIME-RS, SHAP, PRINCE, ACCENT, LXR, GREASE) and specific graph-based explainers originally proposed for GNNs. Here, a unified benchmarking framework is proposed to assess explainers along three dimensions: explanation format (implicit vs. explicit), evaluation level (item-level vs. list-level), and perturbation scope (user interaction vectors vs. user-item interaction graphs). Our evaluation protocol includes effectiveness, sparsity, and computational complexity metrics, and extends existing item-level assessments to top-K list-level explanations. Through extensive experiments on three real-world datasets and six representative recommender models, we analyze how well previously reported strengths of CE methods generalize across diverse setups. We observe that the trade-off between effectiveness and sparsity depends strongly on the specific method and evaluation setting, particularly under the explicit format; in addition, explainer performance remains largely consistent across item level and list level evaluations, and several graph-based explainers exhibit notable scalability limitations on large recommender graphs. Our results refine and challenge earlier conclusions about the robustness and practicality of CE generation methods in recommender systems: https://github.com/L2R-UET/CFExpRec. Khac-Manh Thai, Duc-Hoang Pham, Huy-Son Nguyen, Cam-Van Thi Nguyen, Masoud Mansoury, Duc-Trong Le, Hoang-Quynh Le |
SIGIR | 8 |
| 2026 | Divide and Refine: Enhancing Multimodal Representation and Explainability for Emotion Recognition in ConversationabstractMultimodal emotion recognition in conversation (MERC) requires representations that effectively integrate signals from multiple modalities. These signals include modality-specific cues, information shared across modalities, and interactions that emerge only when modalities are combined. In information-theoretic terms, these correspond to unique, redundant, and synergistic contributions. An ideal representation should leverage all three, yet achieving such balance remains challenging. Recent advances in contrastive learning and augmentation-based methods have made progress, but they often overlook the role of data preparation in preserving these components. In particular, applying augmentations directly to raw inputs or fused embeddings can blur the boundaries between modality-unique and cross-modal signals. To address this challenge, we propose a two-phase framework Divide and Refine (DnR). In the Divide phase, each modality is explicitly decomposed into uniqueness, pairwise redundancy, and synergy. In the Refine phase, tailored objectives enhance the informativeness of these components while maintaining their distinct roles. The refined representations are plug-and-play compatible with diverse multimodal pipelines. Extensive experiments on IEMOCAP and MELD demonstrate consistent improvements across multiple MERC backbones. These results highlight the effectiveness of explicitly dividing, refining, and recombining multimodal representations as a principled strategy for advancing emotion recognition. Our implementation is available at https://github.com/mattam301/DnR-WACV2026 Anh-Tuan Mai, Cam-Van Thi Nguyen, Duc-Trong Le |
WACV | 3 |
| 2026 | Enhancing Graph-based Recommendations with Majority-Voting LLM-Rerank Augmentation
Minh-Anh Nguyen, Bao Nguyen, Ha Lan N. T., Tuan-Anh Hoang, Duc-Trong Le, Dung D. Le |
Mach. Learn. | 5 |
| 2026 | Leveraging self-paced curriculum learning for enhanced modality balance in multimodal conversational emotion recognition
Phuong-Anh Nguyen, The-Son Le, Duc-Trong Le, Cam-Van Thi Nguyen |
Neural Comput. Appl. | 3 |
| 2026 | MODE: A model-agnostic framework for object detection under adverse weather conditions
Tuan-Duc Nguyen, Duc-Trong Le |
Pattern Recognit. | 2 |
| 2025 | Multi-modal Adaptive Mixture of Experts for Cold-start RecommendationabstractRecommendation systems have faced significant challenges in cold-start scenarios, where new items with a limited history of interaction need to be effectively recommended to users. Though multimodal data (e.g., images, text, audio, etc.) offer rich information to address this issue, existing approaches often employ simplistic integration methods such as concatenation, average pooling, or fixed weighting schemes, which fail to capture the complex relationships between modalities. Our study proposes a novel Mixture of Experts framework for multimodal cold-start recommendation (MAMEX), which dynamically leverages latent representation from different modalities. MAMEX utilizes modality-specific expert networks and introduces a learnable gating mechanism that adaptively weights the contribution of each modality based on its content characteristics. This approach enables MAMEX to emphasize the most informative modalities for each item while maintaining robustness when certain modalities are less relevant or missing. Extensive experiments on benchmark datasets show that MAMEX outperforms state-of-the-art models with superior accuracy and adaptability. Van-Khang Nguyen 0005, Duc-Hoang Pham, Huy-Son Nguyen, Cam-Van Thi Nguyen, Hoang-Quynh Le, Duc-Trong Le |
CIKM | 6 |
| 2025 | BRIDGE: Bundle Recommendation via Instruction-Driven GenerationabstractBundle recommendation aims to suggest a set of interconnected items to users. However, diverse interaction types and sparse interaction matrices often pose challenges for previous approaches in accurately predicting user-bundle adoptions. Inspired by the distant supervision strategy and generative paradigm, we propose BRIDGE, a novel framework for bundle recommendation. It consists of two main components, namely the item-sensitive instruction generation and the pseudo bundle generation modules. Inspired by the distant supervision approach, the former is to generate more auxiliary information, e.g., sampled item-sensitive instruction, for training without using external data. This information is subsequently aggregated with collaborative signals from user historical interactions to create pseudo ‘ideal’ bundles. This capability allows BRIDGE to explore all aspects of bundles, rather than being limited to existing real-world bundles. It effectively bridging the gap between user imagination and predefined bundles, hence improving the bundle recommendation performance. Experimental results and analyses validate the superiority of BRIDGE over state-of-the-art methods across four benchmark datasets. Our implementation is available at https://github.com/Rec4Fun/BRIDGE. Tuan-Nghia Bui, Huy-Son Nguyen, Cam-Van Thi Nguyen, Hoang-Quynh Le, Duc-Trong Le |
ECAI | 5 |
| 2025 | RaMen: Multi-Strategy Multi-Modal Learning for Bundle ConstructionabstractExisting studies on bundle construction have relied merely on user feedback via bipartite graphs or enhanced item representations using semantic information. These approaches fail to capture elaborate relations hidden in real-world bundle structures, resulting in suboptimal bundle representations. To overcome this limitation, we propose RaMen, a novel method that provides a holistic multi-strategy approach for bundle construction. RaMen utilizes both intrinsic (characteristics) and extrinsic (collaborative signals) information to model bundle structures through Explicit Strategy-aware Learning (ESL) and Implicit Strategy-aware Learning (ISL). ESL employs task-specific attention mechanisms to encode multi-modal data and direct collaborative relations between items, thereby explicitly capturing essential bundle features. Moreover, ISL computes hyperedge dependencies and hypergraph message passing to uncover shared latent intents among groups of items. Integrating diverse strategies enables RaMen to learn more comprehensive and robust bundle representations. Meanwhile, Multi-strategy Alignment & Discrimination module is employed to facilitate knowledge transfer between learning strategies and ensure discrimination between items/bundles. Extensive experiments demonstrate the effectiveness of RaMen over state-of-the-art models on various domains, justifying valuable insights into complex item set problems. Huy-Son Nguyen, Duc-Hoang Pham, Duc-Trong Le, Hoang-Quynh Le, Padipat Sitkrongwong, Atsuhiro Takasu, Masoud Mansoury |
ECAI | 4 |
| 2025 | Mi-CGA: Cross-modal Graph Attention Network for robust emotion recognition in the presence of incomplete modalities
Cam-Van Thi Nguyen, Hai-Dang Kieu, Quang-Thuy Ha, Xuan-Hieu Phan, Duc-Trong Le |
Neurocomputing | 5 |
| 2025 | Towards efficient pareto-optimal utility-fairness between groups in repeated rankings
Phuong Mai Dinh, Duc-Trong Le, Tuan-Anh Hoang, Dung Duy Le |
Mach. Learn. | 2 |
| 2024 | Towards Robust Continual Learning: A Multi-Head Approach with Online Prototype Equilibrium and Adaptive Prototypical Feedback
Quynh-Trang Pham Thi, Duc-Hung Nguyen, Duc-Trong Le, Tri-Thanh Nguyen, Quang-Thuy Ha |
ACIIDS (2) | 4 |
| 2024 | Curriculum Learning Meets Directed Acyclic Graph for Multimodal Emotion RecognitionabstractEmotion recognition in conversation (ERC) is a crucial task in natural language processing and affective computing. This paper proposes MultiDAG+CL, a novel approach for Multimodal Emotion Recognition in Conversation (ERC) that employs Directed Acyclic Graph (DAG) to integrate textual, acoustic, and visual features within a unified framework. The model is enhanced by Curriculum Learning (CL) to address challenges related to emotional shifts and data imbalance. Curriculum learning facilitates the learning process by gradually presenting training samples in a meaningful order, thereby improving the model’s performance in handling emotional variations and data imbalance. Experimental results on the IEMOCAP and MELD datasets demonstrate that the MultiDAG+CL models outperform baseline models. We release the code for and experiments: https://github.com/vanntc711/MultiDAG-CL. Cam-Van Thi Nguyen, Cao-Bach Nguyen, Duc-Trong Le, Quang-Thuy Ha |
LREC/COLING | 3 |
| 2024 | Collaborative Fair-is-Better Filtering for Implicit FeedbackabstractWith the wide adoption of recommender systems, fairness has increasingly become a critical topic in many applications, such as e-commerce, job search, and online entertainment. Collaborative filtering is susceptible to unfair recommendations for users from sensitive groups due to the non-negligible presence of biases. Whereas recent work mostly concerned fairness with the explicit ratings, we propose fairness-awareness recommendation models, referred to as CFBF, for implicit feedback (e.g., clicks, views, or purchases), which are ubiquitous in today’s context. This paper considers sensitive attributes, such as gender and age, in both disjoined and combined manners to investigate the models’ unfairness. We discuss several fairness metrics for implicit feedback recommendations based on Spearman’s rank correlation and Kendal Tau. Comprehensive experiments on Movielens and LastFM show that CFBF significantly improves user groups’ fairness with comparable and even better ranking performance. Hoang V. Dong, Huu-Quang Nguyen, Hoang D. Nguyen, Duc-Trong Le |
KES | 4 |
| 2024 | Ada2I: Enhancing Modality Balance for Multimodal Conversational Emotion RecognitionabstractMultimodal Emotion Recognition in Conversations (ERC) is a typical multimodal learning task in exploiting various data modalities concurrently. Prior studies on effective multimodal ERC encounter challenges in addressing modality imbalances and optimizing learning across modalities. Dealing with these problems, we present a novel framework named Ada2I, which consists of two inseparable modules namely Adaptive Feature Weighting (AFW) and Adaptive Modality Weighting (AMW) for feature-level and modality-level balancing respectively via leveraging both Inter- and Intra-modal interactions. Additionally, we introduce a refined disparity ratio as part of our training optimization strategy, a simple yet effective measure to assess the overall discrepancy of the model's learning process when handling multiple modalities simultaneously. Experimental results validate the effectiveness of Ada2I with state-of-the-art performance compared to baselines on three benchmark datasets, particularly in addressing modality imbalances. Cam-Van Thi Nguyen, The-Son Le, Anh-Tuan Mai, Duc-Trong Le |
ACM Multimedia | 4 |
| 2023 | Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality InteractionabstractEmotion recognition is a crucial task for human conversation understanding.It becomes more challenging with the notion of multimodal data, e.g., language, voice, and facial expressions.As a typical solution, the globaland the local context information are exploited to predict the emotional label for every single sentence, i.e., utterance, in the dialogue.Specifically, the global representation could be captured via modeling of cross-modal interactions at the conversation level.The local one is often inferred using the temporal information of speakers or emotional shifts, which neglects vital factors at the utterance level.Additionally, most existing approaches take fused features of multiple modalities in an unified input without leveraging modalityspecific representations.Motivating from these problems, we propose the Relational Temporal Graph Neural Network with Auxiliary Cross-Modality Interaction (CORECT), an novel neural network framework that effectively captures conversation-level cross-modality interactions and utterance-level temporal dependencies with the modality-specific manner for conversation understanding.Extensive experiments demonstrate the effectiveness of CORECT via its stateof-the-art results on the IEMOCAP and CMU-MOSEI datasets for the multimodal ERC task. Cam-Van Thi Nguyen, Anh-Tuan Mai, The-Son Le, Hai-Dang Kieu, Duc-Trong Le |
EMNLP | 5 |
| 2023 | Multimodal Machine Learning for Mental Disorder Detection: A Scoping ReviewabstractRecent advancements in machine learning and multimedia technologies have paved new ways for automatic medical diagnosis. In mental health, multimodal inputs such as visual and audible sensing data are promising to investigate the underlying mechanisms of many conditions, such as depression and bipolar disorders. With the increasing burden on healthcare systems, timely diagnosis of mental diseases using multiple modalities might benefit millions of people worldwide. This scoping review provides an exploratory overview of recent multimodal machine learning approaches for mental disorder screening. We also discuss a generalised end-to-end multimodal machine learning pipeline for future research and development of multimodal disease detection. Thuy-Trinh Nguyen, Viet Hoang-Quoc Pham, Duc-Trong Le, Xuan-Son Vu, Fani Deligianni, Hoang D. Nguyen |
KES | 3 |
| 2023 | Reliable Sound Re-labeling Approach for Bird Sound ClassificationabstractPassive acoustic monitoring has become an effective and economically scalable solution for wildlife monitoring in recent years, especially for population monitoring of endangered birds in hardly accessible and high-elevation areas. As sound recorders are deployed on the field for extended periods (months), continuous sound streams are often complex in nature, with noises and intermittent signals. Generally, it is prohibitively expensive to label bird sounds with the exact onset and offset time, thus, during training, the data is often provided with only presence/absence labeling (weak labeling) that states which bird species are present in each long recording without temporal information. During test time, it is, however, desirable to provide fine-grained detection in short audio segments. To bridge the gap between the difference of weak labels used for training and strong labels used for testing, we propose a re-labeling approach with two stages: (1) we train a model with weak labeling; and (2) using the model obtained from stage 1, we generate labels for short audio segments and retrain the model on short audio segments with the newly generated labels. We applied our approach in both classification and sound event detection and achieved consistently good performance across multiple random seeds. In BirdCLEF 2022, our model ranked in the top 1.1%, the 9th best entry out of the total 807 entries in the competition. Hoang Van Truong, Nghia NVN, Alex To, Duc-Trong Le, Hoang D. Nguyen |
KES | 4 |
| 2023 | Self-MI: Efficient Multimodal Fusion via Self-Supervised Multi-Task Learning with Auxiliary Mutual Information Maximization
Cam-Van Nguyen Thi, Ngoc-Hoa Thi Nguyen, Duc-Trong Le, Quang-Thuy Ha |
PACLIC | 3 |
| 2022 | v3MFND: A Deep Multi-domain Multimodal Fake News Detection Model for Vietnamese
Cam-Van Nguyen Thi, Thanh-Toan Vuong, Duc-Trong Le, Quang-Thuy Ha |
ACIIDS (1) | 3 |
| 2021 | Modular Graph Transformer Networks for Multi-Label Image ClassificationabstractWith the recent advances in graph neural networks, there is a rising number of studies on graph-based multi-label classification with the consideration of object dependencies within visual data. Nevertheless, graph representations can become indistinguishable due to the complex nature of label relationships. We propose a multi-label image classification framework based on graph transformer networks to fully exploit inter-label interactions. The paper presents a modular learning scheme to enhance the classification performance by segregating the computational graph into multiple sub-graphs based on modularity. The proposed approach, named Modular Graph Transformer Networks (MGTN), is capable of employing multiple backbones for better information propagation over different sub-graphs guided by graph transformers and convolutions. We validate our framework on MS-COCO and Fashion550K datasets to demonstrate improvements for multi-label image classification. The source code is available at https://github.com/ReML-AI/MGTN. Hoang D. Nguyen, Xuan-Son Vu, Duc-Trong Le |
AAAI | 3 |
| 2020 | Multimodal Review Generation with Privacy and Fairness AwarenessabstractUsers express their opinions towards entities (e.g., restaurants) via online reviews which can be in diverse forms such as text, ratings, and images.Modeling reviews are advantageous for user behavior understanding which, in turn, supports various user-oriented tasks such as recommendation, sentiment analysis, and review generation.In this paper, we propose MG-PriFair, a multimodal neural-based framework, which generates personalized reviews with privacy and fairness awareness.Motivated by the fact that reviews might contain personal information and sentiment bias, we propose a novel differentially private (dp)-embedding model for training privacy guaranteed embeddings and an evaluation approach for sentiment fairness in the food-review domain.Experiments on our novel review dataset show that MG-PriFair is capable of generating plausibly long reviews while controlling the amount of exploited user data and using the least sentimentbiased word embeddings.To the best of our knowledge, we are the first to bring user privacy and sentiment fairness into the review generation task.The dataset and source codes are available at https Xuan-Son Vu, Thanh-Son Nguyen 0001, Duc-Trong Le, Lili Jiang 0002 |
COLING | 3 |
| 2020 | Introducing a New Dataset for Event Detection in Cybersecurity TextsabstractDetecting cybersecurity events is necessary to keep us informed about the fast growing number of such events reported in text.In this work, we focus on the task of event detection (ED) to identify event trigger words for the cybersecurity domain.In particular, to facilitate the future research, we introduce a new dataset for this problem, characterizing the manual annotation for 30 important cybersecurity event types and a large dataset size to develop deep learning models.Comparing to the prior datasets for this task, our dataset involves more event types and supports the modeling of document-level information to improve the performance.We perform extensive evaluation with the current state-of-the-art methods for ED on the proposed dataset.Our experiments reveal the challenges of cybersecurity ED and present many research opportunities in this area for the future work. Hieu Man Duc Trong, Duc-Trong Le, Amir Pouran Ben Veyseh, Thuat Nguyen, Thien Huu Nguyen |
EMNLP (1) | 2 |
| 2020 | Privacy-Preserving Visual Content Tagging using Graph Transformer NetworksabstractWith the rapid growth of Internet media, content tagging has become an important topic with many multimedia understanding applications, including efficient organisation and search. Nevertheless, existing visual tagging approaches are susceptible to inherent privacy risks in which private information may be exposed unintentionally. The use of anonymisation and privacy-protection methods is desirable, but with the expense of task performance. Therefore, this paper proposes an end-to-end framework (SGTN) using Graph Transformer and Convolutional Networks to significantly improve classification and privacy preservation of visual data. Especially, we employ several mechanisms such as differential privacy based graph construction and noise-induced graph transformation to protect the privacy of knowledge graphs. Our approach unveils new state-of-the-art on MS-COCO dataset in various semi-supervised settings. In addition, we showcase a real experiment in the education domain to address the automation of sensitive document tagging. Experimental results show that our approach achieves an excellent balance of model accuracy and privacy preservation on both public and private datasets. Xuan-Son Vu, Duc-Trong Le, Christoffer Edlund, Lili Jiang 0002, Hoang D. Nguyen |
ACM Multimedia | 2 |
| 2019 | Correlation-Sensitive Next-Basket RecommendationabstractItems adopted by a user over time are indicative of the underlying preferences. We are concerned with learning such preferences from observed sequences of adoptions for recommendation. As multiple items are commonly adopted concurrently, e.g., a basket of grocery items or a sitting of media consumption, we deal with a sequence of baskets as input, and seek to recommend the next basket. Intuitively, a basket tends to contain groups of related items that support particular needs. Instead of recommending items independently for the next basket, we hypothesize that incorporating information on pairwise correlations among items would help to arrive at more coherent basket recommendations. Towards this objective, we develop a hierarchical network architecture codenamed Beacon to model basket sequences. Each basket is encoded taking into account the relative importance of items and correlations among item pairs. This encoding is utilized to infer sequential associations along the basket sequence. Extensive experiments on three public real-life datasets showcase the effectiveness of our approach for the next-basket recommendation problem. Duc-Trong Le, Hady Wirawan Lauw, Yuan Fang 0001 |
IJCAI | 1 |
| 2018 | Modeling Contemporaneous Basket Sequences with Twin Networks for Next-Item RecommendationabstractOur interactions with an application frequently leave a heterogeneous and contemporaneous trail of actions and adoptions (e.g., clicks, bookmarks, purchases). Given a sequence of a particular type (e.g., purchases)-- referred to as the target sequence, we seek to predict the next item expected to appear beyond this sequence. This task is known as next-item recommendation. We hypothesize two means for improvement. First, within each time step, a user may interact with multiple items (a basket), with potential latent associations among them. Second, predicting the next item in the target sequence may be helped by also learning from another supporting sequence (e.g., clicks). We develop three twin network structures modeling the generation of both target and support basket sequences. One based on "Siamese networks" facilitates full sharing of parameters between the two sequence types. The other two based on "fraternal networks" facilitate partial sharing of parameters. Experiments on real-world datasets show significant improvements upon baselines relying on one sequence type. Duc-Trong Le, Hady Wirawan Lauw, Yuan Fang 0001 |
IJCAI | 1 |
| 2017 | Basket-Sensitive Personalized Item RecommendationabstractPersonalized item recommendation is useful in narrowing down the list of options provided to a user. In this paper, we address the problem scenario where the user is currently holding a basket of items, and the task is to recommend an item to be added to the basket. Here, we assume that items currently in a basket share some association based on an underlying latent need, e.g., ingredients to prepare some dish, spare parts of some device. Thus, it is important that a recommended item is relevant not only to the user, but also to the existing items in the basket. Towards this goal, we propose two approaches. First, we explore a factorization-based model called BFM that incorporates various types of associations involving the user, the target item to be recommended, and the items currently in the basket. Second, based on our observation that various recommendations towards constructing the same basket should have similar likelihoods, we propose another model called CBFM that further incorporates basket-level constraints. Experiments on three real-life datasets from different domains empirically validate these models against baselines based on matrix factorization and association rules. Duc-Trong Le, Hady Wirawan Lauw, Yuan Fang 0001 |
IJCAI | 1 |
| 2016 | Modeling Sequential Preferences with Dynamic User and Context Factors
Duc-Trong Le, Yuan Fang 0001, Hady Wirawan Lauw |
ECML/PKDD (2) | 1 |
| 2012 | A Model of Vietnamese Person Named Entity Question Answering System
Mai-Vu Tran, Duc-Trong Le, Xuan-Tu Tran, Tien-Tung Nguyen |
PACLIC | 2 |