Cong-Duy Nguyen

dblp:323/7497 · also Cong-Duy T. Nguyen · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-0931-460XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Learning Uncertainty from Sequential Internal Dispersion in Large Language Models
abstract
Uncertainty estimation is a promising approach to detect hallucinations in large language models (LLMs).Recent approaches commonly depend on model internal states to estimate uncertainty.However, they suffer from strict assumptions on how hidden states should evolve across layers, and from information loss by solely focusing on last or mean tokens.To address these issues, we present Sequential Internal Variance Representation (SIVR), a supervised hallucination detection framework that leverages token-wise, layer-wise features derived from hidden states.SIVR adopts a more basic assumption that uncertainty manifests in the degree of dispersion or variance of internal representations across layers, rather than relying on specific assumptions, which makes the method model and task agnostic.It additionally aggregates the full sequence of per-token variance features, learning temporal patterns indicative of factual errors and thereby preventing information loss.Experimental results demonstrate SIVR consistently outperforms strong baselines.Most importantly, SIVR enjoys stronger generalisation and avoids relying on large training sets, highlighting the potential for practical deployment.
Ponhvoan Srey, Xiaobao Wu, Cong-Duy Nguyen, Anh Tuan Luu
ACL (1)3
2025 Multi-Scale Contrastive Learning for Video Temporal Grounding
abstract
Temporal grounding, which localizes video moments related to a natural language query, is a core problem of vision-language learning and video understanding. To encode video moments of varying lengths, recent methods employ a multi-level structure known as a feature pyramid. In this structure, lower levels concentrate on short-range video moments, while higher levels address long-range moments. Because higher levels experience downsampling to accommodate increasing moment length, their capacity to capture information is reduced and consequently leads to degraded information in moment representations. To resolve this problem, we propose a contrastive learning framework to capture salient semantics among video moments. Our key methodology is to leverage samples from the feature space emanating from multiple stages of the video encoder itself requiring neither data augmentation nor online memory banks to obtain positive and negative samples. To enable such an extension, we introduce a sampling process to draw multiple video moments corresponding to a common query. Subsequently, by utilizing these moments' representations across video encoder layers, we instantiate a novel form of multi-scale and cross-scale contrastive learning that links local short-range video moments with global long-range video moments. Extensive experiments demonstrate the effectiveness of our framework for not only long-form but also short-form video grounding.
Thong Thanh Nguyen, Yi Bin, Xiaobao Wu, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu
AAAI5
2025 Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation
abstract
To equip artificial intelligence with a comprehensive understanding towards a temporal world, video and 4D panoptic scene graph generation abstracts visual data into nodes to represent entities and edges to capture temporal relations. Existing methods encode entity masks tracked across temporal dimensions (mask tubes), then predict their relations with temporal pooling operation, which does not fully utilize the motion indicative of the entities' relation. To overcome this limitation, we introduce a contrastive representation learning framework that focuses on motion pattern for temporal scene graph generation. Firstly, our framework encourages the model to learn close representations for mask tubes of similar subject-relation-object triplets. Secondly, we seek to push apart mask tubes from their temporally shuffled versions. Moreover, we also learn distant representations for mask tubes belonging to the same video but different triplets. Extensive experiments show that our motion-aware contrastive framework significantly improves state-of-the-art methods on both video and 4D datasets.
Thong Thanh Nguyen, Xiaobao Wu, Yi Bin, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu
AAAI4
2025 Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual Augmentation
abstract
Cong-Duy T Nguyen, Xiaobao Wu, Thong Thanh Nguyen, Shuai Zhao, Khoi M. Le, Nguyen Viet Anh, Feng Yichao, Anh Tuan Luu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Cong-Duy Nguyen, Xiaobao Wu, Thong Thanh Nguyen, Shuai Zhao 0007, Khoi M. Le, Yichao Feng, Anh Tuan Luu
NAACL (Long Papers)1
2024 READ-PVLA: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling
abstract
Fully fine-tuning pretrained large-scale transformer models has become a popular paradigm for video-language modeling tasks, such as temporal language grounding and video-language summarization. With a growing number of tasks and limited training data, such full fine-tuning approach leads to costly model storage and unstable training. To overcome these shortcomings, we introduce lightweight adapters to the pre-trained model and only update them at fine-tuning time. However, existing adapters fail to capture intrinsic temporal relations among video frames or textual words. Moreover, they neglect the preservation of critical task-related information that flows from the raw video-language input into the adapter’s low-dimensional space. To address these issues, we first propose a novel REcurrent ADapter (READ) that employs recurrent computation to enable temporal modeling capability. Second, we propose Partial Video-Language Alignment (PVLA) objective via the use of partial optimal transport to maintain task-related information flowing into our READ modules. We validate our READ-PVLA framework through extensive experiments where READ-PVLA significantly outperforms all existing fine-tuning strategies on multiple low-resource temporal language grounding and video-language summarization benchmarks.
Thong Nguyen 0003, Xiaobao Wu, Xinshuai Dong, Khoi M. Le, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu
AAAI6
2024 On the Affinity, Rationality, and Diversity of Hierarchical Topic Modeling
abstract
Hierarchical topic modeling aims to discover latent topics from a corpus and organize them into a hierarchy to understand documents with desirable semantic granularity. However, existing work struggles with producing topic hierarchies of low affinity, rationality, and diversity, which hampers document understanding. To overcome these challenges, we in this paper propose Transport Plan and Context-aware Hierarchical Topic Model (TraCo). Instead of early simple topic dependencies, we propose a transport plan dependency method. It constrains dependencies to ensure their sparsity and balance, and also regularizes topic hierarchy building with them. This improves affinity and diversity of hierarchies. We further propose a context-aware disentangled decoder. Rather than previously entangled decoding, it distributes different semantic granularity to topics at different levels by disentangled decoding. This facilitates the rationality of hierarchies. Experiments on benchmark datasets demonstrate that our method surpasses state-of-the-art baselines, effectively improving the affinity, rationality, and diversity of hierarchical topic modeling with better performance on downstream tasks.
Xiaobao Wu, Fengjun Pan, Thong Nguyen 0003, Yichao Feng, Chaoqun Liu, Cong-Duy Nguyen, Anh Tuan Luu
AAAI6
2024 Meta-optimized Angular Margin Contrastive Framework for Video-Language Representation Learning
Thong Nguyen 0003, Yi Bin, Xiaobao Wu, Xinshuai Dong, Khoi Le, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu
ECCV (80)7
2024 Encoding and Controlling Global Semantics for Long-form Video Question Answering
abstract
Seeking answers effectively for long videos is essential to build video question answering (videoQA) systems.Previous methods adaptively select frames and regions from long videos to save computations.However, this fails to reason over the whole sequence of video, leading to sub-optimal performance.To address this problem, we introduce a state space layer (SSL) into multi-modal Transformer to efficiently integrate global semantics of the video, which mitigates the video information loss caused by frame and region selection modules.Our SSL includes a gating unit to enable controllability over the flow of global semantics into visual representations.To further enhance the controllability, we introduce a cross-modal compositional congruence (C 3 ) objective to encourage global semantics aligned with the question.To rigorously evaluate longform videoQA capacity, we construct two new benchmarks Ego-QA and MAD-QA featuring videos of considerably long length, i.e. 17.5 minutes and 1.9 hours, respectively.Extensive experiments demonstrate the superiority of our framework on these new as well as existing datasets.The code, model, and data have been made available at nguyent- thong.github.io/Long_form_VideoQA.
Thong Nguyen 0003, Xiaobao Wu, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu
EMNLP4
2024 Topic Modeling as Multi-Objective Contrastive Optimization
abstract
Recent representation learning approaches enhance neural topic models by optimizing the weighted linear combination of the evidence lower bound (ELBO) of the log-likelihood and the contrastive learning objective that contrasts pairs of input documents. However, document-level contrastive learning might capture low-level mutual information, such as word ratio, which disturbs topic modeling. Moreover, there is a potential conflict between the ELBO loss that memorizes input details for better reconstruction quality, and the contrastive loss which attempts to learn topic representations that generalize among input documents. To address these issues, we first introduce a novel contrastive learning method oriented towards sets of topic vectors to capture useful semantics that are shared among a set of input documents. Secondly, we explicitly cast contrastive topic modeling as a gradient-based multi-objective optimization problem, with the goal of achieving a Pareto stationary solution that balances the trade-off between the ELBO and the contrastive objective. Extensive experiments demonstrate that our framework consistently produces higher-performing neural topic models in terms of topic coherence, topic diversity, and downstream performance.
Thong Thanh Nguyen, Xiaobao Wu, Xinshuai Dong, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu
ICLR4
2024 KDMCSE: Knowledge Distillation Multimodal Sentence Embeddings with Adaptive Angular margin Contrastive Learning
abstract
Cong-Duy Nguyen, Thong Nguyen, Xiaobao Wu, Anh Tuan Luu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Cong-Duy Nguyen, Thong Nguyen 0003, Xiaobao Wu, Anh Tuan Luu
NAACL-HLT1
2024 Supervised learning models for social bot detection: Literature review and benchmark
Hoang-Dung Nguyen, Cong-Duy Nguyen, Phong T. To, Danh H. Nguyen, Huy Nguyen-Gia, Long H. Tran, Anh Q. Tran, An Dang-Hieu, Anh Nguyen-Duc 0001, Thanh Tho Quan
Expert Syst. Appl.3
2023 Expand BERT Representation with Visual Information via Grounded Language Learning with Multimodal Partial Alignment
abstract
Language models have been supervised with both language-only objective and visual grounding in existing studies of visual-grounded language learning. However, due to differences in the distribution and scale of visual-grounded datasets and language corpora, the language model tends to mix up the context of the tokens that occurred in the grounded data with those that do not. As a result, during representation learning, there is a mismatch between the visual information and the contextual meaning of the sentence. To overcome this limitation, we propose GroundedBERT - a grounded language learning method that enhances the BERT representation with visually grounded information. GroundedBERT comprises two components: (i) the original BERT which captures the contextual representation of words learned from the language corpora, and (ii) a visual grounding module which captures visual information learned from visual-grounded datasets. Moreover, we employ Optimal Transport (OT), specifically its partial variant, to solve the fractional alignment problem between the two modalities. Our proposed method significantly outperforms the baseline language models on various language tasks of the GLUE and SQuAD datasets.
Cong-Duy Nguyen, The-Anh Vu-Le, Thong Nguyen 0003, Thanh Tho Quan, Anh Tuan Luu
ACM Multimedia1