VLDB 2026 Research / reviewers in the wild / expert
Thong Nguyen 0003
dblp:29/5255-3
· DBLP profile ↗
11ranked-venue papers
6as first author
11since 2021 · last 2024
0000-0001-7447-7416ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | READ-PVLA: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language ModelingabstractFully fine-tuning pretrained large-scale transformer models has become a popular paradigm for video-language modeling tasks, such as temporal language grounding and video-language summarization. With a growing number of tasks and limited training data, such full fine-tuning approach leads to costly model storage and unstable training. To overcome these shortcomings, we introduce lightweight adapters to the pre-trained model and only update them at fine-tuning time. However, existing adapters fail to capture intrinsic temporal relations among video frames or textual words. Moreover, they neglect the preservation of critical task-related information that flows from the raw video-language input into the adapter’s low-dimensional space. To address these issues, we first propose a novel REcurrent ADapter (READ) that employs recurrent computation to enable temporal modeling capability. Second, we propose Partial Video-Language Alignment (PVLA) objective via the use of partial optimal transport to maintain task-related information flowing into our READ modules. We validate our READ-PVLA framework through extensive experiments where READ-PVLA significantly outperforms all existing fine-tuning strategies on multiple low-resource temporal language grounding and video-language summarization benchmarks. Thong Nguyen 0003, Xiaobao Wu, Xinshuai Dong, Khoi M. Le, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu |
AAAI | 1 |
| 2024 | On the Affinity, Rationality, and Diversity of Hierarchical Topic ModelingabstractHierarchical topic modeling aims to discover latent topics from a corpus and organize them into a hierarchy to understand documents with desirable semantic granularity. However, existing work struggles with producing topic hierarchies of low affinity, rationality, and diversity, which hampers document understanding. To overcome these challenges, we in this paper propose Transport Plan and Context-aware Hierarchical Topic Model (TraCo). Instead of early simple topic dependencies, we propose a transport plan dependency method. It constrains dependencies to ensure their sparsity and balance, and also regularizes topic hierarchy building with them. This improves affinity and diversity of hierarchies. We further propose a context-aware disentangled decoder. Rather than previously entangled decoding, it distributes different semantic granularity to topics at different levels by disentangled decoding. This facilitates the rationality of hierarchies. Experiments on benchmark datasets demonstrate that our method surpasses state-of-the-art baselines, effectively improving the affinity, rationality, and diversity of hierarchical topic modeling with better performance on downstream tasks. Xiaobao Wu, Fengjun Pan, Thong Nguyen 0003, Yichao Feng, Chaoqun Liu, Cong-Duy Nguyen, Anh Tuan Luu |
AAAI | 3 |
| 2024 | Meta-optimized Angular Margin Contrastive Framework for Video-Language Representation Learning
Thong Nguyen 0003, Yi Bin, Xiaobao Wu, Xinshuai Dong, Khoi Le, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu |
ECCV (80) | 1 |
| 2024 | Encoding and Controlling Global Semantics for Long-form Video Question AnsweringabstractSeeking answers effectively for long videos is essential to build video question answering (videoQA) systems.Previous methods adaptively select frames and regions from long videos to save computations.However, this fails to reason over the whole sequence of video, leading to sub-optimal performance.To address this problem, we introduce a state space layer (SSL) into multi-modal Transformer to efficiently integrate global semantics of the video, which mitigates the video information loss caused by frame and region selection modules.Our SSL includes a gating unit to enable controllability over the flow of global semantics into visual representations.To further enhance the controllability, we introduce a cross-modal compositional congruence (C 3 ) objective to encourage global semantics aligned with the question.To rigorously evaluate longform videoQA capacity, we construct two new benchmarks Ego-QA and MAD-QA featuring videos of considerably long length, i.e. 17.5 minutes and 1.9 hours, respectively.Extensive experiments demonstrate the superiority of our framework on these new as well as existing datasets.The code, model, and data have been made available at nguyent- thong.github.io/Long_form_VideoQA. Thong Nguyen 0003, Xiaobao Wu, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu |
EMNLP | 1 |
| 2024 | KDMCSE: Knowledge Distillation Multimodal Sentence Embeddings with Adaptive Angular margin Contrastive LearningabstractCong-Duy Nguyen, Thong Nguyen, Xiaobao Wu, Anh Tuan Luu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Cong-Duy Nguyen, Thong Nguyen 0003, Xiaobao Wu, Anh Tuan Luu |
NAACL-HLT | 2 |
| 2024 | FASTopic: Pretrained Transformer is a Fast, Adaptive, Stable, and Transferable Topic ModelabstractTopic models have been evolving rapidly over the years, from conventional to recent neural models. However, existing topic models generally struggle with either effectiveness, efficiency, or stability, highly impeding their practical applications. In this paper, we propose FASTopic, a fast, adaptive, stable, and transferable topic model. FASTopic follows a new paradigm: Dual Semantic-relation Reconstruction (DSR). Instead of previous conventional, VAE-based, or clustering-based methods, DSR directly models the semantic relations among document embeddings from a pretrained Transformer and learnable topic and word embeddings. By reconstructing through these semantic relations, DSR discovers latent topics. This brings about a neat and efficient topic modeling framework. We further propose a novel Embedding Transport Plan (ETP) method. Rather than early straightforward approaches, ETP explicitly regularizes the semantic relations as optimal transport plans. This addresses the relation bias issue and thus leads to effective topic modeling. Extensive experiments on benchmark datasets demonstrate that our FASTopic shows superior effectiveness, efficiency, adaptivity, stability, and transferability, compared to state-of-the-art baselines across various scenarios. Xiaobao Wu, Thong Nguyen 0003, Delvin Zhang, William Yang Wang, Anh Tuan Luu |
NeurIPS | 2 |
| 2023 | InfoCTM: A Mutual Information Maximization Perspective of Cross-Lingual Topic ModelingabstractCross-lingual topic models have been prevalent for cross-lingual text analysis by revealing aligned latent topics. However, most existing methods suffer from producing repetitive topics that hinder further analysis and performance decline caused by low-coverage dictionaries. In this paper, we propose the Cross-lingual Topic Modeling with Mutual Information (InfoCTM). Instead of the direct alignment in previous work, we propose a topic alignment with mutual information method. This works as a regularization to properly align topics and prevent degenerate topic representations of words, which mitigates the repetitive topic issue. To address the low-coverage dictionary issue, we further propose a cross-lingual vocabulary linking method that finds more linked cross-lingual words for topic alignment beyond the translations of a given dictionary. Extensive experiments on English, Chinese, and Japanese datasets demonstrate that our method outperforms state-of-the-art baselines, producing more coherent, diverse, and well-aligned topics and showing better transferability for cross-lingual classification tasks. Xiaobao Wu, Xinshuai Dong, Thong Nguyen 0003, Chaoqun Liu, Liangming Pan, Anh Tuan Luu |
AAAI | 3 |
| 2023 | Expand BERT Representation with Visual Information via Grounded Language Learning with Multimodal Partial AlignmentabstractLanguage models have been supervised with both language-only objective and visual grounding in existing studies of visual-grounded language learning. However, due to differences in the distribution and scale of visual-grounded datasets and language corpora, the language model tends to mix up the context of the tokens that occurred in the grounded data with those that do not. As a result, during representation learning, there is a mismatch between the visual information and the contextual meaning of the sentence. To overcome this limitation, we propose GroundedBERT - a grounded language learning method that enhances the BERT representation with visually grounded information. GroundedBERT comprises two components: (i) the original BERT which captures the contextual representation of words learned from the language corpora, and (ii) a visual grounding module which captures visual information learned from visual-grounded datasets. Moreover, we employ Optimal Transport (OT), specifically its partial variant, to solve the fractional alignment problem between the two modalities. Our proposed method significantly outperforms the baseline language models on various language tasks of the GLUE and SQuAD datasets. Cong-Duy Nguyen, The-Anh Vu-Le, Thong Nguyen 0003, Thanh Tho Quan, Anh Tuan Luu |
ACM Multimedia | 3 |
| 2022 | Adaptive Contrastive Learning on Multimodal Transformer for Review Helpfulness PredictionabstractModern Review Helpfulness Prediction systems are dependent upon multiple modalities, typically texts and images.Unfortunately, those contemporary approaches pay scarce attention to polish representations of cross-modal relations and tend to suffer from inferior optimization.This might cause harm to model's predictions in numerous cases.To overcome the aforementioned issues, we propose Multimodal Contrastive Learning for Multimodal Review Helpfulness Prediction (MRHP) problem, concentrating on mutual information between input modalities to explicitly elaborate cross-modal relations.In addition, we introduce Adaptive Weighting scheme for our contrastive learning approach in order to increase flexibility in optimization.Lastly, we propose Multimodal Interaction module to address the unalignment nature of multimodal data, thereby assisting the model in producing more reasonable multimodal representations.Experimental results show that our method outperforms prior baselines and achieves state-of-the-art results on two publicly available benchmark datasets for MRHP problem. Thong Nguyen 0003, Xiaobao Wu, Anh Tuan Luu, Zhen Hai, Lidong Bing |
EMNLP | 1 |
| 2021 | Enriching and Controlling Global Semantics for Text SummarizationabstractRecently, Transformer-based models have been proven effective in the abstractive summarization task by creating fluent and informative summaries.Nevertheless, these models still suffer from the short-range dependency problem, causing them to produce summaries that miss the key points of document.In this paper, we attempt to address this issue by introducing a neural topic model empowered with normalizing flow to capture the global semantics of the document, which are then integrated into the summarization model.In addition, to avoid the overwhelming effect of global semantics on contextualized representation, we introduce a mechanism to control the amount of global semantics supplied to the text generation module.Our method outperforms state-of-the-art summarization models on five common text summarization datasets, namely CNN/DailyMail, XSum, Reddit TIFU, arXiv, and PubMed. Thong Nguyen 0003, Anh Tuan Luu, Truc Lu, Thanh Tho Quan |
EMNLP (1) | 1 |
| 2021 | Contrastive Learning for Neural Topic ModelabstractRecent empirical studies show that adversarial topic models (ATM) can successfully capture semantic patterns of the document by differentiating a document with another dissimilar sample. However, utilizing that discriminative-generative architecture has two important drawbacks: (1) the architecture does not relate similar documents, which has the same document-word distribution of salient words; (2) it restricts the ability to integrate external information, such as sentiments of the document, which has been shown to benefit the training of neural topic model. To address those issues, we revisit the adversarial topic architecture in the view point of mathematical analysis, propose a novel approach to re-formulate discriminative goal as an optimization problem, and design a novel sampling method which facilitates the integration of external variables. The reformulation encourages the model to incorporate the relations among similar samples and enforces the constraint on the similarity among dissimilar ones; while the sampling method, which is based on the internal input and reconstructed output, helps inform the model of salient words contributing to the main topic. Experimental results show that our framework outperforms other state-of-the-art neural topic models in three common benchmark datasets that belong to various domains, vocabulary sizes, and document lengths in terms of topic coherence. Thong Nguyen 0003, Anh Tuan Luu |
NeurIPS | 1 |