EDBT 2026 Demo / reviewers in the wild / expert
Xiaobao Wu
dblp:249/8429
· DBLP profile ↗
32ranked-venue papers
10as first author
29since 2021 · last 2026
0000-0003-0076-3924ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 10 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and FailuresabstractYi Hu, Jiaqi Gu, Ruxin Wang, Zijun Yao, Hao Peng, Xiaobao Wu, Jianhui Chen, Muhan Zhang, Liangming Pan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zijun Yao 0002, Hao Peng 0015, Xiaobao Wu, Muhan Zhang, Liangming Pan |
ACL (1) | 6 |
| 2026 | Learning Uncertainty from Sequential Internal Dispersion in Large Language ModelsabstractUncertainty estimation is a promising approach to detect hallucinations in large language models (LLMs).Recent approaches commonly depend on model internal states to estimate uncertainty.However, they suffer from strict assumptions on how hidden states should evolve across layers, and from information loss by solely focusing on last or mean tokens.To address these issues, we present Sequential Internal Variance Representation (SIVR), a supervised hallucination detection framework that leverages token-wise, layer-wise features derived from hidden states.SIVR adopts a more basic assumption that uncertainty manifests in the degree of dispersion or variance of internal representations across layers, rather than relying on specific assumptions, which makes the method model and task agnostic.It additionally aggregates the full sequence of per-token variance features, learning temporal patterns indicative of factual errors and thereby preventing information loss.Experimental results demonstrate SIVR consistently outperforms strong baselines.Most importantly, SIVR enjoys stronger generalisation and avoids relying on large training sets, highlighting the potential for practical deployment. Ponhvoan Srey, Xiaobao Wu, Cong-Duy Nguyen, Anh Tuan Luu |
ACL (1) | 2 |
| 2026 | MUR: Momentum Uncertainty guided Reasoning for Large Language ModelsabstractHang Yan, Fangzhi Xu, Rongman Xu, Yifei Li, Jian Zhang, Haoran Luo, Xiaobao Wu, Anh Tuan Luu, Haiteng Zhao, Qika Lin, Jun Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hang Yan 0010, Fangzhi Xu, Rongman Xu, Yifei Li 0006, Jian Zhang 0087, Haoran Luo 0001, Xiaobao Wu, Anh Tuan Luu, Haiteng Zhao, Qika Lin, Jun Liu 0002 |
ACL (1) | 7 |
| 2026 | Read as You See: Guiding Unimodal LLMs for Low-Resource Explainable Harmful Meme DetectionabstractDetecting harmful memes is crucial for safeguarding the integrity and harmony of online environments, yet existing detection methods are often resource-intensive, inflexible, and lacking explainability, limiting their applicability in assisting real-world web content moderation. We propose U-CoT+, a resource-efficient framework that prioritizes accessibility, flexibility and transparency in harmful meme detection by fully harnessing the capabilities of lightweight unimodal large language models (LLMs). Instead of directly prompting or fine-tuning large multimodal models (LMMs) as black-box classifiers, we avoid immediate reasoning over complex visual inputs but decouple meme content recognition from meme harmfulness analysis through a high-fidelity meme-to-text pipeline, which collaborates lightweight LMMs and LLMs to convert multimodal memes into natural language descriptions that preserve critical visual information, thus enabling text-only LLMs to "see" memes by "reading". Grounded in textual inputs, we further guide unimodal LLMs' reasoning under zero-shot Chain-of-Thoughts (CoT) prompting with targeted, interpretable, context-aware, and easily obtained human-crafted guidelines, thus providing accountable step-by-step rationales, while enabling flexible and efficient adaptation to diverse sociocultural criteria of harmfulness. Extensive experiments on seven benchmark datasets show that U-CoT+ achieves performance comparable to resource-intensive baselines, highlighting its effectiveness and potential as a scalable, explainable, and low-resource solution to support harmful meme detection. Fengjun Pan, Xiaobao Wu, Thanh Tho Quan, Anh Tuan Luu |
WWW | 2 |
| 2026 | Protecting Your Customized LLM Systems From Backdoored Instructions With Metacognitive Probing
Shuai Zhao 0007, Zhongliang Guo 0001, Xiaobao Wu, Yanhao Jia, Luwei Xiao, Anh Tuan Luu |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Multi-Scale Contrastive Learning for Video Temporal GroundingabstractTemporal grounding, which localizes video moments related to a natural language query, is a core problem of vision-language learning and video understanding. To encode video moments of varying lengths, recent methods employ a multi-level structure known as a feature pyramid. In this structure, lower levels concentrate on short-range video moments, while higher levels address long-range moments. Because higher levels experience downsampling to accommodate increasing moment length, their capacity to capture information is reduced and consequently leads to degraded information in moment representations. To resolve this problem, we propose a contrastive learning framework to capture salient semantics among video moments. Our key methodology is to leverage samples from the feature space emanating from multiple stages of the video encoder itself requiring neither data augmentation nor online memory banks to obtain positive and negative samples. To enable such an extension, we introduce a sampling process to draw multiple video moments corresponding to a common query. Subsequently, by utilizing these moments' representations across video encoder layers, we instantiate a novel form of multi-scale and cross-scale contrastive learning that links local short-range video moments with global long-range video moments. Extensive experiments demonstrate the effectiveness of our framework for not only long-form but also short-form video grounding. Thong Thanh Nguyen, Yi Bin, Xiaobao Wu, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu |
AAAI | 3 |
| 2025 | Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph GenerationabstractTo equip artificial intelligence with a comprehensive understanding towards a temporal world, video and 4D panoptic scene graph generation abstracts visual data into nodes to represent entities and edges to capture temporal relations. Existing methods encode entity masks tracked across temporal dimensions (mask tubes), then predict their relations with temporal pooling operation, which does not fully utilize the motion indicative of the entities' relation. To overcome this limitation, we introduce a contrastive representation learning framework that focuses on motion pattern for temporal scene graph generation. Firstly, our framework encourages the model to learn close representations for mask tubes of similar subject-relation-object triplets. Secondly, we seek to push apart mask tubes from their temporally shuffled versions. Moreover, we also learn distant representations for mask tubes belonging to the same video but different triplets. Extensive experiments show that our motion-aware contrastive framework significantly improves state-of-the-art methods on both video and 4D datasets. Thong Thanh Nguyen, Xiaobao Wu, Yi Bin, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu |
AAAI | 2 |
| 2025 | AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World KnowledgeabstractXiaobao Wu, Liangming Pan, Yuxi Xie, Ruiwen Zhou, Shuai Zhao, Yubo Ma, Mingzhe Du, Rui Mao, Anh Tuan Luu, William Yang Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiaobao Wu, Liangming Pan, Yuxi Xie, Ruiwen Zhou, Shuai Zhao 0007, Yubo Ma, Mingzhe Du, Rui Mao 0010, Anh Tuan Luu, William Yang Wang |
ACL (1) | 1 |
| 2025 | RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World ScenariosabstractRuiwen Zhou, Wenyue Hua, Liangming Pan, Sitao Cheng, Xiaobao Wu, En Yu, William Yang Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ruiwen Zhou, Wenyue Hua, Liangming Pan, Sitao Cheng, Xiaobao Wu, En Yu, William Yang Wang |
ACL (1) | 5 |
| 2025 | Zero-to-Strong Generalization: Eliciting Strong Capabilities of Large Language Models Iteratively without Gold LabelsabstractLarge Language Models (LLMs) have demonstrated remarkable performance through supervised fine-tuning or in-context learning using gold labels. However, this paradigm is limited by the availability of gold labels, while in certain scenarios, LLMs may need to perform tasks that are too complex for humans to provide such labels. To tackle this challenge, this study explores whether solely utilizing unlabeled data can elicit strong model capabilities. We propose a new paradigm termed zero-to-strong generalization. We iteratively prompt LLMs to annotate unlabeled data and retain high-quality labels by filtering. Surprisingly, we obverse that this iterative process gradually unlocks LLMs’ potential on downstream tasks. Our experiments on extensive classification and reasoning tasks confirm the effectiveness of our proposed framework. Our analysis indicates that this paradigm is effective for both in-context learning and fine-tuning, and for various model sizes. Chaoqun Liu, Qin Chao, Wenxuan Zhang 0001, Xiaobao Wu, Boyang Li 0001, Anh Tuan Luu, Lidong Bing |
COLING | 4 |
| 2025 | Unsupervised Hallucination Detection by Inspecting Reasoning ProcessesabstractUnsupervised hallucination detection aims to identify hallucinated content generated by large language models (LLMs) without relying on labeled data.While unsupervised methods have gained popularity by eliminating laborintensive human annotations, they frequently rely on proxy signals unrelated to factual correctness.This misalignment biases detection probes toward superficial or non-truthrelated aspects, limiting generalizability across datasets and scenarios.To overcome these limitations, we propose IRIS, an unsupervised hallucination detection framework, leveraging internal representations intrinsic to factual correctness.IRIS prompts the LLM to carefully verify the truthfulness of a given statement, and obtain its contextualized embedding as informative features for training.Meanwhile, the uncertainty of each response is considered a soft pseudolabel for truthfulness.Experimental results demonstrate that IRIS consistently outperforms existing unsupervised methods.Our approach is fully unsupervised, computationally low cost, and works well even with few training data, making it suitable for real-time detection.1 Ponhvoan Srey, Xiaobao Wu, Anh Tuan Luu |
EMNLP | 2 |
| 2025 | KBQA-o1: Agentic Knowledge Base Question Answering with Monte Carlo Tree SearchabstractKnowledge Base Question Answering (KBQA) aims to answer natural language questions with a large-scale structured knowledge base (KB). Despite advancements with large language models (LLMs), KBQA still faces challenges in weak KB awareness, imbalance between effectiveness and efficiency, and high reliance on annotated data. To address these challenges, we propose KBQA-o1, a novel agentic KBQA method with Monte Carlo Tree Search (MCTS). It introduces a ReAct-based agent process for stepwise logical form generation with KB environment exploration. Moreover, it employs MCTS, a heuristic search method driven by policy and reward models, to balance agentic exploration’s performance and search space. With heuristic exploration, KBQA-o1 generates high-quality annotations for further improvement by incremental fine-tuning. Experimental results show that KBQA-o1 outperforms previous low-resource KBQA methods with limited annotated data, boosting Llama-3.1-8B model’s GrailQA F1 performance to 78.5% compared to 48.5% of the previous sota method with GPT-3.5-turbo. Our code is publicly available. Haoran Luo 0001, Haihong E, Yikai Guo, Qika Lin, Xiaobao Wu, Xinyu Mu, Meina Song, Yifan Zhu 0001, Anh Tuan Luu |
ICML | 5 |
| 2025 | Aspect-Based Summarization with Self-Aspect Retrieval Enhanced GenerationabstractAspect-based summarization aims to generate summaries tailored to specific aspects, addressing the resource constraints and limited generalizability of traditional summarization approaches. Recently, large language models have shown promise in this task without the need for training. However, they rely excessively on prompt engineering and face token limits and hallucination challenges, especially with in-context learning. To address these challenges, in this paper, we propose a novel framework for aspect-based summarization: Self-Aspect Retrieval Enhanced Summary Generation. Rather than relying solely on in-context learning, given an aspect, we employ an embedding-driven retrieval mechanism to identify its relevant text segments. This approach extracts the pertinent content while avoiding unnecessary details, thereby mitigating the challenge of token limits. Moreover, our framework optimizes token usage by deleting unrelated parts of the text and ensuring that the model generates output strictly based on the given aspect. With extensive experiments on benchmark datasets, we demonstrate that our framework not only achieves superior performance but also effectively mitigates the token limitation problem. Yichao Feng, Shuai Zhao 0007, Yueqiu Li, Luwei Xiao, Xiaobao Wu, Anh Tuan Luu |
IJCNN | 5 |
| 2025 | Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual AugmentationabstractCong-Duy T Nguyen, Xiaobao Wu, Thong Thanh Nguyen, Shuai Zhao, Khoi M. Le, Nguyen Viet Anh, Feng Yichao, Anh Tuan Luu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Cong-Duy Nguyen, Xiaobao Wu, Thong Thanh Nguyen, Shuai Zhao 0007, Khoi M. Le, Yichao Feng, Anh Tuan Luu |
NAACL (Long Papers) | 2 |
| 2025 | HyperGraphRAG: Retrieval-Augmented Generation via Hypergraph-Structured Knowledge RepresentationabstractStandard Retrieval-Augmented Generation (RAG) relies on chunk-based retrieval, whereas GraphRAG advances this approach by graph-based knowledge representation. However, existing graph-based RAG approaches are constrained by binary relations, as each edge in an ordinary graph connects only two entities, limiting their ability to represent the n-ary relations (n >= 2) in real-world knowledge. In this work, we propose HyperGraphRAG, the first hypergraph-based RAG method that represents n-ary relational facts via hyperedges. HyperGraphRAG consists of a comprehensive pipeline, including knowledge hypergraph construction, retrieval, and generation. Experiments across medicine, agriculture, computer science, and law demonstrate that HyperGraphRAG outperforms both standard RAG and previous graph-based RAG methods in answer accuracy, retrieval efficiency, and generation quality. Haoran Luo 0001, Haihong E, Guanting Chen 0004, Yandan Zheng, Xiaobao Wu, Yikai Guo, Qika Lin, Yu Feng 0015, Zemin Kuang, Meina Song, Yifan Zhu 0001, Anh Tuan Luu |
NeurIPS | 5 |
| 2024 | READ-PVLA: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language ModelingabstractFully fine-tuning pretrained large-scale transformer models has become a popular paradigm for video-language modeling tasks, such as temporal language grounding and video-language summarization. With a growing number of tasks and limited training data, such full fine-tuning approach leads to costly model storage and unstable training. To overcome these shortcomings, we introduce lightweight adapters to the pre-trained model and only update them at fine-tuning time. However, existing adapters fail to capture intrinsic temporal relations among video frames or textual words. Moreover, they neglect the preservation of critical task-related information that flows from the raw video-language input into the adapter’s low-dimensional space. To address these issues, we first propose a novel REcurrent ADapter (READ) that employs recurrent computation to enable temporal modeling capability. Second, we propose Partial Video-Language Alignment (PVLA) objective via the use of partial optimal transport to maintain task-related information flowing into our READ modules. We validate our READ-PVLA framework through extensive experiments where READ-PVLA significantly outperforms all existing fine-tuning strategies on multiple low-resource temporal language grounding and video-language summarization benchmarks. Thong Nguyen 0003, Xiaobao Wu, Xinshuai Dong, Khoi M. Le, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu |
AAAI | 2 |
| 2024 | On the Affinity, Rationality, and Diversity of Hierarchical Topic ModelingabstractHierarchical topic modeling aims to discover latent topics from a corpus and organize them into a hierarchy to understand documents with desirable semantic granularity. However, existing work struggles with producing topic hierarchies of low affinity, rationality, and diversity, which hampers document understanding. To overcome these challenges, we in this paper propose Transport Plan and Context-aware Hierarchical Topic Model (TraCo). Instead of early simple topic dependencies, we propose a transport plan dependency method. It constrains dependencies to ensure their sparsity and balance, and also regularizes topic hierarchy building with them. This improves affinity and diversity of hierarchies. We further propose a context-aware disentangled decoder. Rather than previously entangled decoding, it distributes different semantic granularity to topics at different levels by disentangled decoding. This facilitates the rationality of hierarchies. Experiments on benchmark datasets demonstrate that our method surpasses state-of-the-art baselines, effectively improving the affinity, rationality, and diversity of hierarchical topic modeling with better performance on downstream tasks. Xiaobao Wu, Fengjun Pan, Thong Nguyen 0003, Yichao Feng, Chaoqun Liu, Cong-Duy Nguyen, Anh Tuan Luu |
AAAI | 1 |
| 2024 | Meta-optimized Angular Margin Contrastive Framework for Video-Language Representation Learning
Thong Nguyen 0003, Yi Bin, Xiaobao Wu, Xinshuai Dong, Khoi Le, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu |
ECCV (80) | 3 |
| 2024 | Encoding and Controlling Global Semantics for Long-form Video Question AnsweringabstractSeeking answers effectively for long videos is essential to build video question answering (videoQA) systems.Previous methods adaptively select frames and regions from long videos to save computations.However, this fails to reason over the whole sequence of video, leading to sub-optimal performance.To address this problem, we introduce a state space layer (SSL) into multi-modal Transformer to efficiently integrate global semantics of the video, which mitigates the video information loss caused by frame and region selection modules.Our SSL includes a gating unit to enable controllability over the flow of global semantics into visual representations.To further enhance the controllability, we introduce a cross-modal compositional congruence (C 3 ) objective to encourage global semantics aligned with the question.To rigorously evaluate longform videoQA capacity, we construct two new benchmarks Ego-QA and MAD-QA featuring videos of considerably long length, i.e. 17.5 minutes and 1.9 hours, respectively.Extensive experiments demonstrate the superiority of our framework on these new as well as existing datasets.The code, model, and data have been made available at nguyent- thong.github.io/Long_form_VideoQA. Thong Nguyen 0003, Xiaobao Wu, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu |
EMNLP | 3 |
| 2024 | Are LLMs Good Zero-Shot Fallacy Classifiers?abstractFallacies are defective arguments with faulty reasoning.Detecting and classifying them is a crucial NLP task to prevent misinformation, manipulative claims, and biased decisions.However, existing fallacy classifiers are limited by the requirement for sufficient labeled data for training, which hinders their out-of-distribution (OOD) generalization abilities.In this paper, we focus on leveraging Large Language Models (LLMs) for zero-shot fallacy classification.To elicit fallacy-related knowledge and reasoning abilities of LLMs, we propose diverse single-round and multi-round prompting schemes, applying different taskspecific instructions such as extraction, summarization, and Chain-of-Thought reasoning.With comprehensive experiments on benchmark datasets, we suggest that LLMs could be potential zero-shot fallacy classifiers.In general, LLMs under single-round prompting schemes have achieved acceptable zeroshot performances compared to the best fullshot baselines and can outperform them in all OOD inference scenarios and some opendomain tasks.Our novel multi-round prompting schemes can effectively bring about more improvements, especially for small LLMs.Our analysis further underlines the future research on zero-shot fallacy classification.Codes and data are available at: https://github.com/ panFJCharlotte98/Fallacy_Detection. Fengjun Pan, Xiaobao Wu, Zongrui Li 0001, Anh Tuan Luu |
EMNLP | 2 |
| 2024 | AKEW: Assessing Knowledge Editing in the WildabstractKnowledge editing injects knowledge updates into language models to keep them correct and up-to-date.However, its current evaluations deviate significantly from practice: their knowledge updates solely consist of structured facts derived from meticulously crafted datasets, instead of practical sources-unstructured texts like news articles, and they often overlook practical real-world knowledge updates.To address these issues, in this paper we propose AKEW (Assessing Knowledge Editing in the Wild), a new practical benchmark for knowledge editing.AKEW fully covers three editing settings of knowledge updates: structured facts, unstructured texts as facts, and extracted triplets.It further introduces new datasets featuring both counterfactual and real-world knowledge updates.Through extensive experiments, we demonstrate the considerable gap between state-of-the-art knowledge-editing methods and practical scenarios.Our analyses further highlight key insights to motivate future research for practical knowledge editing 1 . Xiaobao Wu, Liangming Pan, William Yang Wang, Anh Tuan Luu |
EMNLP | 1 |
| 2024 | Topic Modeling as Multi-Objective Contrastive OptimizationabstractRecent representation learning approaches enhance neural topic models by optimizing the weighted linear combination of the evidence lower bound (ELBO) of the log-likelihood and the contrastive learning objective that contrasts pairs of input documents. However, document-level contrastive learning might capture low-level mutual information, such as word ratio, which disturbs topic modeling. Moreover, there is a potential conflict between the ELBO loss that memorizes input details for better reconstruction quality, and the contrastive loss which attempts to learn topic representations that generalize among input documents. To address these issues, we first introduce a novel contrastive learning method oriented towards sets of topic vectors to capture useful semantics that are shared among a set of input documents. Secondly, we explicitly cast contrastive topic modeling as a gradient-based multi-objective optimization problem, with the goal of achieving a Pareto stationary solution that balances the trade-off between the ELBO and the contrastive objective. Extensive experiments demonstrate that our framework consistently produces higher-performing neural topic models in terms of topic coherence, topic diversity, and downstream performance. Thong Thanh Nguyen, Xiaobao Wu, Xinshuai Dong, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu |
ICLR | 2 |
| 2024 | KDMCSE: Knowledge Distillation Multimodal Sentence Embeddings with Adaptive Angular margin Contrastive LearningabstractCong-Duy Nguyen, Thong Nguyen, Xiaobao Wu, Anh Tuan Luu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Cong-Duy Nguyen, Thong Nguyen 0003, Xiaobao Wu, Anh Tuan Luu |
NAACL-HLT | 3 |
| 2024 | FASTopic: Pretrained Transformer is a Fast, Adaptive, Stable, and Transferable Topic ModelabstractTopic models have been evolving rapidly over the years, from conventional to recent neural models. However, existing topic models generally struggle with either effectiveness, efficiency, or stability, highly impeding their practical applications. In this paper, we propose FASTopic, a fast, adaptive, stable, and transferable topic model. FASTopic follows a new paradigm: Dual Semantic-relation Reconstruction (DSR). Instead of previous conventional, VAE-based, or clustering-based methods, DSR directly models the semantic relations among document embeddings from a pretrained Transformer and learnable topic and word embeddings. By reconstructing through these semantic relations, DSR discovers latent topics. This brings about a neat and efficient topic modeling framework. We further propose a novel Embedding Transport Plan (ETP) method. Rather than early straightforward approaches, ETP explicitly regularizes the semantic relations as optimal transport plans. This addresses the relation bias issue and thus leads to effective topic modeling. Extensive experiments on benchmark datasets demonstrate that our FASTopic shows superior effectiveness, efficiency, adaptivity, stability, and transferability, compared to state-of-the-art baselines across various scenarios. Xiaobao Wu, Thong Nguyen 0003, Delvin Zhang, William Yang Wang, Anh Tuan Luu |
NeurIPS | 1 |
| 2023 | InfoCTM: A Mutual Information Maximization Perspective of Cross-Lingual Topic ModelingabstractCross-lingual topic models have been prevalent for cross-lingual text analysis by revealing aligned latent topics. However, most existing methods suffer from producing repetitive topics that hinder further analysis and performance decline caused by low-coverage dictionaries. In this paper, we propose the Cross-lingual Topic Modeling with Mutual Information (InfoCTM). Instead of the direct alignment in previous work, we propose a topic alignment with mutual information method. This works as a regularization to properly align topics and prevent degenerate topic representations of words, which mitigates the repetitive topic issue. To address the low-coverage dictionary issue, we further propose a cross-lingual vocabulary linking method that finds more linked cross-lingual words for topic alignment beyond the translations of a given dictionary. Extensive experiments on English, Chinese, and Japanese datasets demonstrate that our method outperforms state-of-the-art baselines, producing more coherent, diverse, and well-aligned topics and showing better transferability for cross-lingual classification tasks. Xiaobao Wu, Xinshuai Dong, Thong Nguyen 0003, Chaoqun Liu, Liangming Pan, Anh Tuan Luu |
AAAI | 1 |
| 2023 | Fact-Checking Complex Claims with Program-Guided ReasoningabstractLiangming Pan, Xiaobao Wu, Xinyuan Lu, Anh Tuan Luu, William Yang Wang, Min-Yen Kan, Preslav Nakov. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Liangming Pan, Xiaobao Wu, Xinyuan Lu, Anh Tuan Luu, William Yang Wang, Min-Yen Kan, Preslav Nakov |
ACL (1) | 2 |
| 2023 | Effective Neural Topic Modeling with Embedding Clustering RegularizationabstractTopic models have been prevalent for decades with various applications. However, existing topic models commonly suffer from the notorious topic collapsing: discovered topics semantically collapse towards each other, leading to highly repetitive topics, insufficient topic discovery, and damaged model interpretability. In this paper, we propose a new neural topic model, Embedding Clustering Regularization Topic Model (ECRTM). Besides the existing reconstruction error, we propose a novel Embedding Clustering Regularization (ECR), which forces each topic embedding to be the center of a separately aggregated word embedding cluster in the semantic space. This enables each produced topic to contain distinct word semantics, which alleviates topic collapsing. Regularized by ECR, our ECRTM generates diverse and coherent topics together with high-quality topic distributions of documents. Extensive experiments on benchmark datasets demonstrate that ECRTM effectively addresses the topic collapsing issue and consistently surpasses state-of-the-art baselines in terms of topic quality, topic distributions of documents, and downstream classification tasks. Xiaobao Wu, Xinshuai Dong, Thong Thanh Nguyen, Anh Tuan Luu |
ICML | 1 |
| 2022 | Adaptive Contrastive Learning on Multimodal Transformer for Review Helpfulness PredictionabstractModern Review Helpfulness Prediction systems are dependent upon multiple modalities, typically texts and images.Unfortunately, those contemporary approaches pay scarce attention to polish representations of cross-modal relations and tend to suffer from inferior optimization.This might cause harm to model's predictions in numerous cases.To overcome the aforementioned issues, we propose Multimodal Contrastive Learning for Multimodal Review Helpfulness Prediction (MRHP) problem, concentrating on mutual information between input modalities to explicitly elaborate cross-modal relations.In addition, we introduce Adaptive Weighting scheme for our contrastive learning approach in order to increase flexibility in optimization.Lastly, we propose Multimodal Interaction module to address the unalignment nature of multimodal data, thereby assisting the model in producing more reasonable multimodal representations.Experimental results show that our method outperforms prior baselines and achieves state-of-the-art results on two publicly available benchmark datasets for MRHP problem. Thong Nguyen 0003, Xiaobao Wu, Anh Tuan Luu, Zhen Hai, Lidong Bing |
EMNLP | 2 |
| 2022 | Mitigating Data Sparsity for Short Text Topic Modeling by Topic-Semantic Contrastive LearningabstractTo overcome the data sparsity issue in short text topic modeling, existing methods commonly rely on data augmentation or the data characteristic of short texts to introduce more word co-occurrence information.However, most of them do not make full use of the augmented data or the data characteristic: they insufficiently learn the relations among samples in data, leading to dissimilar topic distributions of semantically similar text pairs.To better address data sparsity, in this paper we propose a novel short text topic modeling framework, Topic-Semantic Contrastive Topic Model (TSCTM).To sufficiently model the relations among samples, we employ a new contrastive learning method with efficient positive and negative sampling strategies based on topic semantics.This contrastive learning method refines the representations, enriches the learning signals, and thus mitigates the sparsity issue.Extensive experimental results show that our TSCTM outperforms state-ofthe-art baselines regardless of the data augmentation availability, producing high-quality topics and topic distributions. 1 Xiaobao Wu, Anh Tuan Luu, Xinshuai Dong |
EMNLP | 1 |
| 2020 | Short Text Topic Modeling with Topic Distribution Quantization and Negative Sampling DecoderabstractTopic models have been prevailing for many years on discovering latent semantics while modeling long documents.However, for short texts they generally suffer from data sparsity because of extremely limited word cooccurrences; thus tend to yield repetitive or trivial topics with low quality.In this paper, to address this issue, we propose a novel neural topic model in the framework of autoencoding with a new topic distribution quantization approach generating peakier distributions that are more appropriate for modeling short texts.Besides the encoding, to tackle this issue in terms of decoding, we further propose a novel negative sampling decoder learning from negative samples to avoid yielding repetitive topics.We observe that our model can highly improve short text topic modeling performance.Through extensive experiments on real-world datasets, we demonstrate our model can outperform both strong traditional and neural baselines under extreme data sparsity scenes, producing high-quality topics. Xiaobao Wu, Chunping Li, Yan Zhu 0007, Yishu Miao |
EMNLP (1) | 1 |
| 2020 | Learning Multilingual Topics with Neural Variational Inference
Xiaobao Wu, Chunping Li, Yan Zhu 0007, Yishu Miao |
NLPCC (1) | 1 |
| 2019 | Short Text Topic Modeling with Flexible Word PatternsabstractSince effective semantic representations are utilized in many practical applications, inferring discriminative and coherent latent topics from short texts is a critical and basic task. Traditional topic models like Probabilistic Latent Semantic Analysis (PLSA) and Latent Dirichlet Allocation (LDA) behave not well on short texts due to data sparsity problem. One novel model called Biterm Topic Model (BTM) which models unordered word-pairs (i.e., biterms) from whole corpus was proposed to solve this problem. However, both the performance and efficiency of BTM are reduced because of many irrelevant and useless biterms. In this paper, we propose a Multiterm Topic Model (MTM) for short text topic modeling. MTM extracts variable-length and more correlative word patterns (i.e., multiterms) from the whole corpus. By directly modeling the generative process of multiterms, MTM can infer the word distributions of each topic and the topic distribution of each short text to alleviate the sparsity problem in short text modeling. With the the proper amount of flexible multiterms, learning process of MTM is enhanced. Through extensive experiments on two real-world short text collections, we show that MTM is more efficient and outperforms the baseline models in terms of topic coherence and text classification. Xiaobao Wu, Chunping Li |
IJCNN | 1 |