Anh Tuan Luu

dblp:81/8329 · also Luu Anh Tuan · DBLP profile ↗
← Back
121ranked-venue papers
7as first author
85since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 102 · 5 first-author · 72 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 1 first-author · 17 since 2021Databases, data management, data science and information retrieval · 15 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Stimuli to Minds: Enhancing Psychological Reasoning in LLMs via Bilateral Reinforcement Learning
abstract
Large Language Models show promise in emotion understanding, social reasoning, and empathy, yet struggle with psychologically grounded tasks requiring inference of implicit mental states in complex, socially and contextually ambiguous settings. These limitations stem from lacking theory-aligned supervision and difficulty capturing nuanced mental processes in real-world narratives. To bridge this gap, we leverage expert-labeled scenarios and propose a trajectory-aware reinforcement learning framework imitating expert psychological reasoning. By integrating real-world stimuli with structured reasoning guidance, our approach enables compact models to internalize social-cognitive principles, perform nuanced inference, and support continual self-improvement. Experiments across benchmarks show expert-level interpretive capability across psychological tasks.
Yichao Feng, Haoran Luo 0001, Lang Feng 0007, Shuai Zhao 0007, Anh Tuan Luu
AAAI5
2026 Towards Fast and Accurate Modeling for Cross-Lingual Label Projection
abstract
Information extraction (IE) systems rely on structured data for training, but such annotated data is highly imbalanced across languages, with low-resource languages receiving little attention.Label projection techniques aim to bridge this gap by transferring structured annotations from high-resource to low-resource languages.However, existing methods are either inaccurate or too slow for large-scale use.This work aims to address this problem by developing a more effective method that remains sufficiently efficient for large-scale projection.In particular, we propose to synthesize alignment sequence pairs and fine-tune an encoder model with span alignment objective, while controlling data influence during training.Experimental results across 50+ languages show that our framework consistently outperforms previous state-of-the-art methods while maintaining fast inference speed.In addition, we introduce EXP -the first benchmark for explicit evaluation of label projection, thereby reducing confounders and non-determinism in method assessment.
Thang Le, Huy Huu Nguyen, Anh Tuan Luu, Thamar Solorio, Thien Huu Nguyen
ACL (1)3
2026 Learning Uncertainty from Sequential Internal Dispersion in Large Language Models
abstract
Uncertainty estimation is a promising approach to detect hallucinations in large language models (LLMs).Recent approaches commonly depend on model internal states to estimate uncertainty.However, they suffer from strict assumptions on how hidden states should evolve across layers, and from information loss by solely focusing on last or mean tokens.To address these issues, we present Sequential Internal Variance Representation (SIVR), a supervised hallucination detection framework that leverages token-wise, layer-wise features derived from hidden states.SIVR adopts a more basic assumption that uncertainty manifests in the degree of dispersion or variance of internal representations across layers, rather than relying on specific assumptions, which makes the method model and task agnostic.It additionally aggregates the full sequence of per-token variance features, learning temporal patterns indicative of factual errors and thereby preventing information loss.Experimental results demonstrate SIVR consistently outperforms strong baselines.Most importantly, SIVR enjoys stronger generalisation and avoids relying on large training sets, highlighting the potential for practical deployment.
Ponhvoan Srey, Xiaobao Wu, Cong-Duy Nguyen, Anh Tuan Luu
ACL (1)4
2026 MUR: Momentum Uncertainty guided Reasoning for Large Language Models
abstract
Hang Yan, Fangzhi Xu, Rongman Xu, Yifei Li, Jian Zhang, Haoran Luo, Xiaobao Wu, Anh Tuan Luu, Haiteng Zhao, Qika Lin, Jun Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hang Yan 0010, Fangzhi Xu, Rongman Xu, Yifei Li 0006, Jian Zhang 0087, Haoran Luo 0001, Xiaobao Wu, Anh Tuan Luu, Haiteng Zhao, Qika Lin, Jun Liu 0002
ACL (1)8
2026 Read as You See: Guiding Unimodal LLMs for Low-Resource Explainable Harmful Meme Detection
abstract
Detecting harmful memes is crucial for safeguarding the integrity and harmony of online environments, yet existing detection methods are often resource-intensive, inflexible, and lacking explainability, limiting their applicability in assisting real-world web content moderation. We propose U-CoT+, a resource-efficient framework that prioritizes accessibility, flexibility and transparency in harmful meme detection by fully harnessing the capabilities of lightweight unimodal large language models (LLMs). Instead of directly prompting or fine-tuning large multimodal models (LMMs) as black-box classifiers, we avoid immediate reasoning over complex visual inputs but decouple meme content recognition from meme harmfulness analysis through a high-fidelity meme-to-text pipeline, which collaborates lightweight LMMs and LLMs to convert multimodal memes into natural language descriptions that preserve critical visual information, thus enabling text-only LLMs to "see" memes by "reading". Grounded in textual inputs, we further guide unimodal LLMs' reasoning under zero-shot Chain-of-Thoughts (CoT) prompting with targeted, interpretable, context-aware, and easily obtained human-crafted guidelines, thus providing accountable step-by-step rationales, while enabling flexible and efficient adaptation to diverse sociocultural criteria of harmfulness. Extensive experiments on seven benchmark datasets show that U-CoT+ achieves performance comparable to resource-intensive baselines, highlighting its effectiveness and potential as a scalable, explainable, and low-resource solution to support harmful meme detection.
Fengjun Pan, Xiaobao Wu, Thanh Tho Quan, Anh Tuan Luu
WWW4
2026 ADGT: Enhancing 3D human pose estimation with attention-driven graph-transformers
abstract
2D-to-3D lifting is a fundamental approach in 3D human pose estimation (3DHPE). This task is crucial in applications, including motion analysis and virtual reality. While Graph Convolutional Networks (GCNs) have demonstrated effectiveness in capturing spatial relationships in human skeletons, they suffer from over-smoothing and limited receptive fields. Transformer-based models provide global context but struggle with local feature extraction and computational efficiency. To address these challenges, we propose ADGT, a novel parallel GCN-transformer architecture combining the strengths of both approaches. Our method introduces three key innovations: Hop-Wise Scalable Adaptive GCN to refine local feature extraction, Attention-Based Local Feature Extractor to enhance the integration of local and global representations, and Register-Based Transformer Enhancement to improve feature separation. Extensive experiments on Human3.6M and MPI-INF-3DHP datasets demonstrate ADGT achieves state-of-the-art performance among frame-based methods while maintaining computational efficiency. These results highlight the potential of ADGT for real-time applications requiring accurate and efficient 3DHPE. The code is available at https://github.com/sYANGunique1111/ADGT .
Anh Tuan Luu, Xuan Son Nguyen, Aymeric Histace, Bart Jansen 0001, Hichem Sahli
J. Vis. Commun. Image Represent.2
2026 UniFLE: Uniform Fusion of Multiple LoRA Experts for Backdoor Defense in Large Language Models
abstract
Large language models (LLMs), which serve as a bridge between pre-training and task-specific adaptation, achieve state-of-the-art performance across several downstream tasks through full-parameter fine-tuning (FPFT). However, with the continuous growth of model parameter scales, FPFT requires substantial computational resources, which limits its practicality. Consequently, there is a growing shift toward parameter-efficient fine-tuning (PEFT) methods that update only a limited subset of model parameters, markedly reducing resource consumption. Although this paradigm fosters accelerated research advancements, it also introduces security risks. Empirical studies demonstrate that if LLM weights are backdoored, the backdoors can still be activated even after fine-tuning leveraging PEFT algorithms. To address the aforementioned issue, in this paper, we introduce a novel Uniform Fusion of multiple LoRA Experts algorithm, named UniFLE, designed to defend against backdoor attacks. Specifically, the UniFLE algorithm pioneers the insertion of multiple LoRA experts into the MLP blocks to expand the updatable feature subspace during fine-tuning, and fuses all experts to decouple backdoor features. Additionally, to enhance the diversity of LoRA experts, we introduce a diversity regularization loss that constrains correlations between different experts and encourages them to learn more novel features. The theoretical analysis demonstrates that the UniFLE algorithm effectively decreases the mutual information between the model's intermediate representations and the backdoor representations. To validate the effectiveness of the UniFLE algorithm, we conduct experiments across four tasks, four state-of-the-art LLMs, and three backdoor attack methods. The results consistently demonstrate that our UniFLE algorithm effectively defends against backdoor attacks while preserving model performance. We aspire for our method to enhance model security and contribute to the advancement of the LLM community.
Shuai Zhao 0007, Qika Lin, Yanhao Jia, Anh Tuan Luu
IEEE Trans. Dependable Secur. Comput.6
2026 Protecting Your Customized LLM Systems From Backdoored Instructions With Metacognitive Probing
Shuai Zhao 0007, Zhongliang Guo 0001, Xiaobao Wu, Yanhao Jia, Luwei Xiao, Anh Tuan Luu
IEEE Trans. Inf. Forensics Secur.7
2026 Introduction to the Special Issue on Transformers
Feng Xia 0001, Tyler Derr, Anh Tuan Luu, Richa Singh 0001, Aline Villavicencio
ACM Trans. Intell. Syst. Technol.3
2025 Multi-Scale Contrastive Learning for Video Temporal Grounding
abstract
Temporal grounding, which localizes video moments related to a natural language query, is a core problem of vision-language learning and video understanding. To encode video moments of varying lengths, recent methods employ a multi-level structure known as a feature pyramid. In this structure, lower levels concentrate on short-range video moments, while higher levels address long-range moments. Because higher levels experience downsampling to accommodate increasing moment length, their capacity to capture information is reduced and consequently leads to degraded information in moment representations. To resolve this problem, we propose a contrastive learning framework to capture salient semantics among video moments. Our key methodology is to leverage samples from the feature space emanating from multiple stages of the video encoder itself requiring neither data augmentation nor online memory banks to obtain positive and negative samples. To enable such an extension, we introduce a sampling process to draw multiple video moments corresponding to a common query. Subsequently, by utilizing these moments' representations across video encoder layers, we instantiate a novel form of multi-scale and cross-scale contrastive learning that links local short-range video moments with global long-range video moments. Extensive experiments demonstrate the effectiveness of our framework for not only long-form but also short-form video grounding.
Thong Thanh Nguyen, Yi Bin, Xiaobao Wu, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu
AAAI7
2025 Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation
abstract
To equip artificial intelligence with a comprehensive understanding towards a temporal world, video and 4D panoptic scene graph generation abstracts visual data into nodes to represent entities and edges to capture temporal relations. Existing methods encode entity masks tracked across temporal dimensions (mask tubes), then predict their relations with temporal pooling operation, which does not fully utilize the motion indicative of the entities' relation. To overcome this limitation, we introduce a contrastive representation learning framework that focuses on motion pattern for temporal scene graph generation. Firstly, our framework encourages the model to learn close representations for mask tubes of similar subject-relation-object triplets. Secondly, we seek to push apart mask tubes from their temporally shuffled versions. Moreover, we also learn distant representations for mask tubes belonging to the same video but different triplets. Extensive experiments show that our motion-aware contrastive framework significantly improves state-of-the-art methods on both video and 4D datasets.
Thong Thanh Nguyen, Xiaobao Wu, Yi Bin, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu
AAAI6
2025 FineReason: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle Solving
abstract
Guizhen Chen, Weiwen Xu, Hao Zhang, Hou Pong Chan, Chaoqun Liu, Lidong Bing, Deli Zhao, Anh Tuan Luu, Yu Rong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Guizhen Chen, Weiwen Xu, Hao Zhang 0048, Hou Pong Chan, Chaoqun Liu, Lidong Bing, Deli Zhao, Anh Tuan Luu, Yu Rong 0001
ACL (1)8
2025 LongRecipe: Recipe for Efficient Long Context Generalization in Large Language Models
abstract
Large language models (LLMs) face significant challenges in handling long-context tasks because of their limited effective context window size during pretraining, which restricts their ability to generalize over extended sequences. Meanwhile, extending the context window in LLMs through post-pretraining is highly resource-intensive.To address this, we introduce LongRecipe, an efficient training strategy for extending the context window of LLMs, including impactful token analysis, position index transformation, and training optimization strategies. It simulates long-sequence inputs while maintaining training efficiency and significantly improves the model’s understanding of long-range dependencies. Experiments on three types of LLMs show that LongRecipe can utilize long sequences while requiring only 30% of the target context window size, and reduces computational training resource over 85% compared to full sequence training. Furthermore, LongRecipe also preserves the original LLM’s capabilities in general tasks. Ultimately, we can extend effective context window of open-source LLMs from 8k to 128k, achieving performance close to GPT-4 with just one day of dedicated training using a single GPU with 80G memory.Our code is released at https://github.com/zhiyuanhubj/LongRecipe.
Jinman Zhao, Suyuchen Wang, WangYan WangYan, Wei Shen 0005, Qing Gu 0001, Anh Tuan Luu, See-Kiong Ng, Zhiwei Jiang 0001, Bryan Hooi
ACL (1)8
2025 AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge
abstract
Xiaobao Wu, Liangming Pan, Yuxi Xie, Ruiwen Zhou, Shuai Zhao, Yubo Ma, Mingzhe Du, Rui Mao, Anh Tuan Luu, William Yang Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xiaobao Wu, Liangming Pan, Yuxi Xie, Ruiwen Zhou, Shuai Zhao 0007, Yubo Ma, Mingzhe Du, Rui Mao 0010, Anh Tuan Luu, William Yang Wang
ACL (1)9
2025 Zero-to-Strong Generalization: Eliciting Strong Capabilities of Large Language Models Iteratively without Gold Labels
abstract
Large Language Models (LLMs) have demonstrated remarkable performance through supervised fine-tuning or in-context learning using gold labels. However, this paradigm is limited by the availability of gold labels, while in certain scenarios, LLMs may need to perform tasks that are too complex for humans to provide such labels. To tackle this challenge, this study explores whether solely utilizing unlabeled data can elicit strong model capabilities. We propose a new paradigm termed zero-to-strong generalization. We iteratively prompt LLMs to annotate unlabeled data and retain high-quality labels by filtering. Surprisingly, we obverse that this iterative process gradually unlocks LLMs’ potential on downstream tasks. Our experiments on extensive classification and reasoning tasks confirm the effectiveness of our proposed framework. Our analysis indicates that this paradigm is effective for both in-context learning and fine-tuning, and for various model sizes.
Chaoqun Liu, Qin Chao, Wenxuan Zhang 0001, Xiaobao Wu, Boyang Li 0001, Anh Tuan Luu, Lidong Bing
COLING6
2025 Unsupervised Hallucination Detection by Inspecting Reasoning Processes
abstract
Unsupervised hallucination detection aims to identify hallucinated content generated by large language models (LLMs) without relying on labeled data.While unsupervised methods have gained popularity by eliminating laborintensive human annotations, they frequently rely on proxy signals unrelated to factual correctness.This misalignment biases detection probes toward superficial or non-truthrelated aspects, limiting generalizability across datasets and scenarios.To overcome these limitations, we propose IRIS, an unsupervised hallucination detection framework, leveraging internal representations intrinsic to factual correctness.IRIS prompts the LLM to carefully verify the truthfulness of a given statement, and obtain its contextualized embedding as informative features for training.Meanwhile, the uncertainty of each response is considered a soft pseudolabel for truthfulness.Experimental results demonstrate that IRIS consistently outperforms existing unsupervised methods.Our approach is fully unsupervised, computationally low cost, and works well even with few training data, making it suitable for real-time detection.1
Ponhvoan Srey, Xiaobao Wu, Anh Tuan Luu
EMNLP3
2025 Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective
abstract
Large language model (LLM)-based embedding models, benefiting from large scale pretraining and post-training, have begun to surpass BERT and T5-based models on generalpurpose text embedding tasks such as document retrieval.However, a fundamental limitation of LLM embeddings lies in the unidirectional attention used during autoregressive pre-training, which misaligns with the bidirectional nature of text embedding tasks.To this end, we propose adopting diffusion language models for text embeddings, motivated by their inherent bidirectional architecture and recent success in matching or surpassing LLMs especially on reasoning tasks.We present the first systematic study of the diffusion language embedding model, which outperforms the LLM-based embedding model by 20% on long-document retrieval, 8% on reasoning-intensive retrieval, 2% on instruction-following retrieval, and achieve competitive performance on traditional text embedding benchmarks.Our analysis verifies that bidirectional attention is crucial for encoding global context in long and complex text.
Siyue Zhang, Yilun Zhao 0001, Liyuan Geng, Arman Cohan, Anh Tuan Luu, Chen Zhao 0013
EMNLP5
2025 As Simple as Fine-tuning: LLM Alignment via Bidirectional Negative Feedback Loss
abstract
Direct Preference Optimization (DPO) has emerged as a more computationally efficient alternative to Reinforcement Learning from Human Feedback (RLHF) with Proximal Policy Optimization (PPO), eliminating the need for reward models and online sampling. Despite these benefits, DPO and its variants remain sensitive to hyper-parameters and prone to instability, particularly on mathematical datasets. We argue that these issues arise from the unidirectional likelihood-derivative negative feedback inherent in the log-likelihood loss function. To address this, we propose a novel LLM alignment loss that establishes a stable Bidirectional Negative Feedback (BNF) during optimization. Our proposed BNF loss eliminates the need for pairwise contrastive losses and does not require any extra tunable hyper-parameters or pairwise preference data, streamlining the alignment pipeline to be as simple as supervised fine-tuning. We conduct extensive experiments across two challenging QA benchmarks and four reasoning benchmarks. The experimental results show that BNF achieves comparable performance to the best methods on QA benchmarks, while its performance decrease on the four reasoning benchmarks is significantly lower compared to the best methods, thus striking a better balance between value alignment and reasoning ability. In addition, we further validate the performance of BNF on non-pairwise datasets, and conduct in-depth analysis of log-likelihood and logit shifts across different preference optimization methods. We will release all the source code, checkpoints, and datasets on GitHub.
Feng-Lin Li, Ziqi Jin, Wei Zhang 0218, Anh Tuan Luu
ICLR7
2025 KBQA-o1: Agentic Knowledge Base Question Answering with Monte Carlo Tree Search
abstract
Knowledge Base Question Answering (KBQA) aims to answer natural language questions with a large-scale structured knowledge base (KB). Despite advancements with large language models (LLMs), KBQA still faces challenges in weak KB awareness, imbalance between effectiveness and efficiency, and high reliance on annotated data. To address these challenges, we propose KBQA-o1, a novel agentic KBQA method with Monte Carlo Tree Search (MCTS). It introduces a ReAct-based agent process for stepwise logical form generation with KB environment exploration. Moreover, it employs MCTS, a heuristic search method driven by policy and reward models, to balance agentic exploration’s performance and search space. With heuristic exploration, KBQA-o1 generates high-quality annotations for further improvement by incremental fine-tuning. Experimental results show that KBQA-o1 outperforms previous low-resource KBQA methods with limited annotated data, boosting Llama-3.1-8B model’s GrailQA F1 performance to 78.5% compared to 48.5% of the previous sota method with GPT-3.5-turbo. Our code is publicly available.
Haoran Luo 0001, Haihong E, Yikai Guo, Qika Lin, Xiaobao Wu, Xinyu Mu, Meina Song, Yifan Zhu 0001, Anh Tuan Luu
ICML10
2025 Aspect-Based Summarization with Self-Aspect Retrieval Enhanced Generation
abstract
Aspect-based summarization aims to generate summaries tailored to specific aspects, addressing the resource constraints and limited generalizability of traditional summarization approaches. Recently, large language models have shown promise in this task without the need for training. However, they rely excessively on prompt engineering and face token limits and hallucination challenges, especially with in-context learning. To address these challenges, in this paper, we propose a novel framework for aspect-based summarization: Self-Aspect Retrieval Enhanced Summary Generation. Rather than relying solely on in-context learning, given an aspect, we employ an embedding-driven retrieval mechanism to identify its relevant text segments. This approach extracts the pertinent content while avoiding unnecessary details, thereby mitigating the challenge of token limits. Moreover, our framework optimizes token usage by deleting unrelated parts of the text and ensuring that the model generates output strictly based on the given aspect. With extensive experiments on benchmark datasets, we demonstrate that our framework not only achieves superior performance but also effectively mitigates the token limitation problem.
Yichao Feng, Shuai Zhao 0007, Yueqiu Li, Luwei Xiao, Xiaobao Wu, Anh Tuan Luu
IJCNN6
2025 Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models
abstract
Chaoqun Liu, Wenxuan Zhang, Yiran Zhao, Anh Tuan Luu, Lidong Bing. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Chaoqun Liu, Wenxuan Zhang 0001, Yiran Zhao 0006, Anh Tuan Luu, Lidong Bing
NAACL (Long Papers)4
2025 Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual Augmentation
abstract
Cong-Duy T Nguyen, Xiaobao Wu, Thong Thanh Nguyen, Shuai Zhao, Khoi M. Le, Nguyen Viet Anh, Feng Yichao, Anh Tuan Luu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Cong-Duy Nguyen, Xiaobao Wu, Thong Thanh Nguyen, Shuai Zhao 0007, Khoi M. Le, Yichao Feng, Anh Tuan Luu
NAACL (Long Papers)8
2025 Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization
abstract
Large Language Models (LLMs) generate functionally correct solutions but often fall short in code efficiency, a critical bottleneck for real-world deployment. In this paper, we introduce a novel test-time iterative optimization framework to address this, employing a closed-loop system where LLMs iteratively refine code based on empirical performance feedback from an execution sandbox. We explore three training strategies: Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Group Relative Policy Optimization~(GRPO). Experiments on our Venus dataset and the APPS benchmark show that SFT and DPO rapidly saturate in efficiency gains. In contrast, GRPO, using reinforcement learning (RL) with execution feedback, continuously optimizes code performance, significantly boosting both pass@1 (from 47% to 62%) and the likelihood of outperforming human submissions in efficiency (from 31% to 45%). Our work demonstrates effective test-time code efficiency improvement and critically reveals the power of RL in teaching LLMs to truly self-improve code efficiency. We released our code and data at https://github.com/Elfsong/Afterburner.
Mingzhe Du, Anh Tuan Luu, Yue Liu 0008, Yuhao Qing, Dong Huang 0005, Qian Liu 0033, Zejun Ma 0001, See-Kiong Ng
NeurIPS2
2025 HyperGraphRAG: Retrieval-Augmented Generation via Hypergraph-Structured Knowledge Representation
abstract
Standard Retrieval-Augmented Generation (RAG) relies on chunk-based retrieval, whereas GraphRAG advances this approach by graph-based knowledge representation. However, existing graph-based RAG approaches are constrained by binary relations, as each edge in an ordinary graph connects only two entities, limiting their ability to represent the n-ary relations (n >= 2) in real-world knowledge. In this work, we propose HyperGraphRAG, the first hypergraph-based RAG method that represents n-ary relational facts via hyperedges. HyperGraphRAG consists of a comprehensive pipeline, including knowledge hypergraph construction, retrieval, and generation. Experiments across medicine, agriculture, computer science, and law demonstrate that HyperGraphRAG outperforms both standard RAG and previous graph-based RAG methods in answer accuracy, retrieval efficiency, and generation quality.
Haoran Luo 0001, Haihong E, Guanting Chen 0004, Yandan Zheng, Xiaobao Wu, Yikai Guo, Qika Lin, Yu Feng 0015, Zemin Kuang, Meina Song, Yifan Zhu 0001, Anh Tuan Luu
NeurIPS12
2025 EffiBench-X: A Multi-Language Benchmark for Measuring Efficiency of LLM-Generated Code
abstract
Existing code generation benchmarks primarily evaluate functional correctness, with limited attention to code efficiency, and they are often restricted to a single language such as Python. To address this gap, we introduce EffiBench‑X, the first large‑scale multi‑language benchmark specifically designed for robust efficiency evaluation of LLM‑generated code. EffiBench‑X supports Python, C++, Java, JavaScript, Ruby, and Go, and comprises competitive programming tasks paired with human‑expert solutions as efficiency baselines. Evaluating state‑of‑the‑art LLMs on EffiBench‑X reveals that while models frequently generate functionally correct code, they consistently underperform human experts in efficiency. Even the most efficient LLM‑generated solutions (e.g., Qwen3‑32B) achieve only around 62% of human efficiency on average, with significant language‑specific variation: models tend to perform better in Python, Ruby, and JavaScript than in Java, C++, and Go (e.g., DeepSeek‑R1’s Python code is markedly more efficient than its Java code). These findings highlight the need for research into optimization‑oriented methods to improve the efficiency of LLM‑generated code across diverse languages. The dataset and evaluation infrastructure are publicly available at https://github.com/EffiBench/EffiBench-X.git and https://huggingface.co/datasets/EffiBench/effibench-x.
Yuhao Qing, Boyu Zhu, Mingzhe Du, Zhijiang Guo, Terry Yue Zhuo, Qianru Zhang, Jie Zhang 0050, Heming Cui, Siu-Ming Yiu, Dong Huang 0005, See-Kiong Ng, Anh Tuan Luu
NeurIPS12
2025 Clean-label backdoor attack and defense: An examination of language model vulnerability
Shuai Zhao 0007, Luwei Xiao, Jinming Wen, Anh Tuan Luu
Expert Syst. Appl.5
2025 Dynamic task balancing for joint information extraction
Anran Hao, Jian Su 0002, Siu Cheung Hui, Anh Tuan Luu
Neurocomputing5
2024 From Static to Dynamic: Knowledge Metabolism for Large Language Models
abstract
The immense parameter space of Large Language Models (LLMs) endows them with superior knowledge retention capabilities, allowing them to excel in a variety of natural language processing tasks. However, it also instigates difficulties in consistently tuning LMs to incorporate the most recent knowledge, which may further lead LMs to produce inaccurate and fabricated content. To alleviate this issue, we propose a knowledge metabolism framework for LLMs. This framework proactively sustains the credibility of knowledge through an auxiliary external memory component and directly delivers pertinent knowledge for LM inference, thereby suppressing hallucinations caused by obsolete internal knowledge during the LM inference process. Benchmark experiments demonstrate DynaMind's effectiveness in overcoming this challenge. The code and demo of DynaMind are available at: https://github.com/Elfsong/DynaMind.
Mingzhe Du, Anh Tuan Luu, Bin Ji 0002, See-Kiong Ng
AAAI2
2024 PoetryDiffusion: Towards Joint Semantic and Metrical Manipulation in Poetry Generation
abstract
Controllable text generation is a challenging and meaningful field in natural language generation (NLG). Especially, poetry generation is a typical one with well-defined and strict conditions for text generation which is an ideal playground for the assessment of current methodologies. While prior works succeeded in controlling either semantic or metrical aspects of poetry generation, simultaneously addressing both remains a challenge. In this paper, we pioneer the use of the Diffusion model for generating sonnets and Chinese SongCi poetry to tackle such challenges. In terms of semantics, our PoetryDiffusion model, built upon the Diffusion model, generates entire sentences or poetry by comprehensively considering the entirety of sentence information. This approach enhances semantic expression, distinguishing it from autoregressive and large language models (LLMs). For metrical control, its constraint control module which can be trained individually enables us to flexibly incorporate a novel metrical controller to manipulate and evaluate metrics (format and rhythm). The denoising process in PoetryDiffusion allows for the gradual enhancement of semantics and flexible integration of the metrical controller which can calculate and impose penalties on states that stray significantly from the target control distribution. Experimental results on two datasets demonstrate that our model outperforms existing models in terms of automatic evaluation of semantic, metrical, and overall performance as well as human evaluation. Codes are released to https://github.com/ChorlingLau/PoetryDiffusion.
Chumin Liu, Yue Feng 0002, Anh Tuan Luu, Bryan Hooi
AAAI4
2024 LAMPAT: Low-Rank Adaption for Multilingual Paraphrasing Using Adversarial Training
abstract
Paraphrases are texts that convey the same meaning while using different words or sentence structures. It can be used as an automatic data augmentation tool for many Natural Language Processing tasks, especially when dealing with low-resource languages, where data shortage is a significant problem. To generate a paraphrase in multilingual settings, previous studies have leveraged the knowledge from the machine translation field, i.e., forming a paraphrase through zero-shot machine translation in the same language. Despite good performance on human evaluation, those methods still require parallel translation datasets, thus making them inapplicable to languages that do not have parallel corpora. To mitigate that problem, we proposed the first unsupervised multilingual paraphrasing model, LAMPAT (Low-rank Adaptation for Multilingual Paraphrasing using Adversarial Training), by which monolingual dataset is sufficient enough to generate a human-like and diverse sentence. Throughout the experiments, we found out that our method not only works well for English but can generalize on unseen languages as well. Data and code are available at https://github.com/phkhanhtrinh23/LAMPAT.
Khoi M. Le, Trinh Pham, Thanh Tho Quan, Anh Tuan Luu
AAAI4
2024 READ-PVLA: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling
abstract
Fully fine-tuning pretrained large-scale transformer models has become a popular paradigm for video-language modeling tasks, such as temporal language grounding and video-language summarization. With a growing number of tasks and limited training data, such full fine-tuning approach leads to costly model storage and unstable training. To overcome these shortcomings, we introduce lightweight adapters to the pre-trained model and only update them at fine-tuning time. However, existing adapters fail to capture intrinsic temporal relations among video frames or textual words. Moreover, they neglect the preservation of critical task-related information that flows from the raw video-language input into the adapter’s low-dimensional space. To address these issues, we first propose a novel REcurrent ADapter (READ) that employs recurrent computation to enable temporal modeling capability. Second, we propose Partial Video-Language Alignment (PVLA) objective via the use of partial optimal transport to maintain task-related information flowing into our READ modules. We validate our READ-PVLA framework through extensive experiments where READ-PVLA significantly outperforms all existing fine-tuning strategies on multiple low-resource temporal language grounding and video-language summarization benchmarks.
Thong Nguyen 0003, Xiaobao Wu, Xinshuai Dong, Khoi M. Le, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu
AAAI8
2024 On the Affinity, Rationality, and Diversity of Hierarchical Topic Modeling
abstract
Hierarchical topic modeling aims to discover latent topics from a corpus and organize them into a hierarchy to understand documents with desirable semantic granularity. However, existing work struggles with producing topic hierarchies of low affinity, rationality, and diversity, which hampers document understanding. To overcome these challenges, we in this paper propose Transport Plan and Context-aware Hierarchical Topic Model (TraCo). Instead of early simple topic dependencies, we propose a transport plan dependency method. It constrains dependencies to ensure their sparsity and balance, and also regularizes topic hierarchy building with them. This improves affinity and diversity of hierarchies. We further propose a context-aware disentangled decoder. Rather than previously entangled decoding, it distributes different semantic granularity to topics at different levels by disentangled decoding. This facilitates the rationality of hierarchies. Experiments on benchmark datasets demonstrate that our method surpasses state-of-the-art baselines, effectively improving the affinity, rationality, and diversity of hierarchical topic modeling with better performance on downstream tasks.
Xiaobao Wu, Fengjun Pan, Thong Nguyen 0003, Yichao Feng, Chaoqun Liu, Cong-Duy Nguyen, Anh Tuan Luu
AAAI7
2024 Exploring the Potential of Large Language Models in Computational Argumentation
abstract
Computational argumentation has become an essential tool in various domains, including law, public policy, and artificial intelligence.It is an emerging research field in natural language processing that attracts increasing attention.Research on computational argumentation mainly involves two types of tasks: argument mining and argument generation.As large language models (LLMs) have demonstrated impressive capabilities in understanding context and generating natural language, it is worthwhile to evaluate the performance of LLMs on diverse computational argumentation tasks.This work aims to embark on an assessment of LLMs, such as ChatGPT, Flan models, and LLaMA2 models, in both zero-shot and few-shot settings.We organize existing tasks into six main categories and standardize the format of fourteen openly available datasets.In addition, we present a new benchmark dataset on counter speech generation that aims to holistically evaluate the end-to-end performance of LLMs on argument mining and argument generation.Extensive experiments show that LLMs exhibit commendable performance across most of the datasets, demonstrating their capabilities in the field of argumentation.Our analysis offers valuable suggestions for evaluating computational argumentation and its integration with LLMs in future research endeavors.1
Guizhen Chen, Liying Cheng, Anh Tuan Luu, Lidong Bing
ACL (1)3
2024 UniBridge: A Unified Approach to Cross-Lingual Transfer Learning for Low-Resource Languages
abstract
In this paper, we introduce UniBridge (Cross-Lingual Transfer Learning with Optimized Embeddings and Vocabulary), a comprehensive approach developed to improve the effectiveness of Cross-Lingual Transfer Learning, particularly in languages with limited resources.Our approach tackles two essential elements of a language model: the initialization of embeddings and the optimal vocabulary size.Specifically, we propose a novel embedding initialization method that leverages both lexical and semantic alignment for a language.In addition, we present a method for systematically searching for the optimal vocabulary size, ensuring a balance between model complexity and linguistic coverage.Our experiments across multilingual datasets show that our approach greatly improves the F1-Score in several languages.UniBridge is a robust and adaptable solution for cross-lingual systems in various languages, highlighting the significance of initializing embeddings and choosing the right vocabulary size in cross-lingual environments.
Trinh Pham, Khoi Le, Anh Tuan Luu
ACL (1)3
2024 Meta-optimized Angular Margin Contrastive Framework for Video-Language Representation Learning
Thong Nguyen 0003, Yi Bin, Xiaobao Wu, Xinshuai Dong, Khoi Le, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu
ECCV (80)9
2024 Multi-expert Prompting Improves Reliability, Safety and Usefulness of Large Language Models
abstract
We present Multi-expert Prompting 1 , a novel enhancement of ExpertPrompting (Xu et al., 2023), designed to improve the large language model (LLM) generation.Specifically, it guides an LLM to fulfill an input instruction by simulating multiple experts, aggregating their responses, and selecting the best among individual and aggregated responses.This process is performed in a single chain of thoughts through our seven carefully designed subtasks derived from the Nominal Group Technique (Ven and Delbecq, 1974), a well-established decision-making framework.Our evaluations demonstrate that Multi-expert Prompting significantly outperforms ExpertPrompting and comparable baselines in enhancing the truthfulness, factuality, informativeness, and usefulness of responses while reducing toxicity and hurtfulness.It further achieves state-ofthe-art truthfulness by outperforming the best baseline by 8.69% with ChatGPT.Multi-expert Prompting is efficient, explainable, and highly adaptable to diverse scenarios, eliminating the need for manual prompt construction. * Equal contribution.1 Our codes and data will be made publicly available here.
Do Xuan Long, Duong Ngoc Yen, Anh Tuan Luu, Kenji Kawaguchi, Min-Yen Kan, Nancy F. Chen
EMNLP3
2024 Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration
abstract
While Reinforcement Learning from Human Feedback (RLHF) significantly enhances the generation quality of Large Language Models (LLMs), recent studies have raised concerns regarding the complexity and instability associated with the Proximal Policy Optimization (PPO) algorithm, proposing a series of orderbased alignment methods as viable alternatives.This paper delves into existing order-based methods, unifying them into one framework and examining their inefficiencies in utilizing reward values.Building upon these findings, we propose a new Value-based CaliBration (VCB) method to better align LLMs with human preferences.Experimental results demonstrate that VCB surpasses existing alignment methods on AI assistant and summarization datasets, providing impressive generalizability, robustness, and diversity in different settings.
Feng-Lin Li, Wei Zhang 0218, Anh Tuan Luu
EMNLP6
2024 Encoding and Controlling Global Semantics for Long-form Video Question Answering
abstract
Seeking answers effectively for long videos is essential to build video question answering (videoQA) systems.Previous methods adaptively select frames and regions from long videos to save computations.However, this fails to reason over the whole sequence of video, leading to sub-optimal performance.To address this problem, we introduce a state space layer (SSL) into multi-modal Transformer to efficiently integrate global semantics of the video, which mitigates the video information loss caused by frame and region selection modules.Our SSL includes a gating unit to enable controllability over the flow of global semantics into visual representations.To further enhance the controllability, we introduce a cross-modal compositional congruence (C 3 ) objective to encourage global semantics aligned with the question.To rigorously evaluate longform videoQA capacity, we construct two new benchmarks Ego-QA and MAD-QA featuring videos of considerably long length, i.e. 17.5 minutes and 1.9 hours, respectively.Extensive experiments demonstrate the superiority of our framework on these new as well as existing datasets.The code, model, and data have been made available at nguyent- thong.github.io/Long_form_VideoQA.
Thong Nguyen 0003, Xiaobao Wu, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu
EMNLP6
2024 Are LLMs Good Zero-Shot Fallacy Classifiers?
abstract
Fallacies are defective arguments with faulty reasoning.Detecting and classifying them is a crucial NLP task to prevent misinformation, manipulative claims, and biased decisions.However, existing fallacy classifiers are limited by the requirement for sufficient labeled data for training, which hinders their out-of-distribution (OOD) generalization abilities.In this paper, we focus on leveraging Large Language Models (LLMs) for zero-shot fallacy classification.To elicit fallacy-related knowledge and reasoning abilities of LLMs, we propose diverse single-round and multi-round prompting schemes, applying different taskspecific instructions such as extraction, summarization, and Chain-of-Thought reasoning.With comprehensive experiments on benchmark datasets, we suggest that LLMs could be potential zero-shot fallacy classifiers.In general, LLMs under single-round prompting schemes have achieved acceptable zeroshot performances compared to the best fullshot baselines and can outperform them in all OOD inference scenarios and some opendomain tasks.Our novel multi-round prompting schemes can effectively bring about more improvements, especially for small LLMs.Our analysis further underlines the future research on zero-shot fallacy classification.Codes and data are available at: https://github.com/ panFJCharlotte98/Fallacy_Detection.
Fengjun Pan, Xiaobao Wu, Zongrui Li 0001, Anh Tuan Luu
EMNLP4
2024 AKEW: Assessing Knowledge Editing in the Wild
abstract
Knowledge editing injects knowledge updates into language models to keep them correct and up-to-date.However, its current evaluations deviate significantly from practice: their knowledge updates solely consist of structured facts derived from meticulously crafted datasets, instead of practical sources-unstructured texts like news articles, and they often overlook practical real-world knowledge updates.To address these issues, in this paper we propose AKEW (Assessing Knowledge Editing in the Wild), a new practical benchmark for knowledge editing.AKEW fully covers three editing settings of knowledge updates: structured facts, unstructured texts as facts, and extracted triplets.It further introduces new datasets featuring both counterfactual and real-world knowledge updates.Through extensive experiments, we demonstrate the considerable gap between state-of-the-art knowledge-editing methods and practical scenarios.Our analyses further highlight key insights to motivate future research for practical knowledge editing 1 .
Xiaobao Wu, Liangming Pan, William Yang Wang, Anh Tuan Luu
EMNLP4
2024 Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
abstract
In-context learning, a paradigm bridging the gap between pre-training and fine-tuning, has demonstrated high efficacy in several NLP tasks, especially in few-shot settings. Despite being widely applied, in-context learning is vulnerable to malicious attacks. In this work, we raise security concerns regarding this paradigm. Our studies demonstrate that an attacker can manipulate the behavior of large language models by poisoning the demonstration context, without the need for fine-tuning the model. Specifically, we design a new backdoor attack method, named ICLAttack, to target large language models based on in-context learning. Our method encompasses two types of attacks: poisoning demonstration examples and poisoning demonstration prompts, which can make models behave in alignment with predefined intentions. ICLAttack does not require additional fine-tuning to implant a backdoor, thus preserving the model’s generality. Furthermore, the poisoned examples are correctly labeled, enhancing the natural stealth of our attack method. Extensive experimental results across several language models, ranging in size from 1.3B to 180B parameters, demonstrate the effectiveness of our attack method, exemplified by a high average attack success rate of 95.0% across the three datasets on OPT models.
Shuai Zhao 0007, Meihuizi Jia, Anh Tuan Luu, Fengjun Pan, Jinming Wen
EMNLP3
2024 Topic Modeling as Multi-Objective Contrastive Optimization
abstract
Recent representation learning approaches enhance neural topic models by optimizing the weighted linear combination of the evidence lower bound (ELBO) of the log-likelihood and the contrastive learning objective that contrasts pairs of input documents. However, document-level contrastive learning might capture low-level mutual information, such as word ratio, which disturbs topic modeling. Moreover, there is a potential conflict between the ELBO loss that memorizes input details for better reconstruction quality, and the contrastive loss which attempts to learn topic representations that generalize among input documents. To address these issues, we first introduce a novel contrastive learning method oriented towards sets of topic vectors to capture useful semantics that are shared among a set of input documents. Secondly, we explicitly cast contrastive topic modeling as a gradient-based multi-objective optimization problem, with the goal of achieving a Pareto stationary solution that balances the trade-off between the ELBO and the contrastive objective. Extensive experiments demonstrate that our framework consistently produces higher-performing neural topic models in terms of topic coherence, topic diversity, and downstream performance.
Thong Thanh Nguyen, Xiaobao Wu, Xinshuai Dong, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu
ICLR6
2024 SoVAR: Build Generalizable Scenarios from Accident Reports for Autonomous Driving Testing
abstract
Autonomous driving systems (ADSs) have undergone remarkable development and are increasingly employed in safety-critical applications. However, recently reported data on fatal accidents involving ADSs suggests that the desired level of safety has not yet been fully achieved. Consequently, there is a growing need for more comprehensive and targeted testing approaches to ensure safe driving. Scenarios from real-world accident reports provide valuable resources for ADS testing, including critical scenarios and high-quality seeds. However, existing scenario reconstruction methods from accident reports often exhibit limited accuracy in information extraction. Moreover, due to the diversity and complexity of road environments, matching current accident information with the simulation map data for reconstruction poses significant challenges.
An Guo 0002, Yuan Zhou 0005, Haoxiang Tian 0001, Chunrong Fang, Yunjian Sun, Weisong Sun, Anh Tuan Luu, Yang Liu 0003, Zhenyu Chen 0001
ASE8
2024 SemRoDe: Macro Adversarial Training to Learn Representations that are Robust to Word-Level Attacks
abstract
Brian Formento, Wenjie Feng, Chuan-Sheng Foo, Anh Tuan Luu, See-Kiong Ng. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Brian Formento, Wenjie Feng 0001, Chuan-Sheng Foo, Anh Tuan Luu, See-Kiong Ng
NAACL-HLT4
2024 ToXCL: A Unified Framework for Toxic Speech Detection and Explanation
abstract
Nhat M. Hoang, Xuan Long Do, Duc Anh Do, Duc Anh Vu, Luu Anh Tuan. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Nhat M. Hoang, Do Xuan Long, Duc Anh Vu 0002, Anh Tuan Luu
NAACL-HLT5
2024 Extractive Summarization with Text Generator
abstract
Standard extractive systems suffer from the lack of gold training signals since existing corpora solely provide document and humanwritten summary pairs while disregarding extractive labels.As a result, existing methods resort to imperfect pseudo-labels that are both biased and error-prone, thereby hindering the learning process of extractive models.In contrast, text generators which are commonly employed in abstractive summarization can effortlessly overcome this predicament on account of flexible sequence-to-sequence architectures.Motivated to bypass this inherent limitation, we investigate the possibility of conducting extractive summarization with text generators.Through extensive experiments covering six summarization benchmarks, we show that highquality extractive summaries can be assembled via approximating the outputs (abstractive summaries) of these generators.Moreover, we find that the approximate summaries correlate positively with the auxiliary summaries (i.e. a better generator enables the production of better extractive summaries).Our results signify a new paradigm for training extractive summarizers i.e. learning with generation (abstractive) objectives rather than extractive schemes.
Thang Le, Anh Tuan Luu
NAACL-HLT2
2024 KDMCSE: Knowledge Distillation Multimodal Sentence Embeddings with Adaptive Angular margin Contrastive Learning
abstract
Cong-Duy Nguyen, Thong Nguyen, Xiaobao Wu, Anh Tuan Luu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Cong-Duy Nguyen, Thong Nguyen 0003, Xiaobao Wu, Anh Tuan Luu
NAACL-HLT4
2024 Text2NKG: Fine-Grained N-ary Relation Extraction for N-ary relational Knowledge Graph Construction
abstract
Beyond traditional binary relational facts, n-ary relational knowledge graphs (NKGs) are comprised of n-ary relational facts containing more than two entities, which are closer to real-world facts with broader applications. However, the construction of NKGs remains at a coarse-grained level, which is always in a single schema, ignoring the order and variable arity of entities. To address these restrictions, we propose Text2NKG, a novel fine-grained n-ary relation extraction framework for n-ary relational knowledge graph construction. We introduce a span-tuple classification approach with hetero-ordered merging and output merging to accomplish fine-grained n-ary relation extraction in different arity. Furthermore, Text2NKG supports four typical NKG schemas: hyper-relational schema, event-based schema, role-based schema, and hypergraph-based schema, with high flexibility and practicality. The experimental results demonstrate that Text2NKG achieves state-of-the-art performance in F1 scores on the fine-grained n-ary relation extraction benchmark. Our code and datasets are publicly available.
Haoran Luo 0001, Haihong E, Yuhao Yang 0006, Tianyu Yao, Yikai Guo, Zichen Tang, Wentai Zhang 0004, Shiyao Peng, Kaiyang Wan, Meina Song, Yifan Zhu 0001, Anh Tuan Luu
NeurIPS13
2024 Mercury: A Code Efficiency Benchmark for Code Large Language Models
abstract
Amidst the recent strides in evaluating Large Language Models for Code (Code LLMs), existing benchmarks have mainly focused on the functional correctness of generated code, neglecting the importance of their computational efficiency. To fill the gap, we present Mercury, the first code efficiency benchmark for Code LLMs. It comprises 1,889 Python tasks, each accompanied by adequate solutions that serve as real-world efficiency baselines, enabling a comprehensive analysis of the runtime distribution. Based on the distribution, we introduce a new metric Beyond, which computes a runtime-percentile-weighted Pass score to reflect functional correctness and code efficiency simultaneously. On Mercury, leading Code LLMs can achieve 65% on Pass, while less than 50% on Beyond. Given that an ideal Beyond score would be aligned with the Pass score, it indicates that while Code LLMs exhibit impressive capabilities in generating functionally correct code, there remains a notable gap in their efficiency. Finally, our empirical experiments reveal that Direct Preference Optimization (DPO) serves as a robust baseline for enhancing code efficiency compared with Supervised Fine Tuning (SFT), which paves a promising avenue for future exploration of efficient code generation. Our code and data are available on GitHub: https://github.com/Elfsong/Mercury.
Mingzhe Du, Anh Tuan Luu, Bin Ji 0002, Qian Liu 0033, See-Kiong Ng
NeurIPS2
2024 Uncertainty of Thoughts: Uncertainty-Aware Planning Enhances Information Seeking in LLMs
abstract
In the face of uncertainty, the ability to *seek information* is of fundamental importance. In many practical applications, such as medical diagnosis and troubleshooting, the information needed to solve the task is not initially given, and has to be actively sought by asking follow-up questions (for example, a doctor asking a patient for more details about their symptoms). In this work, we introduce **Uncertainty of Thoughts (UoT)**, an algorithm to augment large language models with the ability to actively seek information by asking effective questions. UoT combines: 1. An *uncertainty-aware simulation approach* which enables the model to simulate possible future scenarios and how likely they are to occur, 2. *Uncertainty-based rewards* motivated by information gain which incentivizes the model to seek information, and 3. A *reward propagation scheme* to select the optimal question to ask in a way that maximizes the expected reward. In experiments on medical diagnosis, troubleshooting and the `20 Questions' game, UoT achieves an average performance improvement of 38.1% in the rate of successful task completion across multiple LLMs compared with direct prompting, and also improves efficiency (i.e., the number of questions needed to complete the task).
Chumin Liu, Xidong Feng, Yilun Zhao 0001, See-Kiong Ng, Anh Tuan Luu, Junxian He, Pang Wei W. Koh, Bryan Hooi
NeurIPS6
2024 FASTopic: Pretrained Transformer is a Fast, Adaptive, Stable, and Transferable Topic Model
abstract
Topic models have been evolving rapidly over the years, from conventional to recent neural models. However, existing topic models generally struggle with either effectiveness, efficiency, or stability, highly impeding their practical applications. In this paper, we propose FASTopic, a fast, adaptive, stable, and transferable topic model. FASTopic follows a new paradigm: Dual Semantic-relation Reconstruction (DSR). Instead of previous conventional, VAE-based, or clustering-based methods, DSR directly models the semantic relations among document embeddings from a pretrained Transformer and learnable topic and word embeddings. By reconstructing through these semantic relations, DSR discovers latent topics. This brings about a neat and efficient topic modeling framework. We further propose a novel Embedding Transport Plan (ETP) method. Rather than early straightforward approaches, ETP explicitly regularizes the semantic relations as optimal transport plans. This addresses the relation bias issue and thus leads to effective topic modeling. Extensive experiments on benchmark datasets demonstrate that our FASTopic shows superior effectiveness, efficiency, adaptivity, stability, and transferability, compared to state-of-the-art baselines across various scenarios.
Xiaobao Wu, Thong Nguyen 0003, Delvin Zhang, William Yang Wang, Anh Tuan Luu
NeurIPS5
2024 Learning facial expression and body gesture visual information for video emotion recognition
Guanyu Hu 0003, Xinyu Yang 0001, Anh Tuan Luu, Yizhuo Dong
Expert Syst. Appl.4
2024 Historical Embedding-Guided Efficient Large-Scale Federated Graph Learning
abstract
Graph convolutional networks (GCNs) are promising for graph learning tasks. For privacy-preserving graph learning tasks involving distributed graph datasets, federated learning (FL)-based GCN (FedGCN) training is required. An important open challenge for FedGCN is scaling to large graphs, which typically incurs 1) high computation overhead for handling the explosively-increasing number of neighbors, and 2) high communication overhead of training GCNs involving multiple FL clients. Thus, neighbor sampling is being studied to enhance the scalability of FedGCNs. Existing FedGCN training techniques with neighbor sampling often produce extremely large communication and computation overhead and inaccurate node embeddings, leading to poor model performance. To bridge this gap, we propose the Federated Adaptive Attention-based Sampling (FedAAS) approach. It achieves substantial cost savings by efficiently leveraging historical embedding estimators and focusing the limited communication resources on transmitting the most influential neighbor node embeddings across FL clients. We further design an adaptive embedding synchronization scheme to optimize the efficiency and accuracy of FedAAS on large-scale datasets. Theoretical analysis shows that the approximation error induced by the staleness of historical embedding is upper bounded, and the model is guaranteed to converge in an efficient manner. Extensive experimental evaluation against four state-of-the-art baselines on six real-world graph datasets show that FedAAS achieves up to 5.12% higher test accuracy, while saving communication and computation costs by 95.11% and 94.76%, respectively.
Anran Li 0001, Yuanyuan Chen 0012, Jian Zhang 0087, Mingfei Cheng, Yihao Huang 0001, Yueming Wu 0001, Anh Tuan Luu, Han Yu 0001
Proc. ACM Manag. Data7
2024 Exploring Clean Label Backdoor Attacks and Defense in Language Models
abstract
Despite being widely applied, pre-trained language models have been proven vulnerable to backdoor attacks. Backdoor attacks are designed to introduce targeted vulnerabilities into models by poisoning a subset of training samples through trigger injection and label modification. Traditional textual backdoor attacks suffer several flaws: the triggers lead to abnormal natural language expressions, and poisoned sample labels are mistakenly labeled. These flaws reduce the stealthiness of the attack and can be easily detected by defense models. In this study, we introduce Cbat, a novel and efficient method to perform clean-label backdoor attack with text style, which does not require external trigger, and the poisoned samples are correctly labeled. Specifically, we develop a sentence rewriting model by leveraging the powerful few-shot learning capability of prompt tuning to generate clean label poisoned samples. Cbat then injects text style as an abstract trigger into the victim model through poisoned samples. We also introduce an algorithm for defending against backdoor attacks, named CbatD, which effectively erases the poisoned samples by locating the lowest training loss and calculating feature relevance. The experiments on text classification tasks demonstrate that our Cbat and CbatD show overall competitive performance in textual backdoor attack and defense. It is noteworthy that Cbat attained leading results in the clean-label backdoor attack benchmark without triggers.
Shuai Zhao 0007, Anh Tuan Luu, Jie Fu 0001, Jinming Wen, Weiqi Luo 0002
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Efficient and Privacy-Preserving Feature Importance-Based Vertical Federated Learning
abstract
Vertical Federated Learning (VFL) enables multiple data owners, each holding a different subset of features about a largely overlapping set of data samples, to collaboratively train a global model. The quality of data owners' local features affects the performance of the VFL model, which makes feature selection vitally important. However, existing feature selection methods for VFL either assume the availability of prior knowledge on the number of noisy features or prior knowledge on the post-training threshold of useful features to be selected, making them unsuitable for practical applications. To bridge this gap, we propose the Federated Stochastic Dual-Gate based Feature Selection (FedSDG-FS) approach. It consists of a Gaussian stochastic dual-gate to efficiently approximate the probability of a feature being selected. FedSDG-FS further designs a local embedding perturbation approach to achieve differential privacy for local training data. To reduce overhead, we propose a feature importance initialization method based on Gini impurity, which can accomplish its goals with only two parameter transmissions between the server and the clients. The enhanced version, FedSDG-FS++, protects the privacy for both the clients' training data and the server's labels through Partially Homomorphic Encryption (PHE) without relying on a trusted third-party. Theoretically, we analyze the convergence rate, privacy guarantees and security analysis of our methods. Extensive experiments on both synthetic and real-world datasets show that FedSDG-FS and FedSDG-FS++ significantly outperform existing approaches in terms of achieving more accurate selection of high-quality features as well as improving VFL performance in a privacy-preserving manner.
Anran Li 0001, Ju Jia, Hongyi Peng, Lan Zhang 0002, Anh Tuan Luu, Han Yu 0001, Xiang-Yang Li 0001
IEEE Trans. Mob. Comput.6
2024 Joint Client-and-Sample Selection for Federated Learning via Bi-Level Optimization
abstract
Federated Learning (FL) enables massive local data owners to collaboratively train a deep learning model without disclosing their private data. The importance of local data samples from various data owners to FL models varies widely. This is exacerbated by the presence of noisy data that exhibit large losses similar to important (hard) samples. Currently, there lacks an FL approach that can effectively distinguish hard samples (which are beneficial) from noisy samples (which are harmful). To bridge this gap, we propose the joint Federated Meta-Weighting based Client and Sample Selection (FedMW-CSS) approach to simultaneously mitigate label noise and hard sample selection. It is a bilevel optimization approach for FL client-and-sample selection and global model construction to achieve hard sample-aware noise-robust learning in a privacy preserving manner. It performs meta-learning based online approximation to iteratively update global FL models, select the most positively influential samples and deal with training data noise. To utilize both the instance-level information and class-level information for better performance improvements, FedMW-CSS efficiently learns a class-level weight by manipulating gradients at the class level, e.g., it performs a gradient descent step on class-level weights, which only relies on intermediate gradients. Theoretically, we analyze the privacy guarantees and convergence of FedMW-CSS. Extensive experiments comparison against eight state-of-the-art baselines on six real-world datasets in the presence of data noise and heterogeneity shows that FedMW-CSS achieves up to 28.5% higher test accuracy, while saving communication and computation costs by at least 49.3% and 1.2%, respectively.
Anran Li 0001, Guangjing Wang 0001, Ming Hu 0003, Jianfei Sun, Lan Zhang 0002, Anh Tuan Luu, Han Yu 0001
IEEE Trans. Mob. Comput.6
2024 MuJo-SF: Multimodal Joint Slot Filling for Attribute Value Prediction of E-Commerce Commodities
abstract
Supplementing product attribute information is a critical step for E-commerce platforms, which further benefits various downstream tasks, including product recommendation, product search, and product knowledge graph construction. Intuitively, the visual information available on e-commerce platforms can effectively function as a primary source for certain product attributes. However, existing works either extract attribute values solely from textual product descriptions or leverage limited visual information (e.g., image features or optical character recognition tokens) to assist extraction, without mining the fine-grained visual cues linked with the products effectively. In this paper, we propose a novel task -Multimodal Joint Slot Filling(MuJo-SF) - that aims to combine multimodal information from both product descriptions and their corresponding product images to jointly fill values into the pre-defined product attribute set. To this end, we develop MAVP, a new dataset with 79 k instances of product description-image pairs. Specifically, we present a strategy to fulfill visualized saliency ascription, which aims to distinguish between text-dependent and image-dependent attributes. For those image-dependent attributes, we annotate the corresponding values from images using distant supervision. Then, we design a model for MuJo-SF, which combines multimodal representations and fills image-dependent and text-dependent attributes separately. Finally, we conduct extensive experiments on MAVP and provide rich results for MuJo-SF, which can be used as baselines to facilitate future research.
Meihuizi Jia, Lei Shen 0001, Anh Tuan Luu, Meng Chen 0006, Lejian Liao, Shaozu Yuan, Xiaodong He 0001
IEEE Trans. Multim.3
2023 InfoCTM: A Mutual Information Maximization Perspective of Cross-Lingual Topic Modeling
abstract
Cross-lingual topic models have been prevalent for cross-lingual text analysis by revealing aligned latent topics. However, most existing methods suffer from producing repetitive topics that hinder further analysis and performance decline caused by low-coverage dictionaries. In this paper, we propose the Cross-lingual Topic Modeling with Mutual Information (InfoCTM). Instead of the direct alignment in previous work, we propose a topic alignment with mutual information method. This works as a regularization to properly align topics and prevent degenerate topic representations of words, which mitigates the repetitive topic issue. To address the low-coverage dictionary issue, we further propose a cross-lingual vocabulary linking method that finds more linked cross-lingual words for topic alignment beyond the translations of a given dictionary. Extensive experiments on English, Chinese, and Japanese datasets demonstrate that our method outperforms state-of-the-art baselines, producing more coherent, diverse, and well-aligned topics and showing better transferability for cross-lingual classification tasks.
Xiaobao Wu, Xinshuai Dong, Thong Nguyen 0003, Chaoqun Liu, Liangming Pan, Anh Tuan Luu
AAAI6
2023 Fact-Checking Complex Claims with Program-Guided Reasoning
abstract
Liangming Pan, Xiaobao Wu, Xinyuan Lu, Anh Tuan Luu, William Yang Wang, Min-Yen Kan, Preslav Nakov. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Liangming Pan, Xiaobao Wu, Xinyuan Lu, Anh Tuan Luu, William Yang Wang, Min-Yen Kan, Preslav Nakov
ACL (1)4
2023 Jointprop: Joint Semi-supervised Learning for Entity and Relation Extraction with Heterogeneous Graph-based Propagation
abstract
Semi-supervised learning has been an important approach to address challenges in extracting entities and relations from limited data.However, current semi-supervised works handle the two tasks (i.e., Named Entity Recognition and Relation Extraction) separately and ignore the cross-correlation of entity and relation instances as well as the existence of similar instances across unlabeled data.To alleviate the issues, we propose Jointprop, a Heterogeneous Graph-based Propagation framework for joint semi-supervised entity and relation extraction, which captures the global structure information between individual tasks and exploits interactions within unlabeled data.Specifically, we construct a unified span-based heterogeneous graph from entity and relation candidates and propagate class labels based on confidence scores.We then employ a propagation learning scheme to leverage the affinities between labelled and unlabeled samples.Experiments on benchmark datasets show that our framework outperforms the state-of-the-art semi-supervised approaches on NER and RE tasks.We show that the joint semi-supervised learning of the two tasks benefits from their codependency and validates the importance of utilizing the shared information between unlabeled data.
Yandan Zheng, Anran Hao, Anh Tuan Luu
ACL (1)3
2023 Unlocking the Potential of User Feedback: Leveraging Large Language Model as User Simulators to Enhance Dialogue System
abstract
Dialogue systems and large language models (LLMs) have gained considerable attention. However, the direct utilization of LLMs as task-oriented dialogue (TOD) models has been found to underperform compared to smaller task-specific models. Nonetheless, it is crucial to acknowledge the significant potential of LLMs and explore improved approaches for leveraging their impressive abilities. Motivated by the goal of leveraging LLMs, we propose an alternative approach called User-Guided Response Optimization (UGRO) to combine it with a smaller TOD model. This approach uses LLM as an annotation-free user simulator to assess dialogue responses, combining them with smaller fine-tuned end-to-end TOD models. By utilizing the satisfaction feedback generated by LLMs, UGRO further optimizes the supervised fine-tuned TOD model. Specifically, the TOD model takes the dialogue history as input and, with the assistance of the user simulator's feedback, generates high-satisfaction responses that meet the user's requirements. Through empirical experiments on two TOD benchmarks, we validate the effectiveness of our method. The results demonstrate that our approach outperforms previous state-of-the-art (SOTA) results.
Yue Feng 0002, Anh Tuan Luu, Bryan Hooi, Aldo Lipani
CIKM3
2023 Rethinking Negative Pairs in Code Search
abstract
Recently, contrastive learning has become a key component in fine-tuning code search models for software development efficiency and effectiveness.It pulls together positive code snippets while pushing negative samples away given search queries.Among contrastive learning, InfoNCE is the most widely used loss function due to its better performance.However, the following problems in negative samples of In-foNCE may deteriorate its representation learning: 1) The existence of false negative samples in large code corpora due to duplications.2).The failure to explicitly differentiate between the potential relevance of negative samples.As an example, a bubble sorting algorithm example is less "negative" than a file saving function for the quick sorting algorithm query.In this paper, we tackle the above problems by proposing a simple yet effective Soft-InfoNCE loss that inserts weight terms into InfoNCE.In our proposed loss function, we apply three methods to estimate the weights of negative pairs and show that the vanilla InfoNCE loss is a special case of Soft-InfoNCE.Theoretically, we analyze the effects of Soft-InfoNCE on controlling the distribution of learnt code representations and on deducing a more precise mutual information estimation.We furthermore discuss the superiority of proposed loss functions with other design alternatives.Extensive experiments demonstrate the effectiveness of Soft-InfoNCE and weights estimation methods under state-of-the-art code search models on a large-scale public dataset consisting of six programming languages.
Haochen Li 0009, Xin Zhou 0008, Anh Tuan Luu, Chunyan Miao
EMNLP3
2023 Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
abstract
The prompt-based learning paradigm, which bridges the gap between pre-training and finetuning, achieves state-of-the-art performance on several NLP tasks, particularly in few-shot settings.Despite being widely applied, promptbased learning is vulnerable to backdoor attacks.Textual backdoor attacks are designed to introduce targeted vulnerabilities into models by poisoning a subset of training samples through trigger injection and label modification.However, they suffer from flaws such as abnormal natural language expressions resulting from the trigger and incorrect labeling of poisoned samples.In this study, we propose ProAttack, a novel and efficient method for performing clean-label backdoor attacks based on the prompt, which uses the prompt itself as a trigger.Our method does not require external triggers and ensures correct labeling of poisoned samples, improving the stealthy nature of the backdoor attack.With extensive experiments on rich-resource and few-shot text classification tasks, we empirically validate ProAttack's competitive performance in textual backdoor attacks.Notably, in the rich-resource setting, ProAttack achieves state-of-the-art attack success rates in the clean-label backdoor attack benchmark without external triggers 1 .
Shuai Zhao 0007, Jinming Wen, Anh Tuan Luu, Jie Fu 0001
EMNLP3
2023 Multi-Scale Receptive Field Graph Model for Emotion Recognition in Conversations
abstract
Emotion recognition in conversations (ERC) has gained more attention, where contextual information modeling and multimodal fusion have been the focus and challenges in recent years. In this paper, we proposed a Multi-Scale Receptive Field Graph model (MSRFG) to tackle the challenges of ERC. Specifically, MSRFG constructs multi-scale perception graphs and learns contextual information via parallel multi-scale receptive field paths. To compensate for the deficiency of temporal information learning by the graph network, MSRFG injects temporal dependencies into the graph network to model the temporal relationships between utterances. Moreover, to achieve the effective fusion of multimodal information, MSRFG converges the multi-scale features of each modality separately and performs the learning of attention weights after the integration of converged features. We carried out experiments on IEMOCAP and MELD datasets to validate the effectiveness of the proposed method, and the results proved the superiority of our model over the existing SOTA methods. The code is available at https://github.com/Janie1996/MSRFG1.
Guanyu Hu 0003, Anh Tuan Luu, Xinyu Yang 0001
ICASSP3
2023 Effective Neural Topic Modeling with Embedding Clustering Regularization
abstract
Topic models have been prevalent for decades with various applications. However, existing topic models commonly suffer from the notorious topic collapsing: discovered topics semantically collapse towards each other, leading to highly repetitive topics, insufficient topic discovery, and damaged model interpretability. In this paper, we propose a new neural topic model, Embedding Clustering Regularization Topic Model (ECRTM). Besides the existing reconstruction error, we propose a novel Embedding Clustering Regularization (ECR), which forces each topic embedding to be the center of a separately aggregated word embedding cluster in the semantic space. This enables each produced topic to contain distinct word semantics, which alleviates topic collapsing. Regularized by ECR, our ECRTM generates diverse and coherent topics together with high-quality topic distributions of documents. Extensive experiments on benchmark datasets demonstrate that ECRTM effectively addresses the topic collapsing issue and consistently surpasses state-of-the-art baselines in terms of topic quality, topic distributions of documents, and downstream classification tasks.
Xiaobao Wu, Xinshuai Dong, Thong Thanh Nguyen, Anh Tuan Luu
ICML4
2023 Expand BERT Representation with Visual Information via Grounded Language Learning with Multimodal Partial Alignment
abstract
Language models have been supervised with both language-only objective and visual grounding in existing studies of visual-grounded language learning. However, due to differences in the distribution and scale of visual-grounded datasets and language corpora, the language model tends to mix up the context of the tokens that occurred in the grounded data with those that do not. As a result, during representation learning, there is a mismatch between the visual information and the contextual meaning of the sentence. To overcome this limitation, we propose GroundedBERT - a grounded language learning method that enhances the BERT representation with visually grounded information. GroundedBERT comprises two components: (i) the original BERT which captures the contextual representation of words learned from the language corpora, and (ii) a visual grounding module which captures visual information learned from visual-grounded datasets. Moreover, we employ Optimal Transport (OT), specifically its partial variant, to solve the fractional alignment problem between the two modalities. Our proposed method significantly outperforms the baseline language models on various language tasks of the GLUE and SQuAD datasets.
Cong-Duy Nguyen, The-Anh Vu-Le, Thong Nguyen 0003, Thanh Tho Quan, Anh Tuan Luu
ACM Multimedia5
2023 A General Framework for Blockchain Data Analysis
Anh Tuan Luu, Tuan-Dat Trinh, Van-Thanh Nguyen
RCIS1
2023 Knowledge Bases and Language Models: Complementing Forces
Fabian M. Suchanek, Anh Tuan Luu
RuleML+RR2
2023 A contrastive learning framework for Event Detection via semantic type prototype representation modelling
Anran Hao, Anh Tuan Luu, Siu Cheung Hui, Jian Su 0002
Neurocomputing2
2023 Benchmarking Graph Neural Networks
abstract
In the last few years, graph neural networks (GNNs) have become the standard toolkit for analyzing and learning from data on graphs. This emerging field has witnessed an extensive growth of promising techniques that have been applied with success to computer science, mathematics, biology, physics and chemistry. But for any successful field to become mainstream and reliable, benchmarks must be developed to quantify progress. This led us in March 2020 to release a benchmark framework that i) comprises of a diverse collection of mathematical and real-world graphs, ii) enables fair model comparison with the same parameter budget to identify key architectures, iii) has an open-source, easy-to use and reproducible code infrastructure, and iv) is flexible for researchers to experiment with new theoretical ideas. As of December 2022, the GitHub repository has reached 2,000 stars and 380 forks, which demonstrates the utility of the proposed open-source framework through the wide usage by the GNN community. In this paper, we present an updated version of our benchmark with a concise presentation of the aforementioned framework characteristics, an additional medium-sized molecular dataset AQSOL, similar to the popular ZINC, but with a real-world measured chemical target, and discuss how this framework can be leveraged to explore new GNN designs and insights. As a proof of value of our benchmark, we study the case of graph positional encoding (PE) in GNNs, which was introduced with this benchmark and has since spurred interest of exploring more powerful PE for Transformers and GNNs in a robust experimental setting.
Vijay Prakash Dwivedi, Chaitanya K. Joshi, Anh Tuan Luu, Thomas Laurent 0001, Yoshua Bengio, Xavier Bresson
J. Mach. Learn. Res.3
2022 Improving Neural Cross-Lingual Abstractive Summarization via Employing Optimal Transport Distance for Knowledge Distillation
abstract
Current state-of-the-art cross-lingual summarization models employ multi-task learning paradigm, which works on a shared vocabulary module and relies on the self-attention mechanism to attend among tokens in two languages. However, correlation learned by self-attention is often loose and implicit, inefficient in capturing crucial cross-lingual representations between languages. The matter worsens when performing on languages with separate morphological or structural features, making the cross-lingual alignment more challenging, resulting in the performance drop. To overcome this problem, we propose a novel Knowledge-Distillation-based framework for Cross-Lingual Summarization, seeking to explicitly construct cross-lingual correlation by distilling the knowledge of the monolingual summarization teacher into the cross-lingual summarization student. Since the representations of the teacher and the student lie on two different vector spaces, we further propose a Knowledge Distillation loss using Sinkhorn Divergence, an Optimal-Transport distance, to estimate the discrepancy between those teacher and student representations. Due to the intuitively geometric nature of Sinkhorn Divergence, the student model can productively learn to align its produced cross-lingual hidden states with monolingual hidden states, hence leading to a strong correlation between distant languages. Experiments on cross-lingual summarization datasets in pairs of distant languages demonstrate that our method outperforms state-of-the-art models under both high and low-resourced settings.
Thong Thanh Nguyen, Anh Tuan Luu
AAAI2
2022 Is Discourse Role Important for Emotion Recognition in Conversation?
abstract
A conversation is a sequence of utterances, where each utterance plays a specific discourse role while expressing a particular emotion. This paper proposes a novel method to exploit latent discourse role information of an utterance to determine the emotion it conveys in a conversation. Specifically, we use a variant of the Variational-Autoencoder (VAE) to model the context-aware latent discourse roles of each utterance in an unsupervised way. The latent discourse role representation further equips the utterance representation with a salient clue for more accurate emotion recognition. Our experiments show that our proposed method beats the best-reported performances on three public Emotion Recognition in Conversation datasets. This proves that the discourse role information of an utterance plays an important role in the emotion recognition task, which no previous work has studied.
Donovan Ong, Jian Su 0002, Bin Chen 0027, Anh Tuan Luu, Ashok Narendranath, Shu-Qi Sun, Yingzhan Lin, Haifeng Wang 0001
AAAI4
2022 Textual Manifold-based Defense Against Natural Language Adversarial Examples
abstract
Recent studies on adversarial images have shown that they tend to leave the underlying low-dimensional data manifold, making them significantly more challenging for current models to make correct predictions.This so-called off-manifold conjecture has inspired a novel line of defenses against adversarial attacks on images.In this study, we find a similar phenomenon occurs in the contextualized embedding space induced by pretrained language models, in which adversarial texts tend to have their embeddings diverge from the manifold of natural ones.Based on this finding, we propose Textual Manifold-based Defense (TMD), a defense mechanism that projects text embeddings onto an approximated embedding manifold before classification.It reduces the complexity of potential adversarial examples, which ultimately enhances the robustness of the protected model.Through extensive experiments, our method consistently and significantly outperforms previous defenses under various attack settings without trading off clean accuracy.To the best of our knowledge, this is the first NLP defense that leverages the manifold structure against adversarial attacks.Our code is available at https://github.com/dangne/tmd.
Dang Nguyen Minh, Anh Tuan Luu
EMNLP2
2022 Adaptive Contrastive Learning on Multimodal Transformer for Review Helpfulness Prediction
abstract
Modern Review Helpfulness Prediction systems are dependent upon multiple modalities, typically texts and images.Unfortunately, those contemporary approaches pay scarce attention to polish representations of cross-modal relations and tend to suffer from inferior optimization.This might cause harm to model's predictions in numerous cases.To overcome the aforementioned issues, we propose Multimodal Contrastive Learning for Multimodal Review Helpfulness Prediction (MRHP) problem, concentrating on mutual information between input modalities to explicitly elaborate cross-modal relations.In addition, we introduce Adaptive Weighting scheme for our contrastive learning approach in order to increase flexibility in optimization.Lastly, we propose Multimodal Interaction module to address the unalignment nature of multimodal data, thereby assisting the model in producing more reasonable multimodal representations.Experimental results show that our method outperforms prior baselines and achieves state-of-the-art results on two publicly available benchmark datasets for MRHP problem.
Thong Nguyen 0003, Xiaobao Wu, Anh Tuan Luu, Zhen Hai, Lidong Bing
EMNLP3
2022 Mitigating Data Sparsity for Short Text Topic Modeling by Topic-Semantic Contrastive Learning
abstract
To overcome the data sparsity issue in short text topic modeling, existing methods commonly rely on data augmentation or the data characteristic of short texts to introduce more word co-occurrence information.However, most of them do not make full use of the augmented data or the data characteristic: they insufficiently learn the relations among samples in data, leading to dissimilar topic distributions of semantically similar text pairs.To better address data sparsity, in this paper we propose a novel short text topic modeling framework, Topic-Semantic Contrastive Topic Model (TSCTM).To sufficiently model the relations among samples, we employ a new contrastive learning method with efficient positive and negative sampling strategies based on topic semantics.This contrastive learning method refines the representations, enriches the learning signals, and thus mitigates the sparsity issue.Extensive experimental results show that our TSCTM outperforms state-ofthe-art baselines regardless of the data augmentation availability, producing high-quality topics and topic distributions. 1
Xiaobao Wu, Anh Tuan Luu, Xinshuai Dong
EMNLP2
2022 Graph Neural Networks with Learnable Structural and Positional Representations
Vijay Prakash Dwivedi, Anh Tuan Luu, Thomas Laurent 0001, Yoshua Bengio, Xavier Bresson
ICLR2
2022 Certified Robustness Against Natural Language Attacks by Causal Intervention
abstract
Deep learning models have achieved great success in many fields, yet they are vulnerable to adversarial examples. This paper follows a causal perspective to look into the adversarial vulnerability and proposes Causal Intervention by Semantic Smoothing (CISS), a novel framework towards robustness against natural language attacks. Instead of merely fitting observational data, CISS learns causal effects p(y|do(x)) by smoothing in the latent semantic space to make robust predictions, which scales to deep architectures and avoids tedious construction of noise customized for specific attacks. CISS is provably robust against word substitution attacks, as well as empirically robust even when perturbations are strengthened by unknown attack algorithms. For example, on YELP, CISS surpasses the runner-up by 6.8% in terms of certified robustness against word substitutions, and achieves 80.7% empirical robustness when syntactic attacks are integrated.
Haiteng Zhao, Xinshuai Dong, Anh Tuan Luu, Zhi-Hong Deng 0001, Hanwang Zhang
ICML4
2022 Audio-Visual Domain Adaptation Feature Fusion for Speech Emotion Recognition
Guanyu Hu 0003, Xinyu Yang 0001, Anh Tuan Luu, Yizhuo Dong
INTERSPEECH4
2022 Long Range Graph Benchmark
abstract
Graph Neural Networks (GNNs) that are based on the message passing (MP) paradigm generally exchange information between 1-hop neighbors to build node representations at each layer. In principle, such networks are not able to capture long-range interactions (LRI) that may be desired or necessary for learning a given task on graphs. Recently, there has been an increasing interest in development of Transformer-based methods for graphs that can consider full node connectivity beyond the original sparse structure, thus enabling the modeling of LRI. However, MP-GNNs that simply rely on 1-hop message passing often fare better in several existing graph benchmarks when combined with positional feature representations, among other innovations, hence limiting the perceived utility and ranking of Transformer-like architectures. Here, we present the Long Range Graph Benchmark (LRGB) with 5 graph learning datasets: $\texttt{PascalVOC-SP}$, $\texttt{COCO-SP}$, $\texttt{PCQM-Contact}$, $\texttt{Peptides-func}$ and $\texttt{Peptides-struct}$ that arguably require LRI reasoning to achieve strong performance in a given task. We benchmark both baseline GNNs and Graph Transformer networks to verify that the models which capture long-range dependencies perform significantly better on these tasks. Therefore, these datasets are suitable for benchmarking and exploration of MP GNNs and Graph Transformer architectures that are intended to capture LRI.
Vijay Prakash Dwivedi, Ladislav Rampásek, Michael Galkin, Ali Parviz, Guy Wolf, Anh Tuan Luu, Dominique Beaini
NeurIPS6
2022 Recipe for a General, Powerful, Scalable Graph Transformer
abstract
We propose a recipe on how to build a general, powerful, scalable (GPS) graph Transformer with linear complexity and state-of-the-art results on a diverse set of benchmarks. Graph Transformers (GTs) have gained popularity in the field of graph representation learning with a variety of recent publications but they lack a common foundation about what constitutes a good positional or structural encoding, and what differentiates them. In this paper, we summarize the different types of encodings with a clearer definition and categorize them as being $\textit{local}$, $\textit{global}$ or $\textit{relative}$. The prior GTs are constrained to small graphs with a few hundred nodes, here we propose the first architecture with a complexity linear in the number of nodes and edges $O(N+E)$ by decoupling the local real-edge aggregation from the fully-connected Transformer. We argue that this decoupling does not negatively affect the expressivity, with our architecture being a universal function approximator on graphs. Our GPS recipe consists of choosing 3 main ingredients: (i) positional/structural encoding, (ii) local message-passing mechanism, and (iii) global attention mechanism. We provide a modular framework $\textit{GraphGPS}$ that supports multiple types of encodings and that provides efficiency and scalability both in small and large graphs. We test our architecture on 16 benchmarks and show highly competitive results in all of them, show-casing the empirical benefits gained by the modularity and the combination of different strategies.
Ladislav Rampásek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu, Guy Wolf, Dominique Beaini
NeurIPS4
2021 Enriching and Controlling Global Semantics for Text Summarization
abstract
Recently, Transformer-based models have been proven effective in the abstractive summarization task by creating fluent and informative summaries.Nevertheless, these models still suffer from the short-range dependency problem, causing them to produce summaries that miss the key points of document.In this paper, we attempt to address this issue by introducing a neural topic model empowered with normalizing flow to capture the global semantics of the document, which are then integrated into the summarization model.In addition, to avoid the overwhelming effect of global semantics on contextualized representation, we introduce a mechanism to control the amount of global semantics supplied to the text generation module.Our method outperforms state-of-the-art summarization models on five common text summarization datasets, namely CNN/DailyMail, XSum, Reddit TIFU, arXiv, and PubMed.
Thong Nguyen 0003, Anh Tuan Luu, Truc Lu, Thanh Tho Quan
EMNLP (1)2
2021 Towards Robustness Against Natural Language Word Substitutions
Xinshuai Dong, Anh Tuan Luu, Rongrong Ji, Hong Liu 0009
ICLR2
2021 Beyond Fully-Connected Layers with Quaternions: Parameterization of Hypercomplex Multiplications with 1/n Parameters
Aston Zhang, Yi Tay, Shuai Zhang 0007, Alvin Chan, Anh Tuan Luu, Siu Cheung Hui, Jie Fu 0001
ICLR5
2021 How Should Pre-Trained Language Models Be Fine-Tuned Towards Adversarial Robustness?
abstract
The fine-tuning of pre-trained language models has a great success in many NLP fields. Yet, it is strikingly vulnerable to adversarial examples, e.g., word substitution attacks using only synonyms can easily fool a BERT-based sentiment analysis model. In this paper, we demonstrate that adversarial training, the prevalent defense technique, does not directly fit a conventional fine-tuning scenario, because it suffers severely from catastrophic forgetting: failing to retain the generic and robust linguistic features that have already been captured by the pre-trained model. In this light, we propose Robust Informative Fine-Tuning (RIFT), a novel adversarial fine-tuning method from an information-theoretical perspective. In particular, RIFT encourages an objective model to retain the features learned from the pre-trained model throughout the entire fine-tuning process, whereas a conventional one only uses the pre-trained weights for initialization. Experimental results show that RIFT consistently outperforms the state-of-the-arts on two popular NLP tasks: sentiment analysis and natural language inference, under different attacks across various pre-trained language models.
Xinshuai Dong, Anh Tuan Luu, Shuicheng Yan, Hanwang Zhang
NeurIPS2
2021 Contrastive Learning for Neural Topic Model
abstract
Recent empirical studies show that adversarial topic models (ATM) can successfully capture semantic patterns of the document by differentiating a document with another dissimilar sample. However, utilizing that discriminative-generative architecture has two important drawbacks: (1) the architecture does not relate similar documents, which has the same document-word distribution of salient words; (2) it restricts the ability to integrate external information, such as sentiments of the document, which has been shown to benefit the training of neural topic model. To address those issues, we revisit the adversarial topic architecture in the view point of mathematical analysis, propose a novel approach to re-formulate discriminative goal as an optimization problem, and design a novel sampling method which facilitates the integration of external variables. The reformulation encourages the model to incorporate the relations among similar samples and enforces the constraint on the similarity among dissimilar ones; while the sampling method, which is based on the internal input and reconstructed output, helps inform the model of salient words contributing to the main topic. Experimental results show that our framework outperforms other state-of-the-art neural topic models in three common benchmark datasets that belong to various domains, vocabulary sizes, and document lengths in terms of topic coherence.
Thong Nguyen 0003, Anh Tuan Luu
NeurIPS2
2020 Capturing Greater Context for Question Generation
abstract
Automatic question generation can benefit many applications ranging from dialogue systems to reading comprehension. While questions are often asked with respect to long documents, there are many challenges with modeling such long documents. Many existing techniques generate questions by effectively looking at one sentence at a time, leading to questions that are easy and not reflective of the human process of question generation. Our goal is to incorporate interactions across multiple sentences to generate realistic questions for long documents. In order to link a broad document context to the target answer, we represent the relevant context via a multi-stage attention mechanism, which forms the foundation of a sequence to sequence model. We outperform state-of-the-art methods on question generation on three question-answering datasets - SQuAD, MS MARCO and NewsQA. 1
Anh Tuan Luu, Darsh J. Shah, Regina Barzilay
AAAI1
2020 Would you Rather? A New Benchmark for Learning Machine Alignment with Cultural Values and Social Preferences
abstract
Understanding human preferences, along with cultural and social nuances, lives at the heart of natural language understanding.Concretely, we present a new task and corpus for learning alignments between machine and human preferences.Our newly introduced problem is concerned with predicting the preferable options from two sentences describing scenarios that may involve social and cultural situations.Our problem is framed as a natural language inference task with crowd-sourced preference votes by human players, obtained from a gamified voting platform.We benchmark several state-of-the-art neural models, along with BERT and friends on this task.Our experimental results show that current state-ofthe-art NLP models still leave much room for improvement.
Yi Tay, Donovan Ong, Jie Fu 0001, Alvin Chan, Nancy F. Chen, Anh Tuan Luu, Christopher Joseph Pal
ACL6
2020 Holistic Multi-Modal Memory Network for Movie Question Answering
abstract
Answering questions using multi-modal context is a challenging problem as it requires a deep integration of diverse data sources. Existing approaches only consider a subset of all possible interactions among data sources during one attention hop. In this paper, we present a Holistic Multi-modal Memory Network (HMMN) framework that fully considers interactions between different input sources (multi-modal context, question) at each hop. In addition, to hone in on relevant information, our framework takes answer choices into consideration during the context retrieval stage. Our HMMN framework effectively integrates information from the multi-modal context, question, and answer choices, enabling more informative context to be retrieved for question answering. Experimental results on the MovieQA and TVQA datasets validate the effectiveness of our HMMN framework. Extensive ablation studies show the importance of holistic reasoning and reveal the contributions of different attention strategies to model performance.
Anran Wang 0001, Anh Tuan Luu, Chuan-Sheng Foo, Hongyuan Zhu 0002, Yi Tay, Vijay Chandrasekhar 0001
IEEE Trans. Image Process.2
2019 Holographic Factorization Machines for Recommendation
abstract
Factorization Machines (FMs) are a class of popular algorithms that have been widely adopted for collaborative filtering and recommendation tasks. FMs are characterized by its usage of the inner product of factorized parameters to model pairwise feature interactions, making it highly expressive and powerful. This paper proposes Holographic Factorization Machines (HFM), a new novel method of enhancing the representation capability of FMs without increasing its parameter size. Our approach replaces the inner product in FMs with holographic reduced representations (HRRs), which are theoretically motivated by associative retrieval and compressed outer products. Empirically, we found that this leads to consistent improvements over vanilla FMs by up to 4% improvement in terms of mean squared error, with improvements larger at smaller parameterization. Additionally, we propose a neural adaptation of HFM which enhances its capability to handle nonlinear structures. We conduct extensive experiments on nine publicly available datasets for collaborative filtering with explicit feedback. HFM achieves state-of-theart performance on all nine, outperforming strong competitors such as Attentional Factorization Machines (AFM) and Neural Matrix Factorization (NeuMF).
Yi Tay, Shuai Zhang 0007, Anh Tuan Luu, Siu Cheung Hui, Lina Yao 0001, Tran Dang Quang Vinh
AAAI3
2019 Simple and Effective Curriculum Pointer-Generator Networks for Reading Comprehension over Long Narratives
abstract
Yi Tay, Shuohang Wang, Anh Tuan Luu, Jie Fu, Minh C. Phan, Xingdi Yuan, Jinfeng Rao, Siu Cheung Hui, Aston Zhang. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Yi Tay, Shuohang Wang, Anh Tuan Luu, Jie Fu 0001, Minh C. Phan, Xingdi Yuan, Jinfeng Rao, Siu Cheung Hui, Aston Zhang
ACL (1)3
2019 Lightweight and Efficient Neural Natural Language Processing with Quaternion Networks
abstract
Many state-of-the-art neural models for NLP are heavily parameterized and thus memory inefficient.This paper proposes a series of lightweight and memory efficient neural architectures for a potpourri of natural language processing (NLP) tasks.To this end, our models exploit computation using Quaternion algebra and hypercomplex spaces, enabling not only expressive inter-component interactions but also significantly (75%) reduced parameter size due to lesser degrees of freedom in the Hamilton product.We propose Quaternion variants of models, giving rise to new architectures such as the Quaternion attention Model and Quaternion Transformer.Extensive experiments on a battery of NLP tasks demonstrates the utility of proposed Quaternion-inspired models, enabling up to 75% reduction in parameter size without significant loss in performance.
Yi Tay, Aston Zhang, Anh Tuan Luu, Jinfeng Rao, Shuai Zhang 0007, Shuohang Wang, Jie Fu 0001, Siu Cheung Hui
ACL (1)3
2019 Compositional De-Attention Networks
abstract
Attentional models are distinctly characterized by their ability to learn relative importance, i.e., assigning a different weight to input values. This paper proposes a new quasi-attention that is compositional in nature, i.e., learning whether to \textit{add}, \textit{subtract} or \textit{nullify} a certain vector when learning representations. This is strongly contrasted with vanilla attention, which simply re-weights input tokens. Our proposed \textit{Compositional De-Attention} (CoDA) is fundamentally built upon the intuition of both similarity and dissimilarity (negative affinity) when computing affinity scores, benefiting from a greater extent of expressiveness. We evaluate CoDA on six NLP tasks, i.e. open domain question answering, retrieval/ranking, natural language inference, machine translation, sentiment analysis and text2code generation. We obtain promising experimental results, achieving state-of-the-art performance on several tasks/datasets.
Yi Tay, Anh Tuan Luu, Aston Zhang, Shuohang Wang, Siu Cheung Hui
NeurIPS2
2018 SkipFlow: Incorporating Neural Coherence Features for End-to-End Automatic Text Scoring
abstract
Deep learning has demonstrated tremendous potential for Automatic Text Scoring (ATS) tasks. In this paper, we describe a new neural architecture that enhances vanilla neural network models with auxiliary neural coherence features. Our new method proposes a new SkipFlow mechanism that models relationships between snapshots of the hidden representations of a long short-term memory (LSTM) network as it reads. Subsequently, the semantic relationships between multiple snapshots are used as auxiliary features for prediction. This has two main benefits. Firstly, essays are typically long sequences and therefore the memorization capability of the LSTM network may be insufficient. Implicit access to multiple snapshots can alleviate this problem by acting as a protection against vanishing gradients. The parameters of the SkipFlow mechanism also acts as an auxiliary memory. Secondly, modeling relationships between multiple positions allows our model to learn features that represent and approximate textual coherence. In our model, we call this neural coherence features. Overall, we present a unified deep learning architecture that generates neural coherence features as it reads in an end-to-end fashion. Our approach demonstrates state-of-the-art performance on the benchmark ASAP dataset, outperforming not only feature engineering baselines but also other deep learning models.
Yi Tay, Minh C. Phan, Anh Tuan Luu, Siu Cheung Hui
AAAI3
2018 Cross Temporal Recurrent Networks for Ranking Question Answer Pairs
abstract
Temporal gates play a significant role in modern recurrent-based neural encoders, enabling fine-grained control over recursive compositional operations over time. In recurrent models such as the long short-term memory (LSTM), temporal gates control the amount of information retained or discarded over time, not only playing an important role in influencing the learned representations but also serving as a protection against vanishing gradients. This paper explores the idea of learning temporal gates for sequence pairs (question and answer), jointly influencing the learned representations in a pairwise manner. In our approach, temporal gates are learned via 1D convolutional layers and then subsequently cross applied across question and answer for joint learning. Empirically, we show that this conceptually simple sharing of temporal gates can lead to competitive performance across multiple benchmarks. Intuitively, what our network achieves can be interpreted as learning representations of question and answer pairs that are aware of what each other is remembering or forgetting, i.e., pairwise temporal gating. Via extensive experiments, we show that our proposed model achieves state-of-the-art performance on two community-based QA datasets and competitive performance on one factoid-based QA dataset.
Yi Tay, Anh Tuan Luu, Siu Cheung Hui
AAAI2
2018 Learning to Attend via Word-Aspect Associative Fusion for Aspect-Based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) tries to predict the polarity of a given document with respect to a given aspect entity. While neural network architectures have been successful in predicting the overall polarity of sentences, aspect-specific sentiment analysis still remains as an open problem. In this paper, we propose a novel method for integrating aspect information into the neural model. More specifically, we incorporate aspect information into the neural model by modeling word-aspect relationships. Our novel model, Aspect Fusion LSTM (AF-LSTM) learns to attend based on associative relationships between sentence words and aspect which allows our model to adaptively focus on the correct words given an aspect term. This ameliorates the flaws of other state-of-the-art models that utilize naive concatenations to model word-aspect similarity. Instead, our model adopts circular convolution and circular correlation to model the similarity between aspect and words and elegantly incorporates this within a differentiable neural attention framework. Finally, our model is end-to-end differentiable and highly related to convolution-correlation (holographic like) memories. Our proposed neural model achieves state-of-the-art performance on benchmark datasets, outperforming ATAE-LSTM by 4%-5% on average across multiple datasets.
Yi Tay, Anh Tuan Luu, Siu Cheung Hui
AAAI2
2018 Reasoning with Sarcasm by Reading In-Between
abstract
Sarcasm is a sophisticated speech act which commonly manifests on social communities such as Twitter and Reddit.The prevalence of sarcasm on the social web is highly disruptive to opinion mining systems due to not only its tendency of polarity flipping but also usage of figurative language.Sarcasm commonly manifests with a contrastive theme either between positive-negative sentiments or between literal-figurative scenarios.In this paper, we revisit the notion of modeling contrast in order to reason with sarcasm.More specifically, we propose an attention-based neural model that looks inbetween instead of across, enabling it to explicitly model contrast and incongruity.We conduct extensive experiments on six benchmark datasets from Twitter, Reddit and the Internet Argument Corpus.Our proposed model not only achieves stateof-the-art performance on all datasets but also enjoys improved interpretability.
Yi Tay, Anh Tuan Luu, Siu Cheung Hui, Jian Su 0002
ACL (1)2
2018 Compare, Compress and Propagate: Enhancing Neural Architectures with Alignment Factorization for Natural Language Inference
abstract
This paper presents a new deep learning architecture for Natural Language Inference (NLI).Firstly, we introduce a new architecture where alignment pairs are compared, compressed and then propagated to upper layers for enhanced representation learning.Secondly, we adopt factorization layers for efficient and expressive compression of alignment vectors into scalar features, which are then used to augment the base word representations.The design of our approach is aimed to be conceptually simple, compact and yet powerful.We conduct experiments on three popular benchmarks, SNLI, MultiNLI and SciTail, achieving competitive performance on all.A lightweight parameterization of our model also enjoys a ≈ 3 times reduction in parameter size compared to the existing state-of-the-art models, e.g., ESIM and DIIN, while maintaining competitive performance.Additionally, visual analysis shows that our propagated features are highly interpretable.
Yi Tay, Anh Tuan Luu, Siu Cheung Hui
EMNLP2
2018 Multi-Granular Sequence Encoding via Dilated Compositional Units for Reading Comprehension
abstract
Sequence encoders are crucial components in many neural architectures for learning to read and comprehend.This paper presents a new compositional encoder for reading comprehension (RC).Our proposed encoder is not only aimed at being fast but also expressive.Specifically, the key novelty behind our encoder is that it explicitly models across multiple granularities using a new dilated composition mechanism.In our approach, gating functions are learned by modeling relationships and reasoning over multi-granular sequence information, enabling compositional learning that is aware of both long and short term information.We conduct experiments on three RC datasets, showing that our proposed encoder demonstrates very promising results both as a standalone encoder as well as a complementary building block.Empirical results show that simple Bi-Attentive architectures augmented with our proposed encoder not only achieves state-of-the-art / highly competitive results but is also considerably faster than other published works.
Yi Tay, Anh Tuan Luu, Siu Cheung Hui
EMNLP2
2018 Co-Stack Residual Affinity Networks with Multi-level Attention Refinement for Matching Text Sequences
abstract
Learning a matching function between two text sequences is a long standing problem in NLP research.This task enables many potential applications such as question answering and paraphrase identification.This paper proposes Co-Stack Residual Affinity Networks (CSRAN), a new and universal neural architecture for this problem.CSRAN is a deep architecture, involving stacked (multi-layered) recurrent encoders.Stacked/Deep architectures are traditionally difficult to train, due to the inherent weaknesses such as difficulty with feature propagation and vanishing gradients.CSRAN incorporates two novel components to take advantage of the stacked architecture.Firstly, it introduces a new bidirectional alignment mechanism that learns affinity weights by fusing sequence pairs across stacked hierarchies.Secondly, it leverages a multi-level attention refinement component between stacked recurrent layers.The key intuition is that, by leveraging information across all network hierarchies, we can not only improve gradient flow but also improve overall performance.We conduct extensive experiments on six well-studied text sequence matching datasets, achieving state-of-the-art performance on all.
Yi Tay, Anh Tuan Luu, Siu Cheung Hui
EMNLP2
2018 Attentive Gated Lexicon Reader with Contrastive Contextual Co-Attention for Sentiment Classification
abstract
This paper proposes a new neural architecture that exploits readily available sentiment lexicon resources.The key idea is that that incorporating a word-level prior can aid in the representation learning process, eventually improving model performance.To this end, our model employs two distinctly unique components, i.e., (1) we introduce a lexicon-driven contextual attention mechanism to imbue lexicon words with long-range contextual information and (2), we introduce a contrastive co-attention mechanism that models contrasting polarities between all positive and negative words in a sentence.Via extensive experiments, we show that our approach outperforms many other neural baselines on sentiment classification tasks on multiple benchmark datasets.
Yi Tay, Anh Tuan Luu, Siu Cheung Hui, Jian Su 0002
EMNLP2
2018 CoupleNet: Paying Attention to Couples with Coupled Attention for Relationship Recommendation
Yi Tay, Anh Tuan Luu, Siu Cheung Hui
ICWSM2
2018 Hermitian Co-Attention Networks for Text Matching in Asymmetrical Domains
abstract
Co-Attentions are highly effective attention mechanisms for text matching applications. Co-Attention enables the learning of pairwise attentions, i.e., learning to attend based on computing word-level affinity scores between two documents. However, text matching problems can exist in either symmetrical or asymmetrical domains. For example, paraphrase identification is a symmetrical task while question-answer matching and entailment classification are considered asymmetrical domains. In this paper, we argue that Co-Attention models in asymmetrical domains require different treatment as opposed to symmetrical domains, i.e., a concept of word-level directionality should be incorporated while learning word-level similarity scores. Hence, the standard inner product in real space commonly adopted in co-attention is not suitable. This paper leverages attractive properties of the complex vector space and proposes a co-attention mechanism based on the complex-valued inner product (Hermitian products). Unlike the real dot product, the dot product in complex space is asymmetric because the first item is conjugated. Aside from modeling and encoding directionality, our proposed approach also enhances the representation learning process. Extensive experiments on five text matching benchmark datasets demonstrate the effectiveness of our approach.
Yi Tay, Anh Tuan Luu, Siu Cheung Hui
IJCAI2
2018 Multi-Pointer Co-Attention Networks for Recommendation
abstract
Many recent state-of-the-art recommender systems such as D-ATT, TransNet and DeepCoNN exploit reviews for representation learning. This paper proposes a new neural architecture for recommendation with reviews. Our model operates on a multi-hierarchical paradigm and is based on the intuition that not all reviews are created equal, i.e., only a selected few are important. The importance, however, should be dynamically inferred depending on the current target. To this end, we propose a review-by-review pointer-based learning scheme that extracts important reviews from user and item reviews and subsequently matches them in a word-by-word fashion. This enables not only the most informative reviews to be utilized for prediction but also a deeper word-level interaction. Our pointer-based method operates with a gumbel-softmax based pointer mechanism that enables the incorporation of discrete vectors within differentiable neural architectures. Our pointer mechanism is co-attentive in nature, learning pointers which are co-dependent on user-item relationships. Finally, we propose a multi-pointer learning scheme that learns to combine multiple views of user-item interactions. We demonstrate the effectiveness of our proposed model via extensive experiments on 24 benchmark datasets from Amazon and Yelp. Empirical results show that our approach significantly outperforms existing state-of-the-art models, with up to 19% and 71% relative improvement when compared to TransNet and DeepCoNN respectively. We study the behavior of our multi-pointer learning mechanism, shedding light on 'evidence aggregation' patterns in review-based recommender systems.
Yi Tay, Anh Tuan Luu, Siu Cheung Hui
KDD2
2018 Multi-Cast Attention Networks
abstract
Attention is typically used to select informative sub-phrases that are used for prediction. This paper investigates the novel use of attention as a form of feature augmentation, i.e, casted attention. We propose Multi-Cast Attention Networks (MCAN), a new attention mechanism and general model architecture for a potpourri of ranking tasks in the conversational modeling and question answering domains. Our approach performs a series of soft attention operations, each time casting a scalar feature upon the inner word embeddings. The key idea is to provide a real-valued hint (feature) to a subsequent encoder layer and is targeted at improving the representation learning process. There are several advantages to this design, e.g., it allows an arbitrary number of attention mechanisms to be casted, allowing for multiple attention types (e.g., co-attention, intra-attention) and attention variants (e.g., alignment-pooling, max-pooling, mean-pooling) to be executed simultaneously. This not only eliminates the costly need to tune the nature of the co-attention layer, but also provides greater extents of explainability to practitioners. Via extensive experiments on four well-known benchmark datasets, we show that MCAN achieves state-of-the-art performance. On the Ubuntu Dialogue Corpus, MCAN outperforms existing state-of-the-art models by 9%. MCAN also achieves the best performing score to date on the well-studied TrecQA dataset.
Yi Tay, Anh Tuan Luu, Siu Cheung Hui
KDD2
2018 Recurrently Controlled Recurrent Networks
abstract
Recurrent neural networks (RNNs) such as long short-term memory and gated recurrent units are pivotal building blocks across a broad spectrum of sequence modeling problems. This paper proposes a recurrently controlled recurrent network (RCRN) for expressive and powerful sequence encoding. More concretely, the key idea behind our approach is to learn the recurrent gating functions using recurrent networks. Our architecture is split into two components - a controller cell and a listener cell whereby the recurrent controller actively influences the compositionality of the listener cell. We conduct extensive experiments on a myriad of tasks in the NLP domain such as sentiment analysis (SST, IMDb, Amazon reviews, etc.), question classification (TREC), entailment classification (SNLI, SciTail), answer selection (WikiQA, TrecQA) and reading comprehension (NarrativeQA). Across all 26 datasets, our results demonstrate that RCRN not only consistently outperforms BiLSTMs but also stacked BiLSTMs, suggesting that our controller architecture might be a suitable replacement for the widely adopted stacked architecture. Additionally, RCRN achieves state-of-the-art results on several well-established datasets.
Yi Tay, Anh Tuan Luu, Siu Cheung Hui
NeurIPS2
2018 Densely Connected Attention Propagation for Reading Comprehension
abstract
We propose DecaProp (Densely Connected Attention Propagation), a new densely connected neural architecture for reading comprehension (RC). There are two distinct characteristics of our model. Firstly, our model densely connects all pairwise layers of the network, modeling relationships between passage and query across all hierarchical levels. Secondly, the dense connectors in our network are learned via attention instead of standard residual skip-connectors. To this end, we propose novel Bidirectional Attention Connectors (BAC) for efficiently forging connections throughout the network. We conduct extensive experiments on four challenging RC benchmarks. Our proposed approach achieves state-of-the-art results on all four, outperforming existing baselines by up to 2.6% to 14.2% in absolute F1 score.
Yi Tay, Anh Tuan Luu, Siu Cheung Hui, Jian Su 0002
NeurIPS2
2018 Hyperbolic Representation Learning for Fast and Efficient Neural Question Answering
abstract
The dominant neural architectures in question answer retrieval are based on recurrent or convolutional encoders configured with complex word matching layers. Given that recent architectural innovations are mostly new word interaction layers or attention-based matching mechanisms, it seems to be a well-established fact that these components are mandatory for good performance. Unfortunately, the memory and computation cost incurred by these complex mechanisms are undesirable for practical applications. As such, this paper tackles the question of whether it is possible to achieve competitive performance with simple neural architectures. We propose a simple but novel deep learning architecture for fast and efficient question-answer ranking and retrieval. More specifically, our proposed model, HyperQA, is a parameter efficient neural network that outperforms other parameter intensive models such as Attentive Pooling BiLSTMs and Multi-Perspective CNNs on multiple QA benchmarks. The novelty behind HyperQA is a pairwise ranking objective that models the relationship between question and answer embeddings in Hyperbolic space instead of Euclidean space. This empowers our model with a self-organizing ability and enables automatic discovery of latent hierarchies while learning embeddings of questions and answers. Our model requires no feature engineering, no similarity matrix matching, no complicated attention mechanisms nor over-parameterized layers and yet outperforms and remains competitive to many models that have these functionalities on multiple benchmarks.
Yi Tay, Anh Tuan Luu, Siu Cheung Hui
WSDM2
2018 Latent Relational Metric Learning via Memory-based Attention for Collaborative Ranking
abstract
This paper proposes a new neural architecture for collaborative ranking with implicit feedback. Our model, LRML (Latent Relational Metric Learning) is a novel metric learning approach for recommendation. More specifically, instead of simple push-pull mechanisms between user and item pairs, we propose to learn latent relations that describe each user item interaction. This helps to alleviate the potential geometric inflexibility of existing metric learning approaches. This enables not only better performance but also a greater extent of modeling capability, allowing our model to scale to a larger number of interactions. In order to do so, we employ a augmented memory module and learn to attend over these memory blocks to construct latent relations. The memory-based attention module is controlled by the user-item interaction, making the learned relation vector specific to each user-item pair. Hence, this can be interpreted as learning an exclusive and optimal relational translation for each user-item interaction. The proposed architecture demonstrates the state-of-the-art performance across multiple recommendation benchmarks. LRML outperforms other metric learning models by 6%-7.5% in terms of [email protected] and [email protected] on large datasets such as Netflix and MovieLens20M. Moreover, qualitative studies also demonstrate evidence that our proposed model is able to infer and encode explicit sentiment, temporal and attribute information despite being only trained on implicit feedback. As such, this ascertains the ability of LRML to uncover hidden relational structure within implicit datasets.
Yi Tay, Anh Tuan Luu, Siu Cheung Hui
WWW2
2018 Personalized question recommendation for English grammar learning
abstract
Abstract Learning English grammar is a very challenging task for many students especially for nonnative English speakers. To learn English well, it is important to understand the concepts of the English grammar with lots of practise on exercise questions. Previous recommendation systems for learning English mainly focused on recommending reading materials and vocabulary. Different from reading material and vocabulary recommendations, grammar question recommendation should recommend questions that have similar grammatical structure and usage to the question of interest. The content similarity calculation methods used in existing recommendation methods cannot represent the similarity between grammar questions effectively. In this paper, we propose a content‐based approach for personalized grammar question recommendation, which recommends similar grammatical structure and usage questions for further practising. Specifically, we propose a novel structure namedparse‐key treeto capture the grammatical structure and usage of grammar questions. We then propose 3 measures to compute the similarity between the question query and database questions for grammar question recommendation. Additionally, we incorporated the proposed recommendation method into a Web‐based English grammar learning system and presented its performance evaluation in this paper. The experimental results have shown that the proposed approach outperforms other classical and state‐of‐the‐art methods in recommending relevant grammar questions.
Lanting Fang, Anh Tuan Luu, Siu Cheung Hui, Lenan Wu
Expert Syst. J. Knowl. Eng.2
2018 Syntactic based approach for grammar question retrieval
abstract
With the popularity of online educational platforms, English learners can learn and practice no matter where they are and what they do. English grammar is one of the important components in learning English. To learn English grammar effectively, it requires students to practice questions containing focused grammar knowledge. In this paper, we study a novel problem of retrieving English grammar questions with similar grammatical focus. Since the grammatical focus similarity is different from textual similarity or sentence syntactic similarity, existing approaches cannot be applied directly to our problem. To address this problem, we propose a syntactic based approach for English grammar question retrieval which can retrieve related grammar questions with similar grammatical focus effectively. In the proposed syntactic based approach, we first propose a new syntactic tree, namely parse-key tree, to capture English grammar questions’ grammatical focus. Next, we propose two kernel functions , namely relaxed tree kernel and part-of-speech order kernel, to compute the similarity between two parse-key trees of the query and grammar questions in the collection. Then, the retrieved grammar questions are ranked according to the similarity between the parse-key trees. In addition, if a query is submitted together with answer choices, conceptual similarity and textual similarity are also incorporated to further improve the retrieval accuracy . The performance results have shown that our proposed approach outperforms the state-of-the-art methods based on statistical analysis and syntactic analysis.
Lanting Fang, Anh Tuan Luu, Siu Cheung Hui, Lenan Wu
Inf. Process. Manag.2
2017 Non-Parametric Estimation of Multiple Embeddings for Link Prediction on Dynamic Knowledge Graphs
abstract
Knowledge graphs play a significant role in many intelligent systems such as semantic search and recommendation systems. Recent works in this area of knowledge graph embeddings such as TransE, TransH and TransR have shown extremely competitive and promising results in relational learning. In this paper, we propose a novel extension of the translational embedding model to solve three main problems of the current models. Firstly, translational models are highly sensitive to hyperparameters such as margin and learning rate. Secondly, the translation principle only allows one spot in vector space for each golden triplet. Thus, congestion of entities and relations in vector space may reduce precision. Lastly, the current models are not able to handle dynamic data especially the introduction of new unseen entities/relations or removal of triplets. In this paper, we propose Parallel Universe TransE (puTransE), an adaptable and robust adaptation of the translational model. Our approach non-parametrically estimates the energy score of a triplet from multiple embedding spaces of structurally and semantically aware triplet selection. Our proposed approach is simple, robust and parallelizable. Our experimental results show that our proposed approach outperforms TransE and many other embedding methods for link prediction on knowledge graphs on both public benchmark dataset and a real world dynamic dataset.
Yi Tay, Anh Tuan Luu, Siu Cheung Hui
AAAI2
2017 Dyadic Memory Networks for Aspect-based Sentiment Analysis
abstract
This paper proposes Dyadic Memory Networks (DyMemNN), a novel extension of end-to-end memory networks (memNN) for aspect-based sentiment analysis (ABSA). Originally designed for question answering tasks, memNN operates via a memory selection operation in which relevant memory pieces are adaptively selected based on the input query. In the problem of ABSA, this is analogous to aspects and documents in which the relationship between each word in the document is compared with the aspect vector. In the standard memory networks, simple dot products or feed forward neural networks are used to model the relationship between aspect and words which lacks representation learning capability. As such, our dyadic memory networks ameliorates this weakness by enabling rich dyadic interactions between aspect and word embeddings by integrating either parameterized neural tensor compositions or holographic compositions into the memory selection operation. To this end, we propose two variations of our dyadic memory networks, namely the Tensor DyMemNN and Holo DyMemNN. Overall, our two models are end-to-end neural architectures that enable rich dyadic interaction between aspect and document which intuitively leads to better performance. Via extensive experiments, we show that our proposed models achieve the state-of-the-art performance and outperform many neural architectures across six benchmark datasets.
Yi Tay, Anh Tuan Luu, Siu Cheung Hui
CIKM2
2017 Multi-Task Neural Network for Non-discrete Attribute Prediction in Knowledge Graphs
abstract
Many popular knowledge graphs such as Freebase, YAGO or DBPedia maintain a list of non-discrete attributes for each entity. Intuitively, these attributes such as height, price or population count are able to richly characterize entities in knowledge graphs. This additional source of information may help to alleviate the inherent sparsity and incompleteness problem that are prevalent in knowledge graphs. Unfortunately, many state-of-the-art relational learning models ignore this information due to the challenging nature of dealing with non-discrete data types in the inherently binary-natured knowledge graphs. In this paper, we propose a novel multi-task neural network approach for both encoding and prediction of non-discrete attribute information in a relational setting. Specifically, we train a neural network for triplet prediction along with a separate network for attribute value regression. Via multi-task learning, we are able to learn representations of entities, relations and attributes that encode information about both tasks. Moreover, such attributes are not only central to many predictive tasks as an information source but also as a prediction target. Therefore, models that are able to encode, incorporate and predict such information in a relational learning context are highly attractive as well. We show that our approach outperforms many state-of-the-art methods for the tasks of relational triplet classification and attribute value prediction.
Yi Tay, Anh Tuan Luu, Minh C. Phan, Siu Cheung Hui
CIKM2
2017 A Syntactic Parse-Key Tree-Based Approach for English Grammar Question Retrieval
Lanting Fang, Anh Tuan Luu, Lenan Wu, Siu Cheung Hui
NLDB2
2017 Learning to Rank Question Answer Pairs with Holographic Dual LSTM Architecture
abstract
We describe a new deep learning architecture for learning to rank question answer pairs. Our approach extends the long short-term memory (LSTM) network with holographic composition to model the relationship between question and answer representations. As opposed to the neural tensor layer that has been adopted recently, the holographic composition provides the benefits of scalable and rich representational learning approach without incurring huge parameter costs. Overall, we present Holographic Dual LSTM (HD-LSTM), a unified architecture for both deep sentence modeling and semantic matching. Essentially, our model is trained end-to-end whereby the parameters of the LSTM are optimized in a way that best explains the correlation between question and answer representations. In addition, our proposed deep learning architecture requires no extensive feature engineering. Via extensive experiments, we show that HD-LSTM outperforms many other neural architectures on two popular benchmark QA datasets. Empirical studies confirm the effectiveness of holographic composition over the neural tensor layer.
Yi Tay, Minh C. Phan, Anh Tuan Luu, Siu Cheung Hui
SIGIR3
2017 Random Semantic Tensor Ensemble for Scalable Knowledge Graph Link Prediction
abstract
Link prediction on knowledge graphs is useful in numerous application areas such as semantic search, question answering, entity disambiguation, enterprise decision support, recommender systems and so on. While many of these applications require a reasonably quick response and may operate on data that is constantly changing, existing methods often lack speed and adaptability to cope with these requirements. This is aggravated by the fact that knowledge graphs are often extremely large and may easily contain millions of entities rendering many of these methods impractical. In this paper, we address the weaknesses of current methods by proposing Random Semantic Tensor Ensemble (RSTE), a scalable ensemble-enabled framework based on tensor factorization. Our proposed approach samples a knowledge graph tensor in its graph representation and performs link prediction via ensembles of tensor factorization. Our experiments on both publicly available datasets and real world enterprise/sales knowledge bases have shown that our approach is not only highly scalable, parallelizable and memory efficient, but also able to increase the prediction accuracy significantly across all datasets.
Yi Tay, Anh Tuan Luu, Siu Cheung Hui, Falk Brauer
WSDM2
2016 Learning Term Embeddings for Taxonomic Relation Identification Using Dynamic Weighting Neural Network
abstract
Taxonomic relation identification aims to recognize the 'is-a' relation between two terms.Previous works on identifying taxonomic relations are mostly based on statistical and linguistic approaches, but the accuracy of these approaches is far from satisfactory.In this paper, we propose a novel supervised learning approach for identifying taxonomic relations using term embeddings.For this purpose, we first design a dynamic weighting neural network to learn term embeddings based on not only the hypernym and hyponym terms, but also the contextual information between them.We then apply such embeddings as features to identify taxonomic relations using a supervised method.The experimental results show that our proposed approach significantly outperforms other state-of-the-art methods by 9% to 13% in terms of accuracy for both general and specific domain datasets.Recently, Yu et al. (2015) proposed a super-
Anh Tuan Luu, Yi Tay, Siu Cheung Hui, See-Kiong Ng
EMNLP1
2016 Utilizing Temporal Information for Taxonomy Construction
abstract
Taxonomies play an important role in many applications by organizing domain knowledge into a hierarchy of ‘ is-a’ relations between terms. Previous work on automatic construction of taxonomies from text documents either ignored temporal information or used fixed time periods to discretize the time series of documents. In this paper, we propose a time-aware method to automatically construct and effectively maintain a taxonomy from a given series of documents preclustered for a domain of interest. The method extracts temporal information from the documents and uses a timestamp contribution function to score the temporal relevance of the evidence from source texts when identifying the taxonomic relations for constructing the taxonomy. Experimental results show that our proposed method outperforms the state-of-the-art methods by increasing F-measure up to 7%–20%. Furthermore, the proposed method can incrementally update the taxonomy by adding fresh relations from new data and removing outdated relations using an information decay function. It thus avoids rebuilding the whole taxonomy from scratch for every update and keeps the taxonomy effectively up-to-date in order to track the latest information trends in the rapidly evolving domain.
Anh Tuan Luu, Siu Cheung Hui, See-Kiong Ng
Trans. Assoc. Comput. Linguistics1
2015 Incorporating Trustiness and Collective Synonym/Contrastive Evidence into Taxonomy Construction
abstract
Taxonomy plays an important role in many applications by organizing domain knowledge into a hierarchy of is-a relations between terms.Previous works on the taxonomic relation identification from text corpora lack in two aspects: 1) They do not consider the trustiness of individual source texts, which is important to filter out incorrect relations from unreliable sources.2) They also do not consider collective evidence from synonyms and contrastive terms, where synonyms may provide additional supports to taxonomic relations, while contrastive terms may contradict them.In this paper, we present a method of taxonomic relation identification that incorporates the trustiness of source texts measured with such techniques as PageRank and knowledge-based trust, and the collective evidence of synonyms and contrastive terms identified by linguistic pattern matching and machine learning.The experimental results show that the proposed features can consistently improve performance up to 4%-10% of F-measure.
Anh Tuan Luu, Jung-Jae Kim 0001, See-Kiong Ng
EMNLP1
2014 Taxonomy Construction Using Syntactic Contextual Evidence
abstract
Taxonomies are the backbone of many structured, semantic knowledge resources.Recent works for extracting taxonomic relations from text focused on collecting lexical-syntactic patterns to extract the taxonomic relations by matching the patterns to text.These approaches, however, often show low coverage due to the lack of contextual analysis across sentences.To address this issue, we propose a novel approach that collectively utilizes contextual information of terms in syntactic structures such that if the set of contexts of a term includes most of contexts of another term, a subsumption relation between the two terms is inferred.We apply this method to the task of taxonomy construction from scratch, where we introduce another novel graph-based algorithm for taxonomic structure induction.Our experiment results show that the proposed method is well complementary with previous methods of linguistic pattern matching and significantly improves recall and thus F-measure.
Anh Tuan Luu, Jung-Jae Kim 0001, See-Kiong Ng
EMNLP1
2012 SeVe: automatic tool for verification of security protocols
Anh Tuan Luu, Jun Sun 0001, Yang Liu 0003, Jin Song Dong 0001, Xiaohong Li 0001, Thanh Tho Quan
Frontiers Comput. Sci. China1