VLDB 2026 Research / reviewers in the wild / expert
Thuy-Trang Vu
dblp:228/5538
· DBLP profile ↗
13ranked-venue papers
4as first author
10since 2021 · last 2026
0009-0005-6296-4184ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LiveCultureBench: a Multi-Agent, Multi-Cultural Benchmark for Large Language Models in Dynamic Social SimulationsabstractLarge language models (LLMs) are increasingly deployed as autonomous agents, yet evaluations focus primarily on task success rather than cultural appropriateness or evaluator reliability.We introduce LIVECULTUREBENCH 1 , a multi-cultural, dynamic benchmark that embeds LLMs as agents in a simulated town and evaluates them on both task completion and adherence to socio-cultural norms.The simulation models a small city as a location graph with synthetic residents having diverse demographic and cultural profiles.Each episode assigns one resident a daily goal while others provide social context.An LLM-based verifier generates structured judgments on norm violations and task progress, which we aggregate into metrics capturing task-norm trade-offs and verifier uncertainty.Using LIVECULTUREBENCH across models and cultural profiles, we study (i) cross-cultural robustness of LLM agents, (ii) how they balance effectiveness against norm sensitivity, and (iii) when LLM-as-a-judge evaluation is reliable for automated benchmarking versus when human oversight is needed. Viet Thanh Pham, Lizhen Qu, Thuy-Trang Vu, Gholamreza Haffari, Dinh Q. Phung |
ACL (1) | 3 |
| 2025 | SCAR: Data Selection via Style Consistency-Aware Response Ranking for Efficient Instruction-Tuning of Large Language ModelsabstractZhuang Li, Yuncheng Hua, Thuy-Trang Vu, Haolan Zhan, Lizhen Qu, Gholamreza Haffari. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhuang Li 0001, Yuncheng Hua, Thuy-Trang Vu, Haolan Zhan, Lizhen Qu, Gholamreza Haffari |
ACL (1) | 3 |
| 2025 | Extending LLMs to New Languages: A Case Study of Llama and Persian AdaptationabstractLarge language models (LLMs) have made great progress in classification and text generation tasks. However, they are mainly trained on English data and often struggle with low-resource languages. In this study, we explore adding a new language, i.e., Persian, to Llama (a model with a limited understanding of Persian) using parameter-efficient fine-tuning. We employ a multi-stage approach involving pretraining on monolingual Persian data, aligning representations through bilingual pretraining and instruction datasets, and instruction-tuning with task-specific datasets. We evaluate the model’s performance at each stage on generation and classification tasks. Our findings suggest that incorporating the Persian language, through bilingual data alignment, can enhance classification accuracy for Persian tasks, with no adverse impact and sometimes even improvements on English tasks. Additionally, the results highlight the model’s initial strength as a critical factor when working with limited training data, with cross-lingual alignment offering minimal benefits for the low-resource language. Knowledge transfer from English to Persian has a marginal effect, primarily benefiting simple classification tasks. Samin Mahdizadeh Sani, Pouya Sadeghi, Thuy-Trang Vu, Yadollah Yaghoobzadeh, Gholamreza Haffari |
COLING | 3 |
| 2025 | MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic CorporaabstractContinually updating model-based indexes in generative retrieval with new documents remains challenging, as full retraining is computationally expensive and impractical under resource constraints.We propose MixLoRA-DSI, a novel framework that combines an expandable mixture of Low-Rank Adaptation experts with a layer-wise out-of-distribution (OOD)driven expansion strategy.Instead of allocating new experts for each new corpus, our proposed expansion strategy enables sublinear parameter growth by selectively introducing new experts only when significant number of OOD documents are detected.Experiments on NQ320k and MS MARCO Passage demonstrate that MixLoRA-DSI outperforms full-model update baselines, with minimal parameter overhead and substantially lower training costs. 1 Tuan-Luc Huynh, Thuy-Trang Vu, Weiqing Wang 0001, Trung Le 0001, Dragan Gasevic, Yuan-Fang Li, Thanh-Toan Do |
EMNLP | 2 |
| 2025 | Fantastic Targets for Concept Erasure in Diffusion Models and Where To Find ThemabstractConcept erasure has emerged as a promising technique for mitigating the risk of harmful content generation in diffusion models by selectively unlearning undesirable concepts. The common principle of previous works to remove a specific concept is to map it to a fixed generic concept, such as a neutral concept or just an empty text prompt. In this paper, we demonstrate that this fixed-target strategy is suboptimal, as it fails to account for the impact of erasing one concept on the others. To address this limitation, we model the concept space as a graph and empirically analyze the effects of erasing one concept on the remaining concepts. Our analysis uncovers intriguing geometric properties of the concept space, where the influence of erasing a concept is confined to a local region. Building on this insight, we propose the Adaptive Guided Erasure (AGE) method, which dynamically selects optimal target concepts tailored to each undesirable concept, minimizing unintended side effects. Experimental results show that AGE significantly outperforms state-of-the-art erasure methods on preserving unrelated concepts while maintaining effective erasure performance. Our code is published at {https://github.com/tuananhbui89/Adaptive-Guided-Erasure}. Anh Tuan Bui, Thuy-Trang Vu, Long Tung Vuong, Trung Le 0001, Paul Montague, Tamas Abraham, Junae Kim, Dinh Q. Phung |
ICLR | 2 |
| 2025 | The Best of Both Worlds: Bridging Quality and Diversity in Data Selection with Bipartite GraphabstractThe performance of large language models (LLMs) is strongly influenced by the quality and diversity of data used during supervised fine-tuning (SFT). However, current data selection methods often prioritize one aspect over the other, resulting in suboptimal training outcomes. To address this, we formulate data selection as a set cover problem and present GraphFilter, a novel approach that balances both quality and diversity in data selection. GraphFilter models the dataset as a bipartite graph connecting sentences to their constituent n-grams, then employs a priority function that combines quality and diversity metrics multiplicatively. GraphFilter iteratively selects sentences with the highest priority, removes covered n-grams from the bipartite graph, and recomputes priorities to reflect the changing data landscape. We validate GraphFilter using three model backbones across six widely-used benchmarks, demonstrating that it outperforms nine existing baselines in both model performance and computational efficiency. Further analysis shows that our design choices lead to more effective subset selection, underscores the value of instruction diversity, and provides insights into how quality and diversity interact with different subset sizes. Minghao Wu, Thuy-Trang Vu, Lizhen Qu, Gholamreza Haffari |
ICML | 2 |
| 2025 | SpeechDialogueFactory: A Framework for Natural Speech Dialogue Generation
Minghan Wang, Ye Bai 0002, Thuy-Trang Vu, Ehsan Shareghi, Gholamreza Haffari |
INTERSPEECH | 4 |
| 2025 | PromptDSI: Prompt-Based Rehearsal-Free Continual Learning for Document Retrieval
Tuan-Luc Huynh, Thuy-Trang Vu, Weiqing Wang 0001, Yinwei Wei, Trung Le 0001, Dragan Gasevic, Yuan-Fang Li, Thanh-Toan Do |
ECML/PKDD (7) | 2 |
| 2024 | Mixture-of-Skills: Learning to Optimize Data Usage for Fine-Tuning Large Language ModelsabstractLarge language models (LLMs) are typically fine-tuned on diverse and extensive datasets sourced from various origins to develop a comprehensive range of skills, such as writing, reasoning, chatting, coding, and more.Each skill has unique characteristics, and these datasets are often heterogeneous and imbalanced, making the fine-tuning process highly challenging.Balancing the development of each skill while ensuring the model maintains its overall performance requires sophisticated techniques and careful dataset curation.In this work, we propose a general, model-agnostic, reinforcement learning framework, MIXTURE-OF-SKILLS (MOS), that learns to optimize data usage automatically during the fine-tuning process.This framework ensures the optimal comprehensive skill development of LLMs by dynamically adjusting the focus on different datasets based on their current learning state.To validate the effectiveness of MOS, we conduct extensive experiments using three diverse LLM backbones on two widely used benchmarks and demonstrate that MOS substantially enhances model performance.Building on the success of MOS, we propose MOSPEC, an adaptation for task-specific fine-tuning, which harnesses the utilities of various datasets for a specific purpose.Our work underlines the significance of dataset rebalancing and present MOS as a powerful, general solution for optimizing data usage in the fine-tuning of LLMs for various purposes. Minghao Wu, Thuy-Trang Vu, Lizhen Qu, Reza Haf |
EMNLP | 2 |
| 2021 | Generalised Unsupervised Domain Adaptation of Neural Machine Translation with Cross-Lingual Data SelectionabstractThis paper considers the unsupervised domain adaptation problem for neural machine translation (NMT), where we assume the access to only monolingual text in either the source or target language in the new domain.We propose a cross-lingual data selection method to extract in-domain sentences in the missing language side from a large generic monolingual corpus.Our proposed method trains an adaptive layer on top of multilingual BERT by contrastive learning to align the representation between the source and target language.This then enables the transferability of the domain classifier between the languages in a zero-shot manner.Once the in-domain data is detected by the classifier, the NMT model is then adapted to the new domain by jointly learning translation and domain discrimination tasks.We evaluate our cross-lingual data selection method on NMT across five diverse domains in three language pairs, as well as a real-world scenario of translation for COVID-19.The results show that our proposed method outperforms other selection baselines up to +1.5 BLEU score. Thuy-Trang Vu, Xuanli He, Dinh Q. Phung, Gholamreza Haffari |
EMNLP (1) | 1 |
| 2020 | Effective Unsupervised Domain Adaptation with Adversarially Trained Language ModelsabstractRecent work has shown the importance of adaptation of broad-coverage contextualised embedding models on the domain of the target task of interest.Current self-supervised adaptation methods are simplistic, as the training signal comes from a small percentage of randomly masked-out tokens.In this paper, we show that careful masking strategies can bridge the knowledge gap of masked language models (MLMs) about the domains more effectively by allocating self-supervision where it is needed.Furthermore, we propose an effective training strategy by adversarially masking out those tokens which are harder to reconstruct by the underlying MLM.The adversarial objective leads to a challenging combinatorial optimisation problem over subsets of tokens, which we tackle efficiently through relaxation to a variational lower-bound and dynamic programming.On six unsupervised domain adaptation tasks involving named entity recognition, our method strongly outperforms the random masking strategy and achieves up to +1.64 F1 score improvements. Thuy-Trang Vu, Dinh Q. Phung, Gholamreza Haffari |
EMNLP (1) | 1 |
| 2019 | Learning How to Active Learn by DreamingabstractHeuristic-based active learning (AL) methods are limited when the data distribution of the underlying learning problems vary.Recent data-driven AL policy learning methods are also restricted to learn from closely related domains.We introduce a new sample-efficient method that learns the AL policy directly on the target domain of interest by using wake and dream cycles.Our approach interleaves between querying the annotation of the selected datapoints to update the underlying student learner and improving AL policy using simulation where the current student learner acts as an imperfect annotator.We evaluate our method on cross-domain and cross-lingual text classification and named entity recognition tasks.Experimental results show that our dream-based AL policy training strategy is more effective than applying the pretrained policy without further fine-tuning, and better than the existing strong baseline methods that use heuristics or reinforcement learning. Thuy-Trang Vu, Ming Liu 0028, Dinh Q. Phung, Gholamreza Haffari |
ACL (1) | 1 |
| 2018 | Automatic Post-Editing of Machine Translation: A Neural Programmer-Interpreter ApproachabstractAutomated Post-Editing (PE) is the task of automatically correcting common and repetitive errors found in machine translation (MT) output.In this paper, we present a neural programmer-interpreter approach to this task, resembling the way that humans perform postediting using discrete edit operations, which we refer to as programs.Our model outperforms previous neural models for inducing PE programs on the WMT17 APE task for German-English up to +1 BLEU score and -0.7 TER scores. Thuy-Trang Vu, Gholamreza Haffari |
EMNLP | 1 |