Tu Vu

dblp:186/7716 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 Discord: Enhancing knowledge distillation via cross-chain of thought and optimal transport alignment between models with different tokenizers
Anh Duc Le, Tu Vu, Ngo Van Linh 0001
Knowl. Based Syst.3
2026 MoL: Mixture of Layers in Cross-Tokenizer embedding model distillation
Hai An Vu, Minh-Phuc Truong, Tu Vu, Ngo Van Linh 0001
Knowl. Based Syst.3
2025 Efficient Model Development through Fine-tuning Transfer
abstract
Modern LLMs face a major obstacle: each new pre-trained model version requires expensive and repetitive alignment.We propose a method that transfers fine-tuning updates across model versions.The key idea is to extract the diff vector, which is the difference in parameters induced by fine-tuning, from a source model version and apply it to the base of a different target version.We show that transferring diff vectors significantly improves the target base model, often achieving performance comparable to its fine-tuned counterpart.For example, applying the fine-tuning updates from Llama 3.0 8B to Llama 3.1 8B increases accuracy by 46.9% on IFEval and 15.7% on Live-CodeBench without further training, surpassing Llama 3.1 8B Instruct.In multilingual settings, we also observe accuracy gains relative to Llama 3.1 8B Instruct, including 4.7% for Malagasy and 15.5% for Turkish on Global MMLU.Our controlled experiments reveal that fine-tuning transfer works best when source and target models are linearly connected in parameter space.We also show that this transfer provides a stronger and more efficient starting point for subsequent fine-tuning.Finally, we propose an iterative recycling-then-finetuning approach for continuous model development, which improves both efficiency and effectiveness.Our findings suggest that fine-tuning transfer is a viable strategy to reduce training costs while maintaining model performance.1
Pin-Jie Lin, Rishab Balasubramanian, Nikhil Kandpal, Tu Vu
EMNLP5
2025 EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport Alignments
abstract
Knowledge distillation (KD) is crucial for compressing large text embedding models, but faces challenges when teacher and student models use different tokenizers (Cross-Tokenizer KD -CTKD).Vocabulary mismatches impede the transfer of relational knowledge encoded in deep representations, such as hidden states and attention matrices, which are vital for producing high-quality embeddings.Existing CTKD methods often focus on direct output alignment, neglecting this crucial structural information.We propose a novel framework tailored for CTKD embedding model distillation.We first map tokens one-to-one via Minimum Edit Distance (MinED).Then, we distill intra-model relational knowledge by aligning attention matrix patterns using Centered Kernel Alignment, focusing on the top-m most important tokens of the directly mapped tokens.Simultaneously, we align final hidden states via Optimal Transport with Importance-Scored Mass Assignment, which emphasizes semantically important token representations, based on importance scores derived from attention weights.We evaluate distillation from state-of-the-art embedding models (e.g., LLM2Vec, BGE) to a Bert-base-uncased model on embedding-reliant tasks such as text classification, sentence pair classification, and semantic textual similarity.Our proposed framework significantly outperforms existing CTKD baselines.By preserving attention structure and prioritizing key representations, our approach yields smaller, highfidelity embedding models despite tokenizer differences.
Minh-Phuc Truong, Hai An Vu, Tu Vu, Nguyen Thi Ngoc Diep, Ngo Van Linh 0001, Thien Huu Nguyen, Trung Le 0001
EMNLP3
2024 Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation
abstract
As large language models (LLMs) evolve, evaluating their output reliably becomes increasingly difficult due to the high cost of human evaluation.To address this, we introduce FLAMe, a family of Foundational Large Autorater Models.FLAMe is trained on a diverse set of over 100 quality assessment tasks, incorporating 5M+ human judgments curated from publicly released human evaluations.FLAMe outperforms models like GPT-4 and Claude-3 on various held-out tasks, and serves as a powerful starting point for finetuning, as shown in our reward model evaluation case study (FLAMe-RM).On Reward-Bench, FLAMe-RM-24B achieves 87.8% accuracy, surpassing GPT-4-0125 (85.9%) and GPT-4o (84.7%).Additionally, we introduce FLAMe-Opt-RM, an efficient tail-patch finetuning approach that offers competitive Re-wardBench performance using 25× fewer training datapoints.Our FLAMe variants outperform popular proprietary LLM-as-a-Judge models on 8 of 12 autorater benchmarks, covering 53 quality assessment tasks, including RewardBench and LLM-AggreFact.Finally, our analysis shows that FLAMe is significantly less biased than other LLM-as-a-Judge models on the CoBBLEr autorater bias benchmark.1 * Tu Vu and Kalpesh Krishna contributed equally to the project leadership, design, and implementation of the work.† Work done while at UMass Amherst.‡ Equal contribution as senior advisors.1 The FLAMe collection is available at https:// huggingface.co/datasets/google/flame-collection."""Input format.""" INSTRUCTIONS:"""Task definition and evaluation instructions."""
Tu Vu, Kalpesh Krishna, Salaheddin Alzubi, Chris Tar, Manaal Faruqui, Yun-Hsuan Sung
EMNLP1
2024 Mixture-of-Experts Meets Instruction Tuning: A Winning Combination for Large Language Models
abstract
Sparse Mixture-of-Experts (MoE) is a neural architecture design that adds learnable parameters to Large Language Models (LLMs) without increasing computational complexity (FLOPs). Instruction tuning is a technique for training LLMs to follow instructions. We advocate combining these two approaches, as we find that MoE models benefit more from instruction tuning than dense models. In particular, we conduct empirical studies across three experimental setups: (i) Direct finetuning on individual downstream tasks devoid of instruction tuning; (ii) Instruction tuning followed by in-context few-shot or zero-shot generalization on downstream tasks; and (iii) Instruction tuning supplemented by further finetuning on individual downstream tasks. In the first scenario, MoE models overall underperform dense models of identical computational capacity. This narrative, however, dramatically changes with the introduction of instruction tuning (in the second and third scenarios), used independently or in conjunction with task-specific finetuning. Our most powerful model, FLAN-MoE-32B, surpasses the performance of Flan-PaLM-62B on four benchmark tasks, while using only a third of the FLOPs. The advancements embodied by FLAN-MoE inspire a reevaluation of the design principles of large-scale, high-performance language models in the framework of task-agnostic learning.
Sheng Shen 0001, Le Hou, Yanqi Zhou, Nan Du 0002, Shayne Longpre, Jason Wei, Hyung Won Chung, Barret Zoph, William Fedus, Tu Vu, Yuexin Wu, Wuyang Chen 0001, Albert Webson, Yunxuan Li, Vincent Y. Zhao, Hongkun Yu 0001, Kurt Keutzer, Trevor Darrell, Denny Zhou
ICLR11
2023 Dialect-robust Evaluation of Generated Text
abstract
Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, Sebastian Gehrmann. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, Sebastian Gehrmann
ACL (1)4
2023 The Flan Collection: Designing Data and Methods for Effective Instruction Tuning
abstract
We study the design decision of publicly available instruction tuning methods, by reproducing and breaking down the development of Flan 2022 (Chung et al., 2022). Through careful ablation studies on the Flan Collection of tasks and methods, we tease apart the effect of design decisions which enable Flan-T5 to outperform prior work by 3-17% across evaluation settings. We find task balancing and enrichment techniques are overlooked but critical to effective instruction tuning, and in particular, training with mixed prompt settings (zero-shot, few-shot, chain-of-thought) actually yields equivalent or stronger (2%) performance in all settings. In further experiments we show Flan-T5 requires less finetuning to converge higher and faster than T5 on single downstream tasks -- motivating instruction-tuned models as more computationally-efficient starting checkpoints for new tasks. Finally, to accelerate research on instruction tuning, we make the Flan 2022 collection of datasets, templates, and methods publicly available.
Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V. Le, Barret Zoph, Jason Wei, Adam Roberts
ICML3
2022 SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer
abstract
There has been growing interest in parameter-efficient methods to apply pre-trained language models to downstream tasks. Building on the Prompt Tuning approach of Lester et al. (2021), which learns task-specific soft prompts to condition a frozen pre-trained model to perform different tasks, we propose a novel prompt-based transfer learning approach called SPoT: Soft Prompt Transfer. SPoT first learns a prompt on one or more source tasks and then uses it to initialize the prompt for a target task. We show that SPoT significantly boosts the performance of Prompt Tuning across many tasks. More remarkably, across all model sizes, SPoT matches or outperforms standard Model Tuning (which fine-tunes all model parameters) on the SuperGLUE benchmark, while using up to 27,000× fewer task-specific parameters. To understand where SPoT is most effective, we conduct a large-scale study on task transferability with 26 NLP tasks in 160 combinations, and demonstrate that many tasks can benefit each other via prompt transfer. Finally, we propose an efficient retrieval approach that interprets task prompts as task embeddings to identify similar tasks and predict the most transferable source tasks for a novel target task.
Tu Vu, Brian Lester, Noah Constant, Rami Al-Rfou, Daniel M. Cer
ACL (1)1
2022 Leveraging QA Datasets to Improve Generative Data Augmentation
abstract
The ability of generative language models (GLMs) to generate text has improved considerably in the last few years, enabling their use for generative data augmentation.In this work, we propose CONDA, an approach to further improve GLMs' ability to generate synthetic data by reformulating data generation as context generation for a given question-answer (QA) pair and leveraging QA datasets for training context generators.Then, we cast downstream tasks into the same question answering format and adapt the fine-tuned context generators to the target task domain.Finally, we use the fine-tuned GLM to generate relevant contexts, which are in turn used as synthetic training data for their corresponding tasks.We perform extensive experiments on multiple classification datasets and demonstrate substantial improvements in performance for both few-and zeroshot settings.Our analysis reveals that QA datasets that require high-level reasoning abilities (e.g., abstractive and common-sense QA datasets) tend to give the best boost in performance in both few-shot and zero-shot settings.
Dheeraj Mekala, Tu Vu, Timo Schick, Jingbo Shang
EMNLP2
2022 Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual Generation
abstract
In this paper, we explore the challenging problem of performing a generative task in a target language when labeled data is only available in English, using summarization as a case study.We assume a strict setting with no access to parallel data or machine translation and find that common transfer learning approaches struggle in this setting, as a generative multilingual model fine-tuned purely on English catastrophically forgets how to generate non-English.Given the recent rise of parameter-efficient adaptation techniques, we conduct the first investigation into how one such method, prompt tuning (Lester et al., 2021), can overcome catastrophic forgetting to enable zero-shot cross-lingual generation.Our experiments show that parameter-efficient prompt tuning provides gains over standard fine-tuning when transferring between lessrelated languages, e.g., from English to Thai.However, a significant gap still remains between these methods and fully-supervised baselines.To improve cross-lingual transfer further, we explore several approaches, including: (1) mixing in unlabeled multilingual data, and (2) explicitly factoring prompts into recombinable language and task components.Our approaches can provide further quality gains, suggesting that robust zero-shot crosslingual generation is within reach.ING and standard MODELTUNING for zero-shot cross-lingual generation (XGEN).We show that increasing model scale and decreasing tunable parameter capacity are key for overcoming catastrophic forgetting on XGEN.• We propose WIKILINGUA-0, a challenging XGEN benchmark and an associated SP-ROUGE evaluation metric, which we hope will facilitate future work evaluating multilingual summarization.• We show that mixing in unsupervised multilingual data can boost XGEN performance, and are the first to combine this approach with PROMPTTUNING.• We propose "factorized prompts", a novel approach that can also help PROMPTTUNING overcome severe catastrophic forgetting.• To facilitate future work, we release our data, pretrained models
Tu Vu, Aditya Barua, Brian Lester, Daniel M. Cer, Mohit Iyyer, Noah Constant
EMNLP1
2021 STraTA: Self-Training with Task Augmentation for Better Few-shot Learning
abstract
Despite their recent successes in tackling many NLP tasks, large-scale pre-trained language models do not perform as well in few-shot settings where only a handful of training examples are available.To address this shortcoming, we propose STraTA, which stands for Self-Training with Task Augmentation, an approach that builds on two key ideas for effective leverage of unlabeled data.First, STraTA uses task augmentation, a novel technique that synthesizes a large amount of data for auxiliary-task fine-tuning from target-task unlabeled texts.Second, STraTA performs selftraining by further fine-tuning the strong base model created by task augmentation on a broad distribution of pseudo-labeled data.Our experiments demonstrate that STraTA can substantially improve sample efficiency across 12 fewshot benchmarks.Remarkably, on the SST-2 sentiment dataset, STraTA, with only 8 training examples per class, achieves comparable results to standard fine-tuning with 67K training examples.Our analyses reveal that task augmentation and self-training are both complementary and independently effective.
Tu Vu, Minh-Thang Luong, Quoc V. Le, Grady Simon, Mohit Iyyer
EMNLP (1)1
2020 Exploring and Predicting Transferability across NLP Tasks
abstract
Tu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni, Adam Trischler, Andrew Mattarella-Micke, Subhransu Maji, Mohit Iyyer. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Tu Vu, Tong Wang 0012, Tsendsuren Munkhdalai, Alessandro Sordoni, Adam Trischler, Andrew Mattarella-Micke, Subhransu Maji, Mohit Iyyer
EMNLP (1)1
2019 Encouraging Paragraph Embeddings to Remember Sentence Identity Improves Classification
abstract
While paragraph embedding models are remarkably effective for downstream classification tasks, what they learn and encode into a single vector remains opaque.In this paper, we investigate a state-of-the-art paragraph embedding method proposed by Zhang et al. (2017) and discover that it cannot reliably tell whether a given sentence occurs in the input paragraph or not.We formulate a sentence content task to probe for this basic linguistic property and find that even a much simpler bag-of-words method has no trouble solving it.This result motivates us to replace the reconstructionbased objective of Zhang et al. (2017) with our sentence content probe objective in a semisupervised setting.Despite its simplicity, our objective improves over paragraph reconstruction in terms of (1) downstream classification accuracies on benchmark datasets, (2) faster training, and (3) better generalization ability. 1
Tu Vu, Mohit Iyyer
ACL (1)1
2017 Customer churn prediction in an internet service provider
abstract
Customer retention is regarded as one of the most concerns in any company, since they provide the fundamental source of revenue for business. Losing customers not only loses the profit, but also may put a whole business in danger. In order to increase customer base, businesses need to improve both acquisition and retention of its customers. Therefore, customer churn prediction is becoming the top concerns that many companies devoted their time and resources to address it. This paper presents the customer churn prediction on an extremely imbalanced data in an Internet Service Provider company to identify the users at the risk of leaving the services. It consists of feature engineering and predictive modeling. In the feature engineering, the most essential features are selected from a large number of created candidates. In the predictive modeling, the imbalance between the number of churners and non-churners was reduced using SMOTE oversampling technique before implementing several models such as AdaBoost, Extra Trees, KNN, Neural Network and XGBoost. Comparing between these models in term of precision and recall, the XGBoost model gives the highest performance. Using the dataset with 98% non-churners and 2% churners, precision and recall of the model are 45.71% and 42.06%, respectively.
Duyen Do, Phuc Huynh, Phuong Vo, Tu Vu
IEEE BigData4
2017 DNS graph mining for malicious domain detection
abstract
As a vital component of variety cyber attacks, malicious domain detection becomes a hot topic for cyber security. Several recent techniques are proposed to identify malicious domains through analysis of DNS data because much of global information in DNS data which cannot be affected by the attackers. The attackers always recycle resources, so they frequently change the domain - IP resolutions and create new domains to avoid detection. Therefore, multiple malicious domains are hosted by the same IPs and multiple IPs also host same malicious domains in simultaneously, which create intrinsic association among them. Hence, using the labeled domains which can be traced back from queries history of all domains to verify and figure out the association of them all. Graphs seem the best candidate to represent for this relationship and there are many algorithms developed on graph with high performance. A graph-based interface can be developed and transformed to the graph mining task of inferring graph node's reputation scores using improvements of the belief propagation algorithm. Then higher reputation scores the nodes reveal, the more malicious probabilities they infer. For demonstration, this paper proposes a malicious domain detection technique and evaluates on a real-world dataset. The dataset is collected from DNS data servers which will be used for building a DNS graph. The proposed technique achieves high performance in accuracy rates over 98.3%, precision and recall rates as: 99.1%, 98.6%. Especially, with a small set of labeled domains (legitimate and malicious domains), the technique can discover a large set of potential malicious domains. The results indicate that the method is strongly effective in detecting malicious domains.
Hau Tran, Phuong Vo, Tu Vu
IEEE BigData4