Viet Dac Lai

dblp:251/8546 · DBLP profile ↗
← Back
21ranked-venue papers
8as first author
18since 2021 · last 2026
0009-0008-1651-4619ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 15 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Lizard: An Efficient Linearization Framework for Large Language Models
abstract
Chien Van Nguyen, Huy Huu Nguyen, Ruiyi Zhang, Hanieh Deilamsalehy, Puneet Mathur, Viet Dac Lai, Haoliang Wang, Jayakumar Subramanian, Ryan A. Rossi, Trung Bui, Nikos Vlassis, Franck Dernoncourt, Thien Huu Nguyen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Chien Van Nguyen, Huy Huu Nguyen, Ruiyi Zhang 0002, Hanieh Deilamsalehy, Puneet Mathur, Viet Dac Lai, Jayakumar Subramanian, Ryan Rossi, Trung Bui, Nikos Vlassis, Franck Dernoncourt, Thien Huu Nguyen
ACL (1)6
2026 Understanding Generative AI Capabilities in Everyday Image Editing Tasks
abstract
Generative AI (GenAI) holds significant promise for automating everyday image editing tasks, especially following the recent release of GPT-4o on March 25, 2025. However, what subjects do people most often want edited? What kinds of editing actions do they want to perform (e.g., removing or stylizing the subject)? Do people prefer precise edits with predictable outcomes, or highly creative ones? By understanding the characteristics of real-world requests and the corresponding edits made by freelance photo-editing wizards, can we draw lessons for improving AI-based editors and determine which types of requests can currently be handled successfully by AI editors? In this paper, we present a unique study addressing these questions by analyzing 83k requests with their associated 305k edits from the recent 12 years on the /r/PhotoshopRequest Reddit community. According to human ratings, approximately only 33% of requests can be fulfilled by the best AI editors (including , , ). Interestingly, AI editors perform worse on low-creativity requests that require precise editing than on more open-ended requests. They often struggle to preserve the identity of people and animals, and frequently make non-requested touch-ups. On the other side of the table, VLM judges (e.g., o1) perform differently than human judges and may prefer AI edits over human edits. Code and qualitative examples are available at: https://psrdataset.github.io/.
Brandon Collins, Mohammad Reza Taesiri, Logan Bolton, Viet Dac Lai, Franck Dernoncourt, Trung Bui, Anh Totti Nguyen
WACV4
2025 Language Model Probabilities are Not Calibrated in Numeric Contexts
abstract
Charles Lovering, Michael Krumdick, Viet Dac Lai, Varshini Reddy, Seth Ebner, Nilesh Kumar, Rik Koncel-Kedziorski, Chris Tanner. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Charles Lovering, Michael Krumdick, Viet Dac Lai, Varshini Reddy, Seth Ebner, Nilesh Kumar, Rik Koncel-Kedziorski, Chris Tanner
ACL (1)3
2025 Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization
abstract
Offline reinforcement learning (RL) is a variant of RL where the policy is learned from a previously collected dataset of trajectories and rewards. In our work, we propose a practical approach to offline RL with large language models (LLMs). We recast the problem as reward-weighted fine-tuning, which can be solved using similar techniques to supervised fine-tuning (SFT). To showcase the value of our approach, we apply it to learning short-horizon question-answering policies of a fixed length, where the agent reasons about potential answers or asks clarifying questions. Our work stands in a stark contrast to state-of-the-art methods in this domain, based on SFT and direct preference optimization, which have additional hyper-parameters and do not directly optimize for rewards. We compare to them empirically, and report major gains in both optimized rewards and language quality.
Subhojyoti Mukherjee, Viet Dac Lai, Raghavendra Addanki, Ryan Rossi, Seunghyun Yoon 0002, Trung Bui, Anup B. Rao, Jayakumar Subramanian, Branislav Kveton
NeurIPS2
2025 LUSIFER: Language Universal Space Integration for Enhanced Representation in Multilingual Text Embedding Models
abstract
Recent advancements in large language models (LLMs) based embedding models have established new state-of-the-art benchmarks for text embedding tasks, particularly in dense vector-based retrieval. However, these models predominantly focus on English, leaving multilingual embedding capabilities largely unexplored. To address this limitation, we present LUSIFER, a novel zero-shot approach that adapts LLM-based embedding models for multilingual tasks without requiring multilingual supervision. LUSIFER's architecture combines a multilingual encoder, serving as a language-universal learner, with an LLM-based embedding model optimized for embedding-specific tasks. These components are seamlessly integrated through a minimal set of trainable parameters that act as a connector, effectively transferring the multilingual encoder's language understanding capabilities to the specialized embedding model. Additionally, to comprehensively evaluate multilingual embedding performance, we introduce a new benchmark encompassing 5 primary embedding tasks, 123 diverse datasets, and coverage across 14 languages. Extensive experimental results demonstrate that LUSIFER significantly enhances the multilingual performance across various embedding tasks, particularly for medium and low-resource languages, without requiring explicit multilingual training data. The code and dataset for training are available at: https://github.com/hieum98/lusifer
Hieu Man, Nghia Trung Ngo, Viet Dac Lai, Ryan Rossi, Franck Dernoncourt, Thien Huu Nguyen
SIGIR3
2024 BizBench: A Quantitative Reasoning Benchmark for Business and Finance
abstract
Michael Krumdick, Rik Koncel-Kedziorski, Viet Dac Lai, Varshini Reddy, Charles Lovering, Chris Tanner. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Michael Krumdick, Rik Koncel-Kedziorski, Viet Dac Lai, Varshini Reddy, Charles Lovering, Chris Tanner
ACL (1)3
2024 CAMAL: A Novel Dataset for Multi-label Conversational Argument Move Analysis
abstract
Understanding the discussion moves that teachers and students use to engage in classroom discussions is important to support pre-service teacher learning and teacher educators. This work introduces a novel conversational multi-label corpus of teaching transcripts collected from a simulated classroom environment for Conversational Argument Move AnaLysis (CAMAL). The dataset offers various argumentation moves used by pre-service teachers and students in mathematics and science classroom discussions. The dataset includes 165 transcripts from these discussions that pre-service elementary teachers facilitated in a simulated classroom environment of five student avatars. The discussion transcripts were annotated by education assessment experts for nine argumentation moves (aka. intents) used by the pre-service teachers and students during the discussions. In this paper, we describe the dataset, our annotation framework, and the models we employed to detect argumentation moves. Our experiments with state-of-the-art models demonstrate the complexity of the CAMAL task presented in the dataset. The result reveals that models that combined CNN and LSTM structures with speaker ID graphs improved the F1-score of our baseline models to detect speakers’ intents by a large margin. Given the complexity of the CAMAL task, it creates research opportunities for future studies. We share the dataset, the source code, and the annotation framework publicly at http://github.com/uonlp/camal-dataset.
Viet Dac Lai, Duy Ngoc Pham, Jonathan Steinberg, Jamie Mikeska, Thien Huu Nguyen
LREC/COLING1
2024 CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
abstract
Extensive training datasets represent one of the important factors for the impressive learning capabilities of large language models (LLMs). However, these training datasets for current LLMs, especially the recent state-of-the-art models, are often not fully disclosed. Creating training data for high-performing LLMs involves extensive cleaning and deduplication to ensure the necessary level of quality. The lack of transparency for training data has thus hampered research on attributing and addressing hallucination and bias issues in LLMs, hindering replication efforts and further advancements in the community. These challenges become even more pronounced in multilingual learning scenarios, where the available multilingual text datasets are often inadequately collected and cleaned. Consequently, there is a lack of open-source and readily usable dataset to effectively train LLMs in multiple languages. To overcome this issue, we present CulturaX, a substantial multilingual dataset with 6.3 trillion tokens in 167 languages, tailored for LLM development. Our dataset undergoes meticulous cleaning and deduplication through a rigorous pipeline of multiple stages to accomplish the best quality for model training, including language identification, URL-based filtering, metric-based cleaning, document refinement, and data deduplication. CulturaX is released in Hugging Face facilitate research and advancements in multilingual LLMs: https://huggingface.co/datasets/uonlp/CulturaX.
Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan Rossi, Thien Huu Nguyen
LREC/COLING3
2024 An Analysis of Multilingual FActScore
abstract
FActScore has gained popularity as a metric to estimate the factuality of long-form texts generated by Large Language Models (LLMs) in English.However, there has not been any work in studying the behavior of FActScore in other languages.This paper studies the limitations of each component in the fourcomponent pipeline of FActScore in the multilingual setting.We introduce a new dataset for FActScore on texts generated by strong multilingual LLMs.Our evaluation shows that LLMs exhibit distinct behaviors in both fact extraction and fact scoring tasks.No LLM produces consistent and reliable FActScore across languages with varying levels of resources.We also find that the knowledge source plays an important role in the quality of the estimated FActScore.Using Wikipedia as the knowledge source may hinder the true FActScore of longform text due to its limited coverage in mediumand low-resource languages.We also incorporate three mitigations to our knowledge source that ultimately improve FActScore estimation across all languages.
Vu Trong Kim, Michael Krumdick, Varshini Reddy, Franck Dernoncourt, Viet Dac Lai
EMNLP5
2023 Boosting Punctuation Restoration with Data Generation and Reinforcement Learning
Viet Dac Lai, Abel Salinas, Hao Tan 0002, Trung Bui, Quan Tran, Seunghyun Yoon 0002, Hanieh Deilamsalehy, Franck Dernoncourt, Thien Huu Nguyen
INTERSPEECH1
2022 MECI: A Multilingual Dataset for Event Causality Identification
abstract
Event Causality Identification (ECI) is the task of detecting causal relations between events mentioned in the text. Although this task has been extensively studied for English materials, it is under-explored for many other languages. A major reason for this issue is the lack of multilingual datasets that provide consistent annotations for event causality relations in multiple non-English languages. To address this issue, we introduce a new multilingual dataset for ECI, called MECI. The dataset employs consistent annotation guidelines for five typologically different languages, i.e., English, Danish, Spanish, Turkish, and Urdu. Our dataset thus enable a new research direction on cross-lingual transfer learning for ECI. Our extensive experiments demonstrate high quality for MECI that can provide ample research challenges and directions for future research. We will publicly release MECI to promote research on multilingual ECI.
Viet Dac Lai, Amir Pouran Ben Veyseh, Minh Nguyen 0007, Franck Dernoncourt, Thien Huu Nguyen
COLING1
2022 Event Extraction in Video Transcripts
abstract
Event extraction (EE) is one of the fundamental tasks for information extraction whose goal is to identify mentions of events and their participants in text. Due to its importance, different methods and datasets have been introduced for EE. However, existing EE datasets are limited to formally written documents such as news articles or scientific papers. As such, the challenges of EE in informal and noisy texts are not adequately studied. In particular, video transcripts constitute an important domain that can benefit tremendously from EE systems (e.g., video retrieval), but has not been studied in EE literature due to the lack of necessary datasets. To address this limitation, we propose the first large-scale EE dataset obtained for transcripts of streamed videos on the video hosting platform Behance to promote future research in this area. In addition, we extensively evaluate existing state-of-the-art EE methods on our new dataset. We demonstrate that such systems cannot achieve adequate performance on the proposed dataset, revealing challenges and opportunities for further research effort.
Amir Pouran Ben Veyseh, Viet Dac Lai, Franck Dernoncourt, Thien Huu Nguyen
COLING2
2022 BehanceCC: A ChitChat Detection Dataset For Livestreaming Video Transcripts
abstract
Livestreaming videos have become an effective broadcasting method for both video sharing and educational purposes. However, livestreaming videos contain a considerable amount of off-topic content (i.e., up to 50%) which introduces significant noises and data load to downstream applications. This paper presents BehanceCC, a new human-annotated benchmark dataset for off-topic detection (also called chitchat detection) in livestreaming video transcripts. In addition to describing the challenges of the dataset, our extensive experiments of various baselines reveal the complexity of chitchat detection for livestreaming videos and suggest potential future research directions for this task. The dataset will be made publicly available to foster research in this area.
Viet Dac Lai, Amir Pouran Ben Veyseh, Franck Dernoncourt, Thien Huu Nguyen
LREC1
2022 BehanceQA: A New Dataset for Identifying Question-Answer Pairs in Video Transcripts
abstract
Question-Answer (QA) is one of the effective methods for storing knowledge which can be used for future retrieval. As such, identifying mentions of questions and their answers in text is necessary for a knowledge construction and retrieval systems. In the literature, QA identification has been well studied in the NLP community. However, most of the prior works are restricted to formal written documents such as papers or websites. As such, Questions and Answers that are presented in informal/noisy documents have not been adequately studied. One of the domains that can significantly benefit from QA identification is the domain of livestreaming video transcripts that involve abundant QA pairs to provide valuable knowledge for future users and services. Since video transcripts are often transcribed automatically for scale, they are prone to errors. Combined with the informal nature of discussion in a video, prior QA identification systems might not be able to perform well in this domain. To enable comprehensive research in this domain, we present a large-scale QA identification dataset annotated by human over transcripts of 500 hours of streamed videos. We employ Behance.net to collect the videos and their automatically obtained transcripts. Furthermore, we conduct extensive analysis on the annotated dataset to understand the complexity of QA identification for livestreaming video transcripts. Our experiments show that the annotated dataset presents unique challenges for existing methods and more research is necessary to explore more effective methods. The dataset and the models developed in this work will be publicly released for future research.
Amir Pouran Ben Veyseh, Viet Dac Lai, Franck Dernoncourt, Thien Huu Nguyen
LREC2
2021 Unleash GPT-2 Power for Event Detection
abstract
Amir Pouran Ben Veyseh, Viet Lai, Franck Dernoncourt, Thien Huu Nguyen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Amir Pouran Ben Veyseh, Viet Dac Lai, Franck Dernoncourt, Thien Huu Nguyen
ACL/IJCNLP (1)2
2021 Learning Prototype Representations Across Few-Shot Tasks for Event Detection
abstract
We address the sampling bias and outlier issues in few-shot learning for event detection, a subtask of information extraction.We propose to model the relations between training tasks in episodic few-shot learning by introducing cross-task prototypes.We further propose to enforce prediction consistency among classifiers across tasks to make the model more robust to outliers.Our extensive experiment shows a consistent improvement on three fewshot learning datasets.The findings suggest that our model is more robust when labeled data of novel event types is limited.
Viet Dac Lai, Franck Dernoncourt, Thien Huu Nguyen
EMNLP (1)1
2021 Cross-Task Instance Representation Interactions and Label Dependencies for Joint Information Extraction with Graph Convolutional Networks
abstract
Existing works on information extraction (IE)have mainly solved the four main tasks separately (entity mention recognition, relation extraction, event trigger detection, and argument extraction), thus failing to benefit from inter-dependencies between tasks.This paper presents a novel deep learning model to simultaneously solve the four tasks of IE in a single model (called FourIE).Compared to few prior work on jointly performing four IE tasks, FourIE features two novel contributions to capture inter-dependencies between tasks.First, at the representation level, we introduce an interaction graph between instances of the four tasks that is used to enrich the prediction representation for one instance with those from related instances of other tasks.Second, at the label level, we propose a dependency graph for the information types in the four IE tasks that captures the connections between the types expressed in an input sentence.A new regularization mechanism is introduced to enforce the consistency between the golden and predicted type dependency graphs to improve representation learning.We show that the proposed model achieves the state-of-the-art performance for joint IE on both monolingual and multilingual learning settings with three different languages.
Minh Nguyen 0007, Viet Dac Lai, Thien Huu Nguyen
NAACL-HLT2
2021 Graph Learning Regularization and Transfer Learning for Few-Shot Event Detection
abstract
We address the poor generalization of few-shot learning models for event detection (ED) using transfer learning and representation regularization. In particular, we propose to transfer knowledge from open-domain word sense disambiguation into few-shot learning models for ED to improve their generalization to new event types. We also propose a novel training signal derived from dependency graphs to regularize the representation learning for ED. Moreover, we evaluate few-shot learning models for ED with a large-scale human-annotated ED dataset to obtain more reliable insights for this problem. Our comprehensive experiments demonstrate that the proposed model outperforms state-of-the-art baseline models in the few-shot learning and supervised learning settings for ED. Code and data splits are available at https://github.com/laiviet/ed-fsl.
Viet Dac Lai, Minh Nguyen 0007, Thien Huu Nguyen, Franck Dernoncourt
SIGIR1
2020 Event Detection: Gate Diversity and Syntactic Importance Scores for Graph Convolution Neural Networks
abstract
Recent studies on event detection (ED) have shown that the syntactic dependency graph can be employed in graph convolution neural networks (GCN) to achieve state-of-the-art performance.However, the computation of the hidden vectors in such graph-based models is agnostic to the trigger candidate words, potentially leaving irrelevant information for the trigger candidate for event prediction.In addition, the current models for ED fail to exploit the overall contextual importance scores of the words, which can be obtained via the dependency tree, to boost the performance.In this study, we propose a novel gating mechanism to filter noisy information in the hidden vectors of the GCN models for ED based on the information from the trigger candidate.We also introduce novel mechanisms to achieve the contextual diversity for the gates and the importance score consistency for the graphs and models in ED.The experiments show that the proposed model achieves state-of-the-art performance on two ED datasets.
Viet Dac Lai, Tuan Ngo Nguyen, Thien Huu Nguyen
EMNLP (1)1
2020 Exploiting the Matching Information in the Support Set for Few Shot Event Classification
Viet Dac Lai, Franck Dernoncourt, Thien Huu Nguyen
PAKDD (2)1
2018 TSix: A Human-involved-creation Dataset for Tweet Summarization
Minh-Tien Nguyen, Viet Dac Lai, Minh Le Nguyen 0001
LREC2