VLDB 2026 Research / reviewers in the wild / expert
Ngan Luu-Thuy Nguyen
dblp:174/4407
· DBLP profile ↗
41ranked-venue papers
1as first author
35since 2021 · last 2026
0000-0003-3931-849XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identifying Distress in Vietnamese Narratives: Dataset Construction and Model Benchmarking
Trong Thanh Le, Thu Trung Tran, Giang Son Tran, Kiet Van Nguyen, Dang Van Thin, Duy Dinh Le, Ngan Luu-Thuy Nguyen |
ACIIDS (1) | 7 |
| 2026 | From Slides to Exams: A Multi-agent Human-AI System for Collaborative Assessment Design
Khiem Hoang Truong, Duc-Tuan Luu, Duong Ngoc Hao, Dang Van Thin, Ngan Luu-Thuy Nguyen |
AIED (1) | 5 |
| 2026 | A new benchmark dataset and mixture-of-experts language models for adversarial natural language inference in Vietnamese
Tin Van Huynh, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
Expert Syst. Appl. | 3 |
| 2026 | ViTextVQA: A large-scale visual question answering dataset and a novel multimodal feature fusion method for Vietnamese text comprehension in images
Quan Van Nguyen, Dan Quang Tran, Huy Quang Pham, Thang Kien-Bao Nguyen, Nghia Hieu Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
Expert Syst. Appl. | 7 |
| 2026 | Natural language processing and computational linguistics for Vietnamese: A comprehensive review
Khiem Vinh Tran, Triet Minh Thai, Son T. Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
Expert Syst. Appl. | 5 |
| 2025 | Vietnamese Words Are Not Constructed from Syllables: Rethinking the Role of Word Segmentation in Natural Language Processing for Vietnamese TextsabstractThe definition of words is the fundamental and crucial linguistic concept. Any changes in word definition lead to changes in the theoretical system of the respective language. Traditionally, researchers in Natural Language Processing (NLP) for Vietnamese texts believe Vietnamese words are constructed from syllables. However, their works did not explicitly mention which linguistic theory they followed for this assumption. Although there are no theoretical guarantees, most NLP studies in Vietnamese accept this assumption. Consequently, word segmentation is recognized as one of the essential stages in NLP for Vietnamese texts. In this study, we address the role of word segmentation for Vietnamese texts from linguistic perspectives. Through our extensive experiments, we show that, based on linguistic theories, performing word segmentation is not appropriate for Vietnamese text understanding. Moreover, we present a novel method, Vietnamese Word TransFormer (ViWordFormer), for modeling Vietnamese word formation. Experimental results indicate that our method is appropriate for modeling Vietnamese word formation from both theoretical and experimental aspects and embark on a novel approach to Vietnamese word representation. Nghia Hieu Nguyen, Tien Dat Nguyen, Ngan Luu-Thuy Nguyen |
AAAI | 3 |
| 2025 | Optimizing Legal Document Retrieval in Vietnamese with Semi-hard Negative Mining
Van-Hoang Le, Duc-Vu Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
ICCCI (1) | 4 |
| 2025 | ViPhoVQA: Toward a Phonemic-Based Method for Mitigating Rare and Out-of-Vocab Words in Vietnamese Text-Based Visual Question Answering
Nghia Hieu Nguyen, Duc-Vu Nguyen, Huy Quoc To, Dinh Dien, Thanh Tho Quan, Ngan Luu-Thuy Nguyen |
ICCCI (1) | 6 |
| 2025 | ViTASA: New benchmark and methods for Vietnamese targeted aspect sentiment analysis for multiple textual domains
Quang Phan-Minh Huynh, Oanh Thi-Hong Le, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
Comput. Speech Lang. | 5 |
| 2025 | ViOCRVQA: novel benchmark dataset and VisionReader for visual question answering by understanding Vietnamese text in images
Huy Quang Pham, Thang Kien-Bao Nguyen, Quan Van Nguyen, Dan Quang Tran, Nghia Hieu Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
Multim. Syst. | 7 |
| 2025 | LMCK: pre-trained language models enhanced with contextual knowledge for Vietnamese natural language inference
Ngan Luu-Thuy Nguyen, Khoa Thi-Kim Phan, Tin Van Huynh, Kiet Van Nguyen |
Multim. Tools Appl. | 1 |
| 2024 | VlogQA: Task, Dataset, and Baseline Models for Vietnamese Spoken-Based Machine Reading ComprehensionabstractThis paper presents the development process of a Vietnamese spoken language corpus for machine reading comprehension (MRC) tasks and provides insights into the challenges and opportunities associated with using realworld data for machine reading comprehension tasks.The existing MRC corpora in Vietnamese mainly focus on formal written documents such as Wikipedia articles, online newspapers, or textbooks.In contrast, the VlogQA consists of 10,076 question-answer pairs based on 1,230 transcript documents sourced from YouTube -an extensive source of user-uploaded content, covering the topics of food and travel.By capturing the spoken language of native Vietnamese speakers in natural settings, an obscure corner overlooked in Vietnamese research, the corpus provides a valuable resource for future research in reading comprehension tasks for the Vietnamese language.Regarding performance evaluation, our deep-learning models achieved the highest F1 score of 75.34% on the test set, indicating significant progress in machine reading comprehension for Vietnamese spoken language data.In terms of EM, the highest score we accomplished is 53.97%, which reflects the challenge in processing spoken-based content and highlights the need for further improvement. Thinh Phuoc Ngo, Khoa Tran Anh Dang, Son T. Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
EACL (1) | 5 |
| 2024 | Prompt Engineering with Large Language Models for Vietnamese Sentiment Classification
Dang Van Thin, Duong Ngoc Hao, Ngan Luu-Thuy Nguyen |
PACLIC | 3 |
| 2024 | Coreference Resolution for Vietnamese Narrative Texts
Hieu-Dai Tran, Duc-Vu Nguyen, Ngan Luu-Thuy Nguyen |
PACLIC | 3 |
| 2024 | ViCLEVR: a visual reasoning dataset and hybrid multimodal fusion model for visual question answering in Vietnamese
Khiem Vinh Tran, Hao Phu Phan, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
Multim. Syst. | 4 |
| 2024 | An approach of data augmentation to improve the performance of BERTology models for Vietnamese hate speech detection
Son T. Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
Multim. Tools Appl. | 3 |
| 2023 | ViHOS: Hate Speech Spans Detection for VietnameseabstractThe rise in hateful and offensive language directed at other users is one of the adverse side effects of the increased use of social networking platforms.This could make it difficult for human moderators to review tagged comments filtered by classification systems.To help address this issue, we present the ViHOS (Vietnamese Hate and Offensive Spans) dataset, the first human-annotated corpus containing 26k spans on 11k comments.We also provide definitions of hateful and offensive spans in Vietnamese comments as well as detailed annotation guidelines.Besides, we conduct experiments with various state-of-the-art models.Specifically, XLM-R Large achieved the best F1-scores in Single span detection and All spans detection, while PhoBERT Large obtained the highest in Multiple spans detection.Finally, our error analysis demonstrates the difficulties in detecting specific types of spans in our data for future research.Our dataset is released on GitHub 1 . Phu Gia Hoang, Canh Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
EACL | 5 |
| 2023 | Towards sustainable agriculture: A lightweight hybrid model and cloud-based collection of datasets for efficient leaf disease detection
Huy-Tan Thai, Kim-Hung Le, Ngan Luu-Thuy Nguyen |
Future Gener. Comput. Syst. | 3 |
| 2023 | Vietnamese Sentiment Analysis: An Overview and Comparative Study of Fine-tuning Pretrained Language ModelsabstractSentiment Analysis (SA) is one of the most active research areas in the Natural Language Processing (NLP) field due to its potential for business and society. With the development of language representation models, numerous methods have shown promising efficiency in fine-tuning pre-trained language models in NLP downstream tasks. For Vietnamese, many available pre-trained language models were also released, including the monolingual and multilingual language models. Unfortunately, all of these models were trained on different architectures, pre-trained data, and pre-processing steps; consequently, fine-tuning these models can be expected to yield different effectiveness. In addition, there is no study focusing on evaluating the performance of these models on the same datasets for the SA task up to now. This article presents a fine-tuning approach to investigate the performance of different pre-trained language models for the Vietnamese SA task. The experimental results show the superior performance of the monolingual PhoBERT model and ViT5 model in comparison with previous studies and provide new state-of-the-art performances on five benchmark Vietnamese SA datasets. To the best of our knowledge, our study is the first attempt to investigate the performance of fine-tuning Transformer-based models on five datasets with different domains and sizes for the Vietnamese SA task. Dang Van Thin, Duong Ngoc Hao, Ngan Luu-Thuy Nguyen |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2023 | A Systematic Literature Review on Vietnamese Aspect-based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) is one of the principal tasks in the automatic deep understanding of texts, widely applied in a broad range of real-world applications. Many studies have been performed on different tasks and datasets for other languages (e.g., English, Chinese) to address this topic. For Vietnamese language, this topic has been attracting considerable interest in recent years. However, we found that many studies tend to repeat the research instead of inheriting and extending the previous works. Moreover, previous studies’ methods of comparison or evaluation metrics have not shown consistency and connection. This might restrict the development of future studies on this research topic. To the best of our knowledge, no research has been conducted to overview the existing studies for the ABSA research in Vietnamese language. The primary objective of this study is to provide a systematic and comprehensive review of the current Vietnamese ABSA research. More specifically, we analyze the early approaches, evaluation metrics, and available published benchmark datasets used in the Vietnamese ABSA task. We also discuss the challenge and recommend potential future directions for Vietnamese ABSA. This work is expected to provide readers with a wealth of knowledge, the research gap, and the challenges in the Vietnamese ABSA field. Dang Van Thin, Duong Ngoc Hao, Ngan Luu-Thuy Nguyen |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2022 | XLMRQA: Open-Domain Question Answering on Vietnamese Wikipedia-Based Textual Knowledge Source
Kiet Van Nguyen, Phong Nguyen-Thuan Do, Nhat Duy Nguyen, Tin Van Huynh, Anh Gia-Tuan Nguyen, Ngan Luu-Thuy Nguyen |
ACIIDS (1) | 6 |
| 2022 | A Comparative Study of Question Answering over Knowledge Bases
Khiem Vinh Tran, Hao Phu Phan, Nguyen Duc Khang Quach, Ngan Luu-Thuy Nguyen, Jun Jo 0001, Thanh Tam Nguyen |
ADMA (1) | 4 |
| 2022 | ViNLI: A Vietnamese Corpus for Studies on Open-Domain Natural Language InferenceabstractOver a decade, the research field of computational linguistics has witnessed the growth of corpora and models for natural language inference (NLI) for rich-resource languages such as English and Chinese. A large-scale and high-quality corpus is necessary for studies on NLI for Vietnamese, which can be considered a low-resource language. In this paper, we introduce ViNLI (Vietnamese Natural Language Inference), an open-domain and high-quality corpus for evaluating Vietnamese NLI models, which is created and evaluated with a strict process of quality control. ViNLI comprises over 30,000 human-annotated premise-hypothesis sentence pairs extracted from more than 800 online news articles on 13 distinct topics. In this paper, we introduce the guidelines for corpus creation which take the specific characteristics of the Vietnamese language in expressing entailment and contradiction into account. To evaluate the challenging level of our corpus, we conduct experiments with state-of-the-art deep neural networks and pre-trained models on our dataset. The best system performance is still far from human performance (a 14.20% gap in accuracy). The ViNLI corpus is a challenging corpus to accelerate progress in Vietnamese computational linguistics. Our corpus is available publicly for research purposes. Tin Van Huynh, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
COLING | 3 |
| 2022 | SPBERTQA: A Two-Stage Question Answering System Based on Sentence Transformers for Medical Texts
Nhung Thi-Hong Nguyen, Phuong Phan-Dieu Ha, Luan Thanh Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
KSEM (2) | 5 |
| 2022 | SMTCE: A Social Media Text Classification Evaluation Benchmark and BERTology Models for Vietnamese
Luan Thanh Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
PACLIC | 3 |
| 2022 | Sentiment Analysis in Code-Mixed Vietnamese-English Sentence-level Hotel Reviews
Dang Van Thin, Duong Ngoc Hao, Ngan Luu-Thuy Nguyen |
PACLIC | 3 |
| 2022 | New Vietnamese Corpus for Machine Reading Comprehension of Health News ArticlesabstractMachine reading comprehension is a natural language understanding task where the computing system is required to read a text and then find the answer to a specific question posed by a human. Large-scale and high-quality corpora are necessary for evaluating machine reading comprehension models. Furthermore, machine reading comprehension (MRC) for the health sector has potential for practical applications; nevertheless, MRC research in this domain is currently scarce. This article presents UIT-ViNewsQA, a new corpus for the Vietnamese language to evaluate MRC models for the healthcare textual domain. The corpus consists of 22,057 human-generated question-answer pairs. Crowd-workers create the questions and answers on a collection of 4,416 online Vietnamese healthcare news articles, where the answers are textual spans extracted from the corresponding articles. We introduce a process for creating a high-quality corpus for the Vietnamese machine reading comprehension task. Linguistically, our corpus accommodates diversity in question and answer types. In addition, we conduct experiments and compare the effectiveness of different MRC methods based on the neural networks and transformer architectures. Experimental results on our corpus show that the MRC system based on ALBERT architecture outperforms the neural network architectures and the BERT-based approach, an exact match score of 65.26% and an F1-score of 84.89%. The best machine model achieves about 10.90% F1-score less efficiently than humans, which proves that exploring machine models on UIT-ViNewsQA to surpass humans is challenging for researchers in the future. Our corpus is publicly available on our website: http://nlp.uit.edu.vn/datasets for research purposes. Kiet Van Nguyen, Tin Van Huynh, Duc-Vu Nguyen, Anh Gia-Tuan Nguyen, Ngan Luu-Thuy Nguyen |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2021 | A Large-Scale Dataset for Hate Speech Detection on Vietnamese Social Media Texts
Son T. Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
IEA/AIE (1) | 3 |
| 2021 | Constructive and Toxic Speech Detection for Open-Domain Social Media Comments in VietnameseabstractThe rise of social media has led to the increasing of comments on online forums. However, there still exists invalid comments which are not informative for users. Moreover, those comments are also quite toxic and harmful to people. In this paper, we create a dataset for constructive and toxic speech detection, named UIT-ViCTSD (Vietnamese Constructive and Toxic Speech Detection dataset) with 10,000 human-annotated comments. For these tasks, we propose a system for constructive and toxic speech detection with the state-of-the-art transfer learning model in Vietnamese NLP as PhoBERT. With this system, we obtain F1-scores of 78.59% and 59.40% for classifying constructive and toxic comments, respectively. Besides, we implement various baseline models as traditional Machine Learning and Deep Neural Network-Based models to evaluate the dataset. With the results, we can solve several tasks on the online discussions and develop the framework for identifying constructiveness and toxicity of Vietnamese social media comments automatically. Luan Thanh Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
IEA/AIE (1) | 3 |
| 2021 | Sentence Extraction-Based Machine Reading Comprehension for Vietnamese
Phong Nguyen-Thuan Do, Nhat Duy Nguyen, Tin Van Huynh, Kiet Van Nguyen, Anh Gia-Tuan Nguyen, Ngan Luu-Thuy Nguyen |
KSEM | 6 |
| 2021 | Joint Chinese Word Segmentation and Part-of-speech Tagging via Two-stage Span Labeling
Duc-Vu Nguyen, Linh-Bao Vo, Ngoc-Linh Tran, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
PACLIC | 5 |
| 2021 | Monolingual vs multilingual BERTology for Vietnamese extractive multi-document summarization
Huy Quoc To, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen, Anh Gia-Tuan Nguyen |
PACLIC | 3 |
| 2021 | Span Labeling Approach for Vietnamese and Chinese Word Segmentation
Duc-Vu Nguyen, Linh-Bao Vo, Dang Van Thin, Ngan Luu-Thuy Nguyen |
PRICAI (2) | 4 |
| 2021 | Vietnamese Complaint Detection on E-Commerce WebsitesabstractCustomer product reviews play a role in improving the quality of products and services for business organizations or their brands. Complaining is an attitude that expresses dissatisfaction with an event or a product not meeting customer expectations. In this paper, we build a Vietnamese Open-domain Complaint Detection dataset (UIT-ViOCD), including 5,485 human-annotated reviews on four categories about product reviews on e-commerce sites. After the data collection phase, we proceed to the annotation task and achieve the inter-annotator agreement (Am) of 87%. Then, we present an extensive methodology for the research purposes and achieve 92.16% by F1-score for identifying complaints. With the results, in future, we aim to build a system for open-domain complaint detection on E-commerce websites. Nhung Thi-Hong Nguyen, Phuong Phan-Dieu Ha, Luan Thanh Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
SoMeT | 5 |
| 2021 | Two New Large Corpora for Vietnamese Aspect-based Sentiment Analysis at Sentence LevelabstractAspect-based sentiment analysis has been studied in both research and industrial communities over recent years. For the low-resource languages, the standard benchmark corpora play an important role in the development of methods. In this article, we introduce two benchmark corpora with the largest sizes at sentence-level for two tasks: Aspect Category Detection and Aspect Polarity Classification in Vietnamese. Our corpora are annotated with high inter-annotator agreements for the restaurant and hotel domains. The release of our corpora would push forward the low-resource language processing community. In addition, we deploy and compare the effectiveness of supervised learning methods with a single and multi-task approach based on deep learning architectures. Experimental results on our corpora show that the multi-task approach based on BERT architecture outperforms the neural network architectures and the single approach. Our corpora and source code are published on this footnoted site. 1 Dang Van Thin, Ngan Luu-Thuy Nguyen, Minh Tri Truong, Lac Si Le, Duy-Tin Vo |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2020 | A Vietnamese Dataset for Evaluating Machine Reading ComprehensionabstractOver 97 million people speak Vietnamese as their native language in the world.However, there are few research studies on machine reading comprehension (MRC) for Vietnamese, the task of understanding a text and answering questions related to it.Due to the lack of benchmark datasets for Vietnamese, we present the Vietnamese Question Answering Dataset (UIT-ViQuAD), a new dataset for the low-resource language as Vietnamese to evaluate MRC models.This dataset comprises over 23,000 human-generated question-answer pairs based on 5,109 passages of 174 Vietnamese articles from Wikipedia.In particular, we propose a new process of dataset creation for Vietnamese MRC.Our in-depth analyses illustrate that our dataset requires abilities beyond simple reasoning like word matching and demands single-sentence and multiple-sentence inferences.Besides, we conduct experiments on state-of-the-art MRC methods for English and Chinese as the first experimental models on UIT-ViQuAD.We also estimate human performance on the dataset and compare it to the experimental results of powerful machine learning models.As a result, the substantial differences between human performance and the best model performance on the dataset indicate that improvements can be made on UIT-ViQuAD in future research.Our dataset is freely available on our website 1 to encourage the research community to overcome challenges in Vietnamese MRC. Kiet Van Nguyen, Duc-Vu Nguyen, Anh Gia-Tuan Nguyen, Ngan Luu-Thuy Nguyen |
COLING | 4 |
| 2020 | UIT-ViIC: A Dataset for the First Evaluation on Vietnamese Image Captioning
Quan Hoang Lam, Quang-Duy Le, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
ICCCI | 4 |
| 2020 | A simple and efficient ensemble classifier combining multiple neural network models on social media datasets in Vietnamese
Huy Duc Huynh, Hang Thi-Thuy Do, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
PACLIC | 4 |
| 2020 | Empirical Study of Text Augmentation on Social Media Text in Vietnamese
Son T. Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
PACLIC | 3 |
| 2018 | Exploring alignment-classification methods in the context of professional writing assistance
Mai Duong, Minh-Quoc Nghiem, Ngan Luu-Thuy Nguyen |
Data Knowl. Eng. | 3 |
| 2003 | A hybrid approach to word order transfer in the English-to-Vietnamese machine translationabstractWord Order transfer is a compulsory stage and has a great effect on the translation result of a transfer-based machine translation system. To solve this problem, we can use fixed rules (rule-based) or stochastic methods (corpus-based) which extract word order transfer rules between two languages. However, each approach has its own advantages and disadvantages. In this paper, we present a hybrid approach based on fixed rules and Transformation-Based Learning (or TBL) method. Our purpose is to transfer automatically the English word orders into the Vietnamese ones. The learning process will be trained on the annotated bilingual corpus (named EVC: English-Vietnamese Corpus) that has been automatically word-aligned, phrase-aligned and POS-tagged. This transfer result is being used for the transfer module in the English-Vietnamese transfer-based machine translation system. Dinh Dien, Ngan Luu-Thuy Nguyen, Quang Xuan Do, Nam Van Chi |
MTSummit | 2 |