VLDB 2026 Research / reviewers in the wild / expert
Kiet Van Nguyen
dblp:174/4526
· DBLP profile ↗
52ranked-venue papers
5as first author
48since 2021 · last 2026
0000-0002-8456-2742ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 5 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ViDia2Std: A Parallel Corpus and Methods for Low-Resource Vietnamese Dialect-to-Standard TranslationabstractVietnamese exhibits extensive dialectal variation, posing challenges for NLP systems trained predominantly on standard Vietnamese. Such systems often underperform on dialectal inputs, especially from underrepresented Central and Southern regions. Previous work on dialect normalization has focused narrowly on Central-to-Northern dialect transfer using synthetic data and limited dialectal diversity. These efforts exclude Southern varieties and intra-regional variants within the North. We introduce ViDia2Std, the first manually annotated parallel corpus for dialect-to-standard Vietnamese translation covering all 63 provinces. Unlike prior datasets, ViDia2Std includes diverse dialects from Central, Southern, and non-standard Northern regions often absent from existing resources, making it the most dialectally inclusive corpus to date. The dataset consists of over 13,000 sentence pairs sourced from real-world Facebook comments and annotated by native speakers across all three dialect regions. To assess annotation consistency, we define a semantic mapping agreement metric that accounts for synonymous standard mappings across annotators. Based on this criterion, we report agreement rates of 86% (North), 82% (Central), and 85% (South). We benchmark several sequence-to-sequence models on ViDia2Std. mBART-large-50 achieves the best results (BLEU 0.8166, ROUGE-L 0.9384, METEOR 0.8925), while ViT5-base offers competitive performance with fewer parameters. ViDia2Std demonstrates that dialect normalization substantially improves downstream tasks, highlighting the need for dialect-aware resources in building robust Vietnamese NLP systems. Khoa Anh Ta, Nguyen Van Dinh, Kiet Van Nguyen |
AAAI | 3 |
| 2026 | Identifying Distress in Vietnamese Narratives: Dataset Construction and Model Benchmarking
Trong Thanh Le, Thu Trung Tran, Giang Son Tran, Kiet Van Nguyen, Dang Van Thin, Duy Dinh Le, Ngan Luu-Thuy Nguyen |
ACIIDS (1) | 4 |
| 2026 | ViWikiFC: Fact-Checking for Vietnamese Wikipedia-Based Textual Knowledge Source
Hung Tuan Le, Long Truong To, Manh Trong Nguyen, Kiet Van Nguyen |
LREC | 4 |
| 2026 | ViX-Ray: A Vietnamese Chest X-Ray Dataset for Vision-Language Models
Duy Vu Minh Nguyen, Chinh Thanh Truong, Phuc Hoang Tran, Hung Tuan Le, Dat Van-Thanh Nguyen, Trung Hieu Pham, Kiet Van Nguyen |
LREC | 7 |
| 2026 | A new benchmark dataset and mixture-of-experts language models for adversarial natural language inference in Vietnamese
Tin Van Huynh, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
Expert Syst. Appl. | 2 |
| 2026 | ViTextVQA: A large-scale visual question answering dataset and a novel multimodal feature fusion method for Vietnamese text comprehension in images
Quan Van Nguyen, Dan Quang Tran, Huy Quang Pham, Thang Kien-Bao Nguyen, Nghia Hieu Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
Expert Syst. Appl. | 6 |
| 2026 | Natural language processing and computational linguistics for Vietnamese: A comprehensive review
Khiem Vinh Tran, Triet Minh Thai, Son T. Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
Expert Syst. Appl. | 4 |
| 2025 | ViFactCheck: A New Benchmark Dataset and Methods for Multi-Domain News Fact-Checking In VietnameseabstractThe rapid spread of information in the digital age highlights the critical need for effective fact-checking tools, particularly for languages with limited resources, such as Vietnamese. In response to this challenge, we introduce ViFactCheck, the first publicly available benchmark dataset designed specifically for Vietnamese fact-checking across multiple online news domains. This dataset contains 7,232 human-annotated pairs of claim-evidence combinations sourced from reputable Vietnamese online news, covering 12 diverse topics. It has been subjected to a meticulous annotation process to ensure high quality and reliability, achieving a Fleiss Kappa inter-annotator agreement score of 0.83. Our evaluation leverages state-of-the-art pre-trained and large language models, employing fine-tuning and prompting techniques to assess performance. Notably, the Gemma model demonstrated superior effectiveness, with an impressive macro F1 score of 89.90%, thereby establishing a new standard for fact-checking benchmarks. This result highlights the robust capabilities of Gemma in accurately identifying and verifying facts in Vietnamese. To further promote advances in fact-checking technology and improve the reliability of digital media, we have made the ViFactCheck dataset, model checkpoints, fact-checking pipelines, and source code freely available on GitHub. This initiative aims to inspire further research and enhance the accuracy of information in low-resource languages. Tran Thai Hoa, Tran Quang Duy, Kiet Van Nguyen |
AAAI | 4 |
| 2025 | Optimizing Legal Document Retrieval in Vietnamese with Semi-hard Negative Mining
Van-Hoang Le, Duc-Vu Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
ICCCI (1) | 3 |
| 2025 | ViNumFCR: A Novel Vietnamese Benchmark for Numerical Reasoning Fact Checking on Social Media NewsabstractIn the digital era, the internet provides rapid and convenient access to vast amounts of information. However, much of this information remains unverified, particularly with the increasing prevalence of falsified numerical data, leading to public confusion and negative societal impacts. To address this issue, we developed ViNumFCR, a first dataset dedicated to fact-checking numerical information in Vietnamese. Comprising over 10,000 samples collected and constructed from online newspaper across 12 different topics. We assessed the performance of various fact-checking models, including Pretrained Language Models and Large Language Models, alongside retrieval techniques for gathering supporting evidence. Experimental results demonstrate that the XLM-R_Large model achieved the highest accuracy of 90.05% on the fact-checking task, while the combined SBERT + BM25 model attained a precision of over 97% on the evidence retrieval task. Additionally, we conducted an in-depth analysis of the linguistic features of the dataset to understand the factors influencing the performance models. The ViNumFCR dataset is publicly available to support further research. Nhi Ngoc Phuong Luong, Anh Thi Lan Le, Tin Van Huynh, Kiet Van Nguyen, Ngan Nguyen |
INLG | 4 |
| 2025 | ViMRHP: A Vietnamese Benchmark Dataset for Multimodal Review Helpfulness Prediction via Human-AI Collaborative Annotation
Truc Mai-Thanh Nguyen, Dat Minh Nguyen, Son T. Luu, Kiet Van Nguyen |
NLDB (1) | 4 |
| 2025 | ViTASA: New benchmark and methods for Vietnamese targeted aspect sentiment analysis for multiple textual domains
Quang Phan-Minh Huynh, Oanh Thi-Hong Le, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
Comput. Speech Lang. | 4 |
| 2025 | LiGT: layout-infused generative transformer for visual question answering on Vietnamese receipts
Thanh-Phong Le, Trung Le Chi Phan, Nghia Hieu Nguyen, Kiet Van Nguyen |
Int. J. Document Anal. Recognit. | 4 |
| 2025 | Open-ViTabQA: A novel benchmark for Vietnamese question answering on open domain wikipedia table
Dung Hoang Dao, Ngan Thi-Kim Huynh, Kiet Van Nguyen |
Knowl. Based Syst. | 4 |
| 2025 | New benchmark dataset and fine-grained cross-modal fusion framework for Vietnamese multimodal aspect-category sentiment analysis
Quy Hoang Nguyen, Minh-Van Truong Nguyen, Kiet Van Nguyen |
Multim. Syst. | 3 |
| 2025 | ViOCRVQA: novel benchmark dataset and VisionReader for visual question answering by understanding Vietnamese text in images
Huy Quang Pham, Thang Kien-Bao Nguyen, Quan Van Nguyen, Dan Quang Tran, Nghia Hieu Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
Multim. Syst. | 6 |
| 2025 | LMCK: pre-trained language models enhanced with contextual knowledge for Vietnamese natural language inference
Ngan Luu-Thuy Nguyen, Khoa Thi-Kim Phan, Tin Van Huynh, Kiet Van Nguyen |
Multim. Tools Appl. | 4 |
| 2025 | XGV-BERT: Leveraging contextualized language model and graph neural network for efficient software vulnerability detection
Vu Le Anh Quan, Chau Thuan Phat, Kiet Van Nguyen, Phan The Duy, Van-Hau Pham |
J. Supercomput. | 3 |
| 2024 | VlogQA: Task, Dataset, and Baseline Models for Vietnamese Spoken-Based Machine Reading ComprehensionabstractThis paper presents the development process of a Vietnamese spoken language corpus for machine reading comprehension (MRC) tasks and provides insights into the challenges and opportunities associated with using realworld data for machine reading comprehension tasks.The existing MRC corpora in Vietnamese mainly focus on formal written documents such as Wikipedia articles, online newspapers, or textbooks.In contrast, the VlogQA consists of 10,076 question-answer pairs based on 1,230 transcript documents sourced from YouTube -an extensive source of user-uploaded content, covering the topics of food and travel.By capturing the spoken language of native Vietnamese speakers in natural settings, an obscure corner overlooked in Vietnamese research, the corpus provides a valuable resource for future research in reading comprehension tasks for the Vietnamese language.Regarding performance evaluation, our deep-learning models achieved the highest F1 score of 75.34% on the test set, indicating significant progress in machine reading comprehension for Vietnamese spoken language data.In terms of EM, the highest score we accomplished is 53.97%, which reflects the challenge in processing spoken-based content and highlights the need for further improvement. Thinh Phuoc Ngo, Khoa Tran Anh Dang, Son T. Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
EACL (1) | 4 |
| 2024 | Multi-Dialect Vietnamese: Task, Dataset, Baseline Models and ChallengesabstractVietnamese, a low-resource language, is typically categorized into three primary dialect groups that belong to Northern, Central, and Southern Vietnam.However, each province within these regions exhibits its own distinct pronunciation variations.Despite the existence of various speech recognition datasets, none of them has provided a fine-grained classification of the 63 dialects specific to individual provinces of Vietnam.To address this gap, we introduce Vietnamese Multi-Dialect (ViMD) dataset, a novel comprehensive dataset capturing the rich diversity of 63 provincial dialects spoken across Vietnam.Our dataset comprises 102.56 hours of audio, consisting of approximately 19,000 utterances, and the associated transcripts contain over 1.2 million words.To provide benchmarks and simultaneously demonstrate the challenges of our dataset, we fine-tune state-of-the-art pre-trained models for two downstream tasks: (1) Dialect identification and (2) Speech recognition.The empirical results suggest two implications including the influence of geographical factors on dialects, and the constraints of current approaches in speech recognition tasks involving multidialect speech data.Our dataset is available for research purposes 1 . Nguyen Dinh, Thanh Dang, Luan Thanh Nguyen, Kiet Van Nguyen |
EMNLP | 4 |
| 2024 | XLMR4MD: New Vietnamese dataset and framework for detecting the consistency of description and permission in Android applications using large language models
Qui Ngoc Nguyen, Cam Nguyen Tan, Kiet Van Nguyen |
Comput. Secur. | 3 |
| 2024 | ViCLEVR: a visual reasoning dataset and hybrid multimodal fusion model for visual question answering in Vietnamese
Khiem Vinh Tran, Hao Phu Phan, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
Multim. Syst. | 3 |
| 2024 | An approach of data augmentation to improve the performance of BERTology models for Vietnamese hate speech detection
Son T. Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
Multim. Tools Appl. | 2 |
| 2024 | Numerical reasoning reading comprehension on Vietnamese COVID-19 news: task, corpus, and challenges
Kiet Van Nguyen, Thang Viet Le, Tinh Pham-Phuc Do |
Neural Comput. Appl. | 1 |
| 2023 | Link Prediction for Wikipedia Articles as a Natural Language Inference TaskabstractLink prediction task is vital to automatically understanding the structure of large knowledge bases. In this paper, we present our system to solve this task at the Data Science and Advanced Analytics 2023 Competition “Efficient and Effective Link Prediction” (DSAA-2023 Competition) [1] with a corpus containing 948,233 training and 238,265 for public testing. This paper introduces an approach to link prediction in Wikipedia articles by formulating it as a natural language inference (NLI) task. Drawing inspiration from recent advancements in natural language processing and understanding, we cast link prediction as an NLI task, wherein the presence of a link between two articles is treated as a premise, and the task is to determine whether this premise holds based on the information presented in the articles. We implemented our system based on the Sentence Pair Classification for Link Prediction for the Wikipedia Articles task. Our system achieved 0.99996 Macro F1-score and 1.00000 Macro F1-score for the public and private test sets, respectively. Our team UIT-NLP ranked 3rd in performance on the private test set, equal to the scores of the first and second places. Our code1is publicly for research purposes.1https://github.com/phanchauthang/dsaa-2023-kaggle/ Chau-Thang Phan, Quoc-Nam Nguyen, Kiet Van Nguyen |
DSAA | 3 |
| 2023 | ViHOS: Hate Speech Spans Detection for VietnameseabstractThe rise in hateful and offensive language directed at other users is one of the adverse side effects of the increased use of social networking platforms.This could make it difficult for human moderators to review tagged comments filtered by classification systems.To help address this issue, we present the ViHOS (Vietnamese Hate and Offensive Spans) dataset, the first human-annotated corpus containing 26k spans on 11k comments.We also provide definitions of hateful and offensive spans in Vietnamese comments as well as detailed annotation guidelines.Besides, we conduct experiments with various state-of-the-art models.Specifically, XLM-R Large achieved the best F1-scores in Single span detection and All spans detection, while PhoBERT Large obtained the highest in Multiple spans detection.Finally, our error analysis demonstrates the difficulties in detecting specific types of spans in our data for future research.Our dataset is released on GitHub 1 . Phu Gia Hoang, Canh Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
EACL | 4 |
| 2023 | ViSoBERT: A Pre-Trained Language Model for Vietnamese Social Media Text ProcessingabstractEnglish and Chinese, known as resource-rich languages, have witnessed the strong development of transformer-based language models for natural language processing tasks.Although Vietnam has approximately 100M people speaking Vietnamese, several pre-trained models, e.g., PhoBERT, ViBERT, and vELEC-TRA, performed well on general Vietnamese NLP tasks, including POS tagging and named entity recognition.These pre-trained language models are still limited to Vietnamese social media tasks.In this paper, we present the first monolingual pre-trained language model for Vietnamese social media texts, ViSoBERT, which is pre-trained on a large-scale corpus of high-quality and diverse Vietnamese social media texts using XLM-R architecture.Moreover, we explored our pre-trained model on five important natural language downstream tasks on Vietnamese social media texts: emotion recognition, hate speech detection, sentiment analysis, spam reviews detection, and hate speech spans detection.Our experiments demonstrate that ViSoBERT, with far fewer parameters, surpasses the previous state-of-the-art models on multiple Vietnamese social media tasks.Our ViSoBERT model is available 4 only for research purposes. Thang Phan, Duc-Vu Nguyen, Kiet Van Nguyen |
EMNLP | 4 |
| 2023 | Machine Reading Comprehension for Vietnamese Customer Reviews: Task, Corpus and Baseline Models
Tinh Pham Phuc Do, Ngoc Dinh Duy Cao, Tin Van Huynh, Kiet Van Nguyen |
PACLIC | 5 |
| 2023 | Academic performance warning system based on data driven for higher education
Hong-Hanh Thi Duong, Linh Thi-My Tran, Huy Quoc To, Kiet Van Nguyen |
Neural Comput. Appl. | 4 |
| 2023 | Vietnamese hate and offensive detection using PhoBERT-CNN and social media streaming data
An Trong Nguyen, Phu Gia Hoang, Canh Duc Luu, Trong-Hop Do, Kiet Van Nguyen |
Neural Comput. Appl. | 6 |
| 2022 | XLMRQA: Open-Domain Question Answering on Vietnamese Wikipedia-Based Textual Knowledge Source
Kiet Van Nguyen, Phong Nguyen-Thuan Do, Nhat Duy Nguyen, Tin Van Huynh, Anh Gia-Tuan Nguyen, Ngan Luu-Thuy Nguyen |
ACIIDS (1) | 1 |
| 2022 | Enhancing Vietnamese Question Generation with Reinforcement Learning
Nguyen Vu, Kiet Van Nguyen |
ACIIDS (1) | 2 |
| 2022 | ViNLI: A Vietnamese Corpus for Studies on Open-Domain Natural Language InferenceabstractOver a decade, the research field of computational linguistics has witnessed the growth of corpora and models for natural language inference (NLI) for rich-resource languages such as English and Chinese. A large-scale and high-quality corpus is necessary for studies on NLI for Vietnamese, which can be considered a low-resource language. In this paper, we introduce ViNLI (Vietnamese Natural Language Inference), an open-domain and high-quality corpus for evaluating Vietnamese NLI models, which is created and evaluated with a strict process of quality control. ViNLI comprises over 30,000 human-annotated premise-hypothesis sentence pairs extracted from more than 800 online news articles on 13 distinct topics. In this paper, we introduce the guidelines for corpus creation which take the specific characteristics of the Vietnamese language in expressing entailment and contradiction into account. To evaluate the challenging level of our corpus, we conduct experiments with state-of-the-art deep neural networks and pre-trained models on our dataset. The best system performance is still far from human performance (a 14.20% gap in accuracy). The ViNLI corpus is a challenging corpus to accelerate progress in Vietnamese computational linguistics. Our corpus is available publicly for research purposes. Tin Van Huynh, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
COLING | 2 |
| 2022 | SPBERTQA: A Two-Stage Question Answering System Based on Sentence Transformers for Medical Texts
Nhung Thi-Hong Nguyen, Phuong Phan-Dieu Ha, Luan Thanh Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
KSEM (2) | 4 |
| 2022 | SMTCE: A Social Media Text Classification Evaluation Benchmark and BERTology Models for Vietnamese
Luan Thanh Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
PACLIC | 2 |
| 2022 | New Vietnamese Corpus for Machine Reading Comprehension of Health News ArticlesabstractMachine reading comprehension is a natural language understanding task where the computing system is required to read a text and then find the answer to a specific question posed by a human. Large-scale and high-quality corpora are necessary for evaluating machine reading comprehension models. Furthermore, machine reading comprehension (MRC) for the health sector has potential for practical applications; nevertheless, MRC research in this domain is currently scarce. This article presents UIT-ViNewsQA, a new corpus for the Vietnamese language to evaluate MRC models for the healthcare textual domain. The corpus consists of 22,057 human-generated question-answer pairs. Crowd-workers create the questions and answers on a collection of 4,416 online Vietnamese healthcare news articles, where the answers are textual spans extracted from the corresponding articles. We introduce a process for creating a high-quality corpus for the Vietnamese machine reading comprehension task. Linguistically, our corpus accommodates diversity in question and answer types. In addition, we conduct experiments and compare the effectiveness of different MRC methods based on the neural networks and transformer architectures. Experimental results on our corpus show that the MRC system based on ALBERT architecture outperforms the neural network architectures and the BERT-based approach, an exact match score of 65.26% and an F1-score of 84.89%. The best machine model achieves about 10.90% F1-score less efficiently than humans, which proves that exploring machine models on UIT-ViNewsQA to surpass humans is challenging for researchers in the future. Our corpus is publicly available on our website: http://nlp.uit.edu.vn/datasets for research purposes. Kiet Van Nguyen, Tin Van Huynh, Duc-Vu Nguyen, Anh Gia-Tuan Nguyen, Ngan Luu-Thuy Nguyen |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2021 | A Large-Scale Dataset for Hate Speech Detection on Vietnamese Social Media Texts
Son T. Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
IEA/AIE (1) | 2 |
| 2021 | A Novel Perspective of Text Classification by Prolog-Based Deductive Databases
Kiet Van Nguyen, Tin Van Huynh, Anh Gia-Tuan Nguyen |
IEA/AIE (2) | 1 |
| 2021 | Constructive and Toxic Speech Detection for Open-Domain Social Media Comments in VietnameseabstractThe rise of social media has led to the increasing of comments on online forums. However, there still exists invalid comments which are not informative for users. Moreover, those comments are also quite toxic and harmful to people. In this paper, we create a dataset for constructive and toxic speech detection, named UIT-ViCTSD (Vietnamese Constructive and Toxic Speech Detection dataset) with 10,000 human-annotated comments. For these tasks, we propose a system for constructive and toxic speech detection with the state-of-the-art transfer learning model in Vietnamese NLP as PhoBERT. With this system, we obtain F1-scores of 78.59% and 59.40% for classifying constructive and toxic comments, respectively. Besides, we implement various baseline models as traditional Machine Learning and Deep Neural Network-Based models to evaluate the dataset. With the results, we can solve several tasks on the online discussions and develop the framework for identifying constructiveness and toxicity of Vietnamese social media comments automatically. Luan Thanh Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
IEA/AIE (1) | 2 |
| 2021 | Machine Learning-Based Empirical Investigation for Credit Scoring in Vietnam's Banking
Binh Van Duong, Linh Quang Tran, An Le-Hoai Tran, An Trong Nguyen, Kiet Van Nguyen |
IEA/AIE (2) | 6 |
| 2021 | Sentence Extraction-Based Machine Reading Comprehension for Vietnamese
Phong Nguyen-Thuan Do, Nhat Duy Nguyen, Tin Van Huynh, Kiet Van Nguyen, Anh Gia-Tuan Nguyen, Ngan Luu-Thuy Nguyen |
KSEM | 4 |
| 2021 | SA2SL: From Aspect-Based Sentiment Analysis to Social Listening System for Business Intelligence
Luong Luc Phan, Phuc Huynh Pham, Kim Thi-Thanh Nguyen, Sieu Khai Huynh, Tham Thi Nguyen, Luan Thanh Nguyen, Tin Van Huynh, Kiet Van Nguyen |
KSEM | 8 |
| 2021 | Span Detection for Vietnamese Aspect-Based Sentiment Analysis
Kim Thi-Thanh Nguyen, Sieu Khai Huynh, Phuc Huynh Pham, Luong Luc Phan, Duc-Vu Nguyen, Kiet Van Nguyen |
PACLIC | 6 |
| 2021 | Joint Chinese Word Segmentation and Part-of-speech Tagging via Two-stage Span Labeling
Duc-Vu Nguyen, Linh-Bao Vo, Ngoc-Linh Tran, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
PACLIC | 4 |
| 2021 | Monolingual vs multilingual BERTology for Vietnamese extractive multi-document summarization
Huy Quoc To, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen, Anh Gia-Tuan Nguyen |
PACLIC | 2 |
| 2021 | ViVQA: Vietnamese Visual Question Answering
An Trong Nguyen, An Tran-Hoai Le, Kiet Van Nguyen |
PACLIC | 4 |
| 2021 | Vietnamese Complaint Detection on E-Commerce WebsitesabstractCustomer product reviews play a role in improving the quality of products and services for business organizations or their brands. Complaining is an attitude that expresses dissatisfaction with an event or a product not meeting customer expectations. In this paper, we build a Vietnamese Open-domain Complaint Detection dataset (UIT-ViOCD), including 5,485 human-annotated reviews on four categories about product reviews on e-commerce sites. After the data collection phase, we proceed to the annotation task and achieve the inter-annotator agreement (Am) of 87%. Then, we present an extensive methodology for the research purposes and achieve 92.16% by F1-score for identifying complaints. With the results, in future, we aim to build a system for open-domain complaint detection on E-commerce websites. Nhung Thi-Hong Nguyen, Phuong Phan-Dieu Ha, Luan Thanh Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
SoMeT | 4 |
| 2021 | An Empirical Investigation of Online News Classification on an Open-Domain, Large-Scale and High-Quality Dataset in VietnameseabstractIn this paper, we build a new dataset UIT-ViON (Vietnamese Online Newspaper) collected from well-known online newspapers in Vietnamese. We collect, process, and create the dataset, then experiment with different machine learning models. In particular, we propose an open-domain, large-scale, and high-quality dataset consisting of 260,000 textual data points annotated with multiple labels for evaluating Vietnamese short text classification. In addition, we present the proposed approach using transformer-based learning (PhoBERT) for Vietnamese short text classification on the dataset, which outperforms traditional machine learning (Naive Bayes and Logistic Regression) and deep learning (Text-CNN and LSTM). As a result, the proposed approach achieves the F1-score of 80.62%. This is a positive result and a premise for developing an automatic news classification system. The study is proposed to significantly save time, costs, and human resources and make it easier for readers to find news related to their interesting topics. In future, we will propose solutions to improve the quality of the dataset and improve the performance of classification models. Phap Ngoc Trinh, Khoa Nguyen-Anh Tran, An Tran-Hoai Le, Luan Van Ha, Kiet Van Nguyen |
SoMeT | 6 |
| 2020 | A Vietnamese Dataset for Evaluating Machine Reading ComprehensionabstractOver 97 million people speak Vietnamese as their native language in the world.However, there are few research studies on machine reading comprehension (MRC) for Vietnamese, the task of understanding a text and answering questions related to it.Due to the lack of benchmark datasets for Vietnamese, we present the Vietnamese Question Answering Dataset (UIT-ViQuAD), a new dataset for the low-resource language as Vietnamese to evaluate MRC models.This dataset comprises over 23,000 human-generated question-answer pairs based on 5,109 passages of 174 Vietnamese articles from Wikipedia.In particular, we propose a new process of dataset creation for Vietnamese MRC.Our in-depth analyses illustrate that our dataset requires abilities beyond simple reasoning like word matching and demands single-sentence and multiple-sentence inferences.Besides, we conduct experiments on state-of-the-art MRC methods for English and Chinese as the first experimental models on UIT-ViQuAD.We also estimate human performance on the dataset and compare it to the experimental results of powerful machine learning models.As a result, the substantial differences between human performance and the best model performance on the dataset indicate that improvements can be made on UIT-ViQuAD in future research.Our dataset is freely available on our website 1 to encourage the research community to overcome challenges in Vietnamese MRC. Kiet Van Nguyen, Duc-Vu Nguyen, Anh Gia-Tuan Nguyen, Ngan Luu-Thuy Nguyen |
COLING | 1 |
| 2020 | UIT-ViIC: A Dataset for the First Evaluation on Vietnamese Image Captioning
Quan Hoang Lam, Quang-Duy Le, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
ICCCI | 3 |
| 2020 | A simple and efficient ensemble classifier combining multiple neural network models on social media datasets in Vietnamese
Huy Duc Huynh, Hang Thi-Thuy Do, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
PACLIC | 3 |
| 2020 | Empirical Study of Text Augmentation on Social Media Text in Vietnamese
Son T. Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
PACLIC | 2 |