VLDB 2026 Research / reviewers in the wild / expert
Tin Van Huynh
dblp:255/5794
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0003-4990-2868ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A new benchmark dataset and mixture-of-experts language models for adversarial natural language inference in Vietnamese
Tin Van Huynh, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
Expert Syst. Appl. | 1 |
| 2025 | ViNumFCR: A Novel Vietnamese Benchmark for Numerical Reasoning Fact Checking on Social Media NewsabstractIn the digital era, the internet provides rapid and convenient access to vast amounts of information. However, much of this information remains unverified, particularly with the increasing prevalence of falsified numerical data, leading to public confusion and negative societal impacts. To address this issue, we developed ViNumFCR, a first dataset dedicated to fact-checking numerical information in Vietnamese. Comprising over 10,000 samples collected and constructed from online newspaper across 12 different topics. We assessed the performance of various fact-checking models, including Pretrained Language Models and Large Language Models, alongside retrieval techniques for gathering supporting evidence. Experimental results demonstrate that the XLM-R_Large model achieved the highest accuracy of 90.05% on the fact-checking task, while the combined SBERT + BM25 model attained a precision of over 97% on the evidence retrieval task. Additionally, we conducted an in-depth analysis of the linguistic features of the dataset to understand the factors influencing the performance models. The ViNumFCR dataset is publicly available to support further research. Nhi Ngoc Phuong Luong, Anh Thi Lan Le, Tin Van Huynh, Kiet Van Nguyen, Ngan Nguyen |
INLG | 3 |
| 2025 | LMCK: pre-trained language models enhanced with contextual knowledge for Vietnamese natural language inference
Ngan Luu-Thuy Nguyen, Khoa Thi-Kim Phan, Tin Van Huynh, Kiet Van Nguyen |
Multim. Tools Appl. | 3 |
| 2023 | Machine Reading Comprehension for Vietnamese Customer Reviews: Task, Corpus and Baseline Models
Tinh Pham Phuc Do, Ngoc Dinh Duy Cao, Tin Van Huynh, Kiet Van Nguyen |
PACLIC | 4 |
| 2022 | XLMRQA: Open-Domain Question Answering on Vietnamese Wikipedia-Based Textual Knowledge Source
Kiet Van Nguyen, Phong Nguyen-Thuan Do, Nhat Duy Nguyen, Tin Van Huynh, Anh Gia-Tuan Nguyen, Ngan Luu-Thuy Nguyen |
ACIIDS (1) | 4 |
| 2022 | ViNLI: A Vietnamese Corpus for Studies on Open-Domain Natural Language InferenceabstractOver a decade, the research field of computational linguistics has witnessed the growth of corpora and models for natural language inference (NLI) for rich-resource languages such as English and Chinese. A large-scale and high-quality corpus is necessary for studies on NLI for Vietnamese, which can be considered a low-resource language. In this paper, we introduce ViNLI (Vietnamese Natural Language Inference), an open-domain and high-quality corpus for evaluating Vietnamese NLI models, which is created and evaluated with a strict process of quality control. ViNLI comprises over 30,000 human-annotated premise-hypothesis sentence pairs extracted from more than 800 online news articles on 13 distinct topics. In this paper, we introduce the guidelines for corpus creation which take the specific characteristics of the Vietnamese language in expressing entailment and contradiction into account. To evaluate the challenging level of our corpus, we conduct experiments with state-of-the-art deep neural networks and pre-trained models on our dataset. The best system performance is still far from human performance (a 14.20% gap in accuracy). The ViNLI corpus is a challenging corpus to accelerate progress in Vietnamese computational linguistics. Our corpus is available publicly for research purposes. Tin Van Huynh, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
COLING | 1 |
| 2022 | New Vietnamese Corpus for Machine Reading Comprehension of Health News ArticlesabstractMachine reading comprehension is a natural language understanding task where the computing system is required to read a text and then find the answer to a specific question posed by a human. Large-scale and high-quality corpora are necessary for evaluating machine reading comprehension models. Furthermore, machine reading comprehension (MRC) for the health sector has potential for practical applications; nevertheless, MRC research in this domain is currently scarce. This article presents UIT-ViNewsQA, a new corpus for the Vietnamese language to evaluate MRC models for the healthcare textual domain. The corpus consists of 22,057 human-generated question-answer pairs. Crowd-workers create the questions and answers on a collection of 4,416 online Vietnamese healthcare news articles, where the answers are textual spans extracted from the corresponding articles. We introduce a process for creating a high-quality corpus for the Vietnamese machine reading comprehension task. Linguistically, our corpus accommodates diversity in question and answer types. In addition, we conduct experiments and compare the effectiveness of different MRC methods based on the neural networks and transformer architectures. Experimental results on our corpus show that the MRC system based on ALBERT architecture outperforms the neural network architectures and the BERT-based approach, an exact match score of 65.26% and an F1-score of 84.89%. The best machine model achieves about 10.90% F1-score less efficiently than humans, which proves that exploring machine models on UIT-ViNewsQA to surpass humans is challenging for researchers in the future. Our corpus is publicly available on our website: http://nlp.uit.edu.vn/datasets for research purposes. Kiet Van Nguyen, Tin Van Huynh, Duc-Vu Nguyen, Anh Gia-Tuan Nguyen, Ngan Luu-Thuy Nguyen |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2021 | A Novel Perspective of Text Classification by Prolog-Based Deductive Databases
Kiet Van Nguyen, Tin Van Huynh, Anh Gia-Tuan Nguyen |
IEA/AIE (2) | 2 |
| 2021 | Sentence Extraction-Based Machine Reading Comprehension for Vietnamese
Phong Nguyen-Thuan Do, Nhat Duy Nguyen, Tin Van Huynh, Kiet Van Nguyen, Anh Gia-Tuan Nguyen, Ngan Luu-Thuy Nguyen |
KSEM | 3 |
| 2021 | SA2SL: From Aspect-Based Sentiment Analysis to Social Listening System for Business Intelligence
Luong Luc Phan, Phuc Huynh Pham, Kim Thi-Thanh Nguyen, Sieu Khai Huynh, Tham Thi Nguyen, Luan Thanh Nguyen, Tin Van Huynh, Kiet Van Nguyen |
KSEM | 7 |