VLDB 2026 Research / reviewers in the wild / expert
Dat Quoc Nguyen
dblp:23/9125
· DBLP profile ↗
29ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0001-8214-2878ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AccurateRAG: A Framework for Building Accurate Retrieval-Augmented Question-Answering Applications
Linh The Nguyen, Chi Tran, Dung Ngoc Nguyen, Van-Cuong Pham, Hoang Ngo, Dat Quoc Nguyen |
LREC | 6 |
| 2024 | Improving Vietnamese-English Medical Machine TranslationabstractMachine translation for Vietnamese-English in the medical domain is still an under-explored research area. In this paper, we introduce MedEV—a high-quality Vietnamese-English parallel dataset constructed specifically for the medical domain, comprising approximately 360K sentence pairs. We conduct extensive experiments comparing Google Translate, ChatGPT (gpt-3.5-turbo), state-of-the-art Vietnamese-English neural machine translation models and pre-trained bilingual/multilingual sequence-to-sequence models on our new MedEV dataset. Experimental results show that the best performance is achieved by fine-tuning “vinai-translate” for each translation direction. We publicly release our dataset to promote further research. Nhu Vo, Dat Quoc Nguyen, Dung D. Le, Massimo Piccardi, Wray L. Buntine |
LREC/COLING | 2 |
| 2024 | JPIS: A Joint Model for Profile-Based Intent Detection and Slot Filling with Slot-to-Intent AttentionabstractProfile-based intent detection and slot filling are important tasks aimed at reducing the ambiguity in user utterances by leveraging user-specific supporting profile information [1]. However, research in these two tasks has not been extensively explored. To fill this gap, we propose a joint model, namely JPIS, designed to enhance profile-based intent detection and slot filling. JPIS incorporates the supporting profile information into its encoder and introduces a slot-to-intent attention mechanism to transfer slot information representations to intent detection. Experimental results show that our JPIS substantially outperforms previous profile-based models, establishing a new state-of-the-art performance in overall accuracy on the Chinese benchmark dataset ProSLU [1]. Thinh Pham, Dat Quoc Nguyen |
ICASSP | 2 |
| 2023 | Two-View Graph Neural Networks for Knowledge Graph Completion
Vinh Tong, Dai Quoc Nguyen, Dinh Q. Phung, Dat Quoc Nguyen |
ESWC | 4 |
| 2023 | XPhoneBERT: A Pre-trained Multilingual Model for Phoneme Representations for Text-to-Speech
Linh The Nguyen, Thinh Pham, Dat Quoc Nguyen |
INTERSPEECH | 3 |
| 2022 | From Disfluency Detection to Intent Detection and Slot FillingabstractWe present the first empirical study investigating the influence of disfluency detection on downstream tasks of intent detection and slot filling. We perform this study for Vietnamese -- a low-resource language that has no previous study as well as no public dataset available for disfluency detection. First, we extend the fluent Vietnamese intent detection and slot filling dataset PhoATIS by manually adding contextual disfluencies and annotating them. Then, we conduct experiments using strong baselines for disfluency detection and joint intent detection and slot filling, which are based on pre-trained language models. We find that: (i) disfluencies produce negative effects on the performances of the downstream intent detection and slot filling tasks, and (ii) in the disfluency context, the pre-trained multilingual language model XLM-R helps produce better intent detection and slot filling performances than the pre-trained monolingual language model PhoBERT, and this is opposite to what generally found in the fluency context. Mai Hoang Dao, Hung-Thinh Truong, Dat Quoc Nguyen |
INTERSPEECH | 3 |
| 2022 | A Vietnamese-English Neural Machine Translation System
Tuan-Duy H. Nguyen, Duy Phung, Duy Tran-Cong Nguyen, Hieu Minh Tran, Manh Luong, Tin Duy Vo, Hung Hai Bui, Dinh Q. Phung, Dat Quoc Nguyen |
INTERSPEECH | 9 |
| 2022 | A High-Quality and Large-Scale Dataset for English-Vietnamese Speech TranslationabstractIn this paper, we introduce a high-quality and large-scale benchmark dataset for English-Vietnamese speech translation with 508 audio hours, consisting of 331K triplets of (sentencelengthed audio, English source transcript sentence, Vietnamese target subtitle sentence).We also conduct empirical experiments using strong baselines and find that the traditional "Cascaded" approach still outperforms the modern "End-to-End" approach.To the best of our knowledge, this is the first largescale English-Vietnamese speech translation study.We hope both our publicly available dataset and study can serve as a starting point for future research and applications on English-Vietnamese speech translation. Linh The Nguyen, Nguyen Luong Tran, Long Doan, Manh Luong, Dat Quoc Nguyen |
INTERSPEECH | 5 |
| 2022 | BARTpho: Pre-trained Sequence-to-Sequence Models for VietnameseabstractWe present BARTpho with two versions, BARTpho syllable and BARTpho word , which are the first public large-scale monolingual sequence-to-sequence models pre-trained for Vietnamese.BARTpho uses the "large" architecture and the pre-training scheme of the sequence-to-sequence denoising autoencoder BART, thus it is especially suitable for generative NLP tasks.We conduct experiments to compare our BARTpho with its competitor mBART on a downstream task of Vietnamese text summarization and show that: in both automatic and human evaluations, BARTpho outperforms the strong baseline mBART and improves the state-of-the-art.We further evaluate and compare BARTpho and mBART on the Vietnamese capitalization and punctuation restoration tasks and also find that BARTpho is more effective than mBART on these two tasks.We publicly release BARTpho to facilitate future research and applications of generative Vietnamese NLP tasks. Nguyen Luong Tran, Duong Minh Le, Dat Quoc Nguyen |
INTERSPEECH | 3 |
| 2022 | Node Co-occurrence based Graph Neural Networks for Knowledge Graph Link PredictionabstractWe introduce a novel embedding model, named NoGE, which aims to integrate co-occurrence among entities and relations into graph neural networks to improve knowledge graph completion (i.e., link prediction). Given a knowledge graph, NoGE constructs a single graph considering entities and relations as individual nodes. NoGE then computes weights for edges among nodes based on the co-occurrence of entities and relations. Next, NoGE proposes Dual Quaternion Graph Neural Networks (DualQGNN) and utilizes DualQGNN to update vector representations for entity and relation nodes. NoGE then adopts a score function to produce the triple scores. Comprehensive experimental results show that NoGE obtains state-of-the-art results on three new and difficult benchmark datasets CoDEx for knowledge graph completion. Dai Quoc Nguyen, Vinh Tong, Dinh Q. Phung, Dat Quoc Nguyen |
WSDM | 4 |
| 2021 | PhoMT: A High-Quality and Large-Scale Benchmark Dataset for Vietnamese-English Machine TranslationabstractWe introduce a high-quality and large-scale Vietnamese-English parallel dataset of 3.02M sentence pairs, which is 2.9M pairs larger than the benchmark Vietnamese-English machine translation corpus IWSLT15.We conduct experiments comparing strong neural baselines and well-known automatic translation engines on our dataset and find that in both automatic and human evaluations: the best performance is obtained by fine-tuning the pretrained sequence-to-sequence denoising autoencoder mBART.To our best knowledge, this is the first large-scale Vietnamese-English machine translation study.We hope our publicly available dataset and study can serve as a starting point for future research and applications on Vietnamese-English machine translation.We release our dataset at: https:// github.com/VinAIResearch/PhoMT. Long Doan, Linh The Nguyen, Nguyen Luong Tran, Thai Hoang, Dat Quoc Nguyen |
EMNLP (1) | 5 |
| 2021 | Intent Detection and Slot Filling for VietnameseabstractIntent detection and slot filling are important tasks in spoken and natural language understanding. However, Vietnamese is a low-resource language in these research topics. In this paper, we present the first public intent detection and slot filling dataset for Vietnamese. In addition, we also propose a joint model for intent detection and slot filling, that extends the recent state-of-the-art JointBERT+CRF model with an intent-slot attention layer to explicitly incorporate intent context information into slot filling via "soft" intent label embedding. Experimental results on our Vietnamese dataset show that our proposed model significantly outperforms JointBERT+CRF. We publicly release our dataset and the implementation of our model at: https://github.com/VinAIResearch/JointIDSF Mai Hoang Dao, Hung-Thinh Truong, Dat Quoc Nguyen |
Interspeech | 3 |
| 2021 | COVID-19 Named Entity Recognition for VietnameseabstractThe current COVID-19 pandemic has lead to the creation of many corpora that facilitate NLP research and downstream applications to help fight the pandemic.However, most of these corpora are exclusively for English.As the pandemic is a global problem, it is worth creating COVID-19 related datasets for languages other than English.In this paper, we present the first manuallyannotated COVID-19 domain-specific dataset for Vietnamese.Particularly, our dataset is annotated for the named entity recognition (NER) task with newly-defined entity types that can be used in other future epidemics.Our dataset also contains the largest number of entities compared to existing Vietnamese NER datasets.We empirically conduct experiments using strong baselines on our dataset, and find that: automatic Vietnamese word segmentation helps improve the NER results and the highest performances are obtained by finetuning pre-trained language models where the monolingual model PhoBERT for Vietnamese (Nguyen and Nguyen, 2020) produces higher results than the multilingual model XLM-R (Conneau et al., 2020).We Hung-Thinh Truong, Mai Hoang Dao, Dat Quoc Nguyen |
NAACL-HLT | 3 |
| 2020 | A Capsule Network-based Model for Learning Node EmbeddingsabstractIn this paper, we focus on learning low-dimensional embeddings for nodes in graph-structured data. To achieve this, we propose Caps2NE -- a new unsupervised embedding model leveraging a network of two capsule layers. Caps2NE induces a routing process to aggregate feature vectors of context neighbors of a given target node at the first capsule layer, then feed these features into the second capsule layer to infer a plausible embedding for the target node. Experimental results show that our proposed Caps2NE obtains state-of-the-art performances on benchmark datasets for the node classification task. Our code is available at: https://github.com/daiquocnguyen/Caps2NE. Dai Quoc Nguyen, Tu Dinh Nguyen, Dat Quoc Nguyen, Dinh Q. Phung |
CIKM | 3 |
| 2020 | ChEMU: Named Entity Recognition and Event Extraction of Chemical Reactions from Patents
Dat Quoc Nguyen, Zenan Zhai, Hiyori Yoshikawa, Biaoyan Fang, Christian Druckenbrodt, Camilo Thorne, Ralph Hoessel, Saber A. Akhondi, Trevor Cohn, Timothy Baldwin, Karin Verspoor |
ECIR (2) | 1 |
| 2020 | A Label Attention Model for ICD Coding from Clinical TextabstractICD coding is a process of assigning the International Classification of Disease diagnosis codes to clinical/medical notes documented by health professionals (e.g. clinicians). This process requires significant human resources, and thus is costly and prone to error. To handle the problem, machine learning has been utilized for automatic ICD coding. Previous state-of-the-art models were based on convolutional neural networks, using a single/several fixed window sizes. However, the lengths and interdependence between text fragments related to ICD codes in clinical text vary significantly, leading to the difficulty of deciding what the best window sizes are. In this paper, we propose a new label attention model for automatic ICD coding, which can handle both the various lengths and the interdependence of the ICD code related text fragments. Furthermore, as the majority of ICD codes are not frequently used, leading to the extremely imbalanced data issue, we additionally propose a hierarchical joint learning mechanism extending our label attention model to handle the issue, using the hierarchical relationships among the codes. Our label attention model achieves new state-of-the-art results on three benchmark MIMIC datasets, and the joint learning mechanism helps improve the performances for infrequent codes. Dat Quoc Nguyen, Anthony N. Nguyen |
IJCAI | 2 |
| 2019 | End-to-End Neural Relation Extraction Using Deep Biaffine Attention
Dat Quoc Nguyen, Karin Verspoor |
ECIR (1) | 1 |
| 2019 | From POS tagging to dependency parsing for biomedical event extractionabstractBACKGROUND: Given the importance of relation or event extraction from biomedical research publications to support knowledge capture and synthesis, and the strong dependency of approaches to this information extraction task on syntactic information, it is valuable to understand which approaches to syntactic processing of biomedical text have the highest performance. RESULTS: We perform an empirical study comparing state-of-the-art traditional feature-based and neural network-based models for two core natural language processing tasks of part-of-speech (POS) tagging and dependency parsing on two benchmark biomedical corpora, GENIA and CRAFT. To the best of our knowledge, there is no recent work making such comparisons in the biomedical context; specifically no detailed analysis of neural models on this data is available. Experimental results show that in general, the neural models outperform the feature-based models on two benchmark biomedical corpora GENIA and CRAFT. We also perform a task-oriented evaluation to investigate the influences of these models in a downstream application on biomedical event extraction, and show that better intrinsic parsing performance does not always imply better extrinsic event extraction performance. CONCLUSION: We have presented a detailed empirical study comparing traditional feature-based and neural network-based models for POS tagging and dependency parsing in the biomedical context, and also investigated the influence of parser selection for a biomedical event extraction downstream task. AVAILABILITY OF DATA AND MATERIALS: We make the retrained models available at https://github.com/datquocnguyen/BioPosDep . Dat Quoc Nguyen, Karin Verspoor |
BMC Bioinform. | 1 |
| 2018 | A Fast and Accurate Vietnamese Word Segmenter
Dat Quoc Nguyen, Dai Quoc Nguyen, Mark Dras, Mark Johnson 0001 |
LREC | 1 |
| 2017 | Search Personalization with Embeddings
Dat Quoc Nguyen, Mark Johnson 0001, Dawei Song 0001, Alistair Willis |
ECIR | 2 |
| 2016 | Neighborhood Mixture Model for Knowledge Base CompletionabstractKnowledge bases are useful resources for many natural language processing tasks, however, they are far from complete.In this paper, we define a novel entity representation as a mixture of its neighborhood in the knowledge base and apply this technique on TransE-a well-known embedding model for knowledge base completion.Experimental results show that the neighborhood information significantly helps to improve the results of the TransE, leading to better performance than obtained by other state-of-the-art embedding models on three benchmark datasets for triple classification, entity prediction and relation prediction tasks. Dat Quoc Nguyen, Kairit Sirts, Lizhen Qu, Mark Johnson 0001 |
CoNLL | 1 |
| 2016 | STransE: a novel embedding model of entities and relationships in knowledge basesabstractDat Quoc Nguyen, Kairit Sirts, Lizhen Qu, Mark Johnson. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Dat Quoc Nguyen, Kairit Sirts, Lizhen Qu, Mark Johnson 0001 |
HLT-NAACL | 1 |
| 2015 | Improving Topic Models with Latent Feature Word RepresentationsabstractProbabilistic topic models are widely used to discover latent topics in document collections, while latent feature vector representations of words have been used to obtain high performance in many NLP tasks. In this paper, we extend two different Dirichlet multinomial topic models by incorporating latent feature vector representations of words trained on very large corpora to improve the word-topic mapping learnt on a smaller corpus. Experimental results show that by using information from the external corpora, our new models produce significant improvements on topic coherence, document clustering and document classification tasks, especially on datasets with few or short documents. Dat Quoc Nguyen, Richard Billingsley, Lan Du 0002, Mark Johnson 0001 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2014 | RDRPOSTagger: A Ripple Down Rules-based Part-Of-Speech TaggerabstractThis paper describes our robust, easyto-use and language independent toolkit namely RDRPOSTagger which employs an error-driven approach to automatically construct a Single Classification Ripple Down Rules tree of transformation rules for POS tagging task.During the demonstration session, we will run the tagger on data sets in 15 different languages. Dat Quoc Nguyen, Dai Quoc Nguyen, Dang Duc Pham, Son Bao Pham |
EACL | 1 |
| 2014 | From Treebank Conversion to Automatic Dependency Parsing for Vietnamese
Dat Quoc Nguyen, Dai Quoc Nguyen, Son Bao Pham, Phuong-Thai Nguyen, Minh Le Nguyen 0001 |
NLDB | 1 |
| 2013 | A Two-Stage Classifier for Sentiment Analysis
Dai Quoc Nguyen, Dat Quoc Nguyen, Son Bao Pham |
IJCNLP | 2 |
| 2012 | A Semantic Approach for Question Analysis
Dai Quoc Nguyen, Dat Quoc Nguyen, Son Bao Pham |
IEA/AIE | 2 |
| 2012 | A Vietnamese Text-Based Conversational Agent
Dai Quoc Nguyen, Dat Quoc Nguyen, Son Bao Pham |
IEA/AIE | 2 |
| 2011 | Ripple Down Rules for Part-of-Speech Tagging
Dat Quoc Nguyen, Dai Quoc Nguyen, Son Bao Pham, Dang Duc Pham |
CICLing (1) | 1 |