VLDB 2026 Research / reviewers in the wild / expert
Dinh Dien
dblp:09/156 · also Dien Dinh
· DBLP profile ↗
25ranked-venue papers
2as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ChartReLA: A compact vision-language model for comprehensive chart reasoning via relationship modeling
Xuan-Quang Nguyen, Quang-Tan Nguyen, Long H. B. Nguyen, Dinh Dien |
Inf. Process. Manag. | 4 |
| 2025 | ViPhoVQA: Toward a Phonemic-Based Method for Mitigating Rare and Out-of-Vocab Words in Vietnamese Text-Based Visual Question Answering
Nghia Hieu Nguyen, Duc-Vu Nguyen, Huy Quoc To, Dinh Dien, Thanh Tho Quan, Ngan Luu-Thuy Nguyen |
ICCCI (1) | 4 |
| 2025 | Enhancing Named Entity Translation from Classical Chinese to Vietnamese in Traditional Vietnamese Medicine Domain: A Hybrid Masking and Dictionary-Augmented ApproachabstractVietnam’s traditional medical texts were historically written in Classical Chinese using Sino-Vietnamese pronunciations. As the Vietnamese language transitioned to a Latin-based national script and interest in integrating traditional medicine with modern healthcare grows, accurate translation of these texts has become increasingly important. However, the diversity of terminology and the complexity of translating medical entities into modern contexts pose significant challenges. To address this, we propose a method that fine-tunes large language models (LLMs) using augmented data and a Hybrid Entity Masking and Replacement (HEMR) strategy to improve named entity translation. We also introduce a parallel named entity translation dataset specifically curated for traditional Vietnamese medicine. Our evaluation across multiple LLMs shows that the proposed approach achieves a translation accuracy of 71.91%, demonstrating its effectiveness. These results underscore the importance of incorporating named entity awareness into translation systems, particularly in low-resource and domain-specific settings like traditional Vietnamese medicine. Nhu Pham, Uyen Nguyen, Long H. B. Nguyen, Dinh Dien |
INLG | 4 |
| 2024 | A Comparative Study of Chart Summarization
An Chu, Thong Huynh, Long H. B. Nguyen, Dinh Dien |
PACLIC | 4 |
| 2024 | VHE: A New Dataset for Event Extraction from Vietnamese Historical Texts
Truc Hoang, Long H. B. Nguyen, Dinh Dien |
PACLIC | 3 |
| 2024 | Advancing Vietnamese Information Retrieval with Learning Objective and Benchmark
Phu-Vinh Nguyen, Minh-Nam Tran, Long H. B. Nguyen, Dinh Dien |
PACLIC | 4 |
| 2024 | ViHerbQA: A Robust QA Model for Vietnamese Traditional Herbal Medicine
Quyen Truong, Long H. B. Nguyen, Dinh Dien |
PACLIC | 3 |
| 2024 | Multi-mask Prefix Tuning: Applying Multiple Adaptive Masks on Deep Prompt Tuning
Qui Tu, Long H. B. Nguyen, Dinh Dien |
PACLIC | 4 |
| 2024 | MUPO: Unlocking the Potential of Unpaired Data for Single Stage Preference Alignment in Language ModelsabstractPreference alignment has recently been on the forefront of helping language model (LM) aligning model behaviors with human preferences, with the most relevant application being chat bots. With the success of Direct Preference Optimization (DPO), there have been many derivative works of the same paradigm that further better different aspects, in both ease of training and performance. Odds Ratio Preference Optimization (ORPO) provides the next step in the development of preference alignment by combining the training pipeline into a single monolithic reference optimization process. In addition, ORPO also eliminates the need for a reference model to further improve on training efficiency without negatively affecting performance. However, ORPO, and many of its predecessor, requires datasets with pre-defined reference pairs of example which could severely limit the amount of compatible training data and lower LM's capability compares to what could be potentially achieved. Inspired by this limitation, we devise a new approach, Monolithic Unpaired Preference Optimization (MUPO), that is capable of making use of the abundantly available data without paired reference and combining the training process into a single monolithic stage for further efficiency and accessibility. Through our experimentation, we found that MUPO's test performances on Stable LM and Phi-2 stayed competitive with previous methods like SFT, SFT+DPO, and ORPO. Truong-Duy Dang, The-Cuong Quang, Long H. B. Nguyen, Dinh Dien |
SMC | 4 |
| 2024 | XLIT: A Method to Bridge Task Discrepancy in Machine Translation Pre-trainingabstractTransfer learning from pre-trained language models to encoder-decoder translation models faces a challenge due to the mismatch between the tasks of pre-training and fine-tuning. Pre-trained models are not explicitly trained to understand the semantic interactions between different languages. To address this issue, a cross-lingual embedding space is used as an interface during the pre-training phase. This approach enables the decoder inputs to attend to the encoder outputs, similar to the fine-tuning process. However, the effectiveness of this transfer heavily relies on the quality of the pre-trained unsupervised cross-lingual embeddings, which introduces complexity and reduces reproducibility. In this study, we propose a pre-training method called Cross-lingual Interaction Transfer (XLIT), which does not depend on other embedding techniques. XLIT effectively reconciles the task discrepancy in machine translation fine-tuning. We conducted extensive experiments involving four low-resource and six very low-resource translation directions. The results of our experiments demonstrate that our method surpasses randomly initialized models and previous pre-training techniques by up to 9.4 BLEU. Furthermore, we demonstrate that our method achieves comparable performance when pre-trained with large-scale monolingual data from various languages. Khang Pham, Long H. B. Nguyen, Dinh Dien |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2023 | Exploring the Role of Monolingual Data in Cross-Attention Pre-training for Neural Machine Translation
Khang Pham, Long H. B. Nguyen, Dinh Dien |
ICCCI | 3 |
| 2022 | Multi-level Community-awareness Graph Neural Networks for Neural Machine TranslationabstractNeural Machine Translation (NMT) aims to translate the source- to the target-language while preserving the original meaning. Linguistic information such as morphology, syntactic, and semantics shall be grasped in token embeddings to produce a high-quality translation. Recent works have leveraged the powerful Graph Neural Networks (GNNs) to encode such language knowledge into token embeddings. Specifically, they use a trained parser to construct semantic graphs given sentences and then apply GNNs. However, most semantic graphs are tree-shaped and too sparse for GNNs which cause the over-smoothing problem. To alleviate this problem, we propose a novel Multi-level Community-awareness Graph Neural Network (MC-GNN) layer to jointly model local and global relationships between words and their linguistic roles in multiple communities. Intuitively, the MC-GNN layer substitutes a self-attention layer at the encoder side of a transformer-based machine translation model. Extensive experiments on four language-pair datasets with common evaluation metrics show the remarkable improvements of our method while reducing the time complexity in very long sentences. Binh P. Nguyen, Long H. B. Nguyen, Dinh Dien |
COLING | 3 |
| 2022 | Intent Detection and Slot Filling from Dependency Parsing Perspective: A Case Study in Vietnamese
Phu-Thinh Pham, Duy Vu-Tran, Duc Do, An-Vinh Luong, Dinh Dien |
PACLIC | 5 |
| 2022 | Integrating Label Attention into CRF-based Vietnamese Constituency Parser
Duy Vu-Tran, Phu-Thinh Pham, Duc Do, An-Vinh Luong, Dinh Dien |
PACLIC | 5 |
| 2020 | Identifying Authors Based on Stylometric measures of Vietnamese texts
Ho Ngoc Lam, Vo Diep Nhu, Dinh Dien, Nguyen Tuyet Nhung |
PACLIC | 3 |
| 2019 | Low-Resource Machine Transliteration Using Recurrent Neural NetworksabstractGrapheme-to-phoneme models are key components in automatic speech recognition and text-to-speech systems. With low-resource language pairs that do not have available and well-developed pronunciation lexicons, grapheme-to-phoneme models are particularly useful. These models are based on initial alignments between grapheme source and phoneme target sequences. Inspired by sequence-to-sequence recurrent neural network--based translation methods, the current research presents an approach that applies an alignment representation for input sequences and pretrained source and target embeddings to overcome the transliteration problem for a low-resource languages pair. Evaluation and experiments involving French and Vietnamese showed that with only a small bilingual pronunciation dictionary available for training the transliteration models, promising results were obtained with a large increase in BLEU scores and a reduction in Translation Error Rate (TER) and Phoneme Error Rate (PER). Moreover, we compared our proposed neural network--based transliteration approach with a statistical one. Ngoc Tan Le, Fatiha Sadat, Lucie Ménard, Dinh Dien |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2017 | Linguistic-Relationships-Based Approach for Improving Word AlignmentabstractThe unsupervised word alignments (such as GIZA++) are widely used in the phrase-based statistical machine translation. The quality of the model is proportional to the size and the quality of the bilingual corpus. However, for low-resource language pairs such as Chinese and Vietnamese, a result of unsupervised word alignment sometimes is of low quality due to the sparse data. In addition, this model does not take advantage of the linguistic relationships to improve performance of word alignment. Chinese and Vietnamese have the same language type and have close linguistic relationships. In this article, we integrate the characteristics of linguistic relationships into the word alignment model to enhance the quality of Chinese-Vietnamese word alignment. These linguistic relationships are Sino-Vietnamese and content word. The experimental results showed that our method improved the performance of word alignment as well as the quality of machine translation. Phuoc Tran, Dinh Dien, Tan Le, Long H. B. Nguyen |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2016 | An Approach to Construct a Named Entity Annotated English-Vietnamese Bilingual CorpusabstractManually constructing an annotated Named Entity (NE) in a bilingual corpus is a time-consuming, labor--intensive, and expensive process, but this is necessary for natural language processing (NLP) tasks such as cross-lingual information retrieval, cross-lingual information extraction, machine translation, etc. In this article, we present an automatic approach to construct an annotated NE in English-Vietnamese bilingual corpus from a bilingual parallel corpus by proposing an aligned NE method. Basing this corpus on a bilingual corpus in which the initial NEs are extracted from its own language separately, the approach tries to correct unrecognized NEs or incorrectly recognized NEs before aligning the NEs by using a variety of bilingual constraints. The generated corpus not only improves the NE recognition results but also creates alignments between English NEs and Vietnamese NEs, which are necessary for training NE translation models. The experimental results show that the approach outperforms the baseline methods effectively. In the English-Vietnamese NE alignment task, the F-measure increases from 68.58% to 79.77%. Thanks to the improvement of the NE recognition quality, the proposed method also increases significantly: the F-measure goes from 84.85% to 88.66% for the English side and from 75.71% to 85.55% for the Vietnamese side. By providing the additional semantic information for the machine translation systems, the BLEU score increases from 33.04% to 45.11%. Long H. B. Nguyen, Dinh Dien, Phuoc Tran |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2016 | Word Re-Segmentation in Chinese-Vietnamese Machine TranslationabstractIn isolated languages, such as Chinese and Vietnamese, words are not separated by spaces, and a word may be formed by one or more syllables. Therefore, word segmentation (WS) is usually the first process that is implemented in the machine translation process. WS in the source and target languages is based on different training corpora, and WS approaches may not be the same. Therefore, the WS that results in these two languages are not often homologous, and thus word alignment results in many 1-n and n-1 alignment pairs in statistical machine translation, which degrades the performance of machine translation. In this article, we will adjust the WS for both Chinese and Vietnamese in particular and for isolated language pairs in general and make the word boundary of the two languages more symmetric in order to strengthen 1-1 alignments and enhance machine translation performance. We have tested this method on the Computational Linguistics Center’s corpus, which consists of 35,623 sentence pairs. The experimental results show that our method has significantly improved the performance of machine translation compared to the baseline translation system, WS translation system, and anchor language-based WS translation systems. Phuoc Tran, Dinh Dien, Long H. B. Nguyen |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2010 | An ontology-driven system for detecting global health events
Nigel Collier, Reiko Matsuda Goodwin, John P. McCrae, Son Doan, Ai Kawazoe, Mike Conway, Asanee Kawtrakul, Koichi Takeuchi, Dinh Dien |
COLING | 9 |
| 2008 | A Syntactic-based Word Re-ordering for English-Vietnamese Statistical Machine Translation System
Hong-Nhung Nguyen Thi, Dinh Dien |
PRICAI | 2 |
| 2008 | BioCaster: detecting public health rumors with a Web-based text mining systemabstractSUMMARY: BioCaster is an ontology-based text mining system for detecting and tracking the distribution of infectious disease outbreaks from linguistic signals on the Web. The system continuously analyzes documents reported from over 1700 RSS feeds, classifies them for topical relevance and plots them onto a Google map using geocoded information. The background knowledge for bridging the gap between Layman's terms and formal-coding systems is contained in the freely available BioCaster ontology which includes information in eight languages focused on the epidemiological role of pathogens as well as geographical locations with their latitudes/longitudes. The system consists of four main stages: topic classification, named entity recognition (NER), disease/location detection and event recognition. Higher order event analysis is used to detect more precisely specified warning signals that can then be notified to registered users via email alerts. Evaluation of the system for topic recognition and entity identification is conducted on a gold standard corpus of annotated news articles. AVAILABILITY: The BioCaster map and ontology are freely available via a web portal at http://www.biocaster.org. Nigel Collier, Son Doan, Ai Kawazoe, Reiko Matsuda Goodwin, Mike Conway, Yoshio Tateno, Hung Quoc Ngo 0002, Dinh Dien, Asanee Kawtrakul, Koichi Takeuchi, Mika Shigematsu, Kiyosu Taniguchi |
Bioinform. | 8 |
| 2007 | Named entity recognition in Vietnamese using classifier votingabstractNamed entity recognition (NER) is one of the fundamental tasks in natural-language processing (NLP). Though the combination of different classifiers has been widely applied in several well-studied languages, this is the first time this method has been applied to Vietnamese. In this article, we describe how voting techniques can improve the performance of Vietnamese NER. By combining several state-of-the-art machine-learning algorithms using voting strategies, our final result outperforms individual algorithms and gained an F-measure of 89.12. A detailed discussion about the challenges of NER in Vietnamese is also presented. Pham Thi Xuan Thao, Tran Quoc Tri, Dinh Dien, Nigel Collier |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2003 | BTL: a hybrid model for English-Vietnamese machine translationabstractMachine Translation (MT) is the most interesting and difficult task which has been posed since the beginning of computer history. The highest difficulty which computers had to face with, is the built-in ambiguity of Natural Languages. Formerly, a lot of human-devised rules have been used to disambiguate those ambiguities. Building such a complete rule-set is time-consuming and labor-intensive task whilst it doesn’t cover all the cases. Besides, when the scale of system increases, it is very difficult to control that rule-set. In this paper, we present a new model of learning-based MT (entitled BTL: Bitext-Transfer Learning) that learns from bilingual corpus to extract disambiguating rules. This model has been experimented in English-to-Vietnamese MT system (EVT) and it gave encouraging results. Dinh Dien, Kiem Hoang, Eduard H. Hovy |
MTSummit | 1 |
| 2003 | A hybrid approach to word order transfer in the English-to-Vietnamese machine translationabstractWord Order transfer is a compulsory stage and has a great effect on the translation result of a transfer-based machine translation system. To solve this problem, we can use fixed rules (rule-based) or stochastic methods (corpus-based) which extract word order transfer rules between two languages. However, each approach has its own advantages and disadvantages. In this paper, we present a hybrid approach based on fixed rules and Transformation-Based Learning (or TBL) method. Our purpose is to transfer automatically the English word orders into the Vietnamese ones. The learning process will be trained on the annotated bilingual corpus (named EVC: English-Vietnamese Corpus) that has been automatically word-aligned, phrase-aligned and POS-tagged. This transfer result is being used for the transfer module in the English-Vietnamese transfer-based machine translation system. Dinh Dien, Ngan Luu-Thuy Nguyen, Quang Xuan Do, Nam Van Chi |
MTSummit | 1 |