Long H. B. Nguyen

dblp:192/3829 · also Long Hong Buu Nguyen, Long Nguyen 0005 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0002-0884-1635ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ChartReLA: A compact vision-language model for comprehensive chart reasoning via relationship modeling
Xuan-Quang Nguyen, Quang-Tan Nguyen, Long H. B. Nguyen, Dinh Dien
Inf. Process. Manag.3
2025 Enhancing Named Entity Translation from Classical Chinese to Vietnamese in Traditional Vietnamese Medicine Domain: A Hybrid Masking and Dictionary-Augmented Approach
abstract
Vietnam’s traditional medical texts were historically written in Classical Chinese using Sino-Vietnamese pronunciations. As the Vietnamese language transitioned to a Latin-based national script and interest in integrating traditional medicine with modern healthcare grows, accurate translation of these texts has become increasingly important. However, the diversity of terminology and the complexity of translating medical entities into modern contexts pose significant challenges. To address this, we propose a method that fine-tunes large language models (LLMs) using augmented data and a Hybrid Entity Masking and Replacement (HEMR) strategy to improve named entity translation. We also introduce a parallel named entity translation dataset specifically curated for traditional Vietnamese medicine. Our evaluation across multiple LLMs shows that the proposed approach achieves a translation accuracy of 71.91%, demonstrating its effectiveness. These results underscore the importance of incorporating named entity awareness into translation systems, particularly in low-resource and domain-specific settings like traditional Vietnamese medicine.
Nhu Pham, Uyen Nguyen, Long H. B. Nguyen, Dinh Dien
INLG3
2024 A Comparative Study of Chart Summarization
An Chu, Thong Huynh, Long H. B. Nguyen, Dinh Dien
PACLIC3
2024 VHE: A New Dataset for Event Extraction from Vietnamese Historical Texts
Truc Hoang, Long H. B. Nguyen, Dinh Dien
PACLIC2
2024 Advancing Vietnamese Information Retrieval with Learning Objective and Benchmark
Phu-Vinh Nguyen, Minh-Nam Tran, Long H. B. Nguyen, Dinh Dien
PACLIC3
2024 ViHerbQA: A Robust QA Model for Vietnamese Traditional Herbal Medicine
Quyen Truong, Long H. B. Nguyen, Dinh Dien
PACLIC2
2024 Multi-mask Prefix Tuning: Applying Multiple Adaptive Masks on Deep Prompt Tuning
Qui Tu, Long H. B. Nguyen, Dinh Dien
PACLIC3
2024 MUPO: Unlocking the Potential of Unpaired Data for Single Stage Preference Alignment in Language Models
abstract
Preference alignment has recently been on the forefront of helping language model (LM) aligning model behaviors with human preferences, with the most relevant application being chat bots. With the success of Direct Preference Optimization (DPO), there have been many derivative works of the same paradigm that further better different aspects, in both ease of training and performance. Odds Ratio Preference Optimization (ORPO) provides the next step in the development of preference alignment by combining the training pipeline into a single monolithic reference optimization process. In addition, ORPO also eliminates the need for a reference model to further improve on training efficiency without negatively affecting performance. However, ORPO, and many of its predecessor, requires datasets with pre-defined reference pairs of example which could severely limit the amount of compatible training data and lower LM's capability compares to what could be potentially achieved. Inspired by this limitation, we devise a new approach, Monolithic Unpaired Preference Optimization (MUPO), that is capable of making use of the abundantly available data without paired reference and combining the training process into a single monolithic stage for further efficiency and accessibility. Through our experimentation, we found that MUPO's test performances on Stable LM and Phi-2 stayed competitive with previous methods like SFT, SFT+DPO, and ORPO.
Truong-Duy Dang, The-Cuong Quang, Long H. B. Nguyen, Dinh Dien
SMC3
2024 XLIT: A Method to Bridge Task Discrepancy in Machine Translation Pre-training
abstract
Transfer learning from pre-trained language models to encoder-decoder translation models faces a challenge due to the mismatch between the tasks of pre-training and fine-tuning. Pre-trained models are not explicitly trained to understand the semantic interactions between different languages. To address this issue, a cross-lingual embedding space is used as an interface during the pre-training phase. This approach enables the decoder inputs to attend to the encoder outputs, similar to the fine-tuning process. However, the effectiveness of this transfer heavily relies on the quality of the pre-trained unsupervised cross-lingual embeddings, which introduces complexity and reduces reproducibility. In this study, we propose a pre-training method called Cross-lingual Interaction Transfer (XLIT), which does not depend on other embedding techniques. XLIT effectively reconciles the task discrepancy in machine translation fine-tuning. We conducted extensive experiments involving four low-resource and six very low-resource translation directions. The results of our experiments demonstrate that our method surpasses randomly initialized models and previous pre-training techniques by up to 9.4 BLEU. Furthermore, we demonstrate that our method achieves comparable performance when pre-trained with large-scale monolingual data from various languages.
Khang Pham, Long H. B. Nguyen, Dinh Dien
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2023 Exploring the Role of Monolingual Data in Cross-Attention Pre-training for Neural Machine Translation
Khang Pham, Long H. B. Nguyen, Dinh Dien
ICCCI2
2022 Multi-level Community-awareness Graph Neural Networks for Neural Machine Translation
abstract
Neural Machine Translation (NMT) aims to translate the source- to the target-language while preserving the original meaning. Linguistic information such as morphology, syntactic, and semantics shall be grasped in token embeddings to produce a high-quality translation. Recent works have leveraged the powerful Graph Neural Networks (GNNs) to encode such language knowledge into token embeddings. Specifically, they use a trained parser to construct semantic graphs given sentences and then apply GNNs. However, most semantic graphs are tree-shaped and too sparse for GNNs which cause the over-smoothing problem. To alleviate this problem, we propose a novel Multi-level Community-awareness Graph Neural Network (MC-GNN) layer to jointly model local and global relationships between words and their linguistic roles in multiple communities. Intuitively, the MC-GNN layer substitutes a self-attention layer at the encoder side of a transformer-based machine translation model. Extensive experiments on four language-pair datasets with common evaluation metrics show the remarkable improvements of our method while reducing the time complexity in very long sentences.
Binh P. Nguyen, Long H. B. Nguyen, Dinh Dien
COLING2
2017 Linguistic-Relationships-Based Approach for Improving Word Alignment
abstract
The unsupervised word alignments (such as GIZA++) are widely used in the phrase-based statistical machine translation. The quality of the model is proportional to the size and the quality of the bilingual corpus. However, for low-resource language pairs such as Chinese and Vietnamese, a result of unsupervised word alignment sometimes is of low quality due to the sparse data. In addition, this model does not take advantage of the linguistic relationships to improve performance of word alignment. Chinese and Vietnamese have the same language type and have close linguistic relationships. In this article, we integrate the characteristics of linguistic relationships into the word alignment model to enhance the quality of Chinese-Vietnamese word alignment. These linguistic relationships are Sino-Vietnamese and content word. The experimental results showed that our method improved the performance of word alignment as well as the quality of machine translation.
Phuoc Tran, Dinh Dien, Tan Le, Long H. B. Nguyen
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2016 An Approach to Construct a Named Entity Annotated English-Vietnamese Bilingual Corpus
abstract
Manually constructing an annotated Named Entity (NE) in a bilingual corpus is a time-consuming, labor--intensive, and expensive process, but this is necessary for natural language processing (NLP) tasks such as cross-lingual information retrieval, cross-lingual information extraction, machine translation, etc. In this article, we present an automatic approach to construct an annotated NE in English-Vietnamese bilingual corpus from a bilingual parallel corpus by proposing an aligned NE method. Basing this corpus on a bilingual corpus in which the initial NEs are extracted from its own language separately, the approach tries to correct unrecognized NEs or incorrectly recognized NEs before aligning the NEs by using a variety of bilingual constraints. The generated corpus not only improves the NE recognition results but also creates alignments between English NEs and Vietnamese NEs, which are necessary for training NE translation models. The experimental results show that the approach outperforms the baseline methods effectively. In the English-Vietnamese NE alignment task, the F-measure increases from 68.58% to 79.77%. Thanks to the improvement of the NE recognition quality, the proposed method also increases significantly: the F-measure goes from 84.85% to 88.66% for the English side and from 75.71% to 85.55% for the Vietnamese side. By providing the additional semantic information for the machine translation systems, the BLEU score increases from 33.04% to 45.11%.
Long H. B. Nguyen, Dinh Dien, Phuoc Tran
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2016 Word Re-Segmentation in Chinese-Vietnamese Machine Translation
abstract
In isolated languages, such as Chinese and Vietnamese, words are not separated by spaces, and a word may be formed by one or more syllables. Therefore, word segmentation (WS) is usually the first process that is implemented in the machine translation process. WS in the source and target languages is based on different training corpora, and WS approaches may not be the same. Therefore, the WS that results in these two languages are not often homologous, and thus word alignment results in many 1-n and n-1 alignment pairs in statistical machine translation, which degrades the performance of machine translation. In this article, we will adjust the WS for both Chinese and Vietnamese in particular and for isolated language pairs in general and make the word boundary of the two languages more symmetric in order to strengthen 1-1 alignments and enhance machine translation performance. We have tested this method on the Computational Linguistics Center’s corpus, which consists of 35,623 sentence pairs. The experimental results show that our method has significantly improved the performance of machine translation compared to the baseline translation system, WS translation system, and anchor language-based WS translation systems.
Phuoc Tran, Dinh Dien, Long H. B. Nguyen
ACM Trans. Asian Low Resour. Lang. Inf. Process.3