Hieu Hoang

dblp:46/3446 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 42% Machine translation · 39% Transfer learning and domain adaptation · 18%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
1.122025
Adapters for Altering LLM Vocabularies: What Languages Benefit the Most? · ICLR 2025
X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at Scale · ICLR 2025
Natural language and speech › Machine translation
low-resource machine translation
0.912025
X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at Scale · ICLR 2025
Natural language and speech › Language models and text generation
multilingual language models
0.912025
Adapters for Altering LLM Vocabularies: What Languages Benefit the Most? · ICLR 2025
Natural language and speech › Language models and text generation › multilingual language models
multilingual large language model
0.912025
X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at Scale · ICLR 2025
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation
0.912025
X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at Scale · ICLR 2025
Natural language and speech › Language models and text generation › tokenization
vocabulary adaptation
0.912025
Adapters for Altering LLM Vocabularies: What Languages Benefit the Most? · ICLR 2025
Natural language and speech › Machine translation
parallel corpus mining
0.412020
ParaCrawl: Web-Scale Acquisition of Parallel Corpora · ACL 2020
Information retrieval
cross-language information retrieval
0.412020
ParaCrawl: Web-Scale Acquisition of Parallel Corpora · ACL 2020
Natural language and speech › Machine translation
statistical machine translation
0.122007
Factored Translation Models · EMNLP-CoNLL 2007
Moses: Open Source Toolkit for Statistical Machine Translation · ACL 2007
Natural language and speech › Machine translation › computer-assisted translation
interactive machine translation
0.112008
Improving Interactive Machine Translation via Mouse Actions · EMNLP 2008

Methods — techniques the papers use, named apart from their topics

language-specific adapter modules · 0.9embedding linear combination · 0.9adaptive rejection preference optimization · 0.9adapter modules · 0.9web crawling · 0.9automatic alignment · 0.9statistical machine translation · 0.1
YearPublicationVenuePosition
2025 Adapters for Altering LLM Vocabularies: What Languages Benefit the Most?
abstract
Vocabulary adaptation, which integrates new vocabulary into pre-trained language models, enables expansion to new languages and mitigates token over-fragmentation. However, existing approaches are limited by their reliance on heuristics or external embeddings. We propose VocADT, a novel method for vocabulary adaptation using adapter modules that are trained to learn the optimal linear combination of existing embeddings while keeping the model’s weights fixed. VocADT offers a flexible and scalable solution without depending on external resources or language constraints. Across 11 languages—with diverse scripts, resource availability, and fragmentation—we demonstrate that VocADT outperforms the original Mistral model (Jiang et al., 2023) and other baselines across various multilingual tasks including natural language understanding and machine translation. We find that Latin-script languages and highly fragmented languages benefit the most from vocabulary adaptation. We further fine-tune the adapted model on the generative task of machine translation and find that vocabulary adaptation is still beneficial after fine-tuning and that VocADT is the most effective.
HyoJung Han 0001, Akiko Eriguchi, Hieu Hoang, Marine Carpuat, Huda Khayrallah
ICLR4
2025 X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at Scale
abstract
Large language models (LLMs) have achieved remarkable success across various NLP tasks with a focus on English due to English-centric pre-training and limited multilingual data. In this work, we focus on the problem of translation, and while some multilingual LLMs claim to support for hundreds of languages, models often fail to provide high-quality responses for mid- and low-resource languages, leading to imbalanced performance heavily skewed in favor of high-resource languages. We introduce **X-ALMA**, a model designed to ensure top-tier performance across 50 diverse languages, regardless of their resource levels. X-ALMA surpasses state-of-the-art open-source multilingual LLMs, such as Aya-101 and Aya-23, in every single translation direction on the FLORES-200 and WMT'23 test datasets according to COMET-22. This is achieved by plug-and-play language-specific module architecture to prevent language conflicts during training and a carefully designed training regimen with novel optimization methods to maximize the translation performance. After the final stage of training regimen, our proposed **A**daptive **R**ejection **P**reference **O**ptimization (**ARPO**) surpasses existing preference optimization methods in translation tasks.
Kenton Murray, Philipp Koehn, Hieu Hoang, Akiko Eriguchi, Huda Khayrallah
ICLR4
2023 A Design Method for an Intelligent Tutoring System with Algorithms Visualization
Hien D. Nguyen 0002, Hieu Hoang, Triet H. M. Nguyen, Khai Truong, Anh T. Huynh, Trong T. Le, Sang Vu
IEA/AIE (1)2
2020 ParaCrawl: Web-Scale Acquisition of Parallel Corpora
abstract
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, William Waites, Dion Wiggins, Jaume Zaragoza. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz-Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strong, Brian Thompson 0001, William Waites, Dion Wiggins, Jaume Zaragoza
ACL5
2019 ParaCrawl: Web-scale parallel corpora for the languages of the EU
Miquel Esplà-Gomis, Mikel L. Forcada, Gema Ramírez-Sánchez, Hieu Hoang
MTSummit (2)4
2014 Integrating an Unsupervised Transliteration Model into Statistical Machine Translation
abstract
We investigate three methods for integrating an unsupervised transliteration model into an end-to-end SMT system.We induce a transliteration model from parallel data and use it to translate OOV words.Our approach is fully unsupervised and language independent.In the methods to integrate transliterations, we observed improvements from 0.23-0.75(∆ 0.41) BLEU points across 7 language pairs.We also show that our mined transliteration corpora provide better rule coverage and translation quality compared to the gold standard transliteration corpora.
Nadir Durrani, Hassan Sajjad 0001, Hieu Hoang, Philipp Koehn
EACL3
2009 Improving Mid-Range Re-Ordering Using Templates of Factors
Hieu Hoang, Philipp Koehn
EACL1
2008 Improving Interactive Machine Translation via Mouse Actions
Germán Sanchis-Trilles, Daniel Ortiz-Martínez, Jorge Civera, Francisco Casacuberta, Enrique Vidal 0001, Hieu Hoang
EMNLP6
2007 Moses: Open Source Toolkit for Statistical Machine Translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondrej Bojar, Alexandra Constantin, Evan Herbst
ACL2
2007 Factored Translation Models
Philipp Koehn, Hieu Hoang
EMNLP-CoNLL2