Tu Anh Dinh

dblp:297/4574 · DBLP profile ↗
← Back
6ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0001-7651-820XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Machine translation · 65% Language models and text generation · 35%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation › machine translation evaluation
translation quality estimation
1.922026
Sigmoid Head for Quality Estimation under Language Ambiguity · ACL (1) 2026
Are Generative Models Underconfident? Better Quality Estimation with Boosted Model Probability · EMNLP 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.812024
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading · EMNLP 2024
Computing education
automated assessment
0.212024
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

human expert grading · 1.5automatic grading · 1.5sigmoid activation · 1.0negative sampling · 1.0probability calibration · 0.9ensemble-based quality estimation · 0.9
YearPublicationVenuePosition
2026 Sigmoid Head for Quality Estimation under Language Ambiguity
abstract
Language model (LM) probability is not a reliable quality estimator, as natural language is ambiguous.When multiple output options are valid, the model's probability distribution is spread across them, which can misleadingly indicate low output quality.This issue is caused by two reasons: (1) LMs' final output activation is softmax, which does not allow multiple correct options to receive high probabilities simultaneuously and (2) LMs' training data is single, one-hot encoded references, indicating that there is only one correct option at each output step.We propose training a module for Quality Estimation on top of pre-trained LMs to address these limitations.The module, called Sigmoid Head, is an extra unembedding head with sigmoid activation to tackle the first limitation.To tackle the second limitation, during the negative sampling process to train the Sigmoid Head, we use a heuristic to avoid selecting potentially alternative correct tokens.Our Sigmoid Head is computationally efficient during training and inference.The probability from Sigmoid Head is notably better quality signal compared to the original softmax head.As the Sigmoid Head does not rely on humanannotated quality data, it is more robust to outof-domain settings compared to supervised QE.
Tu Anh Dinh, Jan Niehues
ACL (1)1
2025 Are Generative Models Underconfident? Better Quality Estimation with Boosted Model Probability
abstract
Quality Estimation (QE) is estimating the quality of the model output during inference when the ground truth is not available.Deriving output quality from the models' output probability is the most trivial and low-effort way.However, we show that the output probability of text-generation models can appear underconfident.At each output step, there can be multiple correct options, making the probability distribution spread out more.Thus, lower probability does not necessarily mean lower output quality.Due to this observation, we propose a QE approach called BOOSTEDPROB 1 , which boosts the model's confidence in cases where there are multiple viable output options.With no increase in complexity, BOOSTEDPROB is notably better than raw model probability in different settings, achieving on average +0.194 improvement in Pearson correlation to groundtruth quality.It also comes close to or outperforms more costly approaches like supervised or ensemble-based QE in certain settings.
Tu Anh Dinh, Jan Niehues
EMNLP1
2024 Quality Estimation with k-nearest Neighbors and Automatic Evaluation for Model-specific Quality Estimation
abstract
Providing quality scores along with Machine Translation (MT) output, so-called reference-free Quality Estimation (QE), is crucial to inform users about the reliability of the translation. We propose a model-specific, unsupervised QE approach, termed kNN-QE, that extracts information from the MT model’s training data using k-nearest neighbors. Measuring the performance of model-specific QE is not straightforward, since they provide quality scores on their own MT output, thus cannot be evaluated using benchmark QE test sets containing human quality scores on premade MT output. Therefore, we propose an automatic evaluation method that uses quality scores from reference-based metrics as gold standard instead of human-generated ones. We are the first to conduct detailed analyses and conclude that this automatic method is sufficient, and the reference-based MetricX-23 is best for the task.
Tu Anh Dinh, Tobias Palzer, Jan Niehues
EAMT (1)1
2024 SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
abstract
Tu Anh Dinh, Carlos Mullov, Leonard Bärmann, Zhaolin Li, Danni Liu, Simon Reiß, Jueun Lee, Nathan Lerzer, Jianfeng Gao, Fabian Peller-Konrad, Tobias Röddiger, Alexander Waibel, Tamim Asfour, Michael Beigl, Rainer Stiefelhagen, Carsten Dachsbacher, Klemens Böhm, Jan Niehues. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Tu Anh Dinh, Carlos Mullov, Leonard Bärmann, Zhaolin Li, Simon Reiß, Jueun Lee, Nathan Lerzer, Jianfeng Gao 0002, Fabian Tërnava, Tobias Röddiger, Alex Waibel, Tamim Asfour, Michael Beigl, Rainer Stiefelhagen, Carsten Dachsbacher, Klemens Böhm, Jan Niehues
EMNLP1
2023 Perturbation-based QE: An Explainable, Unsupervised Word-level Quality Estimation Method for Blackbox Machine Translation
abstract
Quality Estimation (QE) is the task of predicting the quality of Machine Translation (MT) system output, without using any gold-standard translation references. State-of-the-art QE models are supervised: they require human-labeled quality of some MT system output on some datasets for training, making them domain-dependent and MT-system-dependent. There has been research on unsupervised QE, which requires glass-box access to the MT systems, or parallel MT data to generate synthetic errors for training QE models. In this paper, we present Perturbation-based QE - a word-level Quality Estimation approach that works simply by analyzing MT system output on perturbed input source sentences. Our approach is unsupervised, explainable, and can evaluate any type of blackbox MT systems, including the currently prominent large language models (LLMs) with opaque internal processes. For language directions with no labeled QE data, our approach has similar or better performance than the zero-shot supervised approach on the WMT21 shared task. Our approach is better at detecting gender bias and word-sense-disambiguation errors in translation than supervised QE, indicating its robustness to out-of-domain usage. The performance gap is larger when detecting errors on a nontraditional translation-prompting LLM, indicating that our approach is more generalizable to different MT systems. We give examples demonstrating our approach’s explainability power, where it shows which input source words have influence on a certain MT output word.
Tu Anh Dinh, Jan Niehues
MTSummit (1)1
2022 Tackling Data Scarcity in Speech Translation Using Zero-Shot Multilingual Machine Translation Techniques
abstract
Recently, end-to-end speech translation (ST) has gained significant attention as it avoids error propagation. However, the approach suffers from data scarcity. It heavily depends on direct ST data and is less efficient in making use of speech transcription and text translation data, which is often more easily available. In the related field of multilingual text translation, several techniques have been proposed for zero-shot translation. A main idea is to increase the similarity of semantically similar sentences in different languages. We investigate whether these ideas can be applied to speech translation, by building ST models trained on speech transcription and text translation data. We investigate the effects of data augmentation and auxiliary loss function. The techniques were successfully applied to few-shot ST using limited ST data, with improvements of up to +12.9 BLEU points compared to direct end-to-end ST and +3.1 BLEU points compared to ST models fine-tuned from ASR model.
Tu Anh Dinh, Jan Niehues
ICASSP1