EDBT 2026 Demo / reviewers in the wild / expert
Daniel Deutsch
dblp:222/9395
· DBLP profile ↗
19ranked-venue papers
10as first author
15since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 10 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine TranslationabstractParker Riley, Daniel Deutsch, Mara Finkelstein, Colten DiIanni, Juraj Juraska, Markus Freitag. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Parker Riley, Daniel Deutsch, Mara Finkelstein, Colten DiIanni, Juraj Juraska, Markus Freitag |
ACL (1) | 2 |
| 2025 | Enhancing Human Evaluation in Machine Translation with Comparative JudgementabstractHuman evaluation is crucial for assessing rapidly evolving language models but is influenced by annotator proficiency and task design. This study explores the integration of comparative judgment into human annotation for machine translation (MT) and evaluates three annotation setups—point-wise Multidimensional Quality Metrics (MQM), side-by-side (S×S) MQM, and its simplified version S×S relative ranking (RR). In MQM, annotators mark error spans with categories and severity levels. S×S MQM extends MQM to pairwise error annotation for two translations of the same input, while S×S RR focuses on selecting the better output without labeling errors.Key findings are: (1) the S×S settings achieve higher inter-annotator agreement than MQM; (2) S×S MQM enhances inter-translation error marking consistency compared to MQM by, on average, 38.5% for explicitly compared MT systems and 19.5% for others; (3) all annotation settings return stable system rankings, with S×S RR offering a more efficient alternative to (S×S) MQM; (4) the S×S settings highlight subtle errors overlooked in MQM without altering absolute system evaluations.To spur further research, we will release the triply annotated datasets comprising 377 ZhEn and 104 EnDe annotation examples, each covering 10 systems. Yixiao Song, Parker Riley, Daniel Deutsch, Markus Freitag |
ACL (1) | 3 |
| 2025 | Don't Sweat the Small Stuff: Segment-Level Meta-Evaluation Based on Pairwise Difference CorrelationabstractThis paper introduces Pairwise Difference Pearson (PDP), a novel segment-level metaevaluation metric for Machine Translation (MT) that address limitations in previous Pearson's ρ-based and and Kendall's τ -based metaevaluation approaches.PDP is a correlationbased metric that utilizes pairwise differences rather than raw scores.It draws on information from all segments for a more robust understanding of score distributions and uses segmentwise pairwise differences to refine Global Pearson to intra-segment score comparisons.Analysis on the WMT'24 shared task shows PDP properly ranks sentinel evaluation metrics and better aligns with human error weightings than previous work.Noise injection analysis demonstrates PDP's robustness to random noise, segment bias, and system bias while highlighting its sensitivity to extreme outliers. Colten DiIanni, Daniel Deutsch |
EMNLP | 2 |
| 2025 | SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?abstractSenyu Li, Jiayi Wang, Felermino D. M. A. Ali, Colin Cherry, Daniel Deutsch, Eleftheria Briakou, Rui Sousa-Silva, Henrique Lopes Cardoso, Pontus Stenetorp, David Ifeoluwa Adelani. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Senyu Li, Jiayi Wang 0010, Felermino D. M. A. Ali, Colin Cherry, Daniel Deutsch, Eleftheria Briakou, Rui Sousa-Silva, Henrique Lopes Cardoso, Pontus Stenetorp, David Ifeoluwa Adelani |
EMNLP | 5 |
| 2025 | From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test SetabstractAs LLMs continue to become more powerful and versatile, human evaluation has become intractable at scale and reliance on automatic metrics has become the norm. Recently, it has been shown that LLMs are themselves state-of-the-art evaluators for many tasks. These *Autoraters* are typically designed so that they generalize to new systems *and* test sets. In practice, however, evaluation is performed on a small set of fixed, canonical test sets, which are carefully curated to measure the capabilities of interest and are not changed frequently. In this work, we design a method which specializes a prompted Autorater to a given test set, by leveraging historical ratings on the test set to construct in-context learning (ICL) examples. We evaluate our *Specialist* method on the task of fine-grained machine translation evaluation, and show that it dramatically outperforms the state-of-the-art XCOMET metric by 54% and 119% on the WMT'23 and WMT'24 test sets, respectively. We perform extensive analyses to understand the representations learned by our Specialist metrics, and how variability in rater behavior affects their performance. We also verify the generalizability and robustness of our Specialist method across different numbers of ICL examples, LLM backbones, systems to evaluate, and evaluation tasks. Mara Finkelstein, Daniel Deutsch, Parker Riley, Juraj Juraska, Geza Kovacs, Markus Freitag |
ICML | 2 |
| 2025 | Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine TranslationabstractData contamination—the accidental consumption of evaluation examples within the pre-training data—can undermine the validity of evaluation benchmarks. In this paper, we present a rigorous analysis of the effects of contamination on language models at 1B and 8B scales on the machine translation task. Starting from a carefully decontaminated train-test split, we systematically introduce contamination at various stages, scales, and data formats to isolate its effect and measure its impact on performance metrics. Our experiments reveal that contamination with both source and target substantially inflates BLEU scores, and this inflation is 2.5 times larger (up to 30 BLEU points) for 8B compared to 1B models. In contrast, source-only and target-only contamination generally produce smaller, less consistent over-estimations. Finally, we study how the temporal distribution and frequency of contaminated samples influence performance over-estimation across languages with varying degrees of data resources. Yusuf Kocyigit, Eleftheria Briakou, Daniel Deutsch, Jiaming Luo, Colin Cherry, Markus Freitag |
ICML | 3 |
| 2024 | Finding Replicable Human Evaluations via Stable Ranking ProbabilityabstractParker Riley, Daniel Deutsch, George Foster, Viresh Ratnakar, Ali Dabirmoghaddam, Markus Freitag. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Parker Riley, Daniel Deutsch, George F. Foster, Viresh Ratnakar, Ali Dabirmoghaddam, Markus Freitag |
NAACL-HLT | 2 |
| 2023 | A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for SummarizationabstractLining Zhang, Simon Mille, Yufang Hou, Daniel Deutsch, Elizabeth Clark, Yixin Liu, Saad Mahamood, Sebastian Gehrmann, Miruna Clinciu, Khyathi Raghavi Chandu, João Sedoc. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Lining Zhang, Simon Mille, Yufang Hou 0001, Daniel Deutsch, Elizabeth Clark, Yixin Liu 0003, Saad Mahamood, Sebastian Gehrmann, Miruna-Adriana Clinciu, Khyathi Raghavi Chandu, João Sedoc |
ACL (1) | 4 |
| 2023 | Incorporating Question Answering-Based Signals into Abstractive Summarization via Salient Span SelectionabstractIn this work, we propose a method for incorporating question-answering (QA) signals into a summarization model.Our method identifies salient noun phrases (NPs) in the input document by automatically generating wh-questions that are answered by the NPs and automatically determining whether those questions are answered in the gold summaries.This QA-based signal is incorporated into a two-stage summarization model which first marks salient NPs in the input document using a classification model, then conditionally generates a summary.Our experiments demonstrate that the models trained using QA-based supervision generate higher-quality summaries than baseline methods of identifying salient spans on benchmark summarization datasets.Further, we show that the content of the generated summaries can be controlled based on which NPs are marked in the input document.Finally, we propose a method of augmenting the training data so the gold summaries are more consistent with the marked input spans used during training and show how this results in models which learn to better exclude unmarked document content. 1 Daniel Deutsch, Dan Roth 0001 |
EACL | 1 |
| 2023 | Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie CalibrationabstractKendall's τ is frequently used to meta-evaluate how well machine translation (MT) evaluation metrics score individual translations.Its focus on pairwise score comparisons is intuitive but raises the question of how ties should be handled, a gray area that has motivated different variants in the literature.We demonstrate that, in settings like modern MT meta-evaluation, existing variants have weaknesses arising from their handling of ties, and in some situations can even be gamed.We propose instead to meta-evaluate metrics with a version of pairwise accuracy that gives metrics credit for correctly predicting ties, in combination with a tie calibration procedure that automatically introduces ties into metric scores, enabling fair comparison between metrics that do and do not predict ties.We argue and provide experimental evidence that these modifications lead to fairer ranking-based assessments of metric performance.1 Daniel Deutsch, George F. Foster, Markus Freitag |
EMNLP | 1 |
| 2022 | On the Limitations of Reference-Free Evaluations of Generated TextabstractThere is significant interest in developing evaluation metrics which accurately estimate the quality of generated text without the aid of a human-written reference text, which can be time consuming and expensive to collect or entirely unavailable in online applications.However, in this work, we demonstrate that these reference-free metrics are inherently biased and limited in their ability to evaluate generated text, and we argue that they should not be used to measure progress on tasks like machine translation or summarization.We show how reference-free metrics are equivalent to using one generation model to evaluate another, which has several limitations: (1) the metrics can be optimized at test time to find the approximate best-possible output, (2) they are inherently biased toward models which are more similar to their own, and (3) they can be biased against higher-quality outputs, including those written by humans.Therefore, we recommend that reference-free metrics should be used as diagnostic tools for analyzing and understanding model behavior instead of measures of how well models perform a task, in which the goal is to achieve as high of a score as possible.1 Daniel Deutsch, Rotem Dror, Dan Roth 0001 |
EMNLP | 1 |
| 2022 | Re-Examining System-Level Correlations of Automatic Summarization Evaluation MetricsabstractHow reliably an automatic summarization evaluation metric replicates human judgments of summary quality is quantified by systemlevel correlations.We identify two ways in which the definition of the system-level correlation is inconsistent with how metrics are used to evaluate systems in practice and propose changes to rectify this disconnect.First, we calculate the system score for an automatic metric using the full test set instead of the subset of summaries judged by humans, which is currently standard practice.We demonstrate how this small change leads to more precise estimates of system-level correlations.Second, we propose to calculate correlations only on pairs of systems that are separated by small differences in automatic scores which are commonly observed in practice.This allows us to demonstrate that our best estimate of the correlation of ROUGE to human judgments is near 0 in realistic scenarios.The results from the analyses point to the need to collect more high-quality human judgments and to improve automatic metrics when differences in system scores are small.1 Daniel Deutsch, Rotem Dror, Dan Roth 0001 |
NAACL-HLT | 1 |
| 2021 | Understanding the Extent to which Content Quality Metrics Measure the Information Quality of SummariesabstractReference-based metrics such as ROUGE or BERTScore evaluate the content quality of a summary by comparing the summary to a reference.Ideally, this comparison should measure the summary's information quality by calculating how much information the summaries have in common.In this work, we analyze the token alignments used by ROUGE and BERTScore to compare summaries and argue that their scores largely cannot be interpreted as measuring information overlap.Rather, they are better estimates of the extent to which the summaries discuss the same topics.Further, we provide evidence that this result holds true for many other summarization evaluation metrics.The consequence of this result is that the most frequently used summarization evaluation metrics do not align with the community's research goal, to generate summaries with high-quality information.However, we conclude by demonstrating that a recently proposed metric, QAEval, which scores summaries using question-answering, appears to better capture information quality than current evaluations, highlighting a direction for future research. Daniel Deutsch, Dan Roth 0001 |
CoNLL | 1 |
| 2021 | Towards Question-Answering as an Automatic Metric for Evaluating the Content Quality of a SummaryabstractAbstract A desirable property of a reference-based evaluation metric that measures the content quality of a summary is that it should estimate how much information that summary has in common with a reference. Traditional text overlap based metrics such as ROUGE fail to achieve this because they are limited to matching tokens, either lexically or via embeddings. In this work, we propose a metric to evaluate the content quality of a summary using question-answering (QA). QA-based methods directly measure a summary’s information overlap with a reference, making them fundamentally different than text overlap metrics. We demonstrate the experimental benefits of QA-based metrics through an analysis of our proposed metric, QAEval. QAEval outperforms current state-of-the-art metrics on most evaluations using benchmark datasets, while being competitive on others due to limitations of state-of-the-art models. Through a careful analysis of each component of QAEval, we identify its performance bottlenecks and estimate that its potential upper-bound performance surpasses all other automatic metrics, approaching that of the gold-standard Pyramid Method.1 Daniel Deutsch, Tania Bedrax-Weiss, Dan Roth 0001 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | A Statistical Analysis of Summarization Evaluation Metrics Using Resampling MethodsabstractAbstract The quality of a summarization evaluation metric is quantified by calculating the correlation between its scores and human annotations across a large number of summaries. Currently, it is unclear how precise these correlation estimates are, nor whether differences between two metrics’ correlations reflect a true difference or if it is due to mere chance. In this work, we address these two problems by proposing methods for calculating confidence intervals and running hypothesis tests for correlations using two resampling methods, bootstrapping and permutation. After evaluating which of the proposed methods is most appropriate for summarization through two simulation experiments, we analyze the results of applying these methods to several different automatic evaluation metrics across three sets of human annotations. We find that the confidence intervals are rather wide, demonstrating high uncertainty in the reliability of automatic metrics. Further, although many metrics fail to show statistical improvements over ROUGE, two recent works, QAEval and BERTScore, do so in some evaluation settings.1 Daniel Deutsch, Rotem Dror, Dan Roth 0001 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | Is Killed More Significant than Fled? A Contextual Model for Salient Event DetectionabstractIdentifying the key events in a document is critical to holistically understanding its important information.Although measuring the salience of events is highly contextual, most previous work has used a limited representation of events that omits essential information.In this work, we propose a highly contextual model of event salience that uses a rich representation of events, incorporates document-level information and allows for interactions between latent event encodings.Our experimental results on an event salience dataset (Liu et al., 2018) demonstrate that our model improves over previous work by an absolute 2-4% on standard metrics, establishing a new state-of-the-art performance for the task.We also propose a new evaluation metric which addresses flaws in previous evaluation methodologies.Finally, we discuss the importance of salient event detection for the downstream task of summarization. 1 Disha Jindal, Daniel Deutsch, Dan Roth 0001 |
COLING | 2 |
| 2019 | A General-Purpose Algorithm for Constrained Sequential InferenceabstractInference in structured prediction involves finding the best output structure for an input, subject to certain constraints.Many current approaches use sequential inference, which constructs the output in a left-to-right manner.However, there is no general framework to specify constraints in these approaches.We present a principled approach for incorporating constraints into sequential inference algorithms.Our approach expresses constraints using an automaton, which is traversed in lockstep during inference, guiding the search to valid outputs.We show that automata can express commonly used constraints and are easily incorporated into sequential inference.When it is more natural to represent constraints as a set of automata, our algorithm uses an active set method for demonstrably fast and efficient inference.We experimentally show the benefits of our algorithm on constituency parsing and semantic role labeling.For parsing, unlike unconstrained approaches, our algorithm always generates valid output, incurring only a small drop in performance.For semantic role labeling, imposing constraints using our algorithm corrects common errors, improving F 1 by 1.5 points.These benefits increase in low-resource settings.Our active set method achieves a 5.2x relative speedup over a naive approach.1 Daniel Deutsch, Shyam Upadhyay, Dan Roth 0001 |
CoNLL | 1 |
| 2019 | Summary Cloze: A New Task for Content Selection in Topic-Focused SummarizationabstractDaniel Deutsch, Dan Roth. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Daniel Deutsch, Dan Roth 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2018 | A Distributional and Orthographic Aggregation Model for English Derivational MorphologyabstractModeling derivational morphology to generate words with particular semantics is useful in many text generation tasks, such as machine translation or abstractive question answering.In this work, we tackle the task of derived word generation.That is, given the word "run," we attempt to generate the word "runner" for "someone who runs."We identify two key problems in generating derived words from root words and transformations: suffix ambiguity and orthographic irregularity.We contribute a novel aggregation model of derived word generation that learns derivational transformations both as orthographic functions using sequence-to-sequence models and as functions in distributional word embedding space.Our best open-vocabulary model, which can generate novel words, and our best closed-vocabulary model, show 22% and 37% relative error reductions over current state-of-the-art systems on the same dataset. Daniel Deutsch, John Hewitt, Dan Roth 0001 |
ACL (1) | 1 |