Steffen Eger

dblp:69/9271 · DBLP profile ↗
← Back
52ranked-venue papers
11as first author
29since 2021 · last 2026
0000-0003-4663-8336ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 50 · 10 first-author · 29 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 CROC: Evaluating and Training T2I Metrics with Pseudo- and Human-Labeled Contrastive Robustness Checks
abstract
Abstract The assessment of evaluation metrics (meta-evaluation) is crucial for determining the suitability of existing metrics in text-to-image (T2I) generation tasks. Human-based meta-evaluation is costly and time-intensive, and automated alternatives are scarce. We address this gap and propose CROC: a scalable framework for automated Contrastive Robustness Checks that systematically probes and quantifies metric robustness by synthesizing contrastive test cases across a comprehensive taxonomy of image properties. With CROC, we generate a pseudo-labeled dataset (CROCsyn) of over 1 million contrastive prompt–image pairs to enable a fine-grained comparison of evaluation metrics. We also use this dataset to train CROCScore, a new metric that achieves state-of-the-art performance among open-source methods, demonstrating an additional key application of our framework. To complement this dataset, we introduce a human-supervised benchmark (CROChum) targeting especially challenging categories. Our results highlight robustness issues in existing metrics: for example, many fail on prompts involving negation, and all tested open-source metrics fail on at least 24% of cases involving correct identification of body parts.1
Christoph Leiter, Yuki Markus Asano, Margret Keuper, Steffen Eger
Trans. Assoc. Comput. Linguistics4
2025 Argument Summarization and its Evaluation in the Era of Large Language Models
abstract
Moritz Altemeyer, Steffen Eger, Johannes Daxenberger, Yanran Chen, Tim Altendorf, Philipp Cimiano, Benjamin Schiller. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Moritz Altemeyer, Steffen Eger, Johannes Daxenberger, Yanran Chen, Tim Altendorf, Philipp Cimiano, Benjamin Schiller
EMNLP2
2025 Graph-Guided Textual Explanation Generation Framework
abstract
Natural language explanations (NLEs) are commonly used to provide plausible free-text explanations of a model's reasoning about its predictions.However, recent work has questioned their faithfulness, as they may not accurately reflect the model's internal reasoning process regarding its predicted answer.In contrast, highlight explanations-input fragments critical for the model's predicted answers-exhibit measurable faithfulness.Building on this foundation, we propose G-TEx, a Graph-Guided Textual Explanation Generation framework designed to enhance the faithfulness of NLEs.Specifically, highlight explanations are first extracted as faithful cues reflecting the model's reasoning logic toward answer prediction.They are subsequently encoded through a graph neural network layer to guide the NLE generation, which aligns the generated explanations with the model's underlying reasoning toward the predicted answer.Experiments on both encoder-decoder and decoder-only models across three reasoning datasets demonstrate that G-TEx improves NLE faithfulness by up to 12.18% compared to baseline methods.Additionally, G-TEx generates NLEs with greater semantic and lexical similarity to human-written ones.Human evaluations show that G-TEx can decrease redundant content and enhance the overall quality of NLEs.Our work presents a novel method for explicitly guiding NLE generation to enhance faithfulness, serving as a foundation for addressing broader criteria in NLE and generated text.
Shuzhou Yuan, Ran Zhang 0013, Michael Färber 0001, Steffen Eger, Pepa Atanasova, Isabelle Augenstein
EMNLP5
2025 LiTransProQA: An LLM-based Literary Translation Evaluation Metric with Professional Question Answering
abstract
The impact of Large Language Models (LLMs) has extended into literary domains.However, existing evaluation metrics for literature prioritize mechanical accuracy over artistic expression and tend to overrate machine translation as being superior to human translation from experienced professionals.In the long run, this bias could result in an irreversible decline in translation quality and cultural authenticity.In response to the urgent need for a specialized literary evaluation metric, we introduce LITRANSPROQA, a novel, referencefree, LLM-based question-answering framework designed for literary translation evaluation.LITRANSPROQA integrates humans in the loop to incorporate insights from professional literary translators and researchers, focusing on critical elements in literary quality assessment such as literary devices, cultural understanding, and authorial voice.Our extensive evaluation shows that while literaryfinetuned XCOMET-XL yields marginal gains, LITRANSPROQA substantially outperforms current metrics, achieving up to 0.07 gain in correlation and surpassing the best state-of-theart metrics by over 15 points in adequacy assessments.Incorporating professional translator insights as weights further improves performance, highlighting the value of translator inputs.Notably, LITRANSPROQA reaches an adequacy performance comparable to trained linguistic student evaluators, though it still falls behind experienced professional translators.LITRANSPROQA shows broad applicability to open-source models like LLaMA3.3-70b and Qwen2.5-32b,indicating its potential as an accessible and training-free tool for evaluating literary translations that require local processing due to copyright or ethical considerations.
Ran Zhang 0013, Lieve Macken, Steffen Eger
EMNLP4
2025 Tikzero: Zero-Shot Text-Guided Graphics Program Synthesis
Jonas Belouadi, Eddy Ilg, Margret Keuper, Hideki Tanaka, Masao Utiyama, Raj Dabre, Steffen Eger, Simone Paolo Ponzetto
ICCV7
2025 ScImage: How good are multimodal large language models at scientific text-to-image generation?
abstract
Multimodal large language models (LLMs) have demonstrated impressive capabilities in generating high-quality images from textual instructions. However, their performance in generating scientific images—a critical application for accelerating scientific progress—remains underexplored. In this work, we address this gap by introducing ScImage, a benchmark designed to evaluate the multimodal capabilities of LLMs in generating scientific images from textual descriptions. ScImage assesses three key dimensions of understanding: spatial, numeric, and attribute comprehension, as well as their combinations, focusing on the relationships between scientific objects (e.g., squares, circles). We evaluate seven models, GPT-4o, Llama, AutomaTikZ, Dall-E, StableDiffusion, GPT-o1 and Qwen2.5-Coder-Instruct using two modes of output generation: code-based outputs (Python, TikZ) and direct raster image generation. Additionally, we examine four different input languages: English, German, Farsi, and Chinese. Our evaluation, conducted with 11 scientists across three criteria (correctness, relevance, and scientific accuracy), reveals that while GPT4-o produces outputs of decent quality for simpler prompts involving individual dimensions such as spatial, numeric, or attribute understanding in isolation, all models face challenges in this task, especially for more complex prompts. ScImage is available: huggingface.co/datasets/casszhao/ScImage
Leixin Zhang 0001, Steffen Eger, Yinjie Cheng, Weihe Zhai, Jonas Belouadi, Fahimeh Moafian, Zhixue Zhao
ICLR2
2025 PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics
abstract
Daniil Larionov, Steffen Eger. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Daniil Larionov, Steffen Eger
NAACL (Long Papers)2
2025 How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
abstract
Ran Zhang, Wei Zhao, Steffen Eger. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Ran Zhang 0013, Steffen Eger
NAACL (Long Papers)3
2024 Dependencies over Times and Tools (DoTT)
abstract
Purpose: Based on the examples of English and German, we investigate to what extent parsers trained on modern variants of these languages can be transferred to older language levels without loss. Methods: We developed a treebank called DoTT (https://github.com/texttechnologylab/DoTT) which covers, roughly, the time period from 1800 until today, in conjunction with the further development of the annotation tool DependencyAnnotator. DoTT consists of a collection of diachronic corpora enriched with dependency annotations using 3 parsers, 6 pre-trained language models, 5 newly trained models for German, and two tag sets (TIGER and Universal Dependencies). To assess how the different parsers perform on texts from different time periods, we created a gold standard sample as a benchmark. Results: We found that the parsers/models perform quite well on modern texts (document-level LAS ranging from 82.89 to 88.54) and slightly worse on older texts, as expected (average document-level LAS 84.60 vs. 86.14), but not significantly. For German texts, the (German) TIGER scheme achieved slightly better results than UD. Conclusion: Overall, this result speaks for the transferability of parsers to past language levels, at least dating back until around 1800. This very transferability, it is however argued, means that studies of language change in the field of dependency syntax can draw on dependency distance but miss out on some grammatical phenomena.
Andy Lücking, Giuseppe Abrami, Leon Hammerla, Marc Rahn, Daniel Baumartz, Steffen Eger, Alexander Mehler
LREC/COLING6
2024 Evaluating Diversity in Automatic Poetry Generation
abstract
Natural Language Generation (NLG), and more generally generative AI, are among the currently most impactful research fields.Creative NLG, such as automatic poetry generation, is a fascinating niche in this area.While most previous research has focused on forms of the Turing test when evaluating automatic poetry generation -can humans distinguish between automatic and human generated poetry -we evaluate the diversity of automatically generated poetry (with a focus on quatrains), by comparing distributions of generated poetry to distributions of human poetry along structural, lexical, semantic and stylistic dimensions, assessing different model types (word vs. character-level, general purpose LLMs vs. poetry-specific models), including the very recent LLaMA3-8B, and types of fine-tuning (conditioned vs. unconditioned).We find that current automatic poetry systems are considerably underdiverse along multiple dimensions -they often do not rhyme sufficiently, are semantically too uniform and even do not match the length distribution of human poetry.Our experiments reveal, however, that style-conditioning and character-level modeling clearly increases diversity across virtually all dimensions we explore.Our identified limitations may serve as the basis for more genuinely diverse future poetry generation models. 1 L 0.62 12 34 20.18 20 2.84 de LLaMA3 con 0.76 10 47 21.69 21 4.14 en HUMAN 1.00 4 67 28.06 28 6.26 en DeepSpeare 0.57 15 33 23.85 24 2.85 en SA 0.92 12 52 27.36 27 5.38 en ByGPT5 S 0.80 12 44 25.30 25 5.09 en ByGPT5 L 0.77 11 47 24.97 25 4.87 en GPT2 S 0.69 13 55 24.11 24 4.48 en GPT2 L 0.72 13 56 24.74 24 4.94 en GPTNeo S 0.55 11 55 22.67 22 3.89 en GPTNeo L 0.48 13 34 21.93 22 3.16 en LLaMA2 S 0.87 15 75 28.60 27 7.52 en LLaMA2 L 0.67 12 54 23.95 24 4.50 en LLaMA3 0.59 14 60 23.20 23 4.23 en ByGPT5 con S 0.85 13 42 26.21 26 4.96 en ByGPT5 con L 0.84 14 42 25.85 25 4.84 en GPT2 con S 0.86 17 61 28.37 27 6.18 en GPT2 con L 0.83 16 70 27.8227 6.15 en GPTNeo con S 0.74 16 49 25.13 24 4.47 en GPTNeo con L 0.53 12 35 22.26 22 3.36 en LLaMA2 con S 0.70 17 74 33.55 32 7.83 en LLaMA2 con L 0.81 15 56 26.92 26 5.80 en LLaMA3 con 0.78 16 65 27.12 26 5.35
Yanran Chen, Hannes Gröner, Sina Zarrieß, Steffen Eger
EMNLP4
2024 Fine-Grained Detection of Solidarity for Women and Migrants in 155 Years of German Parliamentary Debates
abstract
Solidarity is a crucial concept to understand social relations in societies.In this paper, we explore fine-grained solidarity frames to study solidarity towards women and migrants in German parliamentary debates between 1867 and 2022.Using 2,864 manually annotated text snippets (with a cost exceeding 18k Euro), we evaluate large language models (LLMs) like Llama 3, GPT-3.5, and GPT-4.We find that GPT-4 outperforms other LLMs, approaching human annotation quality.Using GPT-4, we automatically annotate more than 18k further instances (with a cost of around 500 Euro) across 155 years and find that solidarity with migrants outweighs anti-solidarity but that frequencies and solidarity types shift over time.Most importantly, group-based notions of (anti-)solidarity fade in favor of compassionate solidarity, focusing on the vulnerability of migrant groups, and exchange-based anti-solidarity, focusing on the lack of (economic) contribution.Our study highlights the interplay of historical events, socio-economic needs, and political ideologies in shaping migration discourse and social cohesion.We also show that powerful LLMs, if carefully prompted, can be costeffective alternatives to human annotation for hard social scientific tasks.
Aida Kostikova, Dominik Beese, Benjamin Paaßen, Ole Pütz, Gregor Wiedemann, Steffen Eger
EMNLP6
2024 xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation Metrics
abstract
State-of-the-art trainable machine translation evaluation metrics like xCOMET achieve high correlation with human judgment but rely on large encoders (up to 10.7B parameters), making them computationally expensive and inaccessible to researchers with limited resources.To address this issue, we investigate whether the knowledge stored in these large encoders can be compressed while maintaining quality.We employ distillation, quantization, and pruning techniques to create efficient xCOMET alternatives and introduce a novel data collection pipeline for efficient black-box distillation.Our experiments show that, using quantization, xCOMET can be compressed up to three times with no quality degradation.Additionally, through distillation, we create an 278M-sized xCOMET-lite metric, which has only 2.6% of xCOMET-XXL parameters, but retains 92.1% of its quality.Besides, it surpasses strong smallscale metrics like COMET-22 and BLEURT-20 on the WMT22 metrics challenge dataset by 6.4%, despite using 50% fewer parameters.All code, dataset, and models are available online.
Daniil Larionov, Mikhail Seleznyov, Vasiliy Viskov, Alexander Panchenko, Steffen Eger
EMNLP5
2024 PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation
abstract
Large language models (LLMS) have revolutionized NLP research.Notably, in-context learning enables their use as evaluation metrics for natural language generation, making them particularly advantageous in low-resource scenarios and time-restricted applications.In this work, we introduce PrExMe, a large-scale Prompt Exploration for Metrics, where we evaluate more than 720 prompt templates for open-source LLM-based metrics on machine translation (MT) and summarization datasets, totalling over 6.6M evaluations.This extensive comparison (1) benchmarks recent open-source LLMS as metrics and (2) explores the stability and variability of different prompting strategies.We discover that, on the one hand, there are scenarios for which prompts are stable.For instance, some LLMS show idiosyncratic preferences and favor to grade generated texts with textual labels while others prefer to return numeric scores.On the other hand, the stability of prompts and model rankings can be susceptible to seemingly innocuous changes.For example, changing the requested output format from "0 to 100" to "-1 to +1" can strongly affect the rankings in our evaluation.Our study contributes to understanding the impact of different prompting approaches on LLM-based metrics for MT and summarization evaluation, highlighting the most stable prompting patterns and potential limitations. 1
Christoph Leiter, Steffen Eger
EMNLP2
2024 AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ
abstract
Generating bitmap graphics from text has gained considerable attention, yet for scientific figures, vector graphics are often preferred. Given that vector graphics are typically encoded using low-level graphics primitives, generating them directly is difficult. To address this, we propose the use of TikZ, a well-known abstract graphics language that can be compiled to vector graphics, as an intermediate representation of scientific figures. TikZ offers human-oriented, high-level commands, thereby facilitating conditional language modeling with any large language model. To this end, we introduce DaTikZ the first large-scale TikZ dataset, consisting of 120k TikZ drawings aligned with captions. We fine-tune LLaMA on DaTikZ, as well as our new model CLiMA, which augments LLaMA with multimodal CLIP embeddings. In both human and automatic evaluation, CLiMA and LLaMA outperform commercial GPT-4 and Claude 2 in terms of similarity to human-created figures, with CLiMA additionally improving text-image alignment. Our detailed analysis shows that all models generalize well and are not susceptible to memorization. GPT-4 and Claude 2, however, tend to generate more simplistic figures compared to both humans and our models. We make our framework, AutomaTikZ, along with model weights and datasets, publicly available.
Jonas Belouadi, Anne Lauscher, Steffen Eger
ICLR3
2024 DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ
abstract
Creating high-quality scientific figures can be time-consuming and challenging, even though sketching ideas on paper is relatively easy. Furthermore, recreating existing figures that are not stored in formats preserving semantic information is equally complex. To tackle this problem, we introduce DeTikZify, a novel multimodal language model that automatically synthesizes scientific figures as semantics-preserving TikZ graphics programs based on sketches and existing figures. To achieve this, we create three new datasets: DaTikZv2, the largest TikZ dataset to date, containing over 360k human-created TikZ graphics; SketchFig, a dataset that pairs hand-drawn sketches with their corresponding scientific figures; and MetaFig, a collection of diverse scientific figures and associated metadata. We train DeTikZify on MetaFig and DaTikZv2, along with synthetically generated sketches learned from SketchFig. We also introduce an MCTS-based inference algorithm that enables DeTikZify to iteratively refine its outputs without the need for additional training. Through both automatic and human evaluation, we demonstrate that DeTikZify outperforms commercial Claude 3 and GPT-4V in synthesizing TikZ programs, with the MCTS algorithm effectively boosting its performance. We make our code, models, and datasets publicly available.
Jonas Belouadi, Simone Paolo Ponzetto, Steffen Eger
NeurIPS3
2024 Cross-lingual Cross-temporal Summarization: Dataset, Models, Evaluation
abstract
Abstract While summarization has been extensively researched in natural language processing (NLP), cross-lingual cross-temporal summarization (CLCTS) is a largely unexplored area that has the potential to improve cross-cultural accessibility and understanding. This article comprehensively addresses the CLCTS task, including dataset creation, modeling, and evaluation. We (1) build the first CLCTS corpus with 328 instances for hDe-En (extended version with 455 instances) and 289 for hEn-De (extended version with 501 instances), leveraging historical fiction texts and Wikipedia summaries in English and German; (2) examine the effectiveness of popular transformer end-to-end models with different intermediate fine-tuning tasks; (3) explore the potential of GPT-3.5 as a summarizer; and (4) report evaluations from humans, GPT-4, and several recent automatic evaluation metrics. Our results indicate that intermediate task fine-tuned end-to-end models generate bad to moderate quality summaries while GPT-3.5, as a zero-shot summarizer, provides moderate to good quality outputs. GPT-3.5 also seems very adept at normalizing historical text. To assess data contamination in GPT-3.5, we design an adversarial attack scheme in which we find that GPT-3.5 performs slightly worse for unseen source documents compared to seen documents. Moreover, it sometimes hallucinates when the source sentences are inverted against its prior knowledge with a summarization accuracy of 0.67 for plot omission, 0.71 for entity swap, and 0.53 for plot negation. Overall, our regression results of model performances suggest that longer, older, and more complex source texts (all of which are more characteristic for historical language variants) are harder to summarize for all models, indicating the difficulty of the CLCTS task. Regarding evaluation, we observe that both the GPT-4 and BERTScore correlate moderately with human evaluations, implicating great potential for future improvement.
Ran Zhang 0013, Jihed Ouni, Steffen Eger
Comput. Linguistics3
2024 Towards Explainable Evaluation Metrics for Machine Translation
abstract
Unlike classical lexical overlap metrics such as BLEU, most current evaluation metrics for machine translation (for example, COMET or BERTScore) are based on black-box large language models. They often achieve strong correlations with human judgments, but recent research indicates that the lower-quality classical metrics remain dominant, one of the potential reasons being that their decision processes are more transparent. To foster more widespread acceptance of novel high-quality metrics, explainability thus becomes crucial. In this concept paper, we identify key properties as well as key goals of explainable machine translation metrics and provide a comprehensive synthesis of recent techniques, relating them to our established goals and properties. In this context, we also discuss the latest state-of-the-art approaches to explainable metrics based on generative models such as ChatGPT and GPT4. Finally, we contribute a vision of next-generation approaches, including natural language explanations. We hope that our work can help catalyze and guide future research on explainable evaluation metrics and, mediately, also contribute to better and more transparent machine translation systems.
Christoph Leiter, Piyawat Lertvittayakumjorn, Marina Fomicheva, Wei Zhao 0033, Yang Gao 0021, Steffen Eger
J. Mach. Learn. Res.6
2023 ByGPT5: End-to-End Style-conditioned Poetry Generation with Token-free Language Models
abstract
State-of-the-art poetry generation systems are often complex.They either consist of taskspecific model pipelines, incorporate prior knowledge in the form of manually created constraints, or both.In contrast, end-to-end models would not suffer from the overhead of having to model prior knowledge and could learn the nuances of poetry from data alone, reducing the degree of human supervision required.In this work, we investigate end-to-end poetry generation conditioned on styles such as rhyme, meter, and alliteration.We identify and address lack of training data and mismatching tokenization algorithms as possible limitations of past attempts.In particular, we successfully pre-train ByGPT5, a new token-free decoder-only language model, and fine-tune it on a large custom corpus of English and German quatrains annotated with our styles.We show that ByGPT5 outperforms other models such as mT5, ByT5, GPT-2 and ChatGPT, while also being more parameter efficient and performing favorably compared to humans.In addition, we analyze its runtime performance and demonstrate that it is not prone to memorization.We make our code, models, and datasets publicly available.1
Jonas Belouadi, Steffen Eger
ACL (1)2
2023 UScore: An Effective Approach to Fully Unsupervised Evaluation Metrics for Machine Translation
abstract
The vast majority of evaluation metrics for machine translation are supervised, i.e., (i) are trained on human scores, (ii) assume the existence of reference translations, or (iii) leverage parallel data.This hinders their applicability to cases where such supervision signals are not available.In this work, we develop fully unsupervised evaluation metrics.To do so, we leverage similarities and synergies between evaluation metric induction, parallel corpus mining, and MT systems.In particular, we use an unsupervised evaluation metric to mine pseudo-parallel data, which we use to remap deficient underlying vector spaces (iteratively) and to induce an unsupervised MT system, which then provides pseudo-references as an additional component in the metric.Finally, we also induce unsupervised multilingual sentence embeddings from pseudo-parallel data.We show that our fully unsupervised metrics are effective, i.e., they beat supervised competitors on four out of five evaluation datasets.We make our code publicly available.1
Jonas Belouadi, Steffen Eger
EACL2
2023 DiscoScore: Evaluating Text Generation with BERT and Discourse Coherence
abstract
Recently, there has been a growing interest in designing text generation systems from a discourse coherence perspective, e.g., modeling the interdependence between sentences.Still, recent BERT-based evaluation metrics are weak in recognizing coherence, and thus are not reliable in a way to spot the discourselevel improvements of those text generation systems.In this work, we introduce DiscoScore, a parametrized discourse metric, which uses BERT to model discourse coherence from different perspectives, driven by Centering theory.Our experiments encompass 16 non-discourse and discourse metrics, including DiscoScore and popular coherence models, evaluated on summarization and document-level machine translation (MT).We find that (i) the majority of BERT-based metrics correlate much worse with human rated coherence than early discourse metrics, invented a decade ago; (ii) the recent state-of-the-art BARTScore is weak when operated at system level-which is particularly problematic as systems are typically compared in this manner.DiscoScore, in contrast, achieves strong system-level correlation with human ratings, not only in coherence but also in factual consistency and other aspects, and surpasses BARTScore by over 10 correlation points on average.Further, aiming to understand DiscoScore, we provide justifications to the importance of discourse coherence for evaluation metrics, and explain the superiority of one variant over another.Our code is available at https://github.com/AIPHES/ DiscoScore.
Wei Zhao 0033, Michael Strube 0001, Steffen Eger
EACL3
2023 MENLI: Robust Evaluation Metrics from Natural Language Inference
abstract
Abstract Recently proposed BERT-based evaluation metrics for text generation perform well on standard benchmarks but are vulnerable to adversarial attacks, e.g., relating to information correctness. We argue that this stems (in part) from the fact that they are models of semantic similarity. In contrast, we develop evaluation metrics based on Natural Language Inference (NLI), which we deem a more appropriate modeling. We design a preference-based adversarial attack framework and show that our NLI based metrics are much more robust to the attacks than the recent BERT-based metrics. On standard benchmarks, our NLI based metrics outperform existing summarization metrics, but perform below SOTA MT metrics. However, when combining existing metrics with our NLI metrics, we obtain both higher adversarial robustness (15%–30%) and higher quality metrics as measured on standard benchmarks (+5% to 30%).
Yanran Chen, Steffen Eger
Trans. Assoc. Comput. Linguistics2
2022 Constrained Density Matching and Modeling for Cross-lingual Alignment of Contextualized Representations
Wei Zhao 0033, Steffen Eger
ACML2
2022 Layer or Representation Space: What Makes BERT-based Evaluation Metrics Robust?
abstract
The evaluation of recent embedding-based evaluation metrics for text generation is primarily based on measuring their correlation with human evaluations on standard benchmarks. However, these benchmarks are mostly from similar domains to those used for pretraining word embeddings. This raises concerns about the (lack of) generalization of embedding-based metrics to new and noisy domains that contain a different vocabulary than the pretraining data. In this paper, we examine the robustness of BERTScore, one of the most popular embedding-based metrics for text generation. We show that (a) an embedding-based metric that has the highest correlation with human evaluations on a standard benchmark can have the lowest correlation if the amount of input noise or unknown tokens increases, (b) taking embeddings from the first layer of pretrained models improves the robustness of all metrics, and (c) the highest robustness is achieved when using character-level embeddings, instead of token-based embeddings, from the first layer of the pretrained model.
Doan Nam Long Vu, Nafise Sadat Moosavi, Steffen Eger
COLING3
2022 Reproducibility Issues for BERT-based Evaluation Metrics
abstract
Reproducibility is of utmost concern in machine learning and natural language processing (NLP).In the field of natural language generation (especially machine translation), the seminal paper of Post (2018) has pointed out problems of reproducibility of the dominant metric, BLEU, at the time of publication.Nowadays, BERT-based evaluation metrics considerably outperform BLEU.In this paper, we ask whether results and claims from four recent BERT-based metrics can be reproduced.We find that reproduction of claims and results often fails because of (i) heavy undocumented preprocessing involved in the metrics, (ii) missing code and (iii) reporting weaker results for the baseline metrics.(iv) In one case, the problem stems from correlating not to human scores but to a wrong column in the csv file, inflating scores by 5 points.Motivated by the impact of preprocessing, we then conduct a second study where we examine its effects more closely (for one of the metrics).We find that preprocessing can have large effects, especially for highly inflectional languages.In this case, the effect of preprocessing may be larger than the effect of the aggregation mechanism (e.g., greedy alignment vs. Word Mover Distance).
Yanran Chen, Jonas Belouadi, Steffen Eger
EMNLP3
2021 Changes in European Solidarity Before and During COVID-19: Evidence from a Large Crowd- and Expert-Annotated Twitter Dataset
abstract
Alexandra Ils, Dan Liu, Daniela Grunow, Steffen Eger. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Alexandra Ils, Daniela Grunow, Steffen Eger
ACL/IJCNLP (1)4
2021 Better than Average: Paired Evaluation of NLP systems
abstract
Maxime Peyrard, Wei Zhao, Steffen Eger, Robert West. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Maxime Peyrard, Wei Zhao 0033, Steffen Eger, Robert West 0001
ACL/IJCNLP (1)3
2021 Global Explainability of BERT-Based Evaluation Metrics by Disentangling along Linguistic Factors
abstract
Evaluation metrics are a key ingredient for progress of text generation systems.In recent years, several BERT-based evaluation metrics have been proposed (including BERTScore, MoverScore, BLEURT, etc.) which correlate much better with human assessment of text generation quality than BLEU or ROUGE, invented two decades ago.However, little is known what these metrics, which are based on black-box language model representations, actually capture (it is typically assumed they model semantic similarity).In this work, we use a simple regression based global explainability technique to disentangle metric scores along linguistic factors, including semantics, syntax, morphology, and lexical overlap.We show that the different metrics capture all aspects to some degree, but that they are all substantially sensitive to lexical overlap, just like BLEU and ROUGE.This exposes limitations of these novelly proposed metrics, which we also highlight in an adversarial test scenario.
Marvin Kaster, Wei Zhao 0033, Steffen Eger
EMNLP (1)3
2021 TUDA-Reproducibility @ ReproGen: Replicability of Human Evaluation of Text-to-Text and Concept-to-Text Generation
abstract
This paper describes our contribution to the Shared Task ReproGen by Belz et al. (2021), which investigates the reproducibility of human evaluations in the context of Natural Language Generation.We selected the paper "Generation of Company descriptions using concept-to-text and text-to-text deep models: data set collection and systems evaluation" (Qader et al., 2018) and aimed to replicate, as closely to the original as possible, the human evaluation and the subsequent comparison between the human judgements and the automatic evaluation metrics.Here, we first outline the text generation task of the paper of Qader et al. (2018).Then, we document how we approached our replication of the paper's human evaluation.We also discuss the difficulties we encountered and which information was missing.Our replication has medium to strong correlation (0.66 Spearman overall) with the original results of Qader et al. (2018), but due to the missing information about how Qader et al. (2018) compared the human judgements with the metric scores, we have refrained from reproducing this comparison.
Yanran Chen, Steffen Eger
INLG3
2021 Graph routing between capsules
Yang Li 0055, Wei Zhao 0033, Erik Cambria, Suhang Wang, Steffen Eger
Neural Networks5
2020 SUPERT: Towards New Frontiers in Unsupervised Evaluation Metrics for Multi-Document Summarization
abstract
We study unsupervised multi-document summarization evaluation metrics, which require neither human-written reference summaries nor human annotations (e.g.preferences, ratings, etc.).We propose SUPERT, which rates the quality of a summary by measuring its semantic similarity with a pseudo reference summary, i.e. selected salient sentences from the source documents, using contextualized embeddings and soft token alignment techniques.Compared to the state-of-theart unsupervised evaluation metrics, SUPERT correlates better with human ratings by 18-39%.Furthermore, we use SUPERT as rewards to guide a neural-based reinforcement learning summarizer, yielding favorable performance compared to the state-of-the-art unsupervised summarizers.
Yang Gao 0021, Wei Zhao 0033, Steffen Eger
ACL3
2020 On the Limitations of Cross-lingual Encoders as Exposed by Reference-Free Machine Translation Evaluation
abstract
Evaluation of cross-lingual encoders is usually performed either via zero-shot cross-lingual transfer in supervised downstream tasks or via unsupervised cross-lingual textual similarity.In this paper, we concern ourselves with reference-free machine translation (MT) evaluation where we directly compare source texts to (sometimes low-quality) system translations, which represents a natural adversarial setup for multilingual encoders.Referencefree evaluation holds the promise of web-scale comparison of MT systems.We systematically investigate a range of metrics based on state-of-the-art cross-lingual semantic representations obtained with pretrained M-BERT and LASER.We find that they perform poorly as semantic encoders for reference-free MT evaluation and identify their two key limitations, namely, (a) a semantic mismatch between representations of mutual translations and, more prominently, (b) the inability to punish "translationese", i.e., low-quality literal translations.We propose two partial remedies: (1) post-hoc re-alignment of the vector spaces and (2) coupling of semantic-similarity based metrics with target-side language modeling.In segment-level MT evaluation, our best metric surpasses reference-based BLEU by 5.7 correlation points.We make our MT evaluation code available.1
Wei Zhao 0033, Goran Glavas, Maxime Peyrard, Yang Gao 0021, Robert West 0001, Steffen Eger
ACL6
2020 Vec2Sent: Probing Sentence Embeddings with Natural Language Generation
abstract
We introspect black-box sentence embeddings by conditionally generating from them with the objective to retrieve the underlying discrete sentence.We perceive of this as a new unsupervised probing task and show that it correlates well with downstream task performance.We also illustrate how the language generated from different encoders differs.We apply our approach to generate sentence analogies from sentence embeddings.
Martin Kerscher, Steffen Eger
COLING2
2020 Probing Multilingual BERT for Genetic and Typological Signals
abstract
We probe the layers in multilingual BERT (mBERT) for phylogenetic and geographic language signals across 100 languages and compute language distances based on the mBERT representations.We 1) employ the language distances to infer and evaluate language trees, finding that they are close to the reference family tree in terms of quartet tree distance, 2) perform distance matrix regression analysis, finding that the language distances can be best explained by phylogenetic and worst by structural factors and 3) present a novel measure for measuring diachronic meaning stability (based on cross-lingual representation variability) which correlates significantly with published ranked lists based on linguistic approaches.Our results contribute to the nascent field of typological interpretability of cross-lingual text representations.
Taraka Rama, Lisa Beinborn, Steffen Eger
COLING3
2020 How to Probe Sentence Embeddings in Low-Resource Languages: On Structural Design Choices for Probing Task Evaluation
abstract
Sentence encoders map sentences to real valued vectors for use in downstream applications.To peek into these representations-e.g., to increase interpretability of their resultsprobing tasks have been designed which query them for linguistic knowledge.However, designing probing tasks for lesser-resourced languages is tricky, because these often lack largescale annotated data or (high-quality) dependency parsers as a prerequisite of probing task design in English.To investigate how to probe sentence embeddings in such cases, we investigate sensitivity of probing task results to structural design choices, conducting the first such large scale study.We show that design choices like size of the annotated probing dataset and type of classifier used for evaluation do (sometimes substantially) influence probing outcomes.We then probe embeddings in a multilingual setup with design choices that lie in a 'stable region', as we identify for English, and find that results on English do not transfer to other languages.Fairer and more comprehensive sentence-level probing evaluation should thus be carried out on multiple languages in the future.
Steffen Eger, Johannes Daxenberger, Iryna Gurevych
CoNLL1
2020 PO-EMO: Conceptualization, Annotation, and Modeling of Aesthetic Emotions in German and English Poetry
abstract
Most approaches to emotion analysis of social media, literature, news, and other domains focus exclusively on basic emotion categories as defined by Ekman or Plutchik. However, art (such as literature) enables engagement in a broader range of more complex and subtle emotions. These have been shown to also include mixed emotional responses. We consider emotions in poetry as they are elicited in the reader, rather than what is expressed in the text or intended by the author. Thus, we conceptualize a set of aesthetic emotions that are predictive of aesthetic appreciation in the reader, and allow the annotation of multiple labels per line to capture mixed emotions within their context. We evaluate this novel setting in an annotation experiment both with carefully trained experts and via crowdsourcing. Our annotation with experts leads to an acceptable agreement of k = .70, resulting in a consistent dataset for future large scale analysis. Finally, we conduct first emotion classification experiments based on BERT, showing that identifying aesthetic emotions is challenging in our data, with up to .52 F1-micro on the German subset. Data and resources are available at https://github.com/tnhaider/poetry-emotion.
Thomas N. Haider, Steffen Eger, Evgeny Kim, Roman Klinger, Winfried Menninghaus
LREC2
2020 DBPal: A Fully Pluggable NL2SQL Training Pipeline
abstract
Natural language is a promising alternative interface to DBMSs because it enables non-technical users to formulate complex questions in a more concise manner than SQL. Recently, deep learning has gained traction for translating natural language to SQL, since similar ideas have been successful in the related domain of machine translation. However, the core problem with existing deep learning approaches is that they require an enormous amount of training data in order to provide accurate translations. This training data is extremely expensive to curate, since it generally requires humans to manually annotate natural language examples with the corresponding SQL queries (or vice versa). Based on these observations, we propose DBPal, a new approach that augments existing deep learning techniques in order to improve the performance of models for natural language to SQL translation. More specifically, we present a novel training pipeline that automatically generates synthetic training data in order to (1) improve overall translation accuracy, (2) increase robustness to linguistic variation, and (3) specialize the model for the target database. As we show, our DBPal training pipeline is able to improve both the accuracy and linguistic robustness of state-of-the-art natural language to SQL translation models.
Nathaniel Weir, Prasetya Ajie Utama, Alex Galakatos, Andrew Crotty, Amir Ilkhechi, Shekar Ramaswamy, Rohin Bhushan, Nadja Geisler, Benjamin Hättasch, Steffen Eger, Ugur Çetintemel, Carsten Binnig
SIGMOD Conference10
2019 Towards Scalable and Reliable Capsule Networks for Challenging NLP Applications
abstract
Obstacles hindering the development of capsule networks for challenging NLP applications include poor scalability to large output spaces and less reliable routing processes.In this paper, we introduce (i) an agreement score to evaluate the performance of routing processes at instance level; (ii) an adaptive optimizer to enhance the reliability of routing; (iii) capsule compression and partial routing to improve the scalability of capsule networks.We validate our approach on two NLP tasks, namely: multi-label text classification and question answering.Experimental results show that our approach considerably improves over strong competitors on both tasks.In addition, we gain the best results in low-resource settings with few training instances.1
Wei Zhao 0033, Haiyun Peng, Steffen Eger, Erik Cambria, Min Yang 0007
ACL (1)3
2019 MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance
abstract
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, Steffen Eger. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Wei Zhao 0033, Maxime Peyrard, Fei Liu 0004, Yang Gao 0021, Christian M. Meyer, Steffen Eger
EMNLP/IJCNLP (1)6
2018 Killing Four Birds with Two Stones: Multi-Task Learning for Non-Literal Language Detection
abstract
Non-literal language phenomena such as idioms or metaphors are commonly studied in isolation from each other in NLP. However, often similar definitions and features are being used for different phenomena, challenging the distinction. Instead, we propose to view the detection problem as a generalized non-literal language classification problem. In this paper we investigate multi-task learning for related non-literal language phenomena. We show that in contrast to simply joining the data of multiple tasks, multi-task learning consistently improves upon four metaphor and idiom detection tasks in two languages, English and German. Comparing two state-of-the-art multi-task learning architectures, we also investigate when soft parameter sharing and learned information flow can be beneficial for our related tasks. We make our adapted code publicly available.
Erik-Lân Do Dinh, Steffen Eger, Iryna Gurevych
COLING2
2018 Cross-lingual Argumentation Mining: Machine Translation (and a bit of Projection) is All You Need!
abstract
Argumentation mining (AM) requires the identification of complex discourse structures and has lately been applied with success monolingually. In this work, we show that the existing resources are, however, not adequate for assessing cross-lingual AM, due to their heterogeneity or lack of complexity. We therefore create suitable parallel corpora by (human and machine) translating a popular AM dataset consisting of persuasive student essays into German, French, Spanish, and Chinese. We then compare (i) annotation projection and (ii) bilingual word embeddings based direct transfer strategies for cross-lingual AM, finding that the former performs considerably better and almost eliminates the loss from cross-lingual transfer. Moreover, we find that annotation projection works equally well when using either costly human or cheap machine translations. Our code and data are available at http://github.com/UKPLab/coling2018-xling_argument_mining.
Steffen Eger, Johannes Daxenberger, Christian Stab, Iryna Gurevych
COLING1
2018 Is it Time to Swish? Comparing Deep Learning Activation Functions Across NLP tasks
abstract
Activation functions play a crucial role in neural networks because they are the nonlinearities which have been attributed to the success story of deep learning. One of the currently most popular activation functions is ReLU, but several competitors have recently been proposed or 'discovered', including LReLU functions and swish. While most works compare newly proposed activation functions on few tasks (usually from image classification) and against few competitors (usually ReLU), we perform the first large-scale comparison of 21 activation functions across eight different NLP tasks. We find that a largely unknown activation function performs most stably across all tasks, the so-called penalized tanh function. We also show that it can successfully replace the sigmoid and tanh gates in LSTM cells, leading to a 2 percentage point (pp) improvement over the standard choices on a challenging NLP task.
Steffen Eger, Paul Youssef, Iryna Gurevych
EMNLP1
2017 Neural End-to-End Learning for Computational Argumentation Mining
abstract
We investigate neural techniques for endto-end computational argumentation mining (AM).We frame AM both as a tokenbased dependency parsing and as a tokenbased sequence tagging problem, including a multi-task learning setup.Contrary to models that operate on the argument component level, we find that framing AM as dependency parsing leads to subpar performance results.In contrast, less complex (local) tagging models based on BiL-STMs perform robustly across classification scenarios, being able to catch longrange dependencies inherent to the AM problem.Moreover, we find that jointly learning 'natural' subtasks, in a multi-task learning setup, improves performance.
Steffen Eger, Johannes Daxenberger, Iryna Gurevych
ACL (1)1
2017 What is the Essence of a Claim? Cross-Domain Claim Identification
abstract
Argument mining has become a popular research area in NLP.It typically includes the identification of argumentative components, e.g.claims, as the central component of an argument.We perform a qualitative analysis across six different datasets and show that these appear to conceptualize claims quite differently.To learn about the consequences of such different conceptualizations of claim for practical applications, we carried out extensive experiments using state-of-the-art featurerich and deep learning systems, to identify claims in a cross-domain fashion.While the divergent conceptualization of claims in different datasets is indeed harmful to cross-domain classification, we show that there are shared properties on the lexical level as well as system configurations that can help to overcome these gaps.
Johannes Daxenberger, Steffen Eger, Ivan Habernal, Christian Stab, Iryna Gurevych
EMNLP2
2016 Language classification from bilingual word embedding graphs
abstract
We study the role of the second language in bilingual word embeddings in monolingual semantic evaluation tasks. We find strongly and weakly positive correlations between down-stream task performance and second language similarity to the target language. Additionally, we show how bilingual word embeddings can be employed for the task of semantic language classification and that joint semantic spaces vary in meaningful ways across second languages. Our results support the hypothesis that semantic language similarity is influenced by both structural similarity as well as geography/contact.
Steffen Eger, Armin Hoenen, Alexander Mehler
COLING1
2016 Still not there? Comparing Traditional Sequence-to-Sequence Models to Encoder-Decoder Neural Networks on Monotone String Translation Tasks
abstract
We analyze the performance of encoder-decoder neural models and compare them with well-known established methods. The latter represent different classes of traditional approaches that are applied to the monotone sequence-to-sequence tasks OCR post-correction, spelling correction, grapheme-to-phoneme conversion, and lemmatization. Such tasks are of practical relevance for various higher-level research fields including digital humanities, automatic text correction, and speech recognition. We investigate how well generic deep-learning approaches adapt to these tasks, and how they perform in comparison with established and more specialized methods, including our own adaptation of pruned CRFs.
Carsten Schnober, Steffen Eger, Erik-Lân Do Dinh, Iryna Gurevych
COLING2
2016 Lemmatization and Morphological Tagging in German and Latin: A Comparison and a Survey of the State-of-the-art
Steffen Eger, Rüdiger Gleim, Alexander Mehler
LREC1
2015 Multiple Many-to-Many Sequence Alignment for Combining String-Valued Variables: A G2P Experiment
abstract
Steffen Eger. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Steffen Eger
ACL (1)1
2015 Do we need bigram alignment models? On the effect of alignment quality on transduction accuracy in G2P
abstract
We investigate the need for bigram alignment models and the benefit of supervised alignment techniques in graphemeto-phoneme (G2P) conversion.Moreover, we quantitatively estimate the relationship between alignment quality and overall G2P system performance.We find that, in English, bigram alignment models do perform better than unigram alignment models on the G2P task.Moreover, we find that supervised alignment techniques may perform considerably better than their unsupervised brethren and that few manually aligned training pairs suffice for them to do so.Finally, we estimate a highly significant impact of alignment quality on overall G2P transcription performance and that this relationship is linear in nature.
Steffen Eger
EMNLP1
2015 Complex Decomposition of the Negative Distance Kernel
abstract
A Support Vector Machine (SVM) has become a very popular machine learning method for text classification. One reason for this relates to the range of existing kernels which allow for classifying data that is not linearly separable. The linear, polynomial and RBF (Gaussian Radial Basis Function) kernel are commonly used and serve as a basis of comparison in our study. We show how to derive the primal form of the quadratic Power Kernel (PK) -- also called the Negative Euclidean Distance Kernel (NDK) -- by means of complex numbers. We exemplify the NDK in the framework of text categorization using the Dewey Document Classification (DDC) as the target scheme. Our evaluation shows that the power kernel produces F-scores that are comparable to the reference kernels, but is -- except for the linear kernel -- faster to compute. Finally, we show how to extend the NDK-approach by including the Mahalanobis distance.
Tim vor der Brück, Steffen Eger, Alexander Mehler
ICMLA2
2015 Improving G2p from wiktionary and other (web) resources
Steffen Eger
INTERSPEECH1
2013 Sequence alignment with arbitrary steps and further generalizations, with applications to alignments in linguistics
Steffen Eger
Inf. Sci.1
2012 S-Restricted Monotone Alignments: Algorithm, Search Space, and Applications
Steffen Eger
COLING1