VLDB 2026 Research / reviewers in the wild / expert
Juri Opitz
dblp:185/5616
· DBLP profile ↗
17ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0001-6892-4574ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLEF HIPE-2026: Evaluating Accurate and Efficient Person-Place Relation Extraction from Multilingual Historical Texts
Juri Opitz, Corina Julia Raclé, Emanuela Boros, Andrianos Michail, Matteo Romanello, Maud Ehrmann, Simon Clematide |
ECIR (4) | 1 |
| 2026 | A Recipe for Adapting Multilingual Embedders to OCR-Error Robustness and Historical Texts
Andrianos Michail, Stylianos Psychias, Juri Opitz, Simon Clematide |
LREC | 3 |
| 2025 | ConLoan: A Contrastive Multilingual Dataset for Evaluating LoanwordsabstractSina Ahmadi, Micha David Hess, Elena Álvarez-Mellado, Alessia Battisti, Cui Ding, Anne Göhring, Yingqiang Gao, Zifan Jiang, Andrianos Michail, Peshmerge Morad, Joel Niklaus, Maria Christina Panagiotopoulou, Stefano Perrella, Juri Opitz, Anastassia Shaitarova, Rico Sennrich. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Sina Ahmadi, Micha David Hess, Elena Álvarez Mellado, Alessia Battisti, Cui Ding, Anne Göhring, Yingqiang Gao, Zifan Jiang, Andrianos Michail, Peshmerge Morad, Joel Niklaus, Maria Christina Panagiotopoulou, Stefano Perrella, Juri Opitz, Anastassia Shaitarova, Rico Sennrich |
ACL (1) | 14 |
| 2025 | PARAPHRASUS: A Comprehensive Benchmark for Evaluating Paraphrase Detection ModelsabstractThe task of determining whether two texts are paraphrases has long been a challenge in NLP. However, the prevailing notion of paraphrase is often quite simplistic, offering only a limited view of the vast spectrum of paraphrase phenomena. Indeed, we find that evaluating models in a paraphrase dataset can leave uncertainty about their true semantic understanding. To alleviate this, we create PARAPHRASUS, a benchmark designed for multi-dimensional assessment, benchmarking and selection of paraphrase detection models. We find that paraphrase detection models under our fine-grained evaluation lens exhibit trade-offs that cannot be captured through a single classification dataset. Furthermore, PARAPHRASUS allows prompt calibration for different use cases, tailoring LLM models to specific strictness levels. PARAPHRASUS includes 3 challenges spanning over 10 datasets, including 8 repurposed and 2 newly annotated; we release it along with a benchmarking library at https://github.com/impresso/paraphrasus Andrianos Michail, Simon Clematide, Juri Opitz |
COLING | 3 |
| 2025 | Sentence Smith: Controllable Edits for Evaluating Text EmbeddingsabstractControllable and transparent text generation has been a long-standing goal in NLP.Almost as long-standing is a general idea for addressing this challenge: Parsing text to a symbolic representation, and generating from it.However, earlier approaches were hindered by parsing and generation insufficiencies.Using modern parsers and a safety supervision mechanism, we show how close current methods come to this goal.Concretely, we propose the SENTENCE-SMITH framework for English, which has three steps: 1. Parsing a sentence into a semantic graph.2. Applying human-designed semantic manipulation rules.3. Generating text from the manipulated graph.A final entailment check (4.) verifies the validity of the applied transformation.To demonstrate our framework's utility, we use it to induce hard negative text pairs that challenge text embedding models.Since the controllable generation makes it possible to clearly isolate different types of semantic shifts, we can evaluate text embedding models in a fine-grained way, also addressing an issue in current benchmarking where linguistic phenomena remain opaque.Human validation confirms that our transparent generation process produces texts of good quality.Notably, our way of generation is very resource-efficient, since it relies only on smaller neural networks. Andrianos Michail, Reto Gubelmann, Simon Clematide, Juri Opitz |
EMNLP | 5 |
| 2025 | Interpretable Text Embeddings and Text Similarity Explanation: A SurveyabstractText embeddings are a fundamental component in many NLP tasks, including classification, regression, clustering, and semantic search.However, despite their ubiquitous application, challenges persist in interpreting embeddings and explaining similarities between them.In this work, we provide a structured overview of methods specializing in inherently interpretable text embeddings and text similarity explanation, an underexplored research area.We characterize the main ideas, approaches, and tradeoffs.We compare means of evaluation, discuss overarching lessons learned and finally identify opportunities and open challenges for future research. Juri Opitz, Lucas Möller, Andrianos Michail, Sebastian Padó, Simon Clematide |
EMNLP | 1 |
| 2024 | Schroedinger's Threshold: When the AUC Doesn't Predict AccuracyabstractThe Area Under Curve measure (AUC) seems apt to evaluate and compare diverse models, possibly without calibration. An important example of AUC application is the evaluation and benchmarking of models that predict faithfulness of generated text. But we show that the AUC yields an academic and optimistic notion of accuracy that can misalign with the actual accuracy observed in application, yielding significant changes in benchmark rankings. To paint a more realistic picture of downstream model performance (and prepare it for actual application), we explore different calibration modes, testing calibration data and method. Juri Opitz |
LREC/COLING | 1 |
| 2024 | A Survey of AMR ApplicationsabstractIn the ten years since the development of the Abstract Meaning Representation (AMR) formalism, substantial progress has been made on AMR-related tasks such as parsing and alignment.Still, the engineering applications of AMR are not fully understood.In this survey, we categorize and characterize more than 100 papers which use AMR for downstream tasksthe first survey of this kind for AMR.Specifically, we highlight (1) the range of applications for which AMR has been harnessed, and (2) the techniques for incorporating AMR into those applications.We also detect broader AMR engineering patterns and outline areas of future work that seem ripe for AMR incorporation.We hope that this survey will be useful to those interested in using AMR and that it sparks discussion on the role of symbolic representations in the age of neural-focused NLP research. Shira Wein, Juri Opitz |
EMNLP | 2 |
| 2024 | A Survey of Meaning Representations - From Theory to Practical UtilityabstractZacchary Sadeddine, Juri Opitz, Fabian Suchanek. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zacchary Sadeddine, Juri Opitz, Fabian M. Suchanek |
NAACL-HLT | 2 |
| 2024 | A Closer Look at Classification Evaluation Metrics and a Critical Reflection of Common Evaluation PracticeabstractAbstract Classification systems are evaluated in a countless number of papers. However, we find that evaluation practice is often nebulous. Frequently, metrics are selected without arguments, and blurry terminology invites misconceptions. For instance, many works use so-called ‘macro’ metrics to rank systems (e.g., ‘macro F1’) but do not clearly specify what they would expect from such a ‘macro’ metric. This is problematic, since picking a metric can affect research findings and thus any clarity in the process should be maximized. Starting from the intuitive concepts of bias and prevalence, we perform an analysis of common evaluation metrics. The analysis helps us understand the metrics’ underlying properties, and how they align with expectations as found expressed in papers. Then we reflect on the practical situation in the field, and survey evaluation practice in recent shared tasks. We find that metric selection is often not supported with convincing arguments, an issue that can make a system ranking seem arbitrary. Our work aims at providing overview and guidance for more informed and transparent metric selection, fostering meaningful evaluation. Juri Opitz |
Trans. Assoc. Comput. Linguistics | 1 |
| 2023 | Similarity-weighted Construction of Contextualized Commonsense Knowledge Graphs for Knowledge-intense Argumentation TasksabstractArguments often do not make explicit how a conclusion follows from its premises.To compensate for this lack, we enrich arguments with structured background knowledge to support knowledge-intense argumentation tasks.We present a new unsupervised method for constructing Contextualized Commonsense Knowledge Graphs (CCKGs) that selects contextually relevant knowledge from large knowledge graphs (KGs) efficiently and at high quality.Our work goes beyond context-insensitive knowledge extraction heuristics by computing semantic similarity between KG triplets and textual arguments.Using these triplet similarities as weights, we extract contextualized knowledge paths that connect a conclusion to its premise, while maximizing similarity to the argument.We combine multiple paths into a CCKG that we optionally prune to reduce noise and raise precision.Intrinsic evaluation of the quality of our graphs shows that our method is effective for (re)constructing human explanation graphs.Manual evaluations in a large-scale knowledge selection setup confirm high recall and precision of implicit CSK in the CCKGs.Finally, we demonstrate the effectiveness of CCKGs in a knowledge-insensitive argument quality rating task, outperforming strong baselines and rivaling a GPT-3 based system. 1 Moritz Plenz, Juri Opitz, Philipp Heinisch, Philipp Cimiano, Anette Frank |
ACL (1) | 2 |
| 2021 | Towards a Decomposable Metric for Explainable Evaluation of Text Generation from AMRabstractSystems that generate natural language text from abstract meaning representations such as AMR are typically evaluated using automatic surface matching metrics that compare the generated texts to reference texts from which the input meaning representations were constructed.We show that besides wellknown issues from which such metrics suffer, an additional problem arises when applying these metrics for AMR-to-text evaluation, since an abstract meaning representation allows for numerous surface realizations.In this work we aim to alleviate these issues by proposing MF β , a decomposable metric that builds on two pillars.The first is the principle of meaning preservation M: it measures to what extent a given AMR can be reconstructed from the generated sentence using SOTA AMR parsers and applying (finegrained) AMR evaluation metrics to measure the distance between the original and the reconstructed AMR.The second pillar builds on a principle of (grammatical) form F that measures the linguistic quality of the generated text, which we implement using SOTA language models.In two extensive pilot studies we show that fulfillment of both principles offers benefits for AMR-to-text evaluation, including explainability of scores.Since MF β does not necessarily rely on gold AMRs, it may extend to other text generation tasks. Juri Opitz, Anette Frank |
EACL | 1 |
| 2021 | Weisfeiler-Leman in the Bamboo: Novel AMR Graph Metrics and a Benchmark for AMR Graph SimilarityabstractAbstract Several metrics have been proposed for assessing the similarity of (abstract) meaning representations (AMRs), but little is known about how they relate to human similarity ratings. Moreover, the current metrics have complementary strengths and weaknesses: Some emphasize speed, while others make the alignment of graph structures explicit, at the price of a costly alignment step. In this work we propose new Weisfeiler-Leman AMR similarity metrics that unify the strengths of previous metrics, while mitigating their weaknesses. Specifically, our new metrics are able to match contextualized substructures and induce n:m alignments between their nodes. Furthermore, we introduce a Benchmark for AMR Metrics based on Overt Objectives (Bamboo), the first benchmark to support empirical assessment of graph-based MR similarity metrics. Bamboo maximizes the interpretability of results by defining multiple overt objectives that range from sentence similarity objectives to stress tests that probe a metric’s robustness against meaning-altering and meaning- preserving graph transformations. We show the benefits of Bamboo by profiling previous metrics and our own metrics. Results indicate that our novel metrics may serve as a strong baseline for future work. Juri Opitz, Angel Daza, Anette Frank |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | Argumentative Relation Classification with Background KnowledgeabstractA common conception is that the understanding of relations that hold between argument units requires knowledge beyond the text. But to date, argument analysis systems that leverage knowledge resources are still very rare. In this paper, we propose an unsupervised graph-based ranking method that extracts relevant multi-hop knowledge from a background knowledge resource. This knowledge is integrated into a neural argumentative relation classifier via an attention-based gating mechanism. In contrast to prior work we emphasize the selection of relevant multi-hop knowledge, and apply methods to automatically enrich the knowledge resource with missing knowledge. We assess model performance on two datasets, showing considerable improvement over strong baselines. Debjit Paul, Juri Opitz, Maria Becker, Jonathan Kobbe, Graeme Hirst, Anette Frank |
COMMA | 2 |
| 2020 | AMR Similarity Metrics from PrinciplesabstractDifferent metrics have been proposed to compare Abstract Meaning Representation (AMR) graphs. The canonical Smatch metric (Cai and Knight, 2013 ) aligns the variables of two graphs and assesses triple matches. The recent SemBleu metric (Song and Gildea, 2019 ) is based on the machine-translation metric Bleu (Papineni et al., 2002 ) and increases computational efficiency by ablating the variable-alignment. In this paper, i) we establish criteria that enable researchers to perform a principled assessment of metrics comparing meaning representations like AMR; ii) we undertake a thorough analysis of Smatch and SemBleu where we show that the latter exhibits some undesirable properties. For example, it does not conform to the identity of indiscernibles rule and introduces biases that are hard to control; and iii) we propose a novel metric S2 match that is more benevolent to only very slight meaning deviations and targets the fulfilment of all established criteria. We assess its suitability and show its advantages over Smatch and SemBleu. Juri Opitz, Anette Frank, Letitia Parcalabescu |
Trans. Assoc. Comput. Linguistics | 1 |
| 2019 | Exploiting Background Knowledge for Argumentative Relation Classification
Jonathan Kobbe, Juri Opitz, Maria Becker, Ioana Hulpus, Heiner Stuckenschmidt, Anette Frank |
LDK | 2 |
| 2017 | A Mention-Ranking Model for Abstract Anaphora ResolutionabstractResolving abstract anaphora is an important, but difficult task for text understanding. Yet, with recent advances in representation learning this task becomes a more tangible aim. A central property of abstract anaphora is that it establishes a relation between the anaphor embedded in the anaphoric sentence and its (typically non-nominal) antecedent. We propose a mention-ranking model that learns how abstract anaphors relate to their antecedents with an LSTM-Siamese Net. We overcome the lack of training data by generating artificial anaphoric sentence–antecedent pairs. Our model outperforms state-of-the-art results on shell noun resolution. We also report first benchmark results on an abstract anaphora subset of the ARRAU corpus. This corpus presents a greater challenge due to a mixture of nominal and pronominal anaphors and a greater range of confounders. We found model variants that outperform the baselines for nominal anaphors, without training on individual anaphor data, but still lag behind for pronominal anaphors. Our model selects syntactically plausible candidates and – if disregarding syntax – discriminates candidates using deeper features. Ana Marasovic, Leo Born, Juri Opitz, Anette Frank |
EMNLP | 3 |