VLDB 2026 Research / reviewers in the wild / expert
Markus Leippold
dblp:53/9476
· DBLP profile ↗
11ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0001-5983-2360ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
Jingwei Ni, Ekaterina Fadeeva, Mubashara Akhtar, Jiaheng Zhang, Elliott Ash, Markus Leippold, Timothy Baldwin, See-Kiong Ng, Artem Shelmanov, Mrinmaya Sachan |
ACL (1) | 7 |
| 2025 | ReportGRI: Automating GRI Alignment and Report AssessmentabstractOrganisations disclose their sustainability performance in corporate sustainability reports (CSRs). CSRs vary widely in structure and depth depending on the reporting framework. Such disparity, together with report complexity and volume, poses significant challenges to transparency, comparability and standardisation. To address this problem, we introduce ReportGRI, an automated system for Global Reporting Initiative (GRI) indexing and qualitative assessment of CSRs. The interactive framework leverages information retrieval techniques and zero-shot prompting to enable GRI disclosure-based report indexing and report coverage assessment by visualising well-covered topics and reporting gaps. The tool facilitates scalable and explainable benchmarking of Environmental, Social and Governance (ESG) reporting quality, enhancing report interpretation, transparency, and corporate accountability. The system is open-sourced on GitHub with an introduction video. Aida Usmanova, Rana Abdullah, Debayan Banerjee, Markus Leippold, Ricardo Usbeck |
CIKM | 4 |
| 2025 | DIRAS: Efficient LLM Annotation of Document Relevance for Retrieval Augmented GenerationabstractJingwei Ni, Tobias Schimanski, Meihong Lin, Mrinmaya Sachan, Elliott Ash, Markus Leippold. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jingwei Ni, Tobias Schimanski, Meihong Lin, Mrinmaya Sachan, Elliott Ash, Markus Leippold |
NAACL (Long Papers) | 6 |
| 2024 | AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM AnnotatorsabstractJingwei Ni, Minjing Shi, Dominik Stammbach, Mrinmaya Sachan, Elliott Ash, Markus Leippold. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jingwei Ni, Minjing Shi, Dominik Stammbach, Mrinmaya Sachan, Elliott Ash, Markus Leippold |
ACL (1) | 6 |
| 2024 | Towards Faithful and Robust LLM Specialists for Evidence-Based Question-AnsweringabstractAdvances towards more faithful and traceable answers of Large Language Models (LLMs) are crucial for various research and practical endeavors.One avenue in reaching this goal is basing the answers on reliable sources.However, this Evidence-Based QA has proven to work insufficiently with LLMs in terms of citing the correct sources (source quality) and truthfully representing the information within sources (answer attributability).In this work, we systematically investigate how to robustly fine-tune LLMs for better source quality and answer attributability.Specifically, we introduce a data generation pipeline with automated data quality filters, which can synthesize diversified high-quality training and testing data at scale.We further introduce four test sets to benchmark the robustness of fine-tuned specialist models.Extensive evaluation shows that fine-tuning on synthetic data improves performance on both in-and out-of-distribution.Furthermore, we show that data quality, which can be drastically improved by proposed quality filters, matters more than quantity in improving Evidence-Based QA. Tobias Schimanski, Jingwei Ni, Mathias Kraus, Elliott Ash, Markus Leippold |
ACL (1) | 5 |
| 2024 | ClimRetrieve: A Benchmarking Dataset for Information Retrieval from Corporate Climate DisclosuresabstractTo handle the vast amounts of qualitative data produced in corporate climate communication, stakeholders increasingly rely on Retrieval Augmented Generation (RAG) systems.However, a significant gap remains in evaluating domain-specific information retrieval -the basis for answer generation.To address this challenge, this work simulates the typical tasks of a sustainability analyst by examining 30 sustainability reports with 16 detailed climate-related questions.As a result, we obtain a dataset with over 8.5K unique question-source-answer pairs labeled by different levels of relevance.Furthermore, we develop a use case with the dataset to investigate the integration of expert knowledge into information retrieval with embeddings.Although we show that incorporating expert knowledge works, we also outline the critical limitations of embeddings in knowledge-intensive downstream domains like climate change communication.121 All the data and code for this project is available on https://github.com/tobischimanski/ClimRetrieve.2 We thank the expert annotators Aysha Emmerson, Emily Hsu, and Capucine Le Meur for their work on this project.3 For example, companies must describe the processes they use to identify, assess, and manage these risks and opportuni- Tobias Schimanski, Jingwei Ni, Roberto Martín, Nicola Ranger, Markus Leippold |
EMNLP | 5 |
| 2024 | Assessing Large Language Models on Climate InformationabstractAs Large Language Models (LLMs) rise in popularity, it is necessary to assess their capability in critically relevant domains. We present a comprehensive evaluation framework, grounded in science communication research, to assess LLM responses to questions about climate change. Our framework emphasizes both presentational and epistemological adequacy, offering a fine-grained analysis of LLM generations spanning 8 dimensions and 30 issues. Our evaluation task is a real-world example of a growing number of challenging problems where AI can complement and lift human performance. We introduce a novel protocol for scalable oversight that relies on AI Assistance and raters with relevant education. We evaluate several recent LLMs on a set of diverse climate questions. Our results point to a significant gap between surface and epistemological qualities of LLMs in the realm of climate communication. Jannis Bulian, Mike S. Schäfer, Afra Amini, Heidi Lam, Massimiliano Ciaramita, Ben Gaiarin, Michelle Chen Huebscher, Christian Buck, Niels Mede, Markus Leippold, Nadine Strauß |
ICML | 10 |
| 2023 | When Does Aggregating Multiple Skills with Multi-Task Learning Work? A Case Study in Financial NLPabstractMulti-task learning (MTL) aims at achieving a better model by leveraging data and knowledge from multiple tasks.However, MTL does not always work -sometimes negative transfer occurs between tasks, especially when aggregating loosely related skills, leaving it an open question when MTL works.Previous studies show that MTL performance can be improved by algorithmic tricks.However, what tasks and skills should be included is less well explored.In this work, we conduct a case study in Financial NLP where multiple datasets exist for skills relevant to the domain, such as numeric reasoning and sentiment analysis.Due to the task difficulty and data scarcity in the Financial NLP domain, we explore when aggregating such diverse skills from multiple datasets with MTL can work.Our findings suggest that the key to MTL success lies in skill diversity, relatedness between tasks, and choice of aggregation size and shared capacity.Specifically, MTL works well when tasks are diverse but related, and when the size of the task aggregation and the shared capacity of the model are balanced to avoid overwhelming certain tasks. 1 Jingwei Ni, Zhijing Jin 0001, Mrinmaya Sachan, Markus Leippold |
ACL (1) | 5 |
| 2023 | ClimateBERT-NetZero: Detecting and Assessing Net Zero and Reduction TargetsabstractPublic and private actors struggle to assess the vast amounts of information about sustainability commitments made by various institutions.To address this problem, we create a novel tool for automatically detecting corporate, national, and regional net zero and reduction targets in three steps.First, we introduce an expertannotated data set with 3.5K text samples.Second, we train and release ClimateBERT-NetZero, a natural language classifier to detect whether a text contains a net zero or reduction target.Third, we showcase its analysis potential with two use cases: We first demonstrate how ClimateBERT-NetZero can be combined with conventional question-answering (Q&A) models to analyze the ambitions displayed in net zero and reduction targets.Furthermore, we employ the ClimateBERT-NetZero model on quarterly earning call transcripts and outline how communication patterns evolve over time.Our experiments demonstrate promising pathways for extracting and analyzing net zero and emission reduction targets at scale.7 See https://huggingface.co/climatebert/ distilroberta-base-climate-detector for the model. Tobias Schimanski, Julia Anna Bingler, Mathias Kraus, Camilla Hyslop, Markus Leippold |
EMNLP | 5 |
| 2022 | Towards Climate Awareness in NLP ResearchabstractThe climate impact of AI, and NLP research in particular, has become a serious issue given the enormous amount of energy that is increasingly being used for training and running computational models.Consequently, increasing focus is placed on efficient NLP.However, this important initiative lacks simple guidelines that would allow for systematic climate reporting of NLP research.We argue that this deficiency is one of the reasons why very few publications in NLP report key figures that would allow a more thorough examination of environmental impact, and present a quantitative survey to demonstrate this.As a remedy, we propose a climate performance model card with the primary purpose of being practically usable with only limited information about experiments and the underlying computer hardware.We describe why this step is essential to increase awareness about the environmental impact of NLP research and, thereby, paving the way for more thorough discussions.1 Daniel Hershcovich, Nicolas Webersinke, Mathias Kraus, Julia Anna Bingler, Markus Leippold |
EMNLP | 5 |
| 2020 | MuSeM: Detecting Incongruent News Headlines using Mutual Attentive Semantic MatchingabstractMeasuring congruence between two texts has several useful applications, such as detecting the prevalent deceptive and misleading news headlines on the web. Many works have proposed machine learning based solutions such as text similarity between the headline and body text to detect the incongruence. Text similarity based methods fail to perform well due to different inherent challenges such as relative length mismatch between the news headline and its body content and non-overlapping vocabulary. On the other hand, more recent works that use headline guided attention to learn a headline derived contextual representation of the news body also result in convoluting overall representation due to the news body's lengthiness. This paper proposes a method that uses inter-mutual attention-based semantic matching between the original and synthetically generated headlines, which utilizes the difference between all pairs of word embeddings of words involved. The paper also investigates two more variations of our method, which use concatenation and dot-products of word embeddings of the words of original and synthetic headlines. We observe that the proposed method outperforms prior-arts significantly for two publicly available datasets. Rahul Mishra 0004, Piyush Yadav, Rémi Calizzano, Markus Leippold |
ICMLA | 4 |