VLDB 2026 Research / reviewers in the wild / expert
Anastasia Shimorina
dblp:19/10057
· DBLP profile ↗
12ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 91% Trustworthy machine learning · 7% Knowledge representation and reasoning · 2% | |
| Software engineering, system software, and programming languages
1 paper |
Requirements engineering and software design · 100% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › text summarization
dialogue summarization |
0.9 | 1 | 2025 | PoSum-Bench: Benchmarking Position Bias in LLM-based Conversational Summarization · EMNLP 2025 |
Natural language and speech › Language models and text generation
text summarization |
0.9 | 1 | 2025 | PoSum-Bench: Benchmarking Position Bias in LLM-based Conversational Summarization · EMNLP 2025 |
Natural language and speech › Language models and text generation › text generation
surface realisation |
0.4 | 1 | 2019 | Surface Realisation Using Full Delexicalisation · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Language models and text generation › text generation › sentence planning
microplanning |
0.3 | 1 | 2017 | Creating Training Corpora for NLG Micro-Planners · ACL (1) 2017 |
Natural language and speech › Language models and text generation › text generation › text simplification
sentence simplification |
0.3 | 1 | 2017 | Split and Rephrase · EMNLP 2017 |
Natural language and speech › Language models and text generation › text generation › text simplification › sentence simplification
split and rephrase |
0.3 | 1 | 2017 | Split and Rephrase · EMNLP 2017 |
Natural language and speech › Language models and text generation
text generation |
0.3 | 1 | 2017 | Creating Training Corpora for NLG Micro-Planners · ACL (1) 2017 |
Requirements engineering and software design
requirements traceability |
0.3 | 1 | 2017 | ModelWriter: text and model-synchronized document engineering platform · ASE 2017 |
Machine learning › Trustworthy machine learning › fairness
fairness evaluation |
0.3 | 1 | 2025 | PoSum-Bench: Benchmarking Position Bias in LLM-based Conversational Summarization · EMNLP 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge base |
0.1 | 1 | 2017 | Creating Training Corpora for NLG Micro-Planners · ACL (1) 2017 |
Natural language and speech › Language models and text generation › language modeling › language model architecture
sequence-to-sequence model |
0.1 | 1 | 2017 | Split and Rephrase · EMNLP 2017 |
Methods — techniques the papers use, named apart from their topics
semantic similarity metric · 0.9reference-free evaluation · 0.9semantic parsing · 0.6first-order relational logic · 0.6finite model finding · 0.6description logic · 0.6delexicalisation · 0.4sequence-to-sequence · 0.3semantic modeling · 0.3data-to-text generation · 0.3corpus generation · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PoSum-Bench: Benchmarking Position Bias in LLM-based Conversational SummarizationabstractLarge language models (LLMs) are increasingly used for zero-shot conversation summarization, but often exhibit positional bias—tending to overemphasize content from the beginning or end of a conversation while neglecting the middle. To address this issue, we introduce PoSum-Bench, a comprehensive benchmark for evaluating positional bias in conversational summarization, featuring diverse English and French conversational datasets spanning formal meetings, casual conversations, and customer service interactions. We propose a novel semantic similarity-based sentence-level metric to quantify the direction and magnitude of positional bias in model-generated summaries, enabling systematic and reference-free evaluation across conversation positions, languages, and conversational contexts.Our benchmark and methodology thus provide the first systematic, cross-lingual framework for reference-free evaluation of positional bias in conversational summarization, laying the groundwork for developing more balanced and unbiased summarization models. Lionel Delphin-Poulat, Christèle Tarnec, Anastasia Shimorina |
EMNLP | 4 |
| 2021 | A Systematic Review of Reproducibility Research in Natural Language ProcessingabstractAgainst the background of what has been termed a reproducibility crisis in science, the NLP field is becoming increasingly interested in, and conscientious about, the reproducibility of its results.The past few years have seen an impressive range of new initiatives, events and active research in the area.However, the field is far from reaching a consensus about how reproducibility should be defined, measured and addressed, with diversity of views currently increasing rather than converging.With this focused contribution, we aim to provide a wideangle, and as near as possible complete, snapshot of current work on reproducibility in NLP, delineating differences and similarities, and providing pointers to common denominators. Anya Belz, Anastasia Shimorina, Ehud Reiter |
EACL | 3 |
| 2021 | The ReproGen Shared Task on Reproducibility of Human Evaluations in NLG: Overview and ResultsabstractThe NLP field has recently seen a substantial increase in work related to reproducibility of results, and more generally in recognition of the importance of having shared definitions and practices relating to evaluation.Much of the work on reproducibility has so far focused on metric scores, with reproducibility of human evaluation results receiving far less attention.As part of a research programme designed to develop theory and practice of reproducibility assessment in NLP, we organised the first shared task on reproducibility of human evaluations, ReproGen 2021.This paper describes the shared task in detail, summarises results from each of the reproduction studies submitted, and provides further comparative analysis of the results.Out of nine initial team registrations, we received submissions from four teams.Meta-analysis of the four reproduction studies revealed varying degrees of reproducibility, and allowed very tentative first conclusions about what types of evaluation tend to have better reproducibility. Anya Belz, Anastasia Shimorina, Ehud Reiter |
INLG | 2 |
| 2021 | An Error Analysis Framework for Shallow Surface RealisationabstractAbstract The metrics standardly used to evaluate Natural Language Generation (NLG) models, such as BLEU or METEOR, fail to provide information on which linguistic factors impact performance. Focusing on Surface Realization (SR), the task of converting an unordered dependency tree into a well-formed sentence, we propose a framework for error analysis which permits identifying which features of the input affect the models’ results. This framework consists of two main components: (i) correlation analyses between a wide range of syntactic metrics and standard performance metrics and (ii) a set of techniques to automatically identify syntactic constructs that often co-occur with low performance scores. We demonstrate the advantages of our framework by performing error analysis on the results of 174 system runs submitted to the Multilingual SR shared tasks; we show that dependency edge accuracy correlate with automatic metrics thereby providing a more interpretable basis for evaluation; and we suggest ways in which our framework could be used to improve models and data. The framework is available in the form of a toolkit which can be used both by campaign organizers to provide detailed, linguistically interpretable feedback on the state of the art in multilingual SR, and by individual researchers to improve models and datasets.1 Anastasia Shimorina, Claire Gardent, Yannick Parmentier 0001 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | ReproGen: Proposal for a Shared Task on Reproducibility of Human Evaluations in NLGabstractAcross NLP, a growing body of work is looking at the issue of reproducibility.However, replicability of human evaluation experiments and reproducibility of their results is currently under-addressed, and this is of particular concern for NLG where human evaluations are the norm.This paper outlines our ideas for a shared task on reproducibility of human evaluations in NLG which aims (i) to shed light on the extent to which past NLG evaluations have been replicable and reproducible, and (ii) to draw conclusions regarding how evaluations can be designed and reported to increase replicability and reproducibility.If the task is run over several years, we hope to be able to document an overall increase in levels of replicability and reproducibility over time. Anya Belz, Anastasia Shimorina, Ehud Reiter |
INLG | 3 |
| 2019 | Surface Realisation Using Full DelexicalisationabstractAnastasia Shimorina, Claire Gardent. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Anastasia Shimorina, Claire Gardent |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Handling Rare Items in Data-to-Text GenerationabstractNeural approaches to data-to-text generation generally handle rare input items using either delexicalisation or a copy mechanism.We investigate the relative impact of these two methods on two datasets (E2E and WebNLG) and using two evaluation settings.We show (i) that rare items strongly impact performance; (ii) that combining delexicalisation and copying yields the strongest improvement; (iii) that copying underperforms for rare and unseen items and (iv) that the impact of these two mechanisms greatly varies depending on how the dataset is constructed and on how it is split into train, dev and test 1 . Anastasia Shimorina, Claire Gardent |
INLG | 1 |
| 2017 | Creating Training Corpora for NLG Micro-PlannersabstractIn this paper, we present a novel framework for semi-automatically creating linguistically challenging microplanning data-to-text corpora from existing Knowledge Bases.Because our method pairs data of varying size and shape with texts ranging from simple clauses to short texts, a dataset created using this framework provides a challenging benchmark for microplanning.Another feature of this framework is that it can be applied to any large scale knowledge base and can therefore be used to train and learn KB verbalisers.We apply our framework to DBpedia data and compare the resulting dataset with Wen et al. (2016)'s.We show that while Wen et al.'s dataset is more than twice larger than ours, it is less diverse both in terms of input and in terms of text.We thus propose our corpus generation framework as a novel method for creating challenging data sets from which NLG models can be learned which are capable of handling the complex interactions occurring during in micro-planning between lexicalisation, aggregation, surface realisation, referring expression generation and sentence segmentation.To encourage researchers to take up this challenge, we recently made available a dataset created using this framework in the context of the WEBNLG shared task. Claire Gardent, Anastasia Shimorina, Shashi Narayan, Laura Perez-Beltrachini |
ACL (1) | 2 |
| 2017 | Split and RephraseabstractWe propose a new sentence simplification task (Split-and-Rephrase) where the aim is to split a complex sentence into a meaning preserving sequence of shorter sentences.Like sentence simplification, splitting-and-rephrasing has the potential of benefiting both natural language processing and societal applications.Because shorter sentences are generally better processed by NLP systems, it could be used as a preprocessing step which facilitates and improves the performance of parsers, semantic role labelers and machine translation systems.It should also be of use for people with reading disabilities because it allows the conversion of longer sentences into shorter ones.This paper makes two contributions towards this new task.First, we create and make available a benchmark consisting of 1,066,115 tuples mapping a single complex sentence to a sequence of sentences expressing the same meaning.1 Second, we propose five models (vanilla sequence-to-sequence to semantically-motivated models) to understand the difficulty of the proposed task. Shashi Narayan, Claire Gardent, Shay B. Cohen, Anastasia Shimorina |
EMNLP | 4 |
| 2017 | Mapping Natural Language to Description Logic
Bikash Gyawali, Anastasia Shimorina, Claire Gardent, Samuel Cruz-Lara, Mariem Mahfoudh |
ESWC (1) | 2 |
| 2017 | The WebNLG Challenge: Generating Text from RDF DataabstractThe WebNLG challenge consists in mapping sets of RDF triples to text.It provides a common benchmark on which to train, evaluate and compare "microplanners", i.e. generation systems that verbalise a given content by making a range of complex interacting choices including referring expression generation, aggregation, lexicalisation, surface realisation and sentence segmentation.In this paper, we introduce the microplanning task, describe data preparation, introduce our evaluation methodology, analyse participant results and provide a brief description of the participating systems.(3) a. LOCATION-COUNTRY-STARTDATE ⇒ Passive-Apposition-Active Claire Gardent, Anastasia Shimorina, Shashi Narayan, Laura Perez-Beltrachini |
INLG | 2 |
| 2017 | ModelWriter: text and model-synchronized document engineering platformabstractThe ModelWriter platform provides a generic framework for automated traceability analysis. In this paper, we demonstrate how this framework can be used to trace the consistency and completeness of technical documents that consist of a set of System Installation Design Principles used by Airbus to ensure the correctness of aircraft system installation. We show in particular, how the platform allows the integration of two types of reasoning: reasoning about the meaning of text using semantic parsing and description logic theorem proving; and reasoning about document structure using first-order relational logic and finite model finding for traceability analysis. Ferhat Erata, Claire Gardent, Bikash Gyawali, Anastasia Shimorina, Yvan Lussaud, Bedir Tekinerdogan, Geylani Kardas, Anne Monceaux |
ASE | 4 |