Anastasia Shimorina

dblp:19/10057 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 91% Trustworthy machine learning · 7% Knowledge representation and reasoning · 2%
Software engineering, system software, and programming languages
1 paper
Requirements engineering and software design · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › text summarization
dialogue summarization
0.912025
PoSum-Bench: Benchmarking Position Bias in LLM-based Conversational Summarization · EMNLP 2025
Natural language and speech › Language models and text generation
text summarization
0.912025
PoSum-Bench: Benchmarking Position Bias in LLM-based Conversational Summarization · EMNLP 2025
Natural language and speech › Language models and text generation › text generation
surface realisation
0.412019
Surface Realisation Using Full Delexicalisation · EMNLP/IJCNLP (1) 2019
Natural language and speech › Language models and text generation › text generation › sentence planning
microplanning
0.312017
Creating Training Corpora for NLG Micro-Planners · ACL (1) 2017
Natural language and speech › Language models and text generation › text generation › text simplification
sentence simplification
0.312017
Split and Rephrase · EMNLP 2017
Natural language and speech › Language models and text generation › text generation › text simplification › sentence simplification
split and rephrase
0.312017
Split and Rephrase · EMNLP 2017
Natural language and speech › Language models and text generation
text generation
0.312017
Creating Training Corpora for NLG Micro-Planners · ACL (1) 2017
Requirements engineering and software design
requirements traceability
0.312017
ModelWriter: text and model-synchronized document engineering platform · ASE 2017
Machine learning › Trustworthy machine learning › fairness
fairness evaluation
0.312025
PoSum-Bench: Benchmarking Position Bias in LLM-based Conversational Summarization · EMNLP 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge base
0.112017
Creating Training Corpora for NLG Micro-Planners · ACL (1) 2017
Natural language and speech › Language models and text generation › language modeling › language model architecture
sequence-to-sequence model
0.112017
Split and Rephrase · EMNLP 2017

Methods — techniques the papers use, named apart from their topics

semantic similarity metric · 0.9reference-free evaluation · 0.9semantic parsing · 0.6first-order relational logic · 0.6finite model finding · 0.6description logic · 0.6delexicalisation · 0.4sequence-to-sequence · 0.3semantic modeling · 0.3data-to-text generation · 0.3corpus generation · 0.3
YearPublicationVenuePosition
2025 PoSum-Bench: Benchmarking Position Bias in LLM-based Conversational Summarization
abstract
Large language models (LLMs) are increasingly used for zero-shot conversation summarization, but often exhibit positional bias—tending to overemphasize content from the beginning or end of a conversation while neglecting the middle. To address this issue, we introduce PoSum-Bench, a comprehensive benchmark for evaluating positional bias in conversational summarization, featuring diverse English and French conversational datasets spanning formal meetings, casual conversations, and customer service interactions. We propose a novel semantic similarity-based sentence-level metric to quantify the direction and magnitude of positional bias in model-generated summaries, enabling systematic and reference-free evaluation across conversation positions, languages, and conversational contexts.Our benchmark and methodology thus provide the first systematic, cross-lingual framework for reference-free evaluation of positional bias in conversational summarization, laying the groundwork for developing more balanced and unbiased summarization models.
Lionel Delphin-Poulat, Christèle Tarnec, Anastasia Shimorina
EMNLP4
2021 A Systematic Review of Reproducibility Research in Natural Language Processing
abstract
Against the background of what has been termed a reproducibility crisis in science, the NLP field is becoming increasingly interested in, and conscientious about, the reproducibility of its results.The past few years have seen an impressive range of new initiatives, events and active research in the area.However, the field is far from reaching a consensus about how reproducibility should be defined, measured and addressed, with diversity of views currently increasing rather than converging.With this focused contribution, we aim to provide a wideangle, and as near as possible complete, snapshot of current work on reproducibility in NLP, delineating differences and similarities, and providing pointers to common denominators.
Anya Belz, Anastasia Shimorina, Ehud Reiter
EACL3
2021 The ReproGen Shared Task on Reproducibility of Human Evaluations in NLG: Overview and Results
abstract
The NLP field has recently seen a substantial increase in work related to reproducibility of results, and more generally in recognition of the importance of having shared definitions and practices relating to evaluation.Much of the work on reproducibility has so far focused on metric scores, with reproducibility of human evaluation results receiving far less attention.As part of a research programme designed to develop theory and practice of reproducibility assessment in NLP, we organised the first shared task on reproducibility of human evaluations, ReproGen 2021.This paper describes the shared task in detail, summarises results from each of the reproduction studies submitted, and provides further comparative analysis of the results.Out of nine initial team registrations, we received submissions from four teams.Meta-analysis of the four reproduction studies revealed varying degrees of reproducibility, and allowed very tentative first conclusions about what types of evaluation tend to have better reproducibility.
Anya Belz, Anastasia Shimorina, Ehud Reiter
INLG2
2021 An Error Analysis Framework for Shallow Surface Realisation
abstract
Abstract The metrics standardly used to evaluate Natural Language Generation (NLG) models, such as BLEU or METEOR, fail to provide information on which linguistic factors impact performance. Focusing on Surface Realization (SR), the task of converting an unordered dependency tree into a well-formed sentence, we propose a framework for error analysis which permits identifying which features of the input affect the models’ results. This framework consists of two main components: (i) correlation analyses between a wide range of syntactic metrics and standard performance metrics and (ii) a set of techniques to automatically identify syntactic constructs that often co-occur with low performance scores. We demonstrate the advantages of our framework by performing error analysis on the results of 174 system runs submitted to the Multilingual SR shared tasks; we show that dependency edge accuracy correlate with automatic metrics thereby providing a more interpretable basis for evaluation; and we suggest ways in which our framework could be used to improve models and data. The framework is available in the form of a toolkit which can be used both by campaign organizers to provide detailed, linguistically interpretable feedback on the state of the art in multilingual SR, and by individual researchers to improve models and datasets.1
Anastasia Shimorina, Claire Gardent, Yannick Parmentier 0001
Trans. Assoc. Comput. Linguistics1
2020 ReproGen: Proposal for a Shared Task on Reproducibility of Human Evaluations in NLG
abstract
Across NLP, a growing body of work is looking at the issue of reproducibility.However, replicability of human evaluation experiments and reproducibility of their results is currently under-addressed, and this is of particular concern for NLG where human evaluations are the norm.This paper outlines our ideas for a shared task on reproducibility of human evaluations in NLG which aims (i) to shed light on the extent to which past NLG evaluations have been replicable and reproducible, and (ii) to draw conclusions regarding how evaluations can be designed and reported to increase replicability and reproducibility.If the task is run over several years, we hope to be able to document an overall increase in levels of replicability and reproducibility over time.
Anya Belz, Anastasia Shimorina, Ehud Reiter
INLG3
2019 Surface Realisation Using Full Delexicalisation
abstract
Anastasia Shimorina, Claire Gardent. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Anastasia Shimorina, Claire Gardent
EMNLP/IJCNLP (1)1
2018 Handling Rare Items in Data-to-Text Generation
abstract
Neural approaches to data-to-text generation generally handle rare input items using either delexicalisation or a copy mechanism.We investigate the relative impact of these two methods on two datasets (E2E and WebNLG) and using two evaluation settings.We show (i) that rare items strongly impact performance; (ii) that combining delexicalisation and copying yields the strongest improvement; (iii) that copying underperforms for rare and unseen items and (iv) that the impact of these two mechanisms greatly varies depending on how the dataset is constructed and on how it is split into train, dev and test 1 .
Anastasia Shimorina, Claire Gardent
INLG1
2017 Creating Training Corpora for NLG Micro-Planners
abstract
In this paper, we present a novel framework for semi-automatically creating linguistically challenging microplanning data-to-text corpora from existing Knowledge Bases.Because our method pairs data of varying size and shape with texts ranging from simple clauses to short texts, a dataset created using this framework provides a challenging benchmark for microplanning.Another feature of this framework is that it can be applied to any large scale knowledge base and can therefore be used to train and learn KB verbalisers.We apply our framework to DBpedia data and compare the resulting dataset with Wen et al. (2016)'s.We show that while Wen et al.'s dataset is more than twice larger than ours, it is less diverse both in terms of input and in terms of text.We thus propose our corpus generation framework as a novel method for creating challenging data sets from which NLG models can be learned which are capable of handling the complex interactions occurring during in micro-planning between lexicalisation, aggregation, surface realisation, referring expression generation and sentence segmentation.To encourage researchers to take up this challenge, we recently made available a dataset created using this framework in the context of the WEBNLG shared task.
Claire Gardent, Anastasia Shimorina, Shashi Narayan, Laura Perez-Beltrachini
ACL (1)2
2017 Split and Rephrase
abstract
We propose a new sentence simplification task (Split-and-Rephrase) where the aim is to split a complex sentence into a meaning preserving sequence of shorter sentences.Like sentence simplification, splitting-and-rephrasing has the potential of benefiting both natural language processing and societal applications.Because shorter sentences are generally better processed by NLP systems, it could be used as a preprocessing step which facilitates and improves the performance of parsers, semantic role labelers and machine translation systems.It should also be of use for people with reading disabilities because it allows the conversion of longer sentences into shorter ones.This paper makes two contributions towards this new task.First, we create and make available a benchmark consisting of 1,066,115 tuples mapping a single complex sentence to a sequence of sentences expressing the same meaning.1 Second, we propose five models (vanilla sequence-to-sequence to semantically-motivated models) to understand the difficulty of the proposed task.
Shashi Narayan, Claire Gardent, Shay B. Cohen, Anastasia Shimorina
EMNLP4
2017 Mapping Natural Language to Description Logic
Bikash Gyawali, Anastasia Shimorina, Claire Gardent, Samuel Cruz-Lara, Mariem Mahfoudh
ESWC (1)2
2017 The WebNLG Challenge: Generating Text from RDF Data
abstract
The WebNLG challenge consists in mapping sets of RDF triples to text.It provides a common benchmark on which to train, evaluate and compare "microplanners", i.e. generation systems that verbalise a given content by making a range of complex interacting choices including referring expression generation, aggregation, lexicalisation, surface realisation and sentence segmentation.In this paper, we introduce the microplanning task, describe data preparation, introduce our evaluation methodology, analyse participant results and provide a brief description of the participating systems.(3) a. LOCATION-COUNTRY-STARTDATE ⇒ Passive-Apposition-Active
Claire Gardent, Anastasia Shimorina, Shashi Narayan, Laura Perez-Beltrachini
INLG2
2017 ModelWriter: text and model-synchronized document engineering platform
abstract
The ModelWriter platform provides a generic framework for automated traceability analysis. In this paper, we demonstrate how this framework can be used to trace the consistency and completeness of technical documents that consist of a set of System Installation Design Principles used by Airbus to ensure the correctness of aircraft system installation. We show in particular, how the platform allows the integration of two types of reasoning: reasoning about the meaning of text using semantic parsing and description logic theorem proving; and reasoning about document structure using first-order relational logic and finite model finding for traceability analysis.
Ferhat Erata, Claire Gardent, Bikash Gyawali, Anastasia Shimorina, Yvan Lussaud, Bedir Tekinerdogan, Geylani Kardas, Anne Monceaux
ASE4