Ilia Kuznetsov

dblp:199/2048 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 5 first-author · 9 since 2021
YearPublicationVenuePosition
2025 STRICTA: Structured Reasoning in Critical Text Assessment for Peer Review and Beyond
abstract
Nils Dycke, Matej Zečević, Ilia Kuznetsov, Beatrix Suess, Kristian Kersting, Iryna Gurevych. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Nils Dycke, Matej Zecevic, Ilia Kuznetsov, Beatrix Suess, Kristian Kersting, Iryna Gurevych
ACL (1)3
2024 Re3: A Holistic Framework and Dataset for Modeling Collaborative Document Revision
abstract
Collaborative review and revision of textual documents is the core of knowledge work and a promising target for empirical analysis and NLP assistance.Yet, a holistic framework that would allow modeling complex relationships between document revisions, reviews and author responses is lacking.To address this gap, we introduce Re3, a framework for joint analysis of collaborative document revision.We instantiate this framework in the scholarly domain, and present Re3-Sci, a large corpus of aligned scientific paper revisions manually labeled according to their action and intent, and supplemented with the respective peer reviews and human-written edit summaries.We use the new data to provide first empirical insights into collaborative document revision in the academic domain, and to assess the capabilities of state-of-the-art LLMs at automating edit analysis and facilitating text-based collaboration.We make our annotation environment and protocols, the resulting data and experimental code publicly available.1
Qian Ruan, Ilia Kuznetsov, Iryna Gurevych
ACL (1)2
2024 Systematic Task Exploration with LLMs: A Study in Citation Text Generation
abstract
Large language models (LLMs) bring unprecedented flexibility in defining and executing complex, creative natural language generation (NLG) tasks.Yet, this flexibility brings new challenges, as it introduces new degrees of freedom in formulating the task inputs and instructions and in evaluating model performance.To facilitate the exploration of creative NLG tasks, we propose a three-component research framework that consists of systematic input manipulation, reference data, and output measurement.We use this framework to explore citation text generation -a popular scholarly NLP task that lacks consensus on the task definition and evaluation metric and has not yet been tackled within the LLM paradigm.Our results highlight the importance of systematically investigating both task instruction and input configuration when prompting LLMs, and reveal non-trivial relationships between different evaluation metrics used for citation text generation.Additional human generation and human evaluation experiments provide new qualitative insights into the task to guide future research in citation text generation.We make our code 1 and data 2 publicly available.
Furkan Sahinuç, Ilia Kuznetsov, Yufang Hou 0001, Iryna Gurevych
ACL (1)2
2024 Document Structure in Long Document Transformers
abstract
Jan Buchmann, Max Eichler, Jan-Micha Bodensohn, Ilia Kuznetsov, Iryna Gurevych. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Jan Buchmann, Max Eichler, Jan-Micha Bodensohn, Ilia Kuznetsov, Iryna Gurevych
EACL (1)4
2024 Are Large Language Models Good Classifiers? A Study on Edit Intent Classification in Scientific Document Revisions
abstract
Classification is a core NLP task architecture with many potential applications.While large language models (LLMs) have brought substantial advancements in text generation, their potential for enhancing classification tasks remains underexplored.To address this gap, we propose a framework for thoroughly investigating fine-tuning LLMs for classification, including both generation-and encoding-based approaches.We instantiate this framework in edit intent classification (EIC), a challenging and underexplored classification task.Our extensive experiments and systematic comparisons with various training approaches and a representative selection of LLMs yield new insights into their application for EIC.We investigate the generalizability of these findings on five further classification tasks.To demonstrate the proposed methods and address the data shortage for empirical edit analysis, we use our bestperforming EIC model to create Re3-Sci2.0,a new large-scale dataset of 1,780 scientific document revisions with over 94k labeled edits.The quality of the dataset is assessed through human evaluation.The new dataset enables an in-depth empirical study of human editing behavior in academic writing.We make our experimental framework 1 , models and data 2 publicly available.
Qian Ruan, Ilia Kuznetsov, Iryna Gurevych
EMNLP2
2023 NLPeer: A Unified Resource for the Computational Study of Peer Review
abstract
Peer review constitutes a core component of scholarly publishing; yet it demands substantial expertise and training, and is susceptible to errors and biases.Various applications of NLP for peer reviewing assistance aim to support reviewers in this complex process, but the lack of clearly licensed datasets and multi-domain corpora prevent the systematic study of NLP for peer review.To remedy this, we introduce NLPEER -the first ethically sourced multidomain corpus of more than 5k papers and 11k review reports from five different venues.In addition to the new datasets of paper drafts, cameraready versions and peer reviews from the NLP community, we establish a unified data representation and augment previous peer review datasets to include parsed and structured paper representations, rich metadata and versioning information.We complement our resource with implementations and analysis of three reviewing assistance tasks, including a novel guided skimming task.Our work paves the path towards systematic, multi-faceted, evidencebased study of peer review in NLP and beyond.The data 1 and code 2 are publicly available.
Nils Dycke, Ilia Kuznetsov, Iryna Gurevych
ACL (1)2
2023 An Inclusive Notion of Text
abstract
Natural language processing (NLP) researchers develop models of grammar, meaning and communication based on written text.Due to task and data differences, what is considered text can vary substantially across studies.A conceptual framework for systematically capturing these differences is lacking.We argue that clarity on the notion of text is crucial for reproducible and generalizable NLP.Towards that goal, we propose common terminology to discuss the production and transformation of textual data, and introduce a two-tier taxonomy of linguistic and non-linguistic elements that are available in textual sources and can be used in NLP modeling.We apply this taxonomy to survey existing work that extends the notion of text beyond the conservative language-centered view.We outline key desiderata and challenges of the emerging inclusive approach to text in NLP, and suggest community-level reporting as a crucial next step to consolidate the discussion.
Ilia Kuznetsov, Iryna Gurevych
ACL (1)1
2023 CiteBench: A Benchmark for Scientific Citation Text Generation
abstract
Science progresses by building upon the prior body of knowledge documented in scientific publications.The acceleration of research makes it hard to stay up-to-date with the recent developments and to summarize the evergrowing body of prior work.To address this, the task of citation text generation aims to produce accurate textual summaries given a set of papers-to-cite and the citing paper context.Due to otherwise rare explicit anchoring of cited documents in the citing paper, citation text generation provides an excellent opportunity to study how humans aggregate and synthesize textual knowledge from sources.Yet, existing studies are based upon widely diverging task definitions, which makes it hard to study this task systematically.To address this challenge, we propose CITEBENCH: a benchmark for citation text generation that unifies multiple diverse datasets and enables standardized evaluation of citation text generation models across task designs and domains.Using the new benchmark, we investigate the performance of multiple strong baselines, test their transferability between the datasets, and deliver new insights into the task definition and evaluation to guide future research in citation text generation.We make the code for CITEBENCH publicly available at https://github.com/ UKPLab/citebench.
Martin Funkquist, Ilia Kuznetsov, Yufang Hou 0001, Iryna Gurevych
EMNLP2
2022 Revise and Resubmit: An Intertextual Model of Text-based Collaboration in Peer Review
abstract
Abstract Peer review is a key component of the publishing process in most fields of science. Increasing submission rates put a strain on reviewing quality and efficiency, motivating the development of applications to support the reviewing and editorial work. While existing NLP studies focus on the analysis of individual texts, editorial assistance often requires modeling interactions between pairs of texts—yet general frameworks and datasets to support this scenario are missing. Relationships between texts are the core object of the intertextuality theory—a family of approaches in literary studies not yet operationalized in NLP. Inspired by prior theoretical work, we propose the first intertextual model of text-based collaboration, which encompasses three major phenomena that make up a full iteration of the review–revise–and–resubmit cycle: pragmatic tagging, linking, and long-document version alignment. While peer review is used across the fields of science and publication formats, existing datasets solely focus on conference-style review in computer science. Addressing this, we instantiate our proposed model in the first annotated multidomain corpus in journal-style post-publication open peer review, and provide detailed insights into the practical aspects of intertextual annotation. Our resource is a major step toward multidomain, fine-grained applications of NLP in editorial support for peer review, and our intertextual framework paves the path for general-purpose modeling of text-based collaboration. We make our corpus, detailed annotation guidelines, and accompanying code publicly available.1
Ilia Kuznetsov, Jan Buchmann, Max Eichler, Iryna Gurevych
Comput. Linguistics1
2020 A matter of framing: The impact of linguistic formalism on probing results
abstract
Deep pre-trained contextualized encoders like BERT (Devlin et al., 2019) demonstrate remarkable performance on a range of downstream tasks.A recent line of research in probing investigates the linguistic knowledge implicitly learned by these models during pretraining.While most work in probing operates on the task level, linguistic tasks are rarely uniform and can be represented in a variety of formalisms.Any linguistics-based probing study thereby inevitably commits to the formalism used to annotate the underlying data.Can the choice of formalism affect probing results?To investigate, we conduct an in-depth cross-formalism layer probing study in role semantics.We find linguistically meaningful differences in the encoding of semantic role-and proto-role information by BERT depending on the formalism and demonstrate that layer probing can detect subtle differences between the implementations of the same linguistic formalism.Our results suggest that linguistic formalism is an important dimension in probing studies and should be investigated along with the commonly used cross-task and cross-lingual experimental settings.
Ilia Kuznetsov, Iryna Gurevych
EMNLP (1)1
2020 LINSPECTOR: Multilingual Probing Tasks for Word Representations
abstract
Despite an ever-growing number of word representation models introduced for a large number of languages, there is a lack of a standardized technique to provide insights into what is captured by these models. Such insights would help the community to get an estimate of the downstream task performance, as well as to design more informed neural architectures, while avoiding extensive experimentation that requires substantial computational resources not all researchers have access to. A recent development in NLP is to use simple classification tasks, also called probing tasks, that test for a single linguistic feature such as part-of-speech. Existing studies mostly focus on exploring the linguistic information encoded by the continuous representations of English text. However, from a typological perspective the morphologically poor English is rather an outlier: The information encoded by the word order and function words in English is often stored on a subword, morphological level in other languages. To address this, we introduce 15 type-level probing tasks such as case marking, possession, word length, morphological tag count, and pseudoword identification for 24 languages. We present a reusable methodology for creation and evaluation of such tests in a multilingual setting, which is challenging because of a lack of resources, lower quality of tools, and differences among languages. We then present experiments on several diverse multilingual word embedding models, in which we relate the probing task performance for a diverse set of languages to a range of five classic NLP tasks: POS-tagging, dependency parsing, semantic role labeling, named entity recognition, and natural language inference. We find that a number of probing tests have significantly high positive correlation to the downstream tasks, especially for morphologically rich languages. We show that our tests can be used to explore word embeddings or black-box neural models for linguistic cues in a multilingual setting. We release the probing data sets and the evaluation suite LINSPECTOR with https://github.com/UKPLab/linspector .
Gözde Gül Sahin, Clara Vania, Ilia Kuznetsov, Iryna Gurevych
Comput. Linguistics3
2018 From Text to Lexicon: Bridging the Gap between Word Embeddings and Lexical Resources
abstract
Distributional word representations (often referred to as word embeddings) are omnipresent in modern NLP. Early work has focused on building representations for word types, and recent studies show that lemmatization and part of speech (POS) disambiguation of targets in isolation improve the performance of word embeddings on a range of downstream tasks. However, the reasons behind these improvements, the qualitative effects of these operations and the combined performance of lemmatized and POS disambiguated targets are less studied. This work aims to close this gap and puts previous findings into a general perspective. We examine the effect of lemmatization and POS typing on word embedding performance in a novel resource-based evaluation scenario, as well as on standard similarity benchmarks. We show that these two operations have complimentary qualitative and vocabulary-level effects and are best used in combination. We find that the improvement is more pronounced for verbs and show how lemmatization and POS typing implicitly target some of the verb-specific issues. We claim that the observed improvement is a result of better conceptual alignment between word embeddings and lexical resources, stressing the need for conceptually plausible modeling of word embedding targets.
Ilia Kuznetsov, Iryna Gurevych
COLING1
2018 Corpus-Driven Thematic Hierarchy Induction
abstract
Thematic role hierarchy is a linguistic tool used to describe interactions between semantic roles and their syntactic realizations.Despite decades of dedicated research and numerous thematic hierarchy suggestions in the literature, this concept has not been used in NLP so far due to incompatibility and limited scope of existing hierarchies.We introduce an empirical framework for thematic hierarchy induction and evaluate several role ranking strategies on English and German corpus data.We hypothesize that inducing a thematic hierarchy is feasible, that a hierarchy can be induced from small amounts of data and that resulting hierarchies apply cross-lingually.We evaluate these assumptions empirically.
Ilia Kuznetsov, Iryna Gurevych
CoNLL1
2017 Out-of-domain FrameNet Semantic Role Labeling
abstract
Silvana Hartmann, Ilia Kuznetsov, Teresa Martin, Iryna Gurevych. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Silvana Hartmann, Ilia Kuznetsov, Teresa Martin, Iryna Gurevych
EACL (1)2