Malvina Nissim

dblp:91/2392 · DBLP profile ↗
← Back
37ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0001-5289-0971ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 6 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Evaluating the Impact of Source Diversity for RAG in Historical Research
abstract
Historical research increasingly benefits from large language models (LLMs). However, LLMs are prone to factual inaccuracy, unreliability, and biased interpretations of data. Retrieval-augmented generation (RAG) approaches have emerged as solutions, but may inadvertently perpetuate biased perspectives embedded in historical archives. This paper investigates how source diversity in RAG impacts perspective variation in historical question answering. We compile a multilingual corpus (English, French, Dutch) of historical documents spanning multiple countries and focus on Napoleon Bonaparte. We evaluate three Qwen3 models across ten questions using a multi-layered framework combining traditional metrics (BERTScore, ROUGE-L), frame semantics analysis, and syntactic profiling. Our results highlight that, while traditional similarity metrics suggest high semantic consistency, frame-semantic analysis exposes substantial perspective shifts. Baseline answers present "flattened" cross-lingual perspectives, whereas RAG introduces diversity. Critically, this diversity manifests differently across languages, demonstrating language-specific patterns. Our findings highlight limitations of traditional evaluation metrics for perspective-sensitive tasks and demonstrate that RAG constitutes active perspective transformation rather than neutral augmentation.
Ruhi Mahadeshwar, Andreas van Cranenburgh, Tommaso Caselli, Malvina Nissim
LREC4
2026 Reading Time in the Wild: An Assessment of Readability Predictors Based on Naturally-Observed Reading Times
Sijbren Van Vaals, Rik van Noord, Malvina Nissim
LREC3
2025 When Harry Meets Superman: The Role of The Interlocutor in Persona-Based Dialogue Generation
abstract
Endowing dialogue agents with persona information has proven to significantly improve the consistency and diversity of their generations.While much focus has been placed on aligning dialogues with provided personas, the adaptation to the interlocutor's profile remains largely underexplored.In this work, we investigate three key aspects: (1) a model's ability to align responses with both the provided persona and the interlocutor's; (2) its robustness when dealing with familiar versus unfamiliar interlocutors and topics, and (3) the impact of additional fine-tuning on specific persona-based dialogues.We evaluate dialogues generated with diverse speaker pairings and topics, framing the evaluation as an author identification task and employing both LLM-as-a-judge and human evaluations.By systematically masking or disclosing information about the interlocutor, we assess its impact on dialogue generation.Results show that access to the interlocutor's persona improves the recognition of the target speaker, while masking it does the opposite.Although models generalise well across topics, they struggle with unfamiliar interlocutors.Finally, we found that in zero-shot settings, LLMs often copy biographical details, facilitating identification but trivialising the task.
Daniela Occhipinti, Marco Guerini, Malvina Nissim
ACL (1)3
2025 Can Model Uncertainty Function as a Proxy for Multiple-Choice Question Item Difficulty?
abstract
Estimating the difficulty of multiple-choice questions would be great help for educators who must spend substantial time creating and piloting stimuli for their tests, and for learners who want to practice. Supervised approaches to difficulty estimation have yielded to date mixed results. In this contribution we leverage an aspect of generative large models which might be seen as a weakness when answering questions, namely their uncertainty. Specifically, we exploit model uncertainty towards exploring correlations between two different metrics of uncertainty, and the actual student response distribution. While we observe some present but weak correlations, we also discover that the models’ behaviour is different in the case of correct vs wrong answers, and that correlations differ substantially according to the different question types which are included in our fine-grained, previously unused dataset of 451 questions from a Biopsychology course. In discussing our findings, we also suggest potential avenues to further leverage model uncertainty as an additional proxy for item difficulty.
Leonidas Zotos, Hedderik van Rijn, Malvina Nissim
COLING3
2025 Are You Doubtful? Oh, It Might Be Difficult Then! Exploring the Use of Model Uncertainty for Question Difficulty Estimation
Leonidas Zotos, Hedderik van Rijn, Malvina Nissim
EDM3
2025 Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement
abstract
Word-level quality estimation (WQE) aims to automatically identify fine-grained error spans in machine-translated outputs and has found many uses, including assisting translators during post-editing.Modern WQE techniques are often expensive, involving prompting of large language models or ad-hoc training on large amounts of human-labeled data.In this work, we investigate efficient alternatives exploiting recent advances in language model interpretability and uncertainty quantification to identify translation errors from the inner workings of translation models.In our evaluation spanning 14 metrics across 12 translation directions, we quantify the impact of human label variation on metric performance by using multiple sets of human labels.Our results highlight the untapped potential of unsupervised metrics, the shortcomings of supervised methods when faced with label uncertainty, and the brittleness of single-annotator evaluation practices.
Gabriele Sarti, Vilém Zouhar, Malvina Nissim, Arianna Bisazza
EMNLP3
2025 QE4PE: Word-level Quality Estimation for Human Post-Editing
Gabriele Sarti, Vilém Zouhar, Grzegorz Chrupala, Ana Guerberof Arenas, Malvina Nissim, Arianna Bisazza
Trans. Assoc. Comput. Linguistics5
2024 mCoT: Multilingual Instruction Tuning for Reasoning Consistency in Language Models
abstract
Large language models (LLMs) with Chainof-thought (CoT) have recently emerged as a powerful technique for eliciting reasoning to improve various downstream tasks.As most research mainly focuses on English, with few explorations in a multilingual context, the question of how reliable this reasoning capability is in different languages is still open.To address it directly, we study multilingual reasoning consistency across multiple languages, using popular open-source LLMs.First, we compile the first large-scale multilingual math reasoning dataset, mCoT-MATH, covering eleven diverse languages.Then, we introduce multilingual CoT instruction tuning to boost reasoning capability across languages, thereby improving model consistency.While existing LLMs show substantial variation across the languages we consider, and especially low performance for lesser resourced languages, our 7B parameter model mCoT achieves impressive consistency across languages, and superior or comparable performance to close-and open-source models even of much larger sizes.
Huiyuan Lai, Malvina Nissim
ACL (1)2
2024 IT5: Text-to-text Pretraining for Italian Language Understanding and Generation
abstract
We introduce IT5, the first family of encoder-decoder transformer models pretrained specifically on Italian. We document and perform a thorough cleaning procedure for a large Italian corpus and use it to pretrain four IT5 model sizes. We then introduce the ItaGen benchmark, which includes a broad range of natural language understanding and generation tasks for Italian, and use it to evaluate the performance of IT5 models and multilingual baselines. We find monolingual IT5 models to provide the best scale-to-performance ratio across tested models, consistently outperforming their multilingual counterparts and setting a new state-of-the-art for Italian language generation.
Gabriele Sarti, Malvina Nissim
LREC/COLING2
2024 Quantifying the Plausibility of Context Reliance in Neural Machine Translation
abstract
Establishing whether language models can use contextual information in a human-plausible way is important to ensure their safe adoption in real-world settings. However, the questions of $\textit{when}$ and $\textit{which parts}$ of the context affect model generations are typically tackled separately, and current plausibility evaluations are practically limited to a handful of artificial benchmarks. To address this, we introduce $\textbf{P}$lausibility $\textbf{E}$valuation of $\textbf{Co}$ntext $\textbf{Re}$liance (PECoRe), an end-to-end interpretability framework designed to quantify context usage in language models' generations. Our approach leverages model internals to (i) contrastively identify context-sensitive target tokens in generated texts and (ii) link them to contextual cues justifying their prediction. We use PECoRe to quantify the plausibility of context-aware machine translation models, comparing model rationales with human annotations across several discourse-level phenomena. Finally, we apply our method to unannotated model translations to identify context-mediated predictions and highlight instances of (im)plausible context usage throughout generation.
Gabriele Sarti, Grzegorz Chrupala, Malvina Nissim, Arianna Bisazza
ICLR3
2023 DUMB: A Dutch Model Benchmark
abstract
We introduce the Dutch Model Benchmark: DUMB.The benchmark includes a diverse set of datasets for low-, medium-and highresource tasks.The total set of nine tasks includes four tasks that were previously not available in Dutch.Instead of relying on a mean score across tasks, we propose Relative Error Reduction (RER), which compares the DUMB performance of language models to a strong baseline which can be referred to in the future even when assessing different sets of language models.Through a comparison of 14 pre-trained language models (monoand multi-lingual, of varying sizes), we assess the internal consistency of the benchmark tasks, as well as the factors that likely enable high performance.Our results indicate that current Dutch monolingual models under-perform and suggest training larger Dutch models with other architectures and pretraining objectives.At present, the highest performance is achieved by DeBERTaV3 large , XLM-R large and mDeBERTaV3 base .In addition to highlighting best strategies for training larger Dutch models, DUMB will foster further research on Dutch.A public leaderboard is available at dumbench.nl. WSD BERTje{1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 67.9 65.9 WSD RobBERT V1 {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 61.3 60.6 WSD RobBERT V2 {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 66.3 64.1 WSD RobBERT 2022 {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 67.0 63.7 WSD mBERT cased {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 68.2 68.5 WSD XLM-R base {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 66.7 66.5 WSD mDeBERTaV3 base {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 69.8 69.6 WSD XLM-R large {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 73.0 73.1 WSD BERT base {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 60.0 58.2 WSD RoBERTa base {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 62.6 61.1 WSD DeBERTaV3 base {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 66.4 64.4 WSD BERT large {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 60.2 57.2 WSD RoBERTa large {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 64.3 59.1 WSD DeBERTaV3 large {1, 3, 5, 10} {0.0, 0.3} {1e-05, 3e-05, 5e-05, 0.0001} {0.0, 0.1} 71.3 70.2
Wietse de Vries, Martijn Wieling 0001, Malvina Nissim
EMNLP3
2023 A text style transfer system for reducing the physician-patient expertise gap: An analysis with automatic and human evaluations
Luca Bacco, Felice Dell'Orletta, Huiyuan Lai, Mario Merone, Malvina Nissim
Expert Syst. Appl.5
2022 Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 Languages
abstract
Cross-lingual transfer learning with large multilingual pre-trained models can be an effective approach for low-resource languages with no labeled training data.Existing evaluations of zero-shot cross-lingual generalisability of large pre-trained models use datasets with English training data, and test data in a selection of target languages.We explore a more extensive transfer learning setup with 65 different source languages and 105 target languages for part-of-speech tagging.Through our analysis, we show that pre-training of both source and target language, as well as matching language families, writing systems, word order systems, and lexical-phonetic distance significantly impact cross-lingual performance.The findings described in this paper can be used as indicators of which factors are important for effective zero-shot cross-lingual transfer to zero-and low-resource languages.
Wietse de Vries, Martijn Wieling 0001, Malvina Nissim
ACL (1)3
2022 Multi-Figurative Language Generation
abstract
Figurative language generation is the task of reformulating a given text in the desired figure of speech while still being faithful to the original context. We take the first step towards multi-figurative language modelling by providing a benchmark for the automatic generation of five common figurative forms in English. We train mFLAG employing a scheme for multi-figurative language pre-training on top of BART, and a mechanism for injecting the target figurative information into the encoder; this enables the generation of text with the target figurative form from another figurative form without parallel figurative-figurative sentence pairs. Our approach outperforms all strong baselines. We also offer some qualitative analysis and reflections on the relationship between the different figures of speech.
Huiyuan Lai, Malvina Nissim
COLING2
2022 AGILe: The First Lemmatizer for Ancient Greek Inscriptions
abstract
To facilitate corpus searches by classicists as well as to reduce data sparsity when training models, we focus on the automatic lemmatization of ancient Greek inscriptions, which have not received as much attention in this sense as literary text data has. We show that existing lemmatizers for ancient Greek, trained on literary data, are not performant on epigraphic data, due to major language differences between the two types of texts. We thus train the first inscription-specific lemmatizer achieving above 80% accuracy, and make both the models and the lemmatized data available to the community. We also provide a detailed error analysis highlighting peculiarities of inscriptions which again highlights the importance of a lemmatizer dedicated to inscriptions.
Evelien de Graaf, Silvia Stopponi, Jasper K. Bos, Saskia Peels, Malvina Nissim
LREC5
2021 Generic resources are what you need: Style transfer tasks without task-specific parallel training data
abstract
Style transfer aims to rewrite a source text in a different target style while preserving its content.We propose a novel approach to this task that leverages generic resources, and without using any task-specific parallel (source-target) data outperforms existing unsupervised approaches on the two most popular style transfer tasks: formality transfer and polarity swap.In practice, we adopt a multistep procedure which builds on a generic pretrained sequence-to-sequence model (BART).First, we strengthen the model's ability to rewrite by further pre-training BART on both an existing collection of generic paraphrases, as well as on synthetic pairs created using a general-purpose lexical resource.Second, through an iterative back-translation approach, we train two models, each in a transfer direction, so that they can provide each other with synthetically generated pairs, dynamically in the training process.Lastly, we let our best resulting model generate static synthetic pairs to be used in a supervised training regime.Besides methodology and state-of-the-art results, a core contribution of this work is a reflection on the nature of the two tasks we address, and how their differences are highlighted by their response to our approach.
Huiyuan Lai, Antonio Toral, Malvina Nissim
EMNLP (1)3
2021 Sentiment Polarity Classification at EVALITA: Lessons Learned and Open Challenges
abstract
Sentiment analysis in social media is a popular task attracting the interest of the research community, also in recent evaluation campaigns of natural language processing tasks in several languages. We report on our experience in the organization of SENTIment POLarity Classification Task (SENTIPOLC), a shared task on sentiment classification of Italian tweets, proposed for the first time in 2014 within the Evalita evaluation campaign. We present the datasets-which include an enriched annotation scheme for dealing with the impact of figurative language on polarity-the evaluation methodology, and discuss the approaches and results of participating systems. We also offer a reflection on the open challenges of state-of-the-art systems for sentiment analysis of microblogging in Italian, as they emerge from a qualitative analysis of misclassified tweets. Finally, we provide an evaluation of the resources we have created, and share the lessons learned by running this task for two consecutive editions.
Valerio Basile, Nicole Novielli, Danilo Croce, Francesco Barbieri, Malvina Nissim, Viviana Patti
IEEE Trans. Affect. Comput.5
2020 MAGPIE: A Large Corpus of Potentially Idiomatic Expressions
abstract
Given the limited size of existing idiom corpora, we aim to enable progress in automatic idiom processing and linguistic analysis by creating the largest-to-date corpus of idioms for English. Using a fixed idiom list, automatic pre-extraction, and a strictly controlled crowdsourced annotation procedure, we show that it is feasible to build a high-quality corpus comprising more than 50K instances, an order of a magnitude larger than previous resources. Crucial ingredients of crowdsourcing were the selection of crowdworkers, clear and comprehensive instructions, and an interface that breaks down the task in small, manageable steps. Analysis of the resulting corpus revealed strong effects of genre on idiom distribution, providing new evidence for existing theories on what influences idiom usage. The corpus also contains rich metadata, and is made publicly available.
Hessel Haagsma, Johan Bos, Malvina Nissim
LREC3
2020 Invisible to People but not to Machines: Evaluation of Style-aware HeadlineGeneration in Absence of Reliable Human Judgment
abstract
We automatically generate headlines that are expected to comply with the specific styles of two different Italian newspapers. Through a data alignment strategy and different training/testing settings, we aim at decoupling content from style and preserve the latter in generation. In order to evaluate the generated headlines’ quality in terms of their specific newspaper-compliance, we devise a fine-grained evaluation strategy based on automatic classification. We observe that our models do indeed learn newspaper-specific style. Importantly, we also observe that humans aren’t reliable judges for this task, since although familiar with the newspapers, they are not able to discern their specific styles even in the original human-written headlines. The utility of automatic evaluation goes therefore beyond saving the costs and hurdles of manual annotation, and deserves particular care in its design.
Lorenzo De Mattei, Michele Cafagna, Felice Dell'Orletta, Malvina Nissim
LREC4
2020 Fair Is Better than Sensational: Man Is to Doctor as Woman Is to Doctor
abstract
Analogies such as man is to king as woman is to X are often used to illustrate the amazing power of word embeddings. Concurrently, they have also been used to expose how strongly human biases are encoded in vector spaces trained on natural language, with examples like man is to computer programmer as woman is to homemaker. Recent work has shown that analogies are in fact not an accurate diagnostic for bias, but this does not mean that they are not used anymore, or that their legacy is fading. Instead of focusing on the intrinsic problems of the analogy task as a bias detection tool, we discuss a series of issues involving implementation as well as subjective choices that might have yielded a distorted picture of bias in word embeddings. We stand by the truth that human biases are present in word embeddings, and, of course, the need to address them. But analogies are not an accurate tool to do so, and the way they have been most often used has exacerbated some possibly non-existing biases and perhaps hidden others. Because they are still widely popular, and some of them have become classics within and outside the NLP community, we deem it important to provide a series of clarifications that should put well-known, and potentially new analogies, into the right perspective.
Malvina Nissim, Rik van Noord, Rob van der Goot
Comput. Linguistics1
2019 You Write like You Eat: Stylistic Variation as a Predictor of Social Stratification
abstract
Inspired by Labov's seminal work on stylistic variation as a function of social stratification, we develop and compare neural models that predict a person's presumed socio-economic status, obtained through distant supervision, from their writing style on social media.The focus of our work is on identifying the most important stylistic parameters to predict socioeconomic group.In particular, we show the effectiveness of morpho-syntactic features as stylistic predictors of socio-economic group, in contrast to lexical features, which are good predictors of topic.
Angelo Basile, Albert Gatt, Malvina Nissim
ACL (1)3
2017 Sharing Is Caring: The Future of Shared Tasks
abstract
Shared tasks are indisputably drivers of progress and interest for problems in NLP. This is reflected by their increasing popularity, as well as by the fact that new shared tasks regularly emerge for under-researched and under-resourced topics, especially at workshops and smaller conferences.The general procedures and conventions for organizing a shared task have arisen organically over time (Paroubek, Chaudiron, and Hirschman, 2007, Section 7). There is no consistent framework that describes how shared tasks should be organized. This is not a harmful thing per se, but we believe that shared tasks, and by extension the field in general, would benefit from some reflection on the existing conventions. This, in turn, could lead to the future harmonization of shared task procedures.Shared tasks revolve around two aspects: research advancement and competition. We see research advancement as the driving force and main goal behind organizing them. Competition is an instrument to encourage and promote participation. However, just because these two forces are intrinsic to shared tasks does not mean that they always act in the same direction: Ensuring that the competition is fair is not a necessary requirement for advancing the field, and might even slow down progress.Our position in this respect is clear: We do believe that (i) advancing the field should be given priority over ensuring fair competition, also because (ii) inequality is partly unsolvable and intrinsic to life. In other words: Equality between competitors is desirable if it does not hinder research advancement.In the recently established workshop on ethics in NLP,1Parra Escartín et al. (2017) raise a set of considerations involving shared tasks, mainly focusing on areas where general ethical concerns regarding good scientific practice intersect with certain aspects of shared tasks. We find that they raise valid concerns, and in this contribution, we address some of them. However, we take a different perspective. Instead of focusing on ethical issues and potential negative effects of the competition aspect, we rather concentrate on how to bolster scientific progress.We make a simple proposal for the improvement of shared tasks, and discuss how it can help to mitigate the problems raised by Parra Escartín et al. (2017), while not necessarily tackling them directly. We start with assessing the concrete impact and significance of such concerns first.Recently, Parra Escartín et al. (2017) drew attention to a list of potential negative effects and ethical issues concerning shared tasks in NLP. In this section, we take this list as a starting point and examine each problem with respect to the main goal of shared tasks—to advance research in the field. Some issues were regarded as potential concerns rather than definite problems, because it is unclear how large their actual impact is. We believe that some of these concerns needed to be quantified in order to be properly assessed.To this end, we reviewed about 100 recent shared tasks from various campaigns (SemEval, EVALITA, CoNLL, WMT, and CLEF) between 2014 and 2016. We focused on several aspects, such as participation of companies, participation of organizers, closed versus open tracks, and the submission of papers by participants. Note that this is not an exhaustive overview of all shared tasks in NLP, but rather an arbitrary sample to investigate general trends in recent times. We use information drawn from this annotation exercise for assessing some of the problems we report in the following sections. The figures that are relevant for the discussion are reported in Table 1. The annotated spreadsheets used to collect this information are publicly available, together with some basic statistics and additional explanations.2Some potential concerns, although being ethically relevant, are not necessarily a problem in terms of research advancement, and fixing them directly should not be a priority. Here, we assess issues raised by Parra Escartín et al. (2017) that we believe fall into this category.Potential Conflicts of Interest. Parra Escartín et al. (2017) state that participation of organizers or annotators in their own shared task raises questions about inequality among participants, as organizers have earlier access to the data than the regular participants. In our survey, we found that in 5.8% of shared tasks, organizers did indeed participate. However, we also observe that this happens in connection with few participants (average 3.5 compared with 12.1, see Table 1), thus typically smaller shared tasks. This indicates that organizers' participation is more common in small, specialized tasks. The low number of participants can also explain why organizers perform better on average compared with non-organizers (see average normalized rank in Table 1).Unequal Playing Field. An unequal playing field mainly reflects the starting point that the different teams have. Parra Escartín et al. (2017) report on the issue of differences in processing power. An extreme example of this issue is the submission of Durrani et al. (2013) at WMT13, in which they reached the highest scores because they were able to boost the BLEU score by approximately 0.8% by making use of 1TB RAM, which was probably unavailable to the other teams at that time. Computational resources are not the only reason for an unequal playing field, though. There are many other causes that could lead to an unequal playing field—for example, some teams might have access to more proprietary data, proprietary software, or research equipment.The competitive nature of shared tasks can be fun, and stimulating for a variety of reasons (visibility, grant applications, beating state of the art, etc.). Such reasons might not necessarily be positively correlated with advancing the field, though. Here we discuss issues also raised by Parra Escartín et al. (2017) that we believe fall into this category.Secretiveness. As a result of the competitive nature of shared tasks, it can be desirable for participating teams to keep their “secret sauce” private, as this could mean an advantage for a re-run of the same task, or a shared task on a related problem. As a possible effect of secretiveness, Parra Escartín et al. (2017) also mention “Unconscious overlooking of ethical concerns,” actually referring to an unacceptable level of vagueness in papers. In other words, participants may unintentionally describe their systems in an abstract and vague way due to a previously established practice in systems' descriptions.Lack of Description of Negative Results. Given that negative results are informative, their under-representation in shared tasks is a concern. A lack of knowledge about negative results might lead to a research redundancy, which is clearly undesirable. The issue of under-represented negative results is a global concern for the entire field, and for science in general. However, shared tasks provide an excellent opportunity for publishing negative results, as the acceptance of papers for publication does not particularly favor positive results.Redundancy and Replicability in the Field. Parra Escartín et al. (2017) raise issues concerning two types of redundancy, (a) when optimal parameter settings of a previous shared task do not carry over to the new version of the task, therefore it is not clear what is learned; and (b) when algorithms are reimplemented for replicability purposes.Regarding (a), we think that differences in used parameter settings are not actually a bad thing; we learn from this that we overfit on the previous task, or that we need to adapt our systems to another data set or domain.Regarding (b), this is a real problem because starting from scratch to reimplement existing systems is unnecessarily time-consuming. In addition, it would always be desirable to be able to directly reproduce the same results of the same model for the same task (Pedersen, 2008; Fokkens et al., 2013).Withdrawal from Competition. Participants may withdraw from a shared task if their ranking in the competition can negatively affect their reputation and/or future funding. For example, Parra Escartín et al. (2017) suggest that companies might prefer to withdraw from the competition if they are not highly ranked, to avoid blemishing their reputation. This is something that we could not quantify in our survey, as in case of withdrawal there would be no evidence of participation in reports. There are two aspects, though, that we can quantify. The first aspect is the number of teams that do not publish their system's description, which amounts to approximately 9%, and could indeed be related to withdrawals. However, exactly because the paper is missing, information on why a team withdrew is not available. The second aspect is the total number of industry participants, which in our sample amounts to 20% (“Some company” and “Only company” in Table 1). Thus, although there is not much that can be done about withdrawal—and this might not be a problem anyway—we believe that, considering the substantial presence and interest of industries so far, their participation should be accommodated.Potentially Gaming the System. Shared tasks are usually bound to data sets and evaluation metrics. This could lead to competition-oriented participants focusing more on tuning their systems on a given data set and metrics rather than finding a scientifically sound and scalable method for solving the problem. This can, in turn, result in an “unfair” ranking or a misleading relation between a methodology and its value with respect to the research problem. A potential negative outcome of the latter is a scenario where “optimal” methods of a shared task do not carry over to related shared tasks. These problems become more severe when system gaming is combined with a secretive attitude. While tackling this issue, we should take into account that participants might be less eager to write about ad hoc solutions, for example tuning pre-processing components or tailoring a system too closely to specifics of the annotation.Our proposal for future shared tasks is not revolutionary. It simply revolves around the key aspects of sharing, not only resources but also experiences, including negative ones. Specifically, with research progress in mind, we believe sharing should be encouraged and even partially enforced. We therefore suggest an explicit setting for shared tasks in NLP, and reflect on the issue of what organizers could do in order to maximize sharing of information regarding participating systems. We also show how such a simple strategy can help to overcome the problems raised that can hinder research advancement.One of the challenges faced by shared tasks is to ensure a level playing field permitting a transparent comparison of the merits of different methods. The problem is that system A might come out on top of system B not because its method is superior, but because, for example, it was trained on more data. This would favor teams with access to more resources, like companies with large quantities of proprietary in-house data.Traditionally, this problem has been mitigated by establishing “closed tracks.” In closed tracks, participating teams are not allowed to use any training data other than that provided by the shared task organizers. The rationale behind this is that if all systems use exactly the same data, the playing field is equal, and the competition results will show the strengths of the different methods. In order to study the effect of additional training data, many shared tasks have a separate competition, the so-called “open track.”However, this open–closed division is increasingly impractical and ineffective. The main problem is that it is only concerned with training data, whereas the performance of systems can crucially depend on other resources. Examples are external components with pre-trained models, such as part-of-speech taggers and dependency parsers, auxiliary data-derived resources like word embeddings, and other influential factors like the availability of computational resources. Because such models are almost always derived from external language data, it is unclear where to draw the line between closed and open. Should such data be disallowed or not? If not, teams still do not really participate on an equal footing.From a research perspective, banning external models is completely impractical and nonsensical, as most state-of-the-art systems now depend on them. Likewise, trying to force all teams to use the same set of external models, and no other, would place a heavy burden on both organizers and participants. Moreover, restricting the resources participants can use is questionable, because, for research to progress quickly, teams should use the best resources available, or the resources best fitting their system.It is therefore unsurprising that the use of closed tracks has declined in shared tasks in general in the last few years, as we have observed during our review of shared tasks. However, the original problem of unequal playing field, and thus a bias in favor of teams with ample resources, remains.We propose, then, to rethink the problem, not in terms of equal training data, but in terms of equal opportunities. This is closely connected to the wider issue of reproducibility and replicability: Like all published research, shared task results should ideally be fully reproducible by anyone (Pedersen, 2008; Fokkens et al., 2013).3 Moreover, it should be easy to build on others' work to try out new variations of a method, without having to reimplement things from scratch. To ensure this, it is desirable that everything needed to reproduce experimental results is publicly and freely available, including code, data, pre-trained models, and so on. Interestingly, at the CoNLL-2013 shared task a similar step was taken, but only in terms of pre-condition: “While all teams in the shared task use the NUCLE corpus, they are also allowed to use additional external resources (both corpora and tools) so long as they are publicly available and not proprietary” (Ng et al., 2013). We would like to take this a step further, by enforcing the sharing of whatever resource teams might choose to use, so as to favor the injection of new resources in the field.Applying this principle to shared tasks in practice, we propose making the primary competition a “public track,” where participants can use any code, data, and pre-trained models they want, as long as others can then freely obtain them. In other words: All resources used to participate in the shared task should be subsequently shared with the community. Although this does not ensure equal access to resources for the current edition, it will still ensure a progressively more equal footing for the future. We believe this is the crucial step to move the field forward, as everyone will have access to the resources used in state-of-the-art systems. To keep participation possible for teams who cannot or will not make all resources available, a secondary, “proprietary track” can be established.Ranking of systems forms a large part of the appeal of shared tasks. However, rankings should not be overemphasized and are far from being the final goal of shared tasks. Research is supposed to teach us about the merits and characteristics of methods, including insights of what does not work, rather than about which team built the system that performed best on the test data.Negative results are very informative for future developments. Although publishing negative results is difficult, shared tasks do provide the ideal context for disclosing and explaining low performance methods and choices. We believe that shared task organizers should explicitly and strongly solicit the inclusion of what did not work in the reports written by participating teams. This could be even solicited via an online form that participants submit after the evaluation phase, where they comment on what worked well (as commonly done), but also provides a separate section to explain what did not work. This information could in turn be valuable data for organizers when compiling the overview report. Moreover, a clear explanation of what did not work, in connection with availability of code, would help to better understand whether something does not work as an idea or because of a specific implementation.More generally, organizers should encourage—and to some point ensure through the reviewing process—that all participants provide exhaustive reports, potentially including ablation/addition tests, so as to have a picture as comprehensive as possible. Because participating in shared tasks directly implies getting a paper accepted for publication, not everyone describes their system to the satisfaction of external reviewers. This should change, and acceptance should be conditional on clarity and exhaustiveness.We stated that the goal of advancing research should be prioritized over competition. This is especially the case when focusing on the competition aspect would encourage undesired practices like secretiveness and gaming the system. We suggest a simple solution based on the principle of maximizing resource- and information-sharing. As a byproduct, some problematic competition-related issues will be overcome, too. Some outstanding ethical issues cannot be solved, as they are intrinsic in human nature and cannot be controlled for by means of specific guidelines.Introducing proprietary and public tracks will stimulate participants to release their systems and resources. This will directly reduce Secretiveness and the issue of Redundancy and Replicability in the Field. It will also partially address the Unequal Playing Field problem, at least in the long run: Even if at the same competition different teams will have access to different resources, all resources will be available to everyone for the next round. Moreover, being able to access and run systems on different data sets will uncover limitations that might have been due to tailoring systems to the specifics of a given shared task (Potential Gaming the System). The presence of a proprietary track still allows for industrial participation (see Withdrawal from Competition in Section 2.2), where distribution of resources might not be as easy as for other teams.Encouraging participants to write comprehensive reports that include negative results will be a valid instrument towards advancing research, at the same time solving some outstanding problems. The reviewers should probably spend extra time in assessing the single reports and accept them conditionally on clarity requirements, but we believe this is worth the effort. Indeed, enforcing that systems are described properly will ensure and the of there are some issues We believe these are issues that cannot or need not be Withdrawal of teams cannot be controlled if of negative results is encouraged and common practice, it is possible that teams will choose to the competition. on progress and sharing rather than will also The of interest is not relevant in our We do believe that organizers should be allowed to and this is especially for shared tasks that might a number of participating teams due to the nature of the As long as these are explicitly reported in both the overview paper and the system the of results and ranking is to the closed tracks also an equal playing will it make it to different methods over the same We do not think Equality will be increasingly by resource sharing, to the that it is as inequality is part of the As we in Section comparison of methods has not been transparent in closed tracks the use of resources is not clarity in reports and release of systems will make it possible for the to assess which methods work and which do not, and to progressively on the state of the that our and will discussion on shared tasks, to make them more for driving progress in and each task will to have their own settings that the organizers will most However, we do believe that participants to release their in terms of resources, and of and should be a common to from long we together and on various with many It is to everyone for their However, we to mention a few who have to into better The discussion we at of the of the Computational at the of was the actual for this are to everyone and in to and for their and valuable has and to the on what does not work with the open versus closed track setting as it We and Parra Escartín for on earlier of this We are also to for
Malvina Nissim, Lasha Abzianidze, Kilian Evang, Rob van der Goot, Hessel Haagsma, Barbara Plank, Martijn Wieling 0001
Comput. Linguistics1
2016 Leveraging Native Data to Correct Preposition Errors in Learners' Dutch
Lennart Kloppenburg, Malvina Nissim
LREC2
2015 Adding Semantics to Data-Driven Paraphrasing
abstract
Ellie Pavlick, Johan Bos, Malvina Nissim, Charley Beller, Benjamin Van Durme, Chris Callison-Burch. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Ellie Pavlick, Johan Bos, Malvina Nissim, Charley Beller, Benjamin Van Durme, Chris Callison-Burch
ACL (1)3
2010 Creation of Lexical Resources for a Characterisation of Multiword Expressions in Italian
Andrea Zaninello, Malvina Nissim
LREC2
2008 The Italian Particle "ne": Corpus Construction and Analysis
Malvina Nissim, Sara Perboni
LREC1
2006 An Empirical Approach to the Interpretation of Superlatives
Johan Bos, Malvina Nissim
EMNLP2
2006 Learning Information Status of Discourse Entities
Malvina Nissim
EMNLP1
2006 The Impact of Annotation on the Performance of Protein Tagging in Biomedical Text
Beatrice Alex, Malvina Nissim, Claire Grover
LREC2
2005 Exploring the boundaries: gene and protein identification in biomedical text
abstract
BACKGROUND: Good automatic information extraction tools offer hope for automatic processing of the exploding biomedical literature, and successful named entity recognition is a key component for such tools. METHODS: We present a maximum-entropy based system incorporating a diverse set of features for identifying gene and protein names in biomedical abstracts. RESULTS: This system was entered in the BioCreative comparative evaluation and achieved a precision of 0.83 and recall of 0.84 in the "open" evaluation and a precision of 0.78 and recall of 0.85 in the "closed" evaluation. CONCLUSION: Central contributions are rich use of features derived from the training data at multiple levels of granularity, a focus on correctly identifying entity boundaries, and the innovative use of several external knowledge sources including full MEDLINE abstracts and web searches.
Jenny Rose Finkel, Shipra Dingare, Christopher D. Manning, Malvina Nissim, Beatrice Alex, Claire Grover
BMC Bioinform.4
2005 Comparing Knowledge Sources for Nominal Anaphora Resolution
abstract
We compare two ways of obtaining lexical knowledge for antecedent selection in other-anaphora and definite noun phrase coreference. Specifically, we compare an algorithm that relies on links encoded in the manually created lexical hierarchy WordNet and an algorithm that mines corpora by means of shallow lexico-semantic patterns. As corpora we use the British National Corpus (BNC), as well as the Web, which has not been previously used for this task. Our results show that (a) the knowledge encoded in WordNet is often insufficient, especially for anaphor' antecedent relations that exploit subjective or context-dependent knowledge; (b) for other-anaphora, the Web-based method outperforms the WordNet-based method; (c) for definite NP coreference, the Web-based method yields results comparable to those obtained using WordNet over the whole data set and outperforms the WordNet-based method on subsets of the data set; (d) in both case studies, the BNC-based method is worse than the other methods because of data sparseness. Thus, in our studies, the Web-based method alleviated the lexical knowledge gap often encountered in anaphora resolution and handled examples with context-dependent relations between anaphor and antecedent. Because it is inexpensive and needs no hand-modeling of lexical knowledge, it is a promising knowledge source to integrate into anaphora resolution systems.
Katja Markert, Malvina Nissim
Comput. Linguistics2
2004 Using the NITE XML Toolkit on the Switchboard Corpus to Study Syntactic Choice: a Case Study
Jean Carletta, Shipra Dingare, Malvina Nissim, Tatiana Nikitina
LREC3
2004 An Annotation Scheme for Information Status in Dialogue
Malvina Nissim, Shipra Dingare, Jean Carletta, Mark Steedman
LREC1
2003 Syntactic Features and Word Similarity for Supervised Metonymy Resolution
abstract
We present a supervised machine learning algorithm for metonymy resolution, which exploits the similarity between examples of conventional metonymy. We show that syntactic head-modifier relations are a high precision feature for metonymy recognition but suffer from data sparseness. We partially overcome this problem by integrating a thesaurus and introducing simpler grammatical features, thereby preserving precision and increasing recall. Our algorithm generalises over two levels of contextual similarity. Resulting inferences exceed the complexity of inferences undertaken in word sense disambiguation. We also compare automatic and manual methods for syntactic feature extraction.
Malvina Nissim, Katja Markert
ACL1
2003 Using the Web in Machine Learning for Other-Anaphora Resolution
Natalia N. Modjeska, Katja Markert, Malvina Nissim
EMNLP3
2002 Metonymy Resolution as a Classification Task
abstract
We reformulate metonymy resolution as a classification task. This is motivated by the regularity of metonymic readings and makes general classification and word sense disambiguation methods available for metonymy resolution. We then present a case study for location names, presenting both a corpus of location names annotated for metonymy as well as experiments with a supervised classification algorithm on this corpus. We especially explore the contribution of features used in word sense disambiguation to metonymy resolution.
Katja Markert, Malvina Nissim
EMNLP2
2002 Towards a Corpus Annotated for Metonymies: the Case of Location Names
Katja Markert, Malvina Nissim
LREC2