VLDB 2026 Research / reviewers in the wild / expert
Nicola Cancedda
dblp:19/2610
· DBLP profile ↗
30ranked-venue papers
4as first author
12since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HalluLens: LLM Hallucination BenchmarkabstractYejin Bang, Ziwei Ji, Alan Schelten, Anthony Hartshorn, Tara Fowler, Cheng Zhang, Nicola Cancedda, Pascale Fung. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yejin Bang, Ziwei Ji 0001, Alan Schelten, Anthony Hartshorn, Tara Fowler, Nicola Cancedda, Pascale Fung |
ACL (1) | 7 |
| 2025 | Calibrating Verbal Uncertainty as a Linear Feature to Reduce HallucinationsabstractZiwei Ji, Lei Yu, Yeskendir Koishekenov, Yejin Bang, Anthony Hartshorn, Alan Schelten, Cheng Zhang, Pascale Fung, Nicola Cancedda. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ziwei Ji 0001, Yeskendir Koishekenov, Yejin Bang, Anthony Hartshorn, Alan Schelten, Pascale Fung, Nicola Cancedda |
EMNLP | 9 |
| 2025 | Robust LLM safeguarding via refusal feature adversarial trainingabstractLarge language models (LLMs) are vulnerable to adversarial attacks that can elicit harmful responses. Defending against such attacks remains challenging due to the opacity of jailbreaking mechanisms and the high computational cost of training LLMs robustly. We demonstrate that adversarial attacks share a universal mechanism for circumventing LLM safeguards that works by ablating a dimension in the residual stream embedding space called the refusal feature. We further show that the operation of refusal feature ablation (RFA) approximates the worst-case perturbation of offsetting model safety. Based on these findings, we propose Refusal Feature Adversarial Training (ReFAT), a novel algorithm that efficiently performs LLM adversarial training by simulating the effect of input-level attacks via RFA. Experiment results show that ReFAT significantly improves the robustness of three popular LLMs against a wide range of adversarial attacks, with considerably less computational overhead compared to existing adversarial training methods. Virginie Do, Karen Hambardzumyan, Nicola Cancedda |
ICLR | 4 |
| 2025 | Combining Code Generating Large Language Models and Self-Play to Iteratively Refine Strategies in GamesabstractWe propose a self-play approach to generating strategies for playing in multi-player games, where strategies are represented as computer code. We use large language models (LLMs) to generate pieces of code to play in the game, which we refer to as generated bots. We engage the LLM generated bots in competitions, designed to generate increasingly stronger strategies. We follow game theoretic principles in organizing these tournaments, and use a Policy Space Response Oracle (PSRO) approach. We start with an initial set of LLM generated bots, and continue in rounds for adding new bots into the population. Each round adds a bot to the population by asking the LLM to produce code for playing against a bot representing the Nash equilibrium mixture over the current population. Our analysis shows that even a few rounds are sufficient to produces strong bots for playing the game. Our demo shows the process for the game of Checkers. We allow users to select initial bots in the population, run the process, inspect how the bots evolve over time, and play against the generated bots. Yoram Bachrach, Edan Toledo, Karen Hambardzumyan, Despoina Magka, Martin Josifoski, Minqi Jiang, Jakob N. Foerster, Roberta Raileanu, Tatiana Shavrina, Nicola Cancedda, Avraham Ruderman, Katie Millican, Andrei Lupu, Rishi Hazra |
IJCAI | 10 |
| 2025 | LLM Unlearning via Neural Activation RedirectionabstractThe ability to selectively remove knowledge from LLMs is highly desirable. However, existing methods often struggle with balancing unlearning efficacy and retain model utility, and lack controllability at inference time to emulate base model behavior as if it had never seen the unlearned data. In this paper, we propose LUNAR, a novel unlearning method grounded in the Linear Representation Hypothesis and operates by redirecting the representations of unlearned data to activation regions that expresses its inability to answer. We show that contrastive features are not a prerequisite for effective activation redirection, and LUNAR achieves state-of-the-art unlearning performance and superior controllability. Specifically, LUNAR achieves between 2.9x and 11.7x improvement in the combined unlearning efficacy and model utility score (Deviation Score) across various base models and generates coherent, contextually appropriate responses post-unlearning. Moreover, LUNAR effectively reduces parameter updates to a single down-projection matrix, a novel design that significantly enhances efficiency by 20x and robustness. Finally, we demonstrate that LUNAR is robust to white-box adversarial attacks and versatile in real-world scenarios, including handling sequential unlearning requests. William F. Shen, Xinchi Qiu, Meghdad Kurmanji, Alex Iacob, Lorenzo Sani, Nicola Cancedda, Nicholas D. Lane |
NeurIPS | 7 |
| 2025 | AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-benchabstractAI research agents are demonstrating great potential to accelerate scientific progress by automating the design, implementation, and training of machine learning models. We focus on methods for improving agents' performance on MLE-bench, a challenging benchmark where agents compete in Kaggle competitions to solve real-world machine learning problems. We formalize AI research agents as search policies that navigate a space of candidate solutions, iteratively modifying them using operators. By designing and systematically varying different operator sets and search policies (Greedy, MCTS, Evolutionary), we show that their interplay is critical for achieving high performance. Our best pairing of search strategy and operator set achieves a state-of-the-art result on MLE-bench lite, increasing the success rate of achieving a Kaggle medal from 39.6% to 47.7%. Our investigation underscores the importance of jointly considering the search strategy, operator design, and evaluation methodology in advancing automated machine learning. Edan Toledo, Karen Hambardzumyan, Martin Josifoski, Rishi Hazra, Nicolas Mario Baldwin, Alexis Audran-Reiss, Michael Kuchnik, Despoina Magka, Minqi Jiang, Alisia Maria Lupidi, Andrei Lupu, Roberta Raileanu, Tatiana Shavrina, Kelvin Niu, Jean-Christophe Gagnon-Audet, Michael Shvartsman, Shagun Sodhani, Alexander H. Miller, Abhishek Charnalia, Derek Dunfield, Carole-Jean Wu, Pontus Stenetorp, Nicola Cancedda, Jakob N. Foerster, Yoram Bachrach |
NeurIPS | 23 |
| 2024 | Spectral Filters, Dark Signals, and Attention SinksabstractProjecting intermediate representations onto the vocabulary is an increasingly popular interpretation tool for transformer-based LLMs, also known as the logit lens (Nostalgebraist).We propose a quantitative extension to this approach and define spectral filters on intermediate representations based on partitioning the singular vectors of the vocabulary embedding and unembedding matrices into bands.We find that the signals exchanged in the tail end of the spectrum, i.e. corresponding to the singular vectors with smallest singular values, are responsible for attention sinking (Xiao et al., 2023), of which we provide an explanation.We find that the negative log-likelihood of pretrained models can be kept low despite suppressing sizeable parts of the embedding spectrum in a layer-dependent way, as long as attention sinking is preserved.Finally, we discover that the representation of tokens that draw attention from many tokens have large projections on the tail end of the spectrum, and likely act as additional attention sinks. Nicola Cancedda |
ACL (1) | 1 |
| 2024 | Know When To Stop: A Study of Semantic Drift in Text GenerationabstractAva Spataru, Eric Hambro, Elena Voita, Nicola Cancedda. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Ava Spataru, Eric Hambro, Elena Voita, Nicola Cancedda |
NAACL-HLT | 4 |
| 2023 | Polar Ducks and Where to Find Them: Enhancing Entity Linking with Duck Typing and Polar Box EmbeddingsabstractMattia Atzeni, Mikhail Plekhanov, Frederic Dreyer, Nora Kassner, Simone Merello, Louis Martin, Nicola Cancedda. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Mattia Atzeni, Mikhail Plekhanov, Frédéric A. Dreyer, Nora Kassner, Simone Merello, Louis Martin, Nicola Cancedda |
EMNLP | 7 |
| 2023 | Toolformer: Language Models Can Teach Themselves to Use ToolsabstractLanguage models (LMs) exhibit remarkable abilities to solve new tasks from just a few examples or textual instructions, especially at scale. They also, paradoxically, struggle with basic functionality, such as arithmetic or factual lookup, where much simpler and smaller specialized models excel. In this paper, we show that LMs can teach themselves to *use external tools* via simple APIs and achieve the best of both worlds. We introduce *Toolformer*, a model trained to decide which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction. This is done in a self-supervised way, requiring nothing more than a handful of demonstrations for each API. We incorporate a range of tools, including a calculator, a Q&A system, a search engine, a translation system, and a calendar. Toolformer achieves substantially improved zero-shot performance across a variety of downstream tasks, often competitive with much larger models, without sacrificing its core language modeling abilities. Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, Thomas Scialom |
NeurIPS | 8 |
| 2022 | EDIN: An End-to-end Benchmark and Pipeline for Unknown Entity Discovery and IndexingabstractExisting work on Entity Linking mostly assumes that the reference knowledge base is complete, and therefore all mentions can be linked.In practice this is hardly ever the case, as knowledge bases are incomplete and because novel concepts arise constantly.We introduce the temporally segmented Unknown Entity Discovery and Indexing (EDIN) -benchmark where unknown entities, that is entities not part of the knowledge base and without descriptions and labeled mentions, have to be integrated into an existing entity linking system.By contrasting EDIN with zero-shot entity linking, we provide insight on the additional challenges it poses.Building on denseretrieval based entity linking, we introduce the end-to-end EDIN-pipeline that detects, clusters, and indexes mentions of unknown entities in context.Experiments show that indexing a single embedding per entity unifying the information of multiple mentions works better than indexing mentions independently. Nora Kassner, Fabio Petroni, Mikhail Plekhanov, Sebastian Riedel 0001, Nicola Cancedda |
EMNLP | 5 |
| 2022 | Multilingual Autoregressive Entity LinkingabstractAbstract We present mGENRE, a sequence-to- sequence system for the Multilingual Entity Linking (MEL) problem—the task of resolving language-specific mentions to a multilingual Knowledge Base (KB). For a mention in a given language, mGENRE predicts the name of the target entity left-to-right, token-by-token in an autoregressive fashion. The autoregressive formulation allows us to effectively cross-encode mention string and entity names to capture more interactions than the standard dot product between mention and entity vectors. It also enables fast search within a large KB even for mentions that do not appear in mention tables and with no need for large-scale vector indices. While prior MEL works use a single representation for each entity, we match against entity names of as many languages as possible, which allows exploiting language connections between source input and target name. Moreover, in a zero-shot setting on languages with no training data at all, mGENRE treats the target language as a latent variable that is marginalized at prediction time. This leads to over 50% improvements in average accuracy. We show the efficacy of our approach through extensive evaluation including experiments on three popular MEL benchmarks where we establish new state-of-the-art results. Source code available at https://github.com/facebookresearch/GENRE. Nicola De Cao, Ledell Wu, Kashyap Popat, Mikel Artetxe, Naman Goyal 0001, Mikhail Plekhanov, Luke Zettlemoyer, Nicola Cancedda, Sebastian Riedel 0001, Fabio Petroni |
Trans. Assoc. Comput. Linguistics | 8 |
| 2017 | Reply With: Proactive Recommendation of Email AttachmentsabstractEmail responses often contain items---such as a file or a hyperlink to an external document---that are attached to or included inline in the body of the message. Analysis of an enterprise email corpus reveals that 35% of the time when users include these items as part of their response, the attachable item is already present in their inbox or sent folder. A modern email client can proactively retrieve relevant attachable items from the user's past emails based on the context of the current conversation, and recommend them for inclusion, to reduce the time and effort involved in composing the response. In this paper, we propose a weakly supervised learning framework for recommending attachable items to the user. As email search systems are commonly available, we constrain the recommendation task to formulating effective search queries from the context of the conversations. The query is submitted to an existing IR system to retrieve relevant items for attachment. We also present a novel strategy for generating labels from an email corpus---without the need for manual annotations---that can be used to train and evaluate the query formulation model. In addition, we describe a deep convolutional neural network that demonstrates satisfactory performance on this query formulation task when evaluated on the publicly available Avocado dataset and a proprietary dataset of internal emails obtained through an employee participation program. Christophe Van Gysel, Bhaskar Mitra 0001, Matteo Venanzi, Roy Rosemarin, Grzegorz Kukla, Piotr Grudzien, Nicola Cancedda |
CIKM | 7 |
| 2014 | Fast Domain Adaptation of SMT models without in-Domain Parallel Data
Prashant Mathur, Sriram Venkatapathy, Nicola Cancedda |
COLING | 3 |
| 2013 | Generation of Compound Words in Statistical Machine Translation into Compounding LanguagesabstractIn this article we investigate statistical machine translation (SMT) into Germanic languages, with a focus on compound processing. Our main goal is to enable the generation of novel compounds that have not been seen in the training data. We adopt a split-merge strategy, where compounds are split before training the SMT system, and merged after the translation step. This approach reduces sparsity in the training data, but runs the risk of placing translations of compound parts in non-consecutive positions. It also requires a postprocessing step of compound merging, where compounds are reconstructed in the translation output. We present a method for increasing the chances that components that should be merged are translated into contiguous positions and in the right order and show that it can lead to improvements both by direct inspection and in terms of standard translation evaluation metrics. We also propose several new methods for compound merging, based on heuristics and machine learning, which outperform previously suggested algorithms. These methods can produce novel compounds and a translation with at least the same overall quality as the baseline. For all subtasks we show that it is useful to include part-of-speech based information in the translation process, in order to handle compounds. Sara Stymne, Nicola Cancedda, Lars Ahrenberg |
Comput. Linguistics | 2 |
| 2012 | Prediction of Learning Curves in Machine Translation
Prasanth Kolachina, Nicola Cancedda, Marc Dymetman, Sriram Venkatapathy |
ACL (1) | 2 |
| 2012 | Task-Driven Linguistic Analysis based on an Underspecified Features Representation
Stasinos Konstantopoulos, Valia Kordoni, Nicola Cancedda, Vangelis Karkaletsis, Dietrich Klakow, Jean-Michel Renders |
LREC | 3 |
| 2010 | Minimum Error Rate Training by Sampling the Translation Lattice
Samidh Chatterjee, Nicola Cancedda |
EMNLP | 2 |
| 2010 | A Dataset for Assessing Machine Translation Evaluation Metrics
Lucia Specia, Nicola Cancedda, Marc Dymetman |
LREC | 2 |
| 2010 | Pushing the frontier of Statistical Machine Translation: Preface
Lucia Specia, Nicola Cancedda |
Mach. Transl. | 2 |
| 2009 | Source-Language Entailment Modeling for Translating Unknown Terms
Shachar Mirkin, Lucia Specia, Nicola Cancedda, Ido Dagan, Marc Dymetman, Idan Szpektor |
ACL/IJCNLP | 3 |
| 2009 | Phrase-Based Statistical Machine Translation as a Traveling Salesman Problem
Mikhail Zaslavskiy, Marc Dymetman, Nicola Cancedda |
ACL/IJCNLP | 3 |
| 2009 | Estimating the Sentence-Level Quality of Machine Translation Systems
Lucia Specia, Marco Turchi, Nicola Cancedda, Nello Cristianini, Marc Dymetman |
EAMT | 3 |
| 2009 | Complexity-Based Phrase-Table Filtering for Statistical Machine Translation
Nadi Tomeh, Nicola Cancedda, Marc Dymetman |
MTSummit | 2 |
| 2009 | Factored sequence kernels
Nicola Cancedda, Pierre Mahé |
Neurocomputing | 1 |
| 2008 | Shaping research from user requirements, and other exotic things
Nicola Cancedda |
EAMT | 1 |
| 2008 | Factored sequence kernels
Pierre Mahé, Nicola Cancedda |
ESANN | 2 |
| 2003 | Word-Sequence Kernels
Nicola Cancedda, Éric Gaussier, Cyril Goutte, Jean-Michel Renders |
J. Mach. Learn. Res. | 1 |
| 2002 | Combining Labelled and Unlabelled Data: A Case Study on Fisher Kernels and Transductive Inference for Biological Entity Recognition
Cyril Goutte, Hervé Déjean, Éric Gaussier, Nicola Cancedda, Jean-Michel Renders |
CoNLL | 4 |
| 2001 | Probabilistic models for terminology extraction and knowledge structuring from documentsabstractThis paper outlines several problems inherent to knowledge extraction and structuring from textual data, and proposes different probabilistic models for solving them. These models are evaluated on different tasks, and several applications are envisaged. Éric Gaussier, Nicola Cancedda |
SMC | 2 |