EDBT 2026 Demo / reviewers in the wild / expert
Simone Filice
dblp:99/11030
· DBLP profile ↗
24ranked-venue papers
8as first author
11since 2021 · last 2026
0009-0002-6735-9950ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LiveRAG: A Diverse Q&A Dataset with Varying Difficulty Level for RAG EvaluationabstractWith Retrieval-Augmented Generation (RAG) becoming more and more prominent in generative AI solutions, there is an emerging need for systematically evaluating its effectiveness. We introduce the LiveRAG benchmark, a publicly available dataset of 895 synthetic questions and answers designed to support systematic evaluation of RAG-based Q&A systems. This synthetic benchmark is derived from the one used during the SIGIR'2025 LiveRAG challenge, where competitors were evaluated under strict time constraints. It is augmented with information that was not made available to competitors during the challenge, such as the ground-truth answers, together with their associated supporting claims which were used for evaluating competitors' answers. In addition, each question is associated with estimated difficulty and discriminability scores, derived from applying an Item Response Theory model to competitors' responses. Our analysis highlights the benchmark's question diversity, the wide range of difficulty levels, and their usefulness in differentiating between system capabilities. The LiveRAG benchmark will hopefully help the community advance RAG research, conduct systematic evaluation, and develop more robust Q&A systems. David Carmel, Simone Filice, Guy Horowitz, Yoelle Maarek, Alex Shtoff, Oren Somekh, Ran Tavory |
SIGIR | 2 |
| 2025 | The Distracting Effect: Understanding Irrelevant Passages in RAGabstractA well-known issue with Retrieval Augmented Generation (RAG) is that retrieved passages that are irrelevant to the query sometimes distract the answer-generating LLM, causing it to provide an incorrect response.In this paper, we shed light on this core issue and formulate the distracting effect of a passage w.r.t. a query (and an LLM).We provide a quantifiable measure of the distracting effect of a passage and demonstrate its robustness across LLMs.Our research introduces novel methods for identifying and using hard distracting passages to improve RAG systems.By fine-tuning LLMs with these carefully selected distracting passages, we achieve up to a 7.5% increase in answering accuracy compared to counterparts fine-tuned on conventional RAG datasets.Our contribution is two-fold: first, we move beyond the simple binary classification of irrelevant passages as either completely unrelated vs. distracting, and second, we develop and analyze multiple methods for finding hard distracting passages.To our knowledge, no other research has provided such a comprehensive framework for identifying and utilizing hard distracting passages. Chen Amiraz, Florin Cuconasu, Simone Filice, Zohar S. Karnin |
ACL (1) | 3 |
| 2025 | Do RAG Systems Really Suffer From Positional Bias?abstractRetrieval Augmented Generation enhances LLM accuracy by adding passages retrieved from an external corpus to the LLM prompt.This paper investigates how positional bias-the tendency of LLMs to weight information differently based on its position in the promptaffects not only the LLM's capability to capitalize on relevant passages, but also its susceptibility to distracting passages.Through extensive experiments on three benchmarks, we show how state-of-the-art retrieval pipelines, while attempting to retrieve relevant passages, systematically bring highly distracting ones to the top ranks, with over 60% of queries containing at least one highly distracting passage among the top-10 retrieved passages.As a result, the impact of the LLM positional bias, which in controlled settings is often reported as very prominent by related works, is actually marginal in real scenarios since both relevant and distracting passages are, in turn, penalized.Indeed, our findings reveal that sophisticated strategies that attempt to rearrange the passages based on LLM positional preferences do not perform better than random shuffling. Florin Cuconasu, Simone Filice, Guy Horowitz, Yoelle Maarek, Fabrizio Silvestri |
EMNLP | 2 |
| 2025 | The LiveRAG Challenge at SIGIR 2025abstractThe LiveRAG Challenge at SIGIR 2025 provides a competitive platform for advancing Retrieval-Augmented Generation (RAG) technologies. Participants from academia and industry have been invited to build a RAG-based question answering system using a fixed corpus (Fineweb-10BT) and a common open-source LLM (Falcon3-10B-Instruct). The goal is to enable fair, focused comparisons on retrieval and prompting strategies. During the Live Challenge Day, the competing teams must provide answers and supportive information to 500 unseen questions within a strict two-hour window. Evaluation is conducted in two stages: automated LLM-as-a-judge scoring mechanism for correctness and faithfulness, followed by a manual review of top ranked submissions. The winners will be announced and prizes awarded during the LiveRAG Workshop at SIGIR 2025 in Padua, Italy. David Carmel, Simone Filice, Guy Horowitz, Yoelle Maarek, Oren Somekh, Ran Tavory |
SIGIR | 2 |
| 2024 | Enhancing Low-Resource LLMs Classification with PEFT and Synthetic DataabstractLarge Language Models (LLMs) operating in 0-shot or few-shot settings achieve competitive results in Text Classification tasks. In-Context Learning (ICL) typically achieves better accuracy than the 0-shot setting, but it pays in terms of efficiency, due to the longer input prompt. In this paper, we propose a strategy to make LLMs as efficient as 0-shot text classifiers, while getting comparable or better accuracy than ICL. Our solution targets the low resource setting, i.e., when only 4 examples per class are available. Using a single LLM and few-shot real data we perform a sequence of generation, filtering and Parameter-Efficient Fine-Tuning steps to create a robust and efficient classifier. Experimental results show that our approach leads to competitive results on multiple text classification datasets. Parth Patwa, Simone Filice, Zhiyu Chen 0001, Giuseppe Castellucci, Oleg Rokhlenko, Shervin Malmasi |
LREC/COLING | 2 |
| 2024 | The Power of Noise: Redefining Retrieval for RAG SystemsabstractRetrieval-Augmented Generation (RAG) has recently emerged as a method to extend beyond the pre-trained knowledge of Large Language Models by augmenting the original prompt with relevant passages or documents retrieved by an Information Retrieval (IR) system. RAG has become increasingly important for Generative AI solutions, especially in enterprise settings or in any domain in which knowledge is constantly refreshed and cannot be memorized in the LLM. We argue here that the retrieval component of RAG systems, be it dense or sparse, deserves increased attention from the research community, and accordingly, we conduct the first comprehensive and systematic examination of the retrieval strategy of RAG systems. We focus, in particular, on the type of passages IR systems within a RAG solution should retrieve. Our analysis considers multiple factors, such as the relevance of the passages included in the prompt context, their position, and their number. One counter-intuitive finding of this work is that the retriever's highest-scoring documents that are not directly relevant to the query (e.g., do not contain the answer) negatively impact the effectiveness of the LLM. Even more surprising, we discovered that adding random documents in the prompt improves the LLM accuracy by up to 35%. These results highlight the need to investigate the appropriate strategies when integrating retrieval with LLMs, thereby laying the groundwork for future research in this area. Florin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice, Cesare Campagnano, Yoelle Maarek, Nicola Tonellotto, Fabrizio Silvestri |
SIGIR | 4 |
| 2023 | Faithful Low-Resource Data-to-Text Generation through Cycle TrainingabstractZhuoer Wang, Marcus Collins, Nikhita Vedula, Simone Filice, Shervin Malmasi, Oleg Rokhlenko. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zhuoer Wang, Marcus D. Collins, Nikhita Vedula, Simone Filice, Shervin Malmasi, Oleg Rokhlenko |
ACL (1) | 4 |
| 2022 | Preventing Catastrophic Forgetting in Continual Learning of New Natural Language TasksabstractMulti-Task Learning (MTL) is widely-accepted in Natural Language Processing as a standard technique for learning multiple related tasks in one model. Training an MTL model requires having the training data for all tasks available at the same time. As systems usually evolve over time, (e.g., to support new functionalities), adding a new task to an existing MTL model usually requires retraining the model from scratch on all the tasks and this can be time-consuming and computationally expensive. Moreover, in some scenarios, the data used to train the original training may be no longer available, for example, due to storage or privacy concerns. Sudipta Kar, Giuseppe Castellucci, Simone Filice, Shervin Malmasi, Oleg Rokhlenko |
KDD | 3 |
| 2022 | Learning to Generate Examples for Semantic Processing TasksabstractDanilo Croce, Simone Filice, Giuseppe Castellucci, Roberto Basili. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Danilo Croce, Simone Filice, Giuseppe Castellucci, Roberto Basili 0001 |
NAACL-HLT | 2 |
| 2021 | Continual Learning for Named Entity RecognitionabstractNamed Entity Recognition (NER) is a vital task in various NLP applications. However, in many real-world scenarios (e.g., voice-enabled assistants) new named entities are frequently introduced, entailing re-training NER models to support these new entities. Re-annotating the original training data for the new entities could be costly or even impossible when storage limitations or security concerns restrict access to that data, and annotating a new dataset for all of the entities becomes impractical and error-prone as the number of entities increases. To tackle this problem, we introduce a novel Continual Learning approach for NER, which requires new training material to be annotated only for the new entities. To preserve the existing knowledge previously learned by the model, we exploit the Knowledge Distillation (KD) framework, where the existing NER model acts as the teacher for a new NER model (i.e., the student), which learns the new entity by using the new training material and retains knowledge of old entities by imitating the teacher's outputs on this new training set. Our experiments show that this approach allows the student model to ``progressively'' learn to identify new entities without forgetting the previously learned ones. We also present a comparison with multiple strong baselines to demonstrate that our approach is superior for continually updating an NER model. Natawut Monaikul, Giuseppe Castellucci, Simone Filice, Oleg Rokhlenko |
AAAI | 3 |
| 2021 | VoiSeR: A New Benchmark for Voice-Based Search RefinementabstractSimone Filice, Giuseppe Castellucci, Marcus Collins, Eugene Agichtein, Oleg Rokhlenko. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Simone Filice, Giuseppe Castellucci, Marcus D. Collins, Eugene Agichtein, Oleg Rokhlenko |
EACL | 1 |
| 2020 | Using Phoneme Representations to Build Predictive Models Robust to ASR ErrorsabstractEven though Automatic Speech Recognition (ASR) systems significantly improved over the last decade, they still introduce a lot of errors when they transcribe voice to text. One of the most common reasons for these errors is phonetic confusion between similar-sounding expressions. As a result, ASR transcriptions often contain "quasi-oronyms", i.e., words or phrases that sound similar to the source ones, but that have completely different semantics (e.g., "win" instead of "when" or "accessible on defecting" instead of "accessible and affecting"). These errors significantly affect the performance of downstream Natural Language Understanding (NLU) models (e.g., intent classification, slot filling, etc.) and impair user experience. To make NLU models more robust to such errors, we propose novel phonetic-aware text representations. Specifically, we represent ASR transcriptions at the phoneme level, aiming to capture pronunciation similarities, which are typically neglected in word-level representations (e.g., word embeddings). To train and evaluate our phoneme representations, we generate noisy ASR transcriptions of four existing datasets - Stanford Sentiment Treebank, SQuAD, TREC Question Classification and Subjectivity Analysis - and show that common neural network architectures exploiting the proposed phoneme representations can effectively handle noisy transcriptions and significantly outperform state-of-the-art baselines. Finally, we confirm these results by testing our models on real utterances spoken to the Alexa virtual assistant. Anjie Fang, Simone Filice, Nut Limsopatham, Oleg Rokhlenko |
SIGIR | 2 |
| 2020 | Voice-based Reformulation of Community AnswersabstractCommunity Question Answering (CQA) websites, such as Stack Exchange1 or Quora2, allow users to freely ask questions and obtain answers from other users, i.e., the community. Personal assistants, such as Amazon Alexa or Google Home, can also exploit CQA data to answer a broader range of questions and increase customers’ engagement. However, the voice-based interaction poses new challenges to the Question Answering scenario. Even assuming that we are able to retrieve a previously asked question that perfectly matches the user’s query, we cannot simply read its answer to the user. A major limitation is the answer length. Reading these answers to the user is cumbersome and boring. Furthermore, many answers contain non-voice-friendly parts, such as images, or URLs. Simone Filice, Nachshon Cohen, David Carmel |
WWW | 1 |
| 2019 | Making sense of kernel spaces in neural learning
Danilo Croce, Simone Filice, Roberto Basili 0001 |
Comput. Speech Lang. | 2 |
| 2017 | Deep Learning in Semantic Kernel SpacesabstractKernel methods enable the direct usage of structured representations of textual data during language learning and inference tasks.Expressive kernels, such as Tree Kernels, achieve excellent performance in NLP.On the other side, deep neural networks have been demonstrated effective in automatically learning feature representations during training.However, their input is tensor data, i.e., they cannot manage rich structured information.In this paper, we show that expressive kernels and deep neural networks can be combined in a common framework in order to (i) explicitly model structured information and (ii) learn non-linear decision functions.We show that the input layer of a deep architecture can be pre-trained through the application of the Nyström low-rank approximation of kernel spaces.The resulting "kernelized" neural network achieves state-of-the-art accuracy in three different tasks. Danilo Croce, Simone Filice, Giuseppe Castellucci, Roberto Basili 0001 |
ACL (1) | 2 |
| 2017 | On the Use of an Intermediate Class in Boolean Crowdsourced Relevance Annotations for Learning to Rank CommentsabstractIn many Information Retrieval tasks, the boundary between classes is not well defined, and assigning a document to a specific class may be complicated, even for humans. For instance, a document which is not directly related to the user's query may still contain relevant information. In this scenario, an option is to define an intermediate class collecting ambiguous instances. Yet some natural questions arise. Is this annotation strategy convenient? how should the intermediate class be treated? To answer these questions, we explored two community question answering datasets whose comments were originally annotated with three classes. We re-annotated a subset of instances considering a binary good vs bad setting. Our main contribution is to show empirically that the inclusion of an intermediate class to assess Boolean relevance is not useful. Moreover, in case the data is already annotated with a 3-class strategy, the instances from the intermediate class can be safely removed at training time. Alberto Barrón-Cedeño, Giovanni Da San Martino, Simone Filice, Alessandro Moschitti |
SIGIR | 3 |
| 2017 | KELP: a Kernel-based Learning Platform
Simone Filice, Giuseppe Castellucci, Giovanni Da San Martino, Alessandro Moschitti, Danilo Croce, Roberto Basili 0001 |
J. Mach. Learn. Res. | 1 |
| 2016 | Learning to Recognize Ancillary Information for Automatic Paraphrase IdentificationabstractPrevious work on Automatic Paraphrase Identification (PI) is mainly based on modeling text similarity between two sentences.In contrast, we study methods for automatically detecting whether a text fragment only appearing in a sentence of the evaluated sentence pair is important or ancillary information with respect to the paraphrase identification task.Engineering features for this new task is rather difficult, thus, we approach the problem by representing text with syntactic structures and applying tree kernels on them.The results show that the accuracy of our automatic Ancillary Text Classifier (ATC) is promising, i.e., 68.6%, and its output can be used to improve the state of the art in PI. Simone Filice, Alessandro Moschitti |
HLT-NAACL | 1 |
| 2015 | A Stratified Strategy for Efficient Kernel-Based LearningabstractIn Kernel-based Learning the targeted phenomenon is summarized by a set of explanatory examples derived from the training set. When the model size grows with the complexity of the task, such approaches are so computationally demanding that the adoption of comprehensive models is not always viable.In this paper, a general framework aimed at minimizing this problem is proposed: multiple classifiers are stratified and dynamically invoked according to increasing levels of complexity corresponding to incrementally more expressive representation spaces.Computationally expensive inferences are thus adopted only when the classification at lower levels is too uncertain over an individual instance. The application of complex functions is thus avoided where possible, with a significant reduction of the overall costs. The proposed strategy has been integrated within two well-known algorithms: Support Vector Machines and Passive-Aggressive Online classifier.A significant cost reduction (up to 90%), with a negligible performance drop, is observed against two Natural Language Processing tasks, i.e. Question Classification and Sentiment Analysis in Twitter. Simone Filice, Danilo Croce, Roberto Basili 0001 |
AAAI | 1 |
| 2015 | Structural Representations for Learning Relations between Pairs of TextsabstractSimone Filice, Giovanni Da San Martino, Alessandro Moschitti. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Simone Filice, Giovanni Da San Martino, Alessandro Moschitti |
ACL (1) | 1 |
| 2015 | Global Thread-level Inference for Comment Classification in Community Question AnsweringabstractShafiq Joty, Alberto Barrón-Cedeño, Giovanni Da San Martino, Simone Filice, Lluís Màrquez, Alessandro Moschitti, Preslav Nakov. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015. Shafiq R. Joty, Alberto Barrón-Cedeño, Giovanni Da San Martino, Simone Filice, Lluís Màrquez, Alessandro Moschitti, Preslav Nakov |
EMNLP | 4 |
| 2014 | Effective Kernelized Online Learning in Language Processing Tasks
Simone Filice, Giuseppe Castellucci, Danilo Croce, Roberto Basili 0001 |
ECIR | 1 |
| 2013 | Linear Online Learning over Structured Data with Distributed Tree KernelsabstractOnline algorithms are an important class of learning machines as they are extremely simple and computationally efficient. Kernel methods versions can handle structured data, such as trees, and achieve state-of-the-art performance. However kernelized versions of Online Learning algorithms slow down when the number of support vectors becomes large. The traditional way to cope with this problem is introducing budgets that set the maximum number of support vectors. In this paper, we investigate Distributed Trees (DT) as an efficient way to use structured data in online learning. DTs effectively embed the huge feature space of the tree fragments into small vectors, so enabling the use of linear versions of kernel machines over tree structured data. We experiment with the Passive-Aggressive (PA) algorithm by comparing the linear and the kernelized version. A massive dataset made with tree structured data is employed: it is originated from a natural language processing task, the Boundary Detection in the context of Semantic Role Labeling over Frame Net. Results on a sample of the final data show that the DTs along with the Linear PA algorithm and the Tree Kernel along with the Bundgeted PA achieve comparable results in terms of f1-measure. Finally, the exploration of the full dataset allows the former to improve the performance on the classification task, with respect to the latter. Simone Filice, Danilo Croce, Roberto Basili 0001, Fabio Massimo Zanzotto |
ICMLA (1) | 1 |
| 2012 | Distributional Models and Lexical Semantics in Convolution Kernels
Danilo Croce, Simone Filice, Roberto Basili 0001 |
CICLing (1) | 2 |