VLDB 2026 Research / reviewers in the wild / expert
Giuseppe Castellucci
dblp:126/2079
· DBLP profile ↗
18ranked-venue papers
2as first author
9since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Generative Explore-Exploit: Training-free Optimization of Generative Recommender Systems using LLM OptimizersabstractLütfi Kerem Senel, Besnik Fetahu, Davis Yoshida, Zhiyu Chen, Giuseppe Castellucci, Nikhita Vedula, Jason Ingyu Choi, Shervin Malmasi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Lutfi Kerem Senel, Besnik Fetahu, Davis Yoshida, Zhiyu Chen 0001, Giuseppe Castellucci, Nikhita Vedula, Jason Ingyu Choi, Shervin Malmasi |
ACL (1) | 5 |
| 2024 | Enhancing Low-Resource LLMs Classification with PEFT and Synthetic DataabstractLarge Language Models (LLMs) operating in 0-shot or few-shot settings achieve competitive results in Text Classification tasks. In-Context Learning (ICL) typically achieves better accuracy than the 0-shot setting, but it pays in terms of efficiency, due to the longer input prompt. In this paper, we propose a strategy to make LLMs as efficient as 0-shot text classifiers, while getting comparable or better accuracy than ICL. Our solution targets the low resource setting, i.e., when only 4 examples per class are available. Using a single LLM and few-shot real data we perform a sequence of generation, filtering and Parameter-Efficient Fine-Tuning steps to create a robust and efficient classifier. Experimental results show that our approach leads to competitive results on multiple text classification datasets. Parth Patwa, Simone Filice, Zhiyu Chen 0001, Giuseppe Castellucci, Oleg Rokhlenko, Shervin Malmasi |
LREC/COLING | 4 |
| 2023 | External Knowledge Acquisition for End-to-End Document-Oriented Dialog SystemsabstractEnd-to-end neural models for conversational AI often assume that a response can be generated by considering only the knowledge acquired by the model during training.Documentoriented conversational models make a similar assumption by conditioning the input on the document and assuming that any other knowledge is captured in the model's weights.However, a conversation may refer to external knowledge sources.In this work, we present EKo-DoC, an architecture for documentoriented conversations with access to external knowledge: we assume that a conversation is centered around a topic document and that external knowledge is needed to produce responses.EKo-DoC includes a dense passage retriever, a re-ranker, and a response generation model.We train the model end-to-end by using silver labels for the retrieval and re-ranking components that we automatically acquire from the attention signals of the response generation model.We demonstrate with automatic and human evaluations that incorporating external knowledge improves response generation in document-oriented conversations.Our architecture achieves new state-of-the-art results on the Wizard of Wikipedia dataset, outperforming a competitive baseline by 10.3% in Recall@1 and 7.4% in ROUGE-L. Tuan Manh Lai, Giuseppe Castellucci, Saar Kuzi, Heng Ji 0001, Oleg Rokhlenko |
EACL | 2 |
| 2023 | Composing Spoken Hints for Follow-on Question Suggestion in Voice Assistants
Pedro Faustini, Besnik Fetahu, Giuseppe Castellucci, Anjie Fang, Oleg Rokhlenko, Shervin Malmasi |
INTERSPEECH | 3 |
| 2022 | Wizard of Tasks: A Novel Conversational Dataset for Solving Real-World Tasks in Conversational SettingsabstractConversational Task Assistants (CTAs) are conversational agents whose goal is to help humans perform real-world tasks. CTAs can help in exploring available tasks, answering task-specific questions and guiding users through step-by-step instructions. In this work, we present Wizard of Tasks, the first corpus of such conversations in two domains: Cooking and Home Improvement. We crowd-sourced a total of 549 conversations (18,077 utterances) with an asynchronous Wizard-of-Oz setup, relying on recipes from WholeFoods Market for the cooking domain, and WikiHow articles for the home improvement domain. We present a detailed data analysis and show that the collected data can be a valuable and challenging resource for CTAs in two tasks: Intent Classification (IC) and Abstractive Question Answering (AQA). While on IC we acquired a high performing model (>85% F1), on AQA the performance is far from being satisfactory (~27% BertScore-F1), suggesting that more work is needed to solve the task of low-resource AQA. Jason Ingyu Choi, Saar Kuzi, Nikhita Vedula, Giuseppe Castellucci, Marcus D. Collins, Shervin Malmasi, Oleg Rokhlenko, Eugene Agichtein |
COLING | 5 |
| 2022 | Preventing Catastrophic Forgetting in Continual Learning of New Natural Language TasksabstractMulti-Task Learning (MTL) is widely-accepted in Natural Language Processing as a standard technique for learning multiple related tasks in one model. Training an MTL model requires having the training data for all tasks available at the same time. As systems usually evolve over time, (e.g., to support new functionalities), adding a new task to an existing MTL model usually requires retraining the model from scratch on all the tasks and this can be time-consuming and computationally expensive. Moreover, in some scenarios, the data used to train the original training may be no longer available, for example, due to storage or privacy concerns. Sudipta Kar, Giuseppe Castellucci, Simone Filice, Shervin Malmasi, Oleg Rokhlenko |
KDD | 2 |
| 2022 | Learning to Generate Examples for Semantic Processing TasksabstractDanilo Croce, Simone Filice, Giuseppe Castellucci, Roberto Basili. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Danilo Croce, Simone Filice, Giuseppe Castellucci, Roberto Basili 0001 |
NAACL-HLT | 3 |
| 2021 | Continual Learning for Named Entity RecognitionabstractNamed Entity Recognition (NER) is a vital task in various NLP applications. However, in many real-world scenarios (e.g., voice-enabled assistants) new named entities are frequently introduced, entailing re-training NER models to support these new entities. Re-annotating the original training data for the new entities could be costly or even impossible when storage limitations or security concerns restrict access to that data, and annotating a new dataset for all of the entities becomes impractical and error-prone as the number of entities increases. To tackle this problem, we introduce a novel Continual Learning approach for NER, which requires new training material to be annotated only for the new entities. To preserve the existing knowledge previously learned by the model, we exploit the Knowledge Distillation (KD) framework, where the existing NER model acts as the teacher for a new NER model (i.e., the student), which learns the new entity by using the new training material and retains knowledge of old entities by imitating the teacher's outputs on this new training set. Our experiments show that this approach allows the student model to ``progressively'' learn to identify new entities without forgetting the previously learned ones. We also present a comparison with multiple strong baselines to demonstrate that our approach is superior for continually updating an NER model. Natawut Monaikul, Giuseppe Castellucci, Simone Filice, Oleg Rokhlenko |
AAAI | 2 |
| 2021 | VoiSeR: A New Benchmark for Voice-Based Search RefinementabstractSimone Filice, Giuseppe Castellucci, Marcus Collins, Eugene Agichtein, Oleg Rokhlenko. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Simone Filice, Giuseppe Castellucci, Marcus D. Collins, Eugene Agichtein, Oleg Rokhlenko |
EACL | 2 |
| 2020 | GAN-BERT: Generative Adversarial Learning for Robust Text Classification with a Bunch of Labeled ExamplesabstractRecent Transformer-based architectures, e.g., BERT, provide impressive results in many Natural Language Processing tasks.However, most of the adopted benchmarks are made of (sometimes hundreds of) thousands of examples.In many real scenarios, obtaining highquality annotated data is expensive and timeconsuming; in contrast, unlabeled examples characterizing the target task can be, in general, easily collected.One promising method to enable semi-supervised learning has been proposed in image processing, based on Semi-Supervised Generative Adversarial Networks.In this paper, we propose GAN-BERT that extends the fine-tuning of BERT-like architectures with unlabeled data in a generative adversarial setting.Experimental results show that the requirement for annotated examples can be drastically reduced (up to only 50-100 annotated examples), still obtaining good performances in several sentence classification tasks. Danilo Croce, Giuseppe Castellucci, Roberto Basili 0001 |
ACL | 2 |
| 2017 | Deep Learning in Semantic Kernel SpacesabstractKernel methods enable the direct usage of structured representations of textual data during language learning and inference tasks.Expressive kernels, such as Tree Kernels, achieve excellent performance in NLP.On the other side, deep neural networks have been demonstrated effective in automatically learning feature representations during training.However, their input is tensor data, i.e., they cannot manage rich structured information.In this paper, we show that expressive kernels and deep neural networks can be combined in a common framework in order to (i) explicitly model structured information and (ii) learn non-linear decision functions.We show that the input layer of a deep architecture can be pre-trained through the application of the Nyström low-rank approximation of kernel spaces.The resulting "kernelized" neural network achieves state-of-the-art accuracy in three different tasks. Danilo Croce, Simone Filice, Giuseppe Castellucci, Roberto Basili 0001 |
ACL (1) | 3 |
| 2017 | KELP: a Kernel-based Learning Platform
Simone Filice, Giuseppe Castellucci, Giovanni Da San Martino, Alessandro Moschitti, Danilo Croce, Roberto Basili 0001 |
J. Mach. Learn. Res. | 2 |
| 2016 | A Language Independent Method for Generating Large Scale Polarity Lexicons
Giuseppe Castellucci, Danilo Croce, Roberto Basili 0001 |
LREC | 1 |
| 2015 | Acquiring a Large Scale Polarity Lexicon Through Unsupervised Distributional Methods
Giuseppe Castellucci, Danilo Croce, Roberto Basili 0001 |
NLDB | 1 |
| 2014 | Effective and Robust Natural Language Understanding for Human-Robot InteractionabstractRobots are slowly becoming part of everyday life, as they are being marketed for commercial applications (viz. telepresence, cleaning or entertainment). Thus, the ability to interact with non-expert users is becoming a key requirement. Even if user utterances can be efficiently recognized and transcribed by Automatic Speech Recognition systems, several issues arise in translating them into suitable robotic actions. In this paper, we will discuss both approaches providing two existing Natural Language Understanding workflows for Human Robot Interaction. First, we discuss a grammar based approach: it is based on grammars thus recognizing a restricted set of commands. Then, a data driven approach, based on a free-from speech recognizer and a statistical semantic parser, is discussed. The main advantages of both approaches are discussed, also from an engineering perspective, i.e. considering the effort of realizing HRI systems, as well as their reusability and robustness. An empirical evaluation of the proposed approaches is carried out on several datasets, in order to understand performances and identify possible improvements towards the design of NLP components in HRI. Emanuele Bastianelli, Giuseppe Castellucci, Danilo Croce, Roberto Basili 0001, Daniele Nardi |
ECAI | 2 |
| 2014 | Effective Kernelized Online Learning in Language Processing Tasks
Simone Filice, Giuseppe Castellucci, Danilo Croce, Roberto Basili 0001 |
ECIR | 2 |
| 2014 | HuRIC: a Human Robot Interaction Corpus
Emanuele Bastianelli, Giuseppe Castellucci, Danilo Croce, Luca Iocchi, Roberto Basili 0001, Daniele Nardi |
LREC | 2 |
| 2014 | RoboCup@Home Spoken Corpus: Using Robotic Competitions for Gathering Datasets
Emanuele Bastianelli, Luca Iocchi, Daniele Nardi, Giuseppe Castellucci, Danilo Croce, Roberto Basili 0001 |
RoboCup | 4 |