Danilo Croce

dblp:38/2048 · DBLP profile ↗
← Back
40ranked-venue papers
13as first author
9since 2021 · last 2026
0000-0001-9111-1950ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 12 first-author · 7 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5
YearPublicationVenuePosition
2026 Integrating AI and IR Paradigms for Sustainable and Trustworthy Accurate Access to Large Scale Biomedical Information
Federico Borazio, Francesco Labbate, Danilo Croce, Roberto Basili 0001
ECIR (3)3
2026 Learning Molecular Structures from Infrared Spectra Through Latent Evidence Prediction
Sergio José Peresson, Danilo Croce, Roberto Basili 0001
IDA2
2026 Sanskrit Travelogue: A Large-Scale Unified and Annotated Corpus of Sanskrit Texts
Giacomo De Luca, Danilo Croce, Roberto Basili 0001
LREC2
2025 Adapting LLMs for Domain-Specific Retrieval: A Case Study in Nuclear Safety
Federico Borazio, Danilo Croce, Roberto Basili 0001
ECIR (5)2
2025 Grounded Semantic Role Labelling from Synthetic Multimodal Data for Situated Robot Commands
abstract
Understanding natural language commands in situated Human-Robot Interaction (HRI) requires linking linguistic input to perceptual context.Traditional symbolic parsers lack the flexibility to operate in complex, dynamic environments.We introduce a novel Multimodal Grounded Semantic Role Labelling (G-SRL) framework that combines frame semantics with perceptual grounding, enabling robots to interpret commands via multimodal logical forms.Our approach leverages modern Vision Language Models (VLMs), which jointly process text and images, and is supported by an automated pipeline that generates high-quality training data.Structured command annotations are converted into photorealistic scenes via LLM-guided prompt engineering and diffusion models, then rigorously validated through object detection and visual question answering.The pipeline produces over 11,000 imagecommand pairs (3,500+ manually validated), while approaching the quality of manually curated datasets at significantly lower cost.
Claudiu D. Hromei, Antonio Scaiella, Danilo Croce, Roberto Basili 0001
EMNLP3
2024 MM-IGLU: Multi-Modal Interactive Grounded Language Understanding
abstract
This paper explores Interactive Grounded Language Understanding (IGLU) challenges within Human-Robot Interaction (HRI). In this setting, a robot interprets user commands related to its environment, aiming to discern whether a specific command can be executed. If faced with ambiguities or incomplete data, the robot poses relevant clarification questions. Drawing from the NeurIPS 2022 IGLU competition, we enrich the dataset by introducing our multi-modal data and natural language descriptions in MM-IGLU: Multi-Modal Interactive Grounded Language Understanding. Utilizing a BART-based model that integrates the user’s statement with the environment’s description, and a cutting-edge Multi-Modal Large Language Model that merges both visual and textual data, we offer a valuable resource for ongoing research in the domain. Additionally, we discuss the evaluation methods for such tasks, highlighting potential limitations imposed by traditional string-match-based evaluations on this intricate multi-modal challenge. Moreover, we provide an evaluation benchmark based on human judgment to address the limits and capabilities of such baseline models. This resource is released on a dedicated GitHub repository at https://github.com/crux82/MM-IGLU.
Claudiu D. Hromei, Daniele Margiotta, Danilo Croce, Roberto Basili 0001
LREC/COLING3
2024 AI-driven transcriptomic encoders: From explainable models to accurate, sample-independent cancer diagnostics
abstract
In the rapidly evolving domain of medical technology, the utilization of sophisticated algorithms for deciphering transcriptional data has emerged as a critical aspect, especially in the oncology sector. these algorithms, drawing upon methodologies from fields such as natural language processing and advanced image analysis, can significantly enhance the accuracy in predicting cancer-related molecular states. notably, transformer models, renowned for their proficiency in handling extensive datasets, are now being adapted for breakthroughs in medical diagnostics or in stratifying patients according to prognostic levels. our study contributes to the field of precision medicine by integrating transformer-based learning, exemplified by the geneformer model, with explainable aI techniques. these techniques are employed to find out the input variables (genes resulting from genomic transcription) most correlated with the decisions of neural network systems. this insight, a key goal in genomic research, aims to select the most relevant gene subset for each specific task in which a neural network is employed. this selection approach has proven to be effective in two classification tasks: cell type classification and breast cancer type classification. such effectiveness has been demonstrated even across various cohorts of patients. when applying geneformer-like architecture analyses solely to the selected gene subsets, the outcomes either maintain their accuracy or significantly improve. this approach, aims not only to contribute to the identification of vital genetic markers in cancer genomics, but also to exemplify the adaptability of aI models to different datasets, marking a significant step towards the development of accurate and universally applicable diagnostic tools for precision medicine.
Danilo Croce, Artem Smirnov, Luigi Tiburzi, Serena Travaglini, Roberta Costa, Armando Calabrese, Roberto Basili 0001, Nathan Levialdi Ghiron, Gerry Melino
Expert Syst. Appl.1
2022 Learning to Generate Examples for Semantic Processing Tasks
abstract
Danilo Croce, Simone Filice, Giuseppe Castellucci, Roberto Basili. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Danilo Croce, Simone Filice, Giuseppe Castellucci, Roberto Basili 0001
NAACL-HLT1
2021 Sentiment Polarity Classification at EVALITA: Lessons Learned and Open Challenges
abstract
Sentiment analysis in social media is a popular task attracting the interest of the research community, also in recent evaluation campaigns of natural language processing tasks in several languages. We report on our experience in the organization of SENTIment POLarity Classification Task (SENTIPOLC), a shared task on sentiment classification of Italian tweets, proposed for the first time in 2014 within the Evalita evaluation campaign. We present the datasets-which include an enriched annotation scheme for dealing with the impact of figurative language on polarity-the evaluation methodology, and discuss the approaches and results of participating systems. We also offer a reflection on the open challenges of state-of-the-art systems for sentiment analysis of microblogging in Italian, as they emerge from a qualitative analysis of misclassified tweets. Finally, we provide an evaluation of the resources we have created, and share the lessons learned by running this task for two consecutive editions.
Valerio Basile, Nicole Novielli, Danilo Croce, Francesco Barbieri, Malvina Nissim, Viviana Patti
IEEE Trans. Affect. Comput.3
2020 Actionable Ethics through Neural Learning
abstract
While AI is going to produce a great impact on society, its alignment with human values and expectations is an essential step towards a correct harnessing of AI potentials for good. There is a corresponding growing need for mature and established technical standards to enable the assessment of an AI application as the evaluation of its graded adherence to formalized ethics. This is clearly dependent on methods to inject ethical awareness at all stages of an AI application development and use. For this reason we introduce the notion of Embedding Principles of ethics by Design (EPbD) as a comprehensive inductive framework. Although extending generic AI applications, it mainly aims at learning the ethical behaviour through numerical optimization, i.e. deep neural models. The core idea is to support ethics by integrating automated reasoning over formal knowledge and induction from ethically enriched training data. A deep neural network is proposed here to model both the functional as well as the ethical conditions characterizing a target decision. In this way, the discovery of latent ethical knowledge is enabled and made available to the learning process. The application of the above framework to a banking application, i.e. AI-driven Digital Lending, is used to show how accurate classification can be achieved without neglecting the ethical dimension. Results over existing datasets demonstrate that the ethical compliance of the sources can be used to output models able to optimally fine tune the balance between business and ethical accuracy.
Daniele Rossini, Danilo Croce, Sara Mancini, Massimo Pellegrino, Roberto Basili 0001
AAAI2
2020 GAN-BERT: Generative Adversarial Learning for Robust Text Classification with a Bunch of Labeled Examples
abstract
Recent Transformer-based architectures, e.g., BERT, provide impressive results in many Natural Language Processing tasks.However, most of the adopted benchmarks are made of (sometimes hundreds of) thousands of examples.In many real scenarios, obtaining highquality annotated data is expensive and timeconsuming; in contrast, unlabeled examples characterizing the target task can be, in general, easily collected.One promising method to enable semi-supervised learning has been proposed in image processing, based on Semi-Supervised Generative Adversarial Networks.In this paper, we propose GAN-BERT that extends the fine-tuning of BERT-like architectures with unlabeled data in a generative adversarial setting.Experimental results show that the requirement for annotated examples can be drastically reduced (up to only 50-100 annotated examples), still obtaining good performances in several sentence classification tasks.
Danilo Croce, Giuseppe Castellucci, Roberto Basili 0001
ACL1
2020 Grounded language interpretation of robotic commands through structured learning
Andrea Vanzo, Danilo Croce, Emanuele Bastianelli, Roberto Basili 0001, Daniele Nardi
Artif. Intell.2
2019 Auditing Deep Learning processes through Kernel-based Explanatory Models
abstract
Danilo Croce, Daniele Rossini, Roberto Basili. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Danilo Croce, Daniele Rossini, Roberto Basili 0001
EMNLP/IJCNLP (1)1
2019 Making sense of kernel spaces in neural learning
Danilo Croce, Simone Filice, Roberto Basili 0001
Comput. Speech Lang.1
2019 Neural embeddings: accurate and readable inferences based on semantic kernels
abstract
Abstract Sentence embeddings are the suitable input vectors for the neural learning of a number of inferences about content and meaning. Similarity estimation, classification, emotional characterization of sentences as well as pragmatic tasks, such as question answering or dialogue, have largely demonstrated the effectiveness of vector embeddings to model semantics. Unfortunately, most of the above decisions are epistemologically opaque as for the limited interpretability of the acquired neural models based on the involved embeddings. We think that any effective approach to meaning representation should be at least epistemologically coherent. In this paper, we concentrate on thereadabilityof neural models, as a core property of any embedding technique consistent and effective in representing sentence meaning. In this perspective, this paper discusses a novel embedding technique (the Nyström methodology) that corresponds to the reconstruction of a sentence in a kernel space, inspired by rich semantic similarity metrics (a semantic kernel) rather than by a language model. In addition to being based on a kernel that captures grammatical and lexical semantic information, the proposed embedding can be used as the input vector of an effective neural learning architecture, calledKernel-based deep architectures(KDA). Finally, it also characterizesby designthe KDA explanatory capability, as the proposed embedding is derived from examples that are both human readable and labeled. This property is obtained by the integration of KDAs with an explanation methodology, calledlayer-wise relevance propagation (LRP), already proposed in image processing. The Nyström embeddings support here the automatic compilation of argumentations in favor or against a KDA inference, in form of an explanation: each decision can in fact be linked through LRP back to the real examples, that is, the landmarks linguistically related to the input instance. The KDA network output is explained via the analogy with the activated landmarks. Quantitative evaluation of the explanations shows that richer explanations based on semantic and syntagmatic structures characterize convincing arguments, as they effectively help the user in assessing whether or not to trust the machine decisions in different tasks, for example, Question Classification or Semantic Role Labeling. This confirms the epistemological benefit that Nyström embeddings may bring, as linguistically rich and meaningful representations for a variety of inference tasks.
Danilo Croce, Daniele Rossini, Roberto Basili 0001
Nat. Lang. Eng.1
2017 Deep Learning in Semantic Kernel Spaces
abstract
Kernel methods enable the direct usage of structured representations of textual data during language learning and inference tasks.Expressive kernels, such as Tree Kernels, achieve excellent performance in NLP.On the other side, deep neural networks have been demonstrated effective in automatically learning feature representations during training.However, their input is tensor data, i.e., they cannot manage rich structured information.In this paper, we show that expressive kernels and deep neural networks can be combined in a common framework in order to (i) explicitly model structured information and (ii) learn non-linear decision functions.We show that the input layer of a deep architecture can be pre-trained through the application of the Nyström low-rank approximation of kernel spaces.The resulting "kernelized" neural network achieves state-of-the-art accuracy in three different tasks.
Danilo Croce, Simone Filice, Giuseppe Castellucci, Roberto Basili 0001
ACL (1)1
2017 KELP: a Kernel-based Learning Platform
Simone Filice, Giuseppe Castellucci, Giovanni Da San Martino, Alessandro Moschitti, Danilo Croce, Roberto Basili 0001
J. Mach. Learn. Res.5
2016 Large-Scale Kernel-Based Language Learning Through the Ensemble Nystr đdoto o ¨ m Methods
Danilo Croce, Roberto Basili 0001
ECIR1
2016 A Discriminative Approach to Grounded Spoken Language Understanding in Interactive Robotics
Emanuele Bastianelli, Danilo Croce, Andrea Vanzo, Roberto Basili 0001, Daniele Nardi
IJCAI2
2016 A Language Independent Method for Generating Large Scale Polarity Lexicons
Giuseppe Castellucci, Danilo Croce, Roberto Basili 0001
LREC2
2015 A Stratified Strategy for Efficient Kernel-Based Learning
abstract
In Kernel-based Learning the targeted phenomenon is summarized by a set of explanatory examples derived from the training set. When the model size grows with the complexity of the task, such approaches are so computationally demanding that the adoption of comprehensive models is not always viable.In this paper, a general framework aimed at minimizing this problem is proposed: multiple classifiers are stratified and dynamically invoked according to increasing levels of complexity corresponding to incrementally more expressive representation spaces.Computationally expensive inferences are thus adopted only when the classification at lower levels is too uncertain over an individual instance. The application of complex functions is thus avoided where possible, with a significant reduction of the overall costs. The proposed strategy has been integrated within two well-known algorithms: Support Vector Machines and Passive-Aggressive Online classifier.A significant cost reduction (up to 90%), with a negligible performance drop, is observed against two Natural Language Processing tasks, i.e. Question Classification and Sentiment Analysis in Twitter.
Simone Filice, Danilo Croce, Roberto Basili 0001
AAAI2
2015 Using semantic maps for robust natural language interaction with robots
abstract
Modern robotic architectures are equipped with sensors enabling a deep analysis of the environment. In this work, we aim at demonstrating that such perceptual information (here modeled through semantic maps) can be effectively used to enhance the language understanding capabilities of the robot. A robust lexical mapping function based on the Distributional Semantics paradigm is here proposed as a basic model of grounding language towards the environment. We show that making such information available to the underlying language understanding algorithms improves the accuracy throughout the entire interpretation process.
Emanuele Bastianelli, Danilo Croce, Roberto Basili 0001, Daniele Nardi
INTERSPEECH2
2015 Acquiring a Large Scale Polarity Lexicon Through Unsupervised Distributional Methods
Giuseppe Castellucci, Danilo Croce, Roberto Basili 0001
NLDB2
2014 Semantic Compositionality in Tree Kernels
abstract
Kernel-based learning has been largely applied to semantic textual inference tasks. In particular, Tree Kernels (TKs) are crucial in the modeling of syntactic similarity between linguistic instances in Question Answering or Information Extraction tasks. At the same time, lexical semantic information has been studied through the adoption of the so-called Distributional Semantics (DS) paradigm, where lexical vectors are acquired automatically from large corpora. Notice how methods to account for compositional linguistic structures (e.g. grammatically typed bi-grams or complex verb or noun phrases) have been proposed recently by defining algebras on lexical vectors. The result is an extended paradigm called Distributional Compositional Semantics (DCS). Although lexical extensions have been already proposed to generalize TKs towards semantic phenomena (e.g. the predicate argument structures as for role labeling), currently studied TKs do not account for compositionality, in general. In this paper, a novel kernel called Compositionally Smoothed Partial Tree Kernel is proposed to integrate DCS operators into the tree kernel evaluation, by acting both over lexical leaves and non-terminal, i.e. complex compositional, nodes. The empirical results obtained on a Question Classification and Paraphrase Identification tasks show that state-of-the-art performances can be achieved, without resorting to manual feature engineering, thus suggesting that a large set of Web and text mining tasks can be handled successfully by the kernel proposed here.
Paolo Annesi, Danilo Croce, Roberto Basili 0001
CIKM2
2014 A context-based model for Sentiment Analysis in Twitter
Andrea Vanzo, Danilo Croce, Roberto Basili 0001
COLING2
2014 Effective and Robust Natural Language Understanding for Human-Robot Interaction
abstract
Robots are slowly becoming part of everyday life, as they are being marketed for commercial applications (viz. telepresence, cleaning or entertainment). Thus, the ability to interact with non-expert users is becoming a key requirement. Even if user utterances can be efficiently recognized and transcribed by Automatic Speech Recognition systems, several issues arise in translating them into suitable robotic actions. In this paper, we will discuss both approaches providing two existing Natural Language Understanding workflows for Human Robot Interaction. First, we discuss a grammar based approach: it is based on grammars thus recognizing a restricted set of commands. Then, a data driven approach, based on a free-from speech recognizer and a statistical semantic parser, is discussed. The main advantages of both approaches are discussed, also from an engineering perspective, i.e. considering the effort of realizing HRI systems, as well as their reusability and robustness. An empirical evaluation of the proposed approaches is carried out on several datasets, in order to understand performances and identify possible improvements towards the design of NLP components in HRI.
Emanuele Bastianelli, Giuseppe Castellucci, Danilo Croce, Roberto Basili 0001, Daniele Nardi
ECAI3
2014 Effective Kernelized Online Learning in Language Processing Tasks
Simone Filice, Giuseppe Castellucci, Danilo Croce, Roberto Basili 0001
ECIR3
2014 HuRIC: a Human Robot Interaction Corpus
Emanuele Bastianelli, Giuseppe Castellucci, Danilo Croce, Luca Iocchi, Roberto Basili 0001, Daniele Nardi
LREC3
2014 RoboCup@Home Spoken Corpus: Using Robotic Competitions for Gathering Datasets
Emanuele Bastianelli, Luca Iocchi, Daniele Nardi, Giuseppe Castellucci, Danilo Croce, Roberto Basili 0001
RoboCup5
2013 Linear Online Learning over Structured Data with Distributed Tree Kernels
abstract
Online algorithms are an important class of learning machines as they are extremely simple and computationally efficient. Kernel methods versions can handle structured data, such as trees, and achieve state-of-the-art performance. However kernelized versions of Online Learning algorithms slow down when the number of support vectors becomes large. The traditional way to cope with this problem is introducing budgets that set the maximum number of support vectors. In this paper, we investigate Distributed Trees (DT) as an efficient way to use structured data in online learning. DTs effectively embed the huge feature space of the tree fragments into small vectors, so enabling the use of linear versions of kernel machines over tree structured data. We experiment with the Passive-Aggressive (PA) algorithm by comparing the linear and the kernelized version. A massive dataset made with tree structured data is employed: it is originated from a natural language processing task, the Boundary Detection in the context of Semantic Role Labeling over Frame Net. Results on a sample of the final data show that the DTs along with the Linear PA algorithm and the Tree Kernel along with the Bundgeted PA achieve comparable results in terms of f1-measure. Finally, the exploration of the full dataset allows the former to improve the performance on the classification task, with respect to the latter.
Simone Filice, Danilo Croce, Roberto Basili 0001, Fabio Massimo Zanzotto
ICMLA (1)2
2012 Verb Classification using Distributional Similarity in Syntactic and Semantic Structures
Danilo Croce, Alessandro Moschitti, Roberto Basili 0001, Martha Palmer
ACL (1)1
2012 Distributional Models and Lexical Semantics in Convolution Kernels
Danilo Croce, Simone Filice, Roberto Basili 0001
CICLing (1)1
2011 Semantic convolution kernels over dependency trees: smoothed partial tree kernel
abstract
In recent years, natural language processing techniques have been used more and more in IR. Among other syntactic and semantic parsing are effective methods for the design of complex applications like for example question answering and sentiment analysis. Unfortunately, extracting feature representations suitable for machine learning algorithms from linguistic structures is typically difficult. In this paper, we describe one of the most advanced piece of technology for automatic engineering of syntactic and semantic patterns. This method merges together convolution dependency tree kernels with lexical similarities. It can efficiently and effectively measure the similarity between dependency structures, whose lexical nodes are in part or completely different. Its use in powerful algorithm such as Support Vector Machines (SVMs) allows for fast design of accurate automatic systems.
Danilo Croce, Alessandro Moschitti, Roberto Basili 0001
CIKM1
2011 Structured Lexical Similarity via Convolution Kernels on Dependency Trees
Danilo Croce, Alessandro Moschitti, Roberto Basili 0001
EMNLP1
2010 Towards Open-Domain Semantic Role Labeling
Danilo Croce, Cristina Giannone, Paolo Annesi, Roberto Basili 0001
ACL1
2010 Acquiring IE Patterns through Distributional Lexical Semantic Models
Roberto Basili 0001, Danilo Croce, Cristina Giannone, Diego De Cao
CICLing2
2010 Extensive Evaluation of a FrameNet-WordNet mapping resource
Diego De Cao, Danilo Croce, Roberto Basili 0001
LREC2
2009 Cross-Language Frame Semantics Transfer in Bilingual Corpora
Roberto Basili 0001, Diego De Cao, Danilo Croce, Bonaventura Coppola, Alessandro Moschitti
CICLing3
2009 Semantic Word Spaces for Robust Role Labeling
abstract
Semantic role labeling systems are often designed as inductive processes over annotated resources. Supervised algorithms based on complex grammatical information achieve state-of-the-art accuracy. However, their generalization on the argument classification task is poorer, as large performance drops in out-of-domain tests showed. In this paper, a robust method based on a minimal set of grammatical features and a distributional model of lexical semantic information is proposed. The achievable generalization ability is studied in several training conditions where negligible performance drops are observed.
Cristina Giannone, Danilo Croce, Roberto Basili 0001
ICMLA2
2008 Automatic induction of FrameNet lexical units
Marco Pennacchiotti, Diego De Cao, Roberto Basili 0001, Danilo Croce, Michael Roth 0001
EMNLP4