VLDB 2026 Research / reviewers in the wild / expert
Luigi Di Caro
dblp:12/5266
· DBLP profile ↗
56ranked-venue papers
8as first author
18since 2021 · last 2027
0000-0002-7570-637XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 23 · 6 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 6 · 4 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Potential and limitations of LLMs for augmenting lexical knowledge basesabstractLexical Knowledge Bases (KBs) are central to NLP but remain costly to maintain, limited in coverage, and slow to adapt to linguistic change. This study explores the potential of large language models (LLMs) to fix these issues through two key research questions: (RQ1) Are LLMs capable of producing correct and novel information suitable for integration into existing KBs? and (RQ2) To what extent can LLM-generated candidates augment an existing KB beyond its current coverage, and what does the gap between automatic overlap and human-validated correctness imply about that augmentation potential? To investigate these, we introduce a three-stage, model-agnostic pipeline that transforms KB entries into structured prompts to generate candidate concepts, applies human-in-the-loop validation to assess novel outputs, and analyzes overlap with existing KB content to probe coverage. Experiments employed eight open-weight LLMs ranging from 7B to 104B parameters, four diverse KBs (ConceptNet, FrameNet, Semagram, MultiAligNet), and variations in prompting strategies (zero-shot/one-shot) and output formats (JSON/comma-separated), producing 384000 prompt-response pairs. Regarding RQ1, human validation confirmed 86.7% of novel candidates as semantically accurate (95% CI: 85.9-87.4%), with per-KB rates highest for MultiAligNet at 93.1%; one-shot prompting and JSON formatting yielded superior results, while model performance showed no correlation with parameter scale. For RQ2, automatic metrics indicated low overlap between LLM outputs and existing KB entries (e.g., F1@10=0.14 for ConceptNet); when combined with the high human approval rate of non-overlapping candidates, this pattern is consistent with coverage gaps in the studied resources, though other factors (e.g., normalization mismatches, generation bias) may also contribute. Recall patterns correlated with KB size and relationship specificity. Combined with HLoop, this gap measures the augmentation surface available from a single open-weight model: most non-overlapping items are valid candidates the KB does not yet contain, and HLoop quantifies the fraction usable for augmentation, shifting part of the human effort from knowledge creation toward verification. Federico Torrielli, Giovanni Siragusa, Vladimiro Lovera Rulfi, Amon Rapp, Luigi Di Caro |
Expert Syst. Appl. | 5 |
| 2026 | HumaniCA: A Benchmark Resource for the Detection of Users' Ascription of Humanness to Conversational Agents
Sabrina Villata, Amon Rapp, Luigi Di Caro, Federica Cena |
LREC | 3 |
| 2026 | From natural language to knowledge: The Role of LLMs in conceptual and data modeling
Luigi Di Caro, Amon Rapp, Vijayan Sugumaran, Farid Meziane |
Data Knowl. Eng. | 1 |
| 2025 | Leveraging RAG for Privacy Violation Detection and ExplainabilityabstractIn today’s digital landscape, users frequently share vast amounts of information, including confidential data, often without full awareness of the associated privacy risks. This scenario highlights the need for automated methods to identify sensitive information and alert users to such risks. Existing algorithmic solutions for detecting sensitive content typically require either human intervention (rule-based approaches) or labeled data (supervised learning), both of which can be costly and limiting. In this paper, we propose a framework based on Retrieval-Augmented Generation (RAG) to classify privacy-sensitive content while providing contextual explanations. We employed the state-of-the-art generative Large Language Model (LLM) GPT-4o, with Information Retrieval models BM25 and FAISS, enhancing both detection accuracy and explainability. Our method utilizes a curated Knowledge Base of scientific literature on privacy and confidentiality to retrieve contextually relevant information, which is then used to guide the classification process and generate explanations. Experimental evaluations on a real-world dataset (Enron Email Dataset) demonstrate that RAG-based approaches significantly outperform the zero-shot baseline, with BM25 showing the highest performance. This tool is designed to serve end-users, by mitigating risks before data sharing, by enabling proactive monitoring of privacy violations. Stefano Locci, Davide Audrito, Giovanni Livraga, Marco Viviani 0001, Luigi Di Caro |
IJCNN | 5 |
| 2025 | Automated Extraction of Judicial Interpretative Formulas in EU Case Law on VATabstractThis paper addresses the extraction of Judicial Interpretative Formulas (JIFs) in decisions of the Court of Justice of the European Union (CJEU) on Value Added Tax (VAT). European case law includes a significant number of JIFs on this subject, which are crucial for the interpretation of VAT. However, extracting such JIFs manually is effortful, and doing that automatically has not been investigated yet in the VAT domain. Our work proposes the first pipeline method for doing so. We start by defining a set of guidelines for annotating legal texts following a principle definition of JIF. By following such guidelines, we obtain a corpus of 21 expert-labeled CJEU decisions. We keep them for validation and testing. For training, we machine-annotate 80 additional decisions using LLMs. Our experiments show that BERT-based architectures trained on such data perform comparably to LLMs. Giulia Grundler, Piera Santin, Alessia Fidelangeli, Rachele Mignone, Federico Galli, Andrea Galassi, Giuseppe Contissa, Luigi Di Caro, Paolo Torroni |
JURIX | 8 |
| 2025 | Revisiting Cross-Modal Knowledge Distillation: A Disentanglement Approach for RGBD Semantic Segmentation
Roger Ferrod, Cássio Fraga Dantas, Luigi Di Caro, Dino Ienco |
ECML/PKDD (4) | 3 |
| 2025 | How Do People Develop Folk Theories of Generative AI Text-to-Image Models? A Qualitative Study on How People Strive to Explain and Make Sense of GenAIabstractGenerative Artificial Intelligence (GenAI) text-to-image models have made significant progress in emulating human-like outputs. However, understanding the inner functioning of these models remains a challenge due to their complexity and black-box nature. It has been observed that individuals naturally develop informal conceptualizations, termed “folk theories,” to explain the behaviors of algorithmic systems. The specific nature of GenAI text-to-image models, which are obscure in their working principles, yet carry out activities that are peculiar to humans, makes it interesting to investigate people’s theorization about this technology. With this aim, we conducted a qualitative interview study with 20 participants and observed how they accounted for the outputs of Stable Diffusion. The study findings show that participants developed a wide spectrum of conceptualizations, including folk theories that appear distinctive of GenAI text-to-image technology, also ascribing to the model a variety of “mental states.” Furthermore, we found that theory building follows different inductive and deductive trajectories, with participants employing diverse strategies to explain the functioning of the technology. Chiara Di Lodovico, Federico Torrielli, Luigi Di Caro, Amon Rapp |
Int. J. Hum. Comput. Interact. | 3 |
| 2025 | How do people react to ChatGPT's unpredictable behavior? Anthropomorphism, uncanniness, and fear of AI: A qualitative study on individuals' perceptions and understandings of LLMs' nonsensical hallucinationsabstract• We conducted a qualitative study on how people perceive LLMs’ unpredictable behaviors • We interviewed 20 participants to gather their feedback on a hallucination dialogue • We found that unpredictable behaviors change how people experience ChatGPT • We show that these behaviors evoke unsettling emotions and fear of AI Large Language Models (LLMs) have shown impressive capabilities in producing texts of quality and fluency that are similar to those created by humans. Despite their increasing use, however, the broader population's experience of many aspects of interaction with LLMs remains underexplored. This study investigates how diverse individuals perceive and account for “nonsensical hallucinations”, namely, an LLM's unpredictable and meaningless behavior provided as a response to a user's request. We asked 20 participants to interact with ChatGPT 3.5 and experience its hallucinations. Through semi-structured interviews, we found that participants with a computer science background or consistent previous use of LLMs interpret unpredictable nonsensical responses as an error, while novices perceive them as model's autonomous behaviors. Moreover, we discovered that such responses produce an abrupt modification of participants’ perceptions and understandings of the LLM's nature. From a soothing and polite entity, ChatGPT becomes either an obscure and unfamiliar “alien”, or a human-like being potentially hostile to humankind, making also emerge unsettling feelings, which may unveil an underlying fear of Artificial Intelligence. The study contributes to literature on how people react to the unfamiliarity of a technology that may be perceived as alien and yet extremely human-like, generating “uncanny effects,” as well as to research on the anthropomorphizing of technology. Amon Rapp, Chiara Di Lodovico, Luigi Di Caro |
Int. J. Hum. Comput. Stud. | 3 |
| 2025 | How do people experience the images created by generative artificial intelligence? An exploration of people's perceptions, appraisals, and emotions related to a Gen-AI text-to-image model and its creationsabstractGenerative Artificial Intelligence (Gen-AI) has rapidly advanced in recent years, potentially producing enormous impacts on industries, societies, and individuals in the near future. In particular, Gen-AI text-to-image models allow people to easily create high-quality images possibly revolutionizing human creative practices. Despite their increasing use, however, the broader population's perceptions and understandings of Gen-AI-generated images remain understudied in the Human-Computer Interaction (HCI) community. This study investigates how individuals, including those unfamiliar with Gen-AI, perceive Gen-AI text-to-image (Stable Diffusion) outputs. Study findings reveal that participants appraise Gen-AI images based on their technical quality and fidelity in representing a subject, often experiencing them as either prototypical or strange: these experiences may raise awareness of societal biases and evoke unsettling feelings that extend to the Gen-AI itself. The study also uncovers several “relational” strategies that participants employ to cope with concerns related to Gen-AI, contributing to the understanding of reactions to uncanny technology and the (de)humanization of intelligent agents. Moreover, the study offers design suggestions on how to use the anthropomorphizing of the text-to-image model as design material, and the Gen-AI images as support for critical design sessions. Amon Rapp, Chiara Di Lodovico, Federico Torrielli, Luigi Di Caro |
Int. J. Hum. Comput. Stud. | 4 |
| 2024 | EcoVerse: An Annotated Twitter Dataset for Eco-Relevance Classification, Environmental Impact Analysis, and Stance DetectionabstractAnthropogenic ecological crisis constitutes a significant challenge that all within the academy must urgently face, including the Natural Language Processing (NLP) community. While recent years have seen increasing work revolving around climate-centric discourse, crucial environmental and ecological topics outside of climate change remain largely unaddressed, despite their prominent importance. Mainstream NLP tasks, such as sentiment analysis, dominate the scene, but there remains an untouched space in the literature involving the analysis of environmental impacts of certain events and practices. To address this gap, this paper presents EcoVerse, an annotated English Twitter dataset of 3,023 tweets spanning a wide spectrum of environmental topics. We propose a three-level annotation scheme designed for Eco-Relevance Classification, Stance Detection, and introducing an original approach for Environmental Impact Analysis. We detail the data collection, filtering, and labeling process that led to the creation of the dataset. Remarkable Inter-Annotator Agreement indicates that the annotation scheme produces consistent annotations of high quality. Subsequent classification experiments using BERT-based models, including ClimateBERT, are presented. These yield encouraging results, while also indicating room for a model specifically tailored for environmental texts. The dataset is made freely available to stimulate further research. Francesca Grasso, Stefano Locci, Giovanni Siragusa, Luigi Di Caro |
LREC/COLING | 4 |
| 2024 | Towards a Multimodal Framework for Remote Sensing Image Change Retrieval and Captioning
Roger Ferrod, Luigi Di Caro, Dino Ienco |
DS (2) | 2 |
| 2024 | Introduction for computer law and security review: special issue "knowledge management for law"
Emilio Sulis, Luigi Di Caro, Rohan Nanda |
Comput. Law Secur. Rev. | 2 |
| 2024 | A tag-based methodology for the detection of user repair strategies in task-oriented conversational agents
Francesca Alloatti, Francesca Grasso, Roger Ferrod, Giovanni Siragusa, Luigi Di Caro, Federica Cena |
Comput. Speech Lang. | 5 |
| 2023 | Enriching Wikipedia Texts through Geographic Information ExtractionabstractGeographic Information Extraction (GIE) involves the extraction of geo-referenced information from a data collection through steps of geoparsing and geocoding. The former is a process that starts from a free textual description of locations with the goal of identifying an unambiguous location, such as specific geographic coordinates expressed as latitude-longitude. Differently, geocoding regards the easier task of translating an exact and well-formatted location such as postal addresses. This paper presents MAWI, i.e. a pipeline that starts from generic texts about cities that first extracts geographic information to automatically detect possible points of interest, then generates textual snippets from their contexts by means of Natural Language Processing (NLP) techniques. The adopted methodology involves several modules, ranging from publicly available geocoding systems to NLP libraries for Named Entity Recognition and text segmentation. The impact of the proposal includes multiple tasks and applications, e.g. i) the enrichment of public platforms of geographic data, ii) the detection of geographic scopes in textual documents, iii) a geo-centric exploration of locations in the tourism domain, and so forth. In this contribution, we present an experimentation of the system with 50 input Wikipedia pages referring different cities, first demonstrating its effectiveness with a running example, then evaluating its power to detect and structure a highly-significant amount of novel geo-referenced information with respect to what currently encoded in Wikipedia. Data and code are publicly available for future research at https://anonymous.4open.science/r/PointOfInterest-8D80/. Laura Ventrice, Luigi Di Caro |
ASONAM | 2 |
| 2023 | How Shall a Machine Call a Thing?
Federico Torrielli, Amon Rapp, Luigi Di Caro |
NLDB | 3 |
| 2022 | MultiAligNet: Cross-lingual Knowledge Bridges Between Words and Senses
Francesca Grasso, Vladimiro Lovera Rulfi, Luigi Di Caro |
EKAW | 3 |
| 2022 | Exploiting co-occurrence networks for classification of implicit inter-relationships in legal texts
Emilio Sulis, Llio Humphreys, Fabiana Vernero, Ilaria Angela Amantea, Davide Audrito, Luigi Di Caro |
Inf. Syst. | 6 |
| 2021 | Structured Semantic Modeling of Scientific Citation Intents
Roger Ferrod, Luigi Di Caro, Claudio Schifanella |
ESWC | 2 |
| 2020 | What2Cite: Unveiling Topics and Citations Dependencies for Scientific Literature Exploration and Recommendation
Davide Giosa, Luigi Di Caro |
EKAW | 2 |
| 2020 | The Role of Vocabulary Mediation to Discover and Represent Relevant Information in Privacy PoliciesabstractTo date, the effort made by existing vocabularies to provide a shared representation of the data protection domain is not fully exploited. Different natural language processing (NLP) techniques have been applied to the text of privacy policies without, however, taking advantage of existing vocabularies to provide those documents with a shared semantic superstructure. In this paper we show how a recently released domain-specific vocabulary, i.e. the Data Privacy Vocabulary (DPV), can be used to discover, in privacy policies, the information that is relevant with respect to the concepts modelled in the vocabulary itself. We also provide a machine-readable representation of this information to bridge the unstructured textual information to the formal taxonomy modelled in it. This is the first approach to the automatic processing of privacy policies that relies on the DPV, fuelling further investigation on the applicability of existing semantic resources to promote the reuse of information and the interoperability between systems in the data protection domain. Valentina Leone, Luigi Di Caro |
JURIX | 2 |
| 2020 | Populating Legal Ontologies using Semantic Role LabelingabstractThis paper is concerned with the goal of maintaining legal information and compliance systems: the ‘resource consumption bottleneck’ of creating semantic technologies manually. The use of automated information extraction techniques could significantly reduce this bottleneck. The research question of this paper is: How to address the resource bottleneck problem of creating specialist knowledge management systems? In particular, how to semi-automate the extraction of norms and their elements to populate legal ontologies? This paper shows that the acquisition paradox can be addressed by combining state-of-the-art general-purpose NLP modules with pre- and post-processing using rules based on domain knowledge. It describes a Semantic Role Labeling based information extraction system to extract norms from legislation and represent them as structured norms in legal ontologies. The output is intended to help make laws more accessible, understandable, and searchable in legal document management systems such as Eunomos (Boella et al., 2016). Llio Humphreys, Guido Boella, Luigi Di Caro, Livio Robaldo, Leon van der Torre, Sepideh Ghanavati, Robert Muthuri |
LREC | 3 |
| 2020 | Building Semantic Grams of Human KnowledgeabstractWord senses are typically defined with textual definitions for human consumption and, in computational lexicons, put in context via lexical-semantic relations such as synonymy, antonymy, hypernymy, etc. In this paper we embrace a radically different paradigm that provides a slot-filler structure, called “semagram”, to define the meaning of words in terms of their prototypical semantic information. We propose a semagram-based knowledge model composed of 26 semantic relationships which integrates features from a range of different sources, such as computational lexicons and property norms. We describe an annotation exercise regarding 50 concepts over 10 different categories and put forward different automated approaches for extending the semagram base to thousands of concepts. We finally evaluated the impact of the proposed resource on a semantic similarity task, showing significant improvements over state-of-the-art word embeddings. Valentina Leone, Giovanni Siragusa, Luigi Di Caro, Roberto Navigli |
LREC | 3 |
| 2020 | What's in a Definition? An Investigation of Semantic Features in Lexical Dictionaries
Luigi Di Caro |
WEBIST | 1 |
| 2019 | Disclosing Citation Meanings for Augmented Research Retrieval and ExplorationabstractIn recent years, new digital technologies are being used to support the navigation and the analysis of scientific publications, justified by the increasing number of articles published every year. For this reason, experts make use of on-line systems to browse thousands of articles in search of relevant information. In this paper, we present a new method that automatically assigns meanings to references on the basis of the citation text through a Natural Language Processing pipeline and a slightly-supervised clustering process. The resulting network of semantically-linked articles allows an informed exploration of the research panorama through semantic paths. The proposed approach has been validated using the ACL Anthology Dataset containing several thousands of papers related to the Computational Linguistics field. A manual evaluation on the extracted citation meanings carried to very high levels of accuracy. Finally, a freely-available web-based application has been developed and published on-line. Roger Ferrod, Claudio Schifanella, Luigi Di Caro, Mario Cataldi |
ESWC | 3 |
| 2019 | Frequent Use Cases Extraction from Legal Texts in the Data Protection DomainabstractBecause of the recent entry into force of the General Data Protection Regulation (GDPR), a growing of documents issued by the European Union institutions and authorities often mention and discuss various use cases to be handled to comply with GDPR principles.This contribution addresses the problem of extracting recurrent use cases from legal documents belonging to the data protection domain by exploiting existing Ontology Design Patterns (ODPs).An analysis of ODPs that could be looked for inside data protection related documents is provided.Moreover, a first insight on how Natural Language Processing techniques could be exploited to identify recurrent ODPs from legal texts is presented.Thus, the proposed approach aims to identify standard use cases in the data protection field at EU level to promote the reuse of existing formalisations of knowledge. Valentina Leone, Luigi Di Caro |
JURIX | 2 |
| 2019 | Real Life Application of a Question Answering System Using BERT Language ModelabstractReal life scenarios are often left untouched by the newest advances in research.They usually require the resolution of some specific task applied to a restricted domain, all the while providing small amounts of data to begin with.In this study we apply one of the newest innovations in Deep Learning to a task of text classification.The goal is to create a question answering system in Italian that provides information about a specific subject, e-invoicing and digital billing.Italy recently introduced a new legislation about e-invoicing and people have some legit doubts, therefore a large share of professionals could benefit from this tool.We gathered few pairs of question and answers; afterwards, we expanded the data, using it as a training corpus for BERT language model.Through a separate test corpus we evaluated the accuracy of the answer provided.Values show that the automatic system alone performs surprisingly well.The demo interface is hosted on Telegram, which makes the system immediately available to test. Francesca Alloatti, Luigi Di Caro, Gianpiero Sportelli |
SIGdial | 2 |
| 2018 | Ontology Development for Competence Assessment in Virtual Communities of Practice
Alice Barana, Luigi Di Caro, Michele Fioravera, Marina Marchisio, Sergio Rabellino |
AIED (2) | 2 |
| 2018 | Towards Adaptive Systems for Automatic Formative Assessment in Virtual Learning CommunitiesabstractThis paper presents a model for structuring shared resources, proposed to activate learners' formative and proactive assessment processes by enhancing the sharing-workflow among instructors. The model is based on the features of a Virtual Learning Community (VLC), and provides for the development and reuse of ontologies. The adoption of the model for the implementation of an adaptive system for the dispatching of materials is discussed on the basis of the results from clustering analyses. Semantic-similarity measures are compared by their application for clustering a collection of mathematical questions for automatic assessment, created by instructors within a national-wide VLC for Secondary Schools. Marina Marchisio, Luigi Di Caro, Michele Fioravera, Sergio Rabellino |
COMPSAC (1) | 2 |
| 2017 | A unifying similarity measure for automated identification of national implementations of european union directivesabstractThis paper presents a unifying text similarity measure (USM) for automated identification of national implementations of European Union (EU) directives. The proposed model retrieves the transposed provisions of national law at a fine-grained level for each article of the directive. USM incorporates methods for matching common words, common sequences of words and approximate string matching. It was used for identifying transpositions on a multilingual corpus of four directives and their corresponding national implementing measures (NIMs) in three different languages : English, French and Italian. We further utilized a corpus of four additional directives and their corresponding NIMs in English language for a thorough test of the USM approach. We evaluated the model by comparing our results with a gold standard consisting of official correlation tables (where available) or correspondences manually identified by domain experts. Our results indicate that USM was able to identify transpositions with average F-score values of 0.808, 0.736 and 0.708 for French, Italian and English Directive-NIM pairs respectively in the multilingual corpus. A comparison with state-of-the-art methods for text similarity illustrates that USM achieves a higher F-score and recall across both the corpora. Rohan Nanda, Luigi Di Caro, Guido Boella, Hristo Konstantinov, Tenyo Tyankov, Daniel Traykov, Hristo Hristov, Francesco Costamagna, Llio Humphreys, Livio Robaldo, Michele Romano |
ICAIL | 2 |
| 2017 | Linking European Case Law: BO-ECLI Parser, an Open Framework for the Automatic Extraction of Legal LinksabstractIn this paper we present the BO-ECLI Parser, an open framework for the extraction of legal references from case-law issued by judicial authorities of European member States. The problem of automatic legal links extraction from texts is tackled for multiple languages and jurisdictions by providing a common stack which is customizable through pluggable extensions in order to cover the linguistic diversity and specific peculiarities of national legal citation practices. The aim is to increase the availability in the public domain of machine readable references metadata for case-law by sharing common services, a guided methodology and efficient solutions to recurrent problems in legal references extraction, that reduce the effort needed by national data providers to develop their own extraction solution. Tommaso Agnoloni, Lorenzo Bacci, Ginevra Peruginelli, Marc van Opijnen, Jos van den Oever, Monica Palmirani, Luca Cervone, Octavian Bujor, Arantxa Arsuaga Lecuona, Alberto Boada García, Luigi Di Caro, Giovanni Siragusa |
JURIX | 11 |
| 2017 | Concept Recognition in European and National LawabstractThis paper presents a concept recognition system for European and national legislation. Current named entity recognition (NER) systems do not focus on identifying concepts which are essential for interpretation and harmonization of European and national law. We utilized the IATE (Inter-Active Terminology for Europe) vocabulary, a state-of-the-art named entity recognition system and Wikipedia to generate an annotated corpus for concept recognition. We applied conditional random fields (CRF) to identify concepts on a corpus of European directives and Statutory Instruments (SIs) of the United Kingdom. The CRF-based concept recognition system achieved an F1 score of 0.71 over the combined corpus of directives and SIs. Our results indicate the usability of a CRF-based learning system over dictionary tagging and state-of-the-art methods. Rohan Nanda, Giovanni Siragusa, Luigi Di Caro, Martin Theobald, Guido Boella, Livio Robaldo, Francesco Costamagna |
JURIX | 3 |
| 2017 | Legalbot: A Deep Learning-Based Conversational Agent in the Legal Domain
Kolawole John Adebayo, Luigi Di Caro, Livio Robaldo, Guido Boella |
NLDB | 2 |
| 2016 | Text Segmentation with Topic Modeling and Entity Coherence
Kolawole John Adebayo, Luigi Di Caro, Guido Boella |
HIS | 2 |
| 2016 | Neural Reasoning for Legal Text UnderstandingabstractWe propose a domain specific Question Answering system. We deviate from approaching this problem as a Textual Entailment task. We implemented a Memory Network-based Question Answering system which test a Machine's understanding of legal text and identifies whether an answer to a question is correct or wrong, given some background knowledge. We also prepared a corpus of real USA MBE Bar exams for this task. We report our initial result and direction for future works. Kolawole John Adebayo, Guido Boella, Luigi Di Caro |
JURIX | 3 |
| 2016 | A Text Similarity Approach for Automated Transposition Detection of European Union DirectivesabstractThis paper investigates the application of text similarity techniques to automatically detect the transposition of European Union (EU) directives into the national law. Currently, the European Commission (EC) resorts to time-consuming and expensive manual methods like conformity checking studies and legal analysis for identifying national transposition measures. We utilize both lexical and semantic similarity techniques and supplement them with knowledge from EuroVoc to identify transpositions. We then evaluate our approach by comparing the results with the correlation tables (gold standard). Our results indicate that both similarity techniques proved to be effective in detecting transpositions. Such systems could be used to identify the transposed provisions by both EC and legal professionals. Rohan Nanda, Luigi Di Caro, Guido Boella |
JURIX | 2 |
| 2016 | Automatic Enrichment of WordNet with Common-Sense Knowledge
Luigi Di Caro, Guido Boella |
LREC | 1 |
| 2016 | Ranking Researchers Through Collaboration Pattern Analysis
Mario Cataldi, Luigi Di Caro, Claudio Schifanella |
ECML/PKDD (3) | 2 |
| 2015 | Linking legal open data: breaking the accessibility and language barrier in european legislation and case lawabstractIn this paper we describe how the EUCases FP7 project is addressing the problem of lifting Legal Open Data to Linked Open Data to develop new applications for the legal information provision market by enriching structurally the documents (first of all with navigable references among legal texts) and semantically (with concepts from ontologies and classification). First we describe the social and economic need for breaking the accessibility barrier in legal information in the EU, then we describe the technological challenges and finally we explain how the EUCases project is addressing them by a combination of Human Language Technologies. Guido Boella, Luigi Di Caro, Michele Graziadei, Loredana Cupi, Carlo Emilio Salaroglio, Llio Humphreys, Hristo Konstantinov, Kornel Marko, Livio Robaldo, Claudio Ruffini, Kiril Ivanov Simov, Andrea Violato, Veli N. Stroetmann |
ICAIL | 2 |
| 2015 | Mapping Recitals to Normative Provisions in EU Legislation to Assist Legal InterpretationabstractThis paper looks at the use of recitals in the interpretation of EU legislation, and mechanisms for connecting them to normative provisions. The purposive approach to the interpretation of EU legislation taken by the European Court of Justice makes frequent references to recitals as helping to establish the purpose of normative provisions. Our research uses a cosine similarity based approach to link articles with relevant provisions to help legal professionals and lay end-users interpret the law. Such support can be used in legal knowledge-based systems. Llio Humphreys, Cristiana Teixeira Santos, Luigi Di Caro, Guido Boella, Leon van der Torre, Livio Robaldo |
JURIX | 3 |
| 2014 | Tell me who your friends are and I'll tell you who you are: Studying the evolution of collaborations in research environmentsabstractNowadays, many tools and systems are available to allow the analysis and the comparison of researchers' scientific production. The reason underlying such interest is evident: promotions, funding allocations, and employments are currently based on the evaluation (and direct comparisons) of publication lists. Existing measures, like H-index, aim at supporting this process by automatic calculations of quality and/or quantity indices. In this work, we propose a demonstration of a web environment, available at http://d-index.di.unito.it, that faces the problem of studying the impact of collaborations in research communities. The presented system allows the estimation of the impact of each scientific collaboration on the production of each researchers indexed by the DBLP bibliographic database by means of a novel time-based modeling of collaborative environments. The proposed application provides several interactions and visualization schemes to deeply discover clear and latent insights, over time, around the work of each researcher. Furthermore, it allows cross-community rankings of the authors depending on similar collaboration patterns and dependences. Mario Cataldi, Myriam Lamolle, Luigi Di Caro, Claudio Schifanella |
ASONAM | 3 |
| 2014 | Compliance with Multiple Regulations
Sepideh Ghanavati, Llio Humphreys, Guido Boella, Luigi Di Caro, Livio Robaldo, Leon van der Torre |
ER | 4 |
| 2014 | Exploiting networks in Law
Livio Robaldo, Guido Boella, Luigi Di Caro, Andrea Violato |
LREC | 3 |
| 2014 | KnowNow: A Serendipity-Based Educational Tool for Learning Time-Linked Knowledge
Luigi Di Caro, Livio Robaldo, Nicoletta Bersia |
ECML/PKDD (3) | 1 |
| 2014 | Learning from syntax generalizations for automatic semantic annotation
Guido Boella, Luigi Di Caro, Alice Ruggeri, Livio Robaldo |
J. Intell. Inf. Syst. | 2 |
| 2013 | A system for classifying multi-label text into EuroVocabstractIn this work we present a working system for automatic classification of text documents into the EuroVoc multilingual thesaurus. EuroVoc contains around 7,000 categories with different levels of specificity. The system relies on a simple approach for the treatment of multi-label texts where each document may have more than one associated category. The classifier is based on the well-known Support Vector Machine algorithm trained using the JRC-Acquis corpus, containing around 23,000 documents labeled with six EuroVoc categories in average. The demonstration scenario will show the ability of the system to classify documents taken on site from the Eur-Lex web portal of the European Union, together with features for visualization and navigation of the texts at different granulatity. Guido Boella, Luigi Di Caro, Daniele Rispoli, Livio Robaldo |
ICAIL | 2 |
| 2013 | Supervised Learning of Syntactic Contexts for Uncovering Definitions and Extracting Hypernym Relations in Text Databases
Guido Boella, Luigi Di Caro |
ECML/PKDD (2) | 2 |
| 2013 | Personalized emerging topic detection based on a term aging modelabstractTwitter is a popular microblogging service that acts as a ground-level information news flashes portal where people with different background, age, and social condition provide information about what is happening in front of their eyes. This characteristic makes Twitter probably the fastest information service in the world. In this article, we recognize this role of Twitter and propose a novel, user-aware topic detection technique that permits to retrieve, in real time, the most emerging topics of discussion expressed by the community within the interests of specific users. First, we analyze the topology of Twitter looking at how the information spreads over the network, taking into account the authority/influence of each active user. Then, we make use of a novel term aging model to compute the burstiness of each term, and provide a graph-based method to retrieve the minimal set of terms that can represent the corresponding topic. Finally, since any user can have topic preferences inferable from the shared content, we leverage such knowledge to highlight the most emerging topics within her foci of interest. As evaluation we then provide several experiments together with a user study proving the validity and reliability of the proposed approach. Mario Cataldi, Luigi Di Caro, Claudio Schifanella |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2012 | Multi-label Classification of Legislative Text into EuroVocabstractIn this paper we present a novel method for the automatic classification of multi-label text documents. In principle, automatic classification of text is usually tackled by supervised Machine Learning techniques like Support Vector Machines (SVM), that typically achieve state-of-the-art accuracy in several domains. Nevertheless, SVM can not handle multi-labeled documents, thus a specific preprocessing of the data is needed. In this paper we present a novel technique for the transformation of multi-label data into mono-label that is able to maintain all the information, allowing the use of standard approaches like SVM. We then evaluate our system using JRC-Acquis-it, a large dataset of italian legislation that has been manually annotated according to EuroVoc, demonstrating the potential of our approach compared to the current state of the art. Guido Boella, Luigi Di Caro, Leonardo Lesmo, Daniele Rispoli, Livio Robaldo |
JURIX | 2 |
| 2012 | D-INDEX: a web environment for analyzing dependences among scientific collaboratorsabstractIn this work, we demonstrate a web application, available at http://d-index.di.unito.it, that permits to analyze the scientific profiles of all the researchers indexed by DBLP by focusing on the collaborations that contributed to define their curricula. The presented application allows the user to analyze the profile of a researcher, her dependence degrees on all the co-authors (along her entire scientific publication history) and to make comparisons among them in terms of dependence patterns. In particular, it is possible to estimate and visualize how much a researcher has benefited from collaboration with another researcher as well as the communities in which she has been involved. Moreover, the application permits to compare, in a single chart, each researcher with all the scientists indexed in DBLP by focusing on their dependences with respect to many other parameters like the total number of papers, the number of collaborations and the length of the scientific careers. Claudio Schifanella, Luigi Di Caro, Mario Cataldi, Marie-Aude Aufaure |
KDD | 2 |
| 2012 | NLP Challenges for Eunomos a Tool to Build and Manage Legal Knowledge
Guido Boella, Luigi Di Caro, Llio Humphreys, Livio Robaldo, Leon van der Torre |
LREC | 2 |
| 2012 | PhC: Multiresolution Visualization and Exploration of Text Corpora with Parallel Hierarchical CoordinatesabstractThe high-dimensional nature of the textual data complicates the design of visualization tools to support exploration of large document corpora. In this article, we first argue that the Parallel Coordinates (PC) technique, which can map multidimensional vectors onto a 2D space in such a way that elements with similar values are represented as similar poly-lines or curves in the visualization space, can be used to help users discern patterns in document collections. The inherent reduction in dimensionality during the mapping from multidimensional points to 2D lines, however, may result in visual complications. For instance, the lines that correspond to clusters of objects that are separate in the multidimensional space may overlap each other in the 2D space; the resulting increase in the number of crossings would make it hard to distinguish the individual document clusters. Such crossings of lines and overly dense regions are significant sources of visual clutter, thus avoiding them may help interpret the visualization. In this article, we note that visual clutter can be significantly reduced by adjusting the resolution of the individual term coordinates by clustering the corresponding values. Such reductions in the resolution of the individual term-coordinates, however, will lead to a certain degree of information loss and thus the appropriate resolution for the term-coordinates has to be selected carefully. Thus, in this article we propose a controlled clutter reduction approach, called Parallel hierarchical Coordinates (or PhC ), for reducing the visual clutter in PC-based visualizations of text corpora. We define visual clutter and information loss measures and provide extensive evaluations that show that the proposed PhC provides significant visual gains (i.e., multiple orders of reductions in visual clutter) with small information loss during visualization and exploration of document collections. K. Selçuk Candan, Luigi Di Caro, Maria Luisa Sapino |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2011 | Unraveling multi-dimensional data using pDViewabstractWe present the pattern development view (pDView) system for multidimensional scientific data visualization. The pDView system relies on a novel pattern development tree (pDTree) structure to unravel patterns in multidimensional data without having to rely on visualizations that require either significant degrees of projections that eliminate certain dimensions at the expense of the others or introduce significant visual overhead due to overly-rich multi-dimensional graphic interfaces. Instead, pDView maps data along all its relevant dimensions onto a pDTree structure, capturing and visualizing the underlying fundamental relationships. The user is able to vary contextual parameters to observe the strength and robustness of these relationships under different situations. Luigi Di Caro, Maria Luisa Sapino, K. Selçuk Candan |
EDBT | 1 |
| 2010 | Analyzing the Role of Dimension Arrangement for Data Visualization in Radviz
Luigi Di Caro, Vanessa Frías-Martínez, Enrique Frías-Martínez |
PAKDD (2) | 1 |
| 2009 | CoSeNa: a context-based search and navigation systemabstractMost of the existing document and web search engines rely on keyword-based queries. To find matches, these queries are processed using retrieval algorithms that rely on word frequencies, topic recentness, document authority, and (in some cases) available ontologies. In this paper, we propose an innovative approach to exploring text collections using a novel keywords-by-concepts (KbC) graph, which supports navigation using domain-specific concepts as well as keywords that are characterizing the text corpus. The KbC graph is a weighted graph, created by tightly integrating keywords extracted from documents and concepts obtained from domain taxonomies. Documents in the corpus are associated to the nodes of the graph based on evidence supporting contextual relevance; thus, the KbC graph supports contextually informed access to these documents. In this paper, we also present CoSeNa (Context-based Search and Navigation) system that leverages the KbC model as the basis for document exploration and retrieval as well as contextually-informed media integration. Mario Cataldi, Claudio Schifanella, K. Selçuk Candan, Maria Luisa Sapino, Luigi Di Caro |
MEDES | 5 |
| 2009 | ClusTR: Exploring Multivariate Cluster Correlations and Topic Trends
Luigi Di Caro, Alejandro Jaimes |
ECML/PKDD (2) | 1 |
| 2008 | Using tagflake for condensing navigable tag hierarchies from tag cloudsabstractWe present the tagFlake system, which supports semantically informed navigation within a tag cloud. tagFlake relies on TMine for organizing tags extracted from textual content in hierarchical organizations, suitable for navigation, visualization, classification, and tracking. TMine extracts the most significant tag/terms from text documents and maps them onto a hierarchy in such a way that descendant terms are contextually dependent on their ancestors within the given corpus of documents. This provides tagFlake with a mechanism for enabling navigation within the tag space and for classification of the text documents based on the contextual structure captured by the created hierarchy. tagFlake is language neutral, since it does not rely on any natural language processing technique and is unsupervised. Luigi Di Caro, K. Selçuk Candan, Maria Luisa Sapino |
KDD | 1 |