Paola Velardi

dblp:12/2807 · DBLP profile ↗
← Back
76ranked-venue papers
16as first author
14since 2021 · last 2026
0000-0003-0884-1499ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 51 · 12 first-author · 7 since 2021Databases, data management, data science and information retrieval · 18 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 2 since 2021Systems, architecture and hardware · 5 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 V-RECS: a NL2Vis Recommender for Chart Generation with Explanations, Captioning, and Suggestions
abstract
NL2Vis (Natural Language to Visualization) is an emerging research area that involves interpreting natural language queries and translating them into visualizations that accurately represent the underlying data. It holds considerable potential for application, as it greatly facilitates data exploration for non-expert users. Following the growing use of generative AI in NL2Vis applications, we present V-RECS, the first LLM-based Visual Recommender augmented with explanations (E), captioning (C), and suggestions (S) to support further data exploration. V-RECS’ visualization narratives facilitate both response verification and data exploration by non-expert users. Furthermore, our proposed solution mitigates computational, controllability, and cost issues associated with using powerful LLMs by leveraging a methodology for effectively fine-tuning small models, such as LLama-2-7B. To generate insightful visualization narratives, we use Chain-of-Thoughts (CoT), a prompt engineering technique that helps LLMs identify and generate the logical steps to produce a correct answer. Since CoT is reported to perform poorly with small LMs, we adopted a strategy in which a large LLM (GPT-4), acting as a Teacher, generates CoT-based instructions to fine-tune a small model, Llama-2-7B, which plays the role of a Student. Extensive experiments - based on a framework for the quantitative evaluation of AI-based visualizations and on a manual assessment by a group of participants - show that V-RECS achieves performance scores comparable to GPT-4 at a much lower cost.
Luca Podo, Paola Velardi, Marco Angelini
AVI2
2026 Seamless monitoring of stress levels leveraging a foundational model for time sequences
abstract
Accurate and continuous monitoring of physiological stress is crucial, especially for patients with neurodegenerative diseases. Traditional monitoring methods, such as Electrocardiogram (ECG), are often invasive and limited in duration, while data from lightweight wearable devices, though more practical for seamless monitoring, typically suffers from significant quality degradation compared to clinical-grade measurements. The challenge lies in developing a robust, long-term, and patient-friendly stress monitoring system that overcomes the limitations of conventional approaches and the accuracy compromises of current wearables. Such a system must also provide actionable, interpretable insights for clinicians and adapt to individual patient variability. This manuscript introduces a methodology for seamless stress level monitoring by leveraging UniTS, a foundational model for time series. Our approach redefines stress detection as an anomaly detection problem, establishing a personalized baseline for each patient’s physiological behavior. Furthermore, to enhance clinical utility and trust, the system integrates a Large Language Model (LLM) to generate human-readable explanations for detected anomalies. The proposed UniTS-based methodology demonstrates superior performance, outperforming 12 top-performing methods on three benchmark datasets. Crucially, it achieves performance comparable to that obtained from more invasive, clinical-grade devices (like ECG) even when utilizing data from lightweight wearable devices, thereby enabling truly seamless monitoring. Furthermore, the system has been successfully tested in a real-world environment, in the context of a project to monitor elderly patients with cognitive disorders in their homes. This work presents an advancement in physiological stress monitoring by offering a personalized, explainable, and continuously adaptive system. We extend and fine-tune UniTS to support contextual anomaly detection and LLM-driven explainability, addressing critical gaps in current healthcare monitoring, fostering enhanced clinician control, improved system predictability, and facilitating long-term, real-world applicability for patients with neurodegenerative conditions. • Introduces a novel, personalized approach to stress detection by reformulating it as an anomaly detection task using a foundational time-series model (UniTS). • Demonstrates that lightweight wearable devices can achieve accuracy comparable to clinical-grade sensors like ECG, enabling seamless, long-term patient monitoring. • Integrates explainability through Large Language Models (LLMs), providing human-readable, clinician-oriented insights into detected anomalies. • Outperforms 12 state-of-the-art models across multiple benchmark datasets. • Validated in a real-world pilot study involving elderly patients with neurodegenerative disorders, the system shows high precision and robustness, supporting its deployment in home-care environments for vulnerable populations.
Davide Gabrielli, Bardh Prenkaj, Paola Velardi
Artif. Intell. Medicine3
2025 AI on the Pulse: Real-Time Health Anomaly Detection with Wearable and Ambient Intelligence
abstract
We introduce AI on the Pulse, a real-world-ready anomaly detection system that continuously monitors patients using a fusion of wearable sensors, ambient intelligence, and advanced AI models. Powered by UniTS, a state-of-the-art (SoTA) universal time-series model, our framework autonomously learns each patient's unique physiological and behavioral patterns, detecting subtle deviations that signal potential health risks. Unlike classification methods that require impractical, continuous labeling in real-world scenarios, our approach uses anomaly detection to provide real-time, personalized alerts for reactive home-care interventions. Our approach outperforms 12 SoTA anomaly detection methods, demonstrating robustness across both high-fidelity medical devices (ECG) and consumer wearables, with a ~22% improvement in F1 score. However, the true impact of AI on the Pulse lies in @HOME, where it has been successfully deployed for continuous, real-world patient monitoring. By operating with non-invasive, lightweight devices like smartwatches, our system proves that high-quality health monitoring is possible without clinical-grade equipment. Beyond detection, we enhance interpretability by integrating LLMs, translating anomaly scores into clinically meaningful insights for healthcare professionals.
Davide Gabrielli, Bardh Prenkaj, Paola Velardi, Stefano Faralli 0001
CIKM3
2025 TRADES: Generating Realistic Market Simulations with Diffusion Models
abstract
Financial markets are complex systems characterized by high statistical noise, nonlinearity, volatility, and constant evolution. Thus, modeling them is extremely hard. Here, we address the task of generating realistic and responsive Limit Order Book (LOB) market simulations, which are fundamental for calibrating and testing trading strategies, performing market impact experiments, and generating synthetic market data. We propose a novel TRAnsformer-based Denoising Diffusion Probabilistic Engine for LOB Simulations (TRADES). TRADES generates realistic order flows as time series conditioned on the state of the market, leveraging a transformer-based architecture that captures the temporal and spatial characteristics of high-frequency market data. There is a notable absence of quantitative metrics for evaluating generative market simulation models in the literature. To tackle this problem, we adapt the predictive score, a metric measured as an MAE, to market data by training a stock price predictive model on synthetic data and testing it on real data. We compare TRADES with previous works on two stocks, reporting a ×3.27 and ×3.48 improvement over SoTA according to the predictive score, demonstrating that we generate useful synthetic market data for financial downstream tasks. Furthermore, we assess TRADES’s market simulation realism and responsiveness, showing that it effectively learns the conditional data distribution and successfully reacts to an experimental agent, giving sprout to possible calibrations and evaluations of trading strategies and market impact experiments. To perform the experiments, we developed DeepMarket, the first open-source Python framework for LOB market simulation with deep learning. In our repository, we include a synthetic LOB dataset composed of TRADES’s generated simulations.
Leonardo Berti, Bardh Prenkaj, Paola Velardi
ECAI3
2025 Agnostic Visual Recommendation Systems: Open Challenges and Future Directions
abstract
Visualization Recommendation Systems (VRSs) are a novel and challenging field of study aiming to help generate insightful visualizations from data and support non-expert users in information discovery. Among the many contributions proposed in this area, some systems embrace the ambitious objective of imitating human analysts to identify relevant relationships in data and make appropriate design choices to represent these relationships with insightful charts. We denote these systems as "agnostic" VRSs since they do not rely on human-provided constraints and rules but try to learn the task autonomously. Despite the high application potential of agnostic VRSs, their progress is hindered by several obstacles, including the absence of standardized datasets to train recommendation algorithms, the difficulty of learning design rules, and defining quantitative criteria for evaluating the perceptual effectiveness of generated plots. This article summarizes the literature on agnostic VRSs and outlines promising future research directions.
Luca Podo, Bardh Prenkaj, Paola Velardi
IEEE Trans. Vis. Comput. Graph.3
2024 Unsupervised Detection of Behavioural Drifts With Dynamic Clustering and Trajectory Analysis
abstract
Real-time monitoring of human behaviours, especially in e-Health applications, has been an active area of research in the past decades. On top of IoT-based sensing environments, anomaly detection algorithms have been proposed for the early detection of abnormalities. Gradual change procedures, commonly referred to as drift anomalies, have received much less attention in the literature because they represent a much more challenging scenario than sudden temporary changes (point anomalies). In this article, we propose, for the first time, a fully unsupervised real-time drift detection algorithm named DynAmo, which can identify drift periods as they are happening. DynAmo comprises a dynamic clustering component to capture the overall trends of monitored behaviours and a trajectory generation component, which extracts features from the densest cluster centroids. Finally, we apply an ensemble of divergence tests on sliding reference and detection windows to detect drift periods in the behavioural sequence.
Bardh Prenkaj, Paola Velardi
IEEE Trans. Knowl. Data Eng.2
2023 A self-supervised algorithm to detect signs of social isolation in the elderly from daily activity sequences
Bardh Prenkaj, Dario Aragona, Alessandro Flaborea, Fabio Galasso, Saverio Gravina, Luca Podo, Emilia Reda, Paola Velardi
Artif. Intell. Medicine8
2023 A Benchmark Study on Knowledge Graphs Enrichment and Pruning Methods in the Presence of Noisy Relationships
abstract
In the past few years, knowledge graphs (KGs), as a form of structured human intelligence, have attracted considerable research attention from academia and industry. In this very active field of study, a widely explored problem is that of link prediction, the task of predicting whether two nodes should be connected, based on node attributes and local or global graph connectivity properties. The state of the art in this area is represented by techniques based on graph embeddings. However, KGs, especially those acquired using automated or partly automated techniques, are often riddled with noise, e.g., wrong relationships, which makes the problem of link deletion as important as that of link prediction. In this paper, we address three main research questions. The first is about the true effectiveness of different knowledge graph embedding models under the presence of an increasing number of wrong links. The second is to asses if methods that can predict unknown relationships effectively, work equally well in recognizing incorrect relations. The third is to verify if there are systems robust enough to maintain primacy in all experimental conditions. To answer these research questions, we performed a systematic benchmark study in which the experimental setting includes ten state-of-the-art models, three common KG datasets with different structural properties and three downstream tasks: the widely explored tasks of link prediction and triple classification, and the less popular task of link deletion. Comparative studies often yield contradictory results, where the same systems score better or worse depending on the experimental context. In our work, in order to facilitate the discovery of clear performance patterns and their interpretation, we select and/or aggregate performance data to highlight each specific comparison dimension: dataset complexity, type of task, category of models, and robustness against noise.
Stefano Faralli 0001, Andrea Lenzi, Paola Velardi
J. Artif. Intell. Res.3
2022 AnomalyByClick: An Interactive Visualization Tool for Monitoring Activities of Daily Living and Anomaly Annotation
abstract
We present AnomalyByClick, an interactive visualization system that allows the monitoring and analysis of anomalies in the behavior of elderly patients during daily activities in a living environment.
Luca Podo, Paola Velardi
AVI2
2022 Plotly.plus, an Improved Dataset for Visualization Recommendation
abstract
Visualization recommendation is a novel and challenging field of study, whose aim is to provide non-expert users with automatic tools for insight discovery from data. Advances in this research area are hindered by the absence of reliable datasets on which to train the recommender systems. To the best of our knowledge, Plotly corpus is the only publicly available dataset, but as complained by many authors and discussed in this article, it contains many labeling errors, which greatly limits its usefulness. We release an improved version of the original dataset, named Plotly.plus, which we obtained through an automated procedure with minimal post-editing. In addition to a manual validation by a group of data science students, we demonstrate that when training two state-of-the-art abstract image classifiers on Plotly.plus, systems' performance improves more than twice as much as when the original dataset is used, showing that Plotly.plus facilitates the discovery of significant perceptual patterns.
Luca Podo, Paola Velardi
CIKM2
2022 A Large Interlinked Knowledge Graph of the Italian Cultural Heritage
abstract
Knowledge is the lifeblood for a plethora of applications such as search, recommender systems and natural language understanding. Thanks to the efforts in the fields of Semantic Web and Linked Open Data a growing number of interlinked knowledge bases are supporting the development of advanced knowledge-based applications. Unfortunately, for a large number of domain-specific applications, these knowledge bases are unavailable. In this paper, we present a resource consisting of a large knowledge graph linking the Italian cultural heritage entities (defined in the ArCo ontology) with the concepts defined on well-known knowledge bases (i.e., DBpedia and the Getty GVP ontology). We describe the methodologies adopted for the semi-automatic resource creation and provide an in-depth analysis of the resulting interlinked graph.
Stefano Faralli 0001, Andrea Lenzi, Paola Velardi
LREC3
2022 Supporting Personalized Health Care With Social Media Analytics: An Application to Hypothyroidism
abstract
Social media analytics can considerably contribute to understanding health conditions beyond clinical practice, by capturing patients’ discussions and feelings about their quality of life in relation to disease treatments. In this article, we propose a methodology to support a detailed analysis of the therapeutic experience in patients affected by a specific disease, as it emerges from health forums. As a use case to test the proposed methodology, we analyze the experience of patients affected by hypothyroidism and their reactions to standard therapies. Our approach is based on a data extraction and filtering pipeline, a novel topic detection model named Generative Text Compression with Agglomerative Clustering Summarization ( GTCACS ), and an in-depth data analytic process. We advance the state of the art on automated detection of adverse drug reactions ( ADRs ) since, rather than simply detecting and classifying positive or negative reactions to a therapy, we are capable of providing a fine characterization of patients along different dimensions, such as co-morbidities, symptoms, and emotional states.
Giorgio Grani, Andrea Lenzi, Paola Velardi
ACM Trans. Comput. Heal.3
2022 A Network-Based Analysis of Disease Modules From a Taxonomic Perspective
abstract
OBJECTIVE: Human-curated diseaseontologies are widely used for diagnostic evaluation, treatment and data comparisons over time, and clinical decision support. The classification principles underlying these ontologies are guided by the analysis of observable pathological similarities between disorders, often based on anatomical or histological principles. Although, thanks to recent advances in molecular biology, disease ontologies are slowly changing to integrate the etiological and genetic origins of diseases, nosology still reflects this "reductionist" perspective. Proximity relationships of disease modules (hereafter DMs) in the human interactome network are now increasingly used in diagnostics, to identify pathobiologically similar diseases and to support drug repurposing and discovery. On the other hand, similarity relations induced from structural proximity of DMs also have several limitations, such as incomplete knowledge of disease-gene relationships and reliability of clinical trials to assess their validity. The purpose of the study described in this paper is to shed more light on disease similarities by analyzing the relationship between categorical proximity of diseases in human-curated ontologies and structural proximity of the related DMs in the interactome. METHOD: We propose a method (and related algorithms) to automatically induce a hierarchical structure from proximity relations between DMs, and to compare this structure with a human-curated disease taxonomy. RESULTS: We demonstrate that the proposed method allows to systematically analyze commonalities and differences among structural and categorical similarity of human diseases, help refine and extend human disease classification systems, and identify promising network areas where new disease-gene interactions can be discovered.
Giorgio Grani, Lorenzo Madeddu, Paola Velardi
IEEE J. Biomed. Health Informatics3
2021 Hidden space deep sequential risk prediction on student trajectories
Bardh Prenkaj, Damiano Distante, Stefano Faralli 0001, Paola Velardi
Future Gener. Comput. Syst.4
2020 A Reproducibility Study of Deep and Surface Machine Learning Methods for Human-related Trajectory Prediction
abstract
In this paper, we compare several deep and surface state-of-the-art machine learning methods for risk prediction in problems that can be modelled as a trajectory of events separated by irregular time intervals. Trajectories are the abstract representation of many real-life data, such as patient records, student e-tivities, online financial transactions, and many others. Given the continuously increasing number of machine learning methods to predict future high-risk events in these contexts, we aim to provide more insight into reproducibility and applicability of these methods when changing datasets, parameters, and evaluation measures. As an additional contribution, we release to the community the implementations of all compared methods.
Bardh Prenkaj, Paola Velardi, Damiano Distante, Stefano Faralli 0001
CIKM2
2020 Multiple Knowledge GraphDB (MKGDB)
abstract
We present MKGDB, a large-scale graph database created as a combination of multiple taxonomy backbones extracted from 5 existing knowledge graphs, namely: ConceptNet, DBpedia, WebIsAGraph, WordNet and the Wikipedia category hierarchy. MKGDB, thanks the versatility of the Neo4j graph database manager technology, is intended to favour and help the development of open-domain natural language processing applications relying on knowledge bases, such as information extraction, hypernymy discovery, topic clustering, and others. Our resource consists of a large hypernymy graph which counts more than 37 million nodes and more than 81 million hypernymy relations.
Stefano Faralli 0001, Paola Velardi, Farid Yusifli
LREC2
2019 Predicting Disease Genes Using Connectivity and Functional Features
abstract
We predict disease-genes relations on the human interactome network using a methodology that jointly learns functional and connectivity patterns surrounding proteins. To exploit at best latent information in the network, we propose an extended version of random walks, named Random Watcher-Walker (RW2), which is shown to perform better than other state-of-the-art algorithms. We also show that performance of RW2and other compared state-of-the-art algorithms is extremely sensitive to the interactome used, and to the adopted disease categorizations, since this influences the ability to capture regularities in presence of sparsity and incompleteness.
Lorenzo Madeddu, Giovanni Stilo, Paola Velardi
BIBM3
2019 The social phenotype: Extracting a patient-centered perspective of diabetes from health-related blogs
Andrea Lenzi, Marianna Maranghi, Giovanni Stilo, Paola Velardi
Artif. Intell. Medicine4
2019 A topic recommender for journalists
Alessandro Cucchiarelli, Christian Morbidoni, Giovanni Stilo, Paola Velardi
Inf. Retr. J.4
2018 Efficient Pruning of Large Knowledge Graphs
abstract
In this paper we present an efficient and highly accurate algorithm to prune noisy or over-ambiguous knowledge graphs given as input an extensional definition of a domain of interest, namely as a set of instances or concepts. Our method climbs the graph in a bottom-up fashion, iteratively layering the graph and pruning nodes and edges in each layer while not compromising the connectivity of the set of input nodes. Iterative layering and protection of pre-defined nodes allow to extract semantically coherent DAG structures from noisy or over-ambiguous cyclic graphs, without loss of information and without incurring in computational bottlenecks, which are the main problem of state-of-the-art methods for cleaning large, i.e., Web-scale, knowledge graphs. We apply our algorithm to the tasks of pruning automatically acquired taxonomies using benchmarking data from a SemEval evaluation exercise, as well as the extraction of a domain-adapted taxonomy from the Wikipedia category hierarchy. The results show the superiority of our approach over state-of-art algorithms in terms of both output quality and computational efficiency.
Stefano Faralli 0001, Irene Finocchi, Simone Paolo Ponzetto, Paola Velardi
IJCAI4
2018 A Large Multilingual and Multi-domain Dataset for Recommender Systems
Giorgia Di Tommaso, Stefano Faralli 0001, Paola Velardi
LREC3
2018 Wiki-MID: A Very Large Multi-domain Interests Dataset of Twitter Users with Mappings to Wikipedia
Giorgia Di Tommaso, Stefano Faralli 0001, Giovanni Stilo, Paola Velardi
ISWC (2)4
2018 CrumbTrail: An efficient methodology to reduce multiple inheritance in knowledge graphs
Stefano Faralli 0001, Irene Finocchi, Simone Paolo Ponzetto, Paola Velardi
Knowl. Based Syst.4
2017 Detecting network leaders in enterprises
abstract
This paper describes an interdisciplinary study aimed at analyzing leadership in less formal collaboration environments, such as enterprise social networks (ESNs). To conduct our research, we defined a measure of network leadership which draws on organization theory and on a computational model based on multiplex networks. This model, along with a social network analysis toolkit developed in the context of the present study, enabled the systematic empirical analysis of a large ESN, as a function of gender, time, roles, and discussed topics.
Giorgia Di Tommaso, Giovanni Stilo, Paola Velardi
CSCWD3
2017 A Gendered Analysis of Leadership in Enterprise Social Networks
Giorgia Di Tommaso, Giovanni Stilo, Paola Velardi
ICWSM3
2017 Hashtag Sense Clustering Based on Temporal Similarity
abstract
Hashtags are creative labels used in micro-blogs to characterize the topic of a message/discussion. Regardless of the use for which they were originally intended, hashtags cannot be used as a means to cluster messages with similar content. First, because hashtags are created in a spontaneous and highly dynamic way by users in multiple languages, the same topic can be associated with different hashtags, and conversely, the same hashtag may refer to different topics in different time periods. Second, contrary to common words, hashtag disambiguation is complicated by the fact that no sense catalogs (e.g., Wikipedia or WordNet) are available; and, furthermore, hashtag labels are difficult to analyze, as they often consist of acronyms, concatenated words, and so forth. A common way to determine the meaning of hashtags has been to analyze their context, but, as we have just pointed out, hashtags can have multiple and variable meanings. In this article, we propose a temporal sense clustering algorithm based on the idea that semantically related hashtags have similar and synchronous usage patterns.
Giovanni Stilo, Paola Velardi
Comput. Linguistics2
2017 Automatic acquisition of a taxonomy of microblogs users' interests
Stefano Faralli 0001, Giovanni Stilo, Paola Velardi
J. Web Semant.3
2016 Efficient temporal mining of micro-blog texts and its application to event discovery
Giovanni Stilo, Paola Velardi
Data Min. Knowl. Discov.2
2015 Large Scale Homophily Analysis in Twitter Using a Twixonomy
Stefano Faralli 0001, Giovanni Stilo, Paola Velardi
IJCAI3
2014 Temporal Semantics: Time-Varying Hashtag Sense Clustering
Giovanni Stilo, Paola Velardi
EKAW2
2014 Twitter mining for fine-grained syndromic surveillance
Paola Velardi, Giovanni Stilo, Alberto Eugenio Tozzi, Francesco Gesualdo
Artif. Intell. Medicine1
2013 OntoLearn Reloaded: A Graph-Based Algorithm for Taxonomy Induction
abstract
In 2004 we published in this journal an article describing OntoLearn, one of the first systems to automatically induce a taxonomy from documents and Web sites. Since then, OntoLearn has continued to be an active area of research in our group and has become a reference work within the community. In this paper we describe our next-generation taxonomy learning methodology, which we name OntoLearn Reloaded. Unlike many taxonomy learning approaches in the literature, our novel algorithm learns both concepts and relations entirely from scratch via the automated extraction of terms, definitions, and hypernyms. This results in a very dense, cyclic and potentially disconnected hypernym graph. The algorithm then induces a taxonomy from this graph via optimal branching and a novel weighting policy. Our experiments show that we obtain high-quality results, both when building brand-new taxonomies and when reconstructing sub-hierarchies of existing taxonomies.
Paola Velardi, Stefano Faralli 0001, Roberto Navigli
Comput. Linguistics1
2012 A New Method for Evaluating Automatically Learned Terminological Taxonomies
Paola Velardi, Roberto Navigli, Stefano Faralli 0001, Juana María Ruiz-Martínez
LREC1
2011 A Graph-Based Algorithm for Inducing Lexical Taxonomies from Scratch
abstract
In this paper we present a graph-based approach aimed at learning a lexical taxonomy automatically starting from a domain corpus and the Web. Unlike many taxonomy learning approaches in the literature, our novel algorithm learns both concepts and relations entirely from scratch via the automated extraction of terms, definitions and hypernyms. This results in a very dense, cyclic and possibly disconnected hypernym graph. The algorithm then induces a taxonomy from the graph. Our experiments show that we obtain high-quality results, both when building brand-new taxonomies and when reconstructing WordNet sub-hierarchies. 1
Roberto Navigli, Paola Velardi, Stefano Faralli 0001
IJCAI2
2010 Learning Word-Class Lattices for Definition and Hypernym Extraction
Roberto Navigli, Paola Velardi
ACL2
2010 An Annotated Dataset for Extracting Definitions and Hypernyms from the Web
Roberto Navigli, Paola Velardi, Juana María Ruiz-Martínez
LREC2
2008 Content-Based Social Network Analysis
abstract
Relationships among actors in traditional social network analysis are modelled as a function of the quantity of relations (co-authorships, business relations, friendship, etc.). In contrast, within a business, social or research community, network analysts are interested in the communicative content exchanged by the community members, not merely in the number of relationships. In order to meet this need, this paper presents a novel social network model, in which the actors are not simply represented through the intensity of their mutual relationships, but also through the analysis and evolution of their shared interests. Text mining and clustering techniques are used to capture the content of communication and to identify the most popular topics.
Paola Velardi, Roberto Navigli, Alessandro Cucchiarelli, Mirco Curzi
ECAI1
2008 Advancing Topic Ontology Learning through Term Extraction
Blaz Fortuna, Nada Lavrac, Paola Velardi
PRICAI3
2008 Modeling Collaborations Content in Social Network Analysis
abstract
This paper presents a methodology and a software application to support the analysis of collaborations and collaboration content in scientific communities. High quality terminology extraction, semantic graphs and clustering techniques are used to identify the relevant research topics. Social analysis tools are then used to study the emergence of interests around certain topics, the evolution of collaborations around these themes, and to identify potential for better cooperation.
Alessandro Cucchiarelli, Paola Velardi, Fulvio D'Antonio, Mirco Curzi
Web Intelligence2
2008 Monitoring the status of a research community through a Knowledge Map
abstract
The NoE (Network of Excellence) INTEROP is an instrument for strengthening the excellence of European research in interoperability of enterprise applications, by bringing together the complementary competences required to develop interoperability in
Paola Velardi, Alessandro Cucchiarelli, Fulvio D'Antonio
Web Intell. Agent Syst.1
2007 Semantic Indexing of a Competence Map to Support Scientific Collaboration in a Research Community
Paola Velardi, Roberto Navigli, Michaël Petit
IJCAI1
2007 A Semantically Enriched Competency Management System to Support the Analysis of a Web-based Research Network
abstract
While it is generally acknowledged that domain ontologies can significantly improve knowledge management systems (KMS) within organizations and among distributed web communities, we have little evidence of operational ontology-based KMS and their practical utility in real settings in the literature. We describe here the INTEROP KMap, a fully implemented, semantically indexed, competency management system, used to facilitate research collaboration and coordination of a Network of Excellence (NoE) on Enterprise Interoperability. Since the main highlighted advantages of ontologies are improved information access and interoperability, our aim in this paper is to give experimental support to these claims. We provide a summary description and usage data on the KMap, as well as experiments to quantify the added value of semantic search wrt traditional document ranking measures.
Paola Velardi, Alessandro Cucchiarelli, Michaël Petit
Web Intelligence1
2007 A Taxonomy Learning Method and Its Application to Characterize a Scientific Web Community
abstract
The need to extract and manage domain-specific taxonomies has become increasingly relevant in recent years. A taxonomy is a form of business intelligence used to integrate information, reduce semantic heterogeneity, describe emergent communities and interest groups, and facilitate communication between information systems. We present a semiautomated strategy to extract domain-specific taxonomies from Web documents and its application to model a network of excellence in the emerging research field of enterprise interoperability
Paola Velardi, Alessandro Cucchiarelli, Michaël Petit
IEEE Trans. Knowl. Data Eng.1
2006 Ontology Enrichment Through Automatic Semantic Annotation of On-Line Glossaries
Roberto Navigli, Paola Velardi
EKAW2
2005 Structural Semantic Interconnections: A Knowledge-Based Approach to Word Sense Disambiguation
abstract
Word Sense Disambiguation (WSD) is traditionally considered an Al-hard problem. A break-through in this field would have a significant impact on many relevant Web-based applications, such as Web information retrieval, improved access to Web services, information extraction, etc. Early approaches to WSD, based on knowledge representation techniques, have been replaced in the past few years by more robust machine learning and statistical techniques. The results of recent comparative evaluations of WSD systems, however, show that these methods have inherent limitations. On the other hand, the increasing availability of large-scale, rich lexical knowledge resources seems to provide new challenges to knowledge-based approaches. In this paper, we present a method, called structural semantic interconnections (SSI), which creates structural specifications of the possible senses for each word in a context and selects the best hypothesis according to a grammar G, describing relations between sense specifications. Sense specifications are created from several available lexical resources that we integrated in part manually, in part with the help of automatic procedures. The SSI algorithm has been applied to different semantic disambiguation problems, like automatic ontology population, disambiguation of sentences in generic texts, disambiguation of words in glossary definitions. Evaluation experiments have been performed on specific knowledge domains (e.g., tourism, computer networks, enterprise interoperability), as well as on standard disambiguation test sets.
Roberto Navigli, Paola Velardi
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Quantitative and Qualitative Evaluation of the OntoLearn Ontology Learning System
Roberto Navigli, Paola Velardi, Alessandro Cucchiarelli, Francesca Neri
COLING2
2004 Automatic Generation of Glosses in the OntoLearn System
Alessandro Cucchiarelli, Roberto Navigli, Francesca Neri, Paola Velardi
LREC4
2004 Learning Domain Ontologies from Document Warehouses and Dedicated Web Sites
abstract
We present a method and a tool, OntoLearn, aimed at the extraction of domain ontologies from Web sites, and more generally from documents shared among the members of virtual organizations. OntoLearn first extracts a domain terminology from available documents. Then, complex domain terms are semantically interpreted and arranged in a hierarchical fashion. Finally, a general-purpose ontology, WordNet, is trimmed and enriched with the detected domain concepts. The major novel aspect of this approach is semantic interpretation, that is, the association of a complex concept with a complex term. This involves finding the appropriate WordNet concept for each word of a terminological string and the appropriate conceptual relations that hold among the concept components. Semantic interpretation is based on a new word sense disambiguation algorithm, called structural semantic interconnections.
Roberto Navigli, Paola Velardi
Comput. Linguistics2
2003 Text Mining Techniques to Automatically Enrich a Domain Ontology
Michele Missikoff, Paola Velardi, Paolo Fabriani
Appl. Intell.2
2002 Automatic Adaptation of WordNet to Domains
Roberto Navigli, Paola Velardi
LREC2
2002 The Usable Ontology: An Environment for Building and Assessing a Domain Ontology
Michele Missikoff, Roberto Navigli, Paola Velardi
ISWC3
2001 Using text processing techniques to automatically enrich a domain ontology
abstract
Abstract- Though the utility of domain Ontologies is now widely acknowledged in an increasing number of domains, several barriers must be overcome before Ontologies become practical and useful tools. A critical issue is the task of identifying, defining, and entering the concept definitions. In case of large and complex application domains this task can be lengthy, costly, and controversial (since different persons may have different points of view about the same concept). To reduce time, cost (and, sometimes, harsh discussions) it is highly advisable to refer, in constructing or updating an ontology, to the documents available in the field. In this paper we describe OntoLearn, a text-mining tool devised to improve human productivity during the process of ontology construction. 1.
Paola Velardi, Paolo Fabriani, Michele Missikoff
FOIS1
2001 Unsupervised Named Entity Recognition Using Syntatic and Semantic Contextual Evidence
abstract
Proper nouns form an open class, making the incompleteness of manually or automatically learned classification rules an obvious problem. The purpose of this paper is twofold: first, to suggest the use of a complementary “backup” method to increase the robustness of any hand-crafted or machine-learning-based NE tagger; and second, to explore the effectiveness of using more fine-grained evidence—namely, syntactic and semantic contextual knowledge—in classifying NEs.
Alessandro Cucchiarelli, Paola Velardi
Comput. Linguistics2
2000 A Theoretical Analysis of Context-based Learning Algorithms or Word Sense Disambiguation
Paola Velardi, Alessandro Cucchiarelli
ECAI1
2000 Will Very Large Corpora Play For Semantic Disambiguation The Role That Massive Computing Power Is Playing For Other AI-Hard Problems?
Alessandro Cucchiarelli, Enrico Faggioli, Paola Velardi
LREC3
2000 Automatic adaptation of proper noun dictionaries through cooperation of machine learning and probabilistic methods
abstract
The recognition of Proper Nouns (PNs) is considered an important task in the area of Information Retrieval and Extraction. However the high performance of most existing PN classifiers heavily depends upon the availability of large dictionaries of domain-specific Proper Nouns, and a certain amount of manual work for rule writing or manual tagging. Though it is not a heavy requirement to rely on some existing PN dictionary (often these resources are available on the web), its coverage of a domain corpus may be rather low, in absence of manual updating. In this paper we propose a technique for the automatic updating of an PN Dictionary through the cooperation of an inductive and a probabilistic classifier. In our experiments we show that, whenever an existing PN Dictionary allows the identification of 50% of the proper nouns within a corpus, our technique allows, without additional manual effort, the successful recognition of about 90% of the remaining 50%.
Georgios Petasis, Alessandro Cucchiarelli, Paola Velardi, Georgios Paliouras, Vangelis Karkaletsis, Constantine D. Spyropoulos
SIGIR3
1999 Semantic tagging of unknown proper nouns
abstract
In this paper, we describe a context-based method to semantically tag unknown proper nouns (U-PNs) in corpora. Like many others, our system relies on a gazetteer and a set of context-dependent heuristics to classify proper nouns. However, proper nouns are an open-end class: when parsing new fragments of a corpus, even in the same language domain, we can expect that several proper nouns cannot be semantically tagged. The algorithm that we propose assigns to an unknown PN an entity type based on the analysis of syntactically and semantically similar contexts already seen in the application corpus. The performance of the algorithm is evaluated not only in terms of precision, following the tradition of MUC conferences, but also in terms of information gain, an information theoretic measure that takes into account the complexity of the classification task.
Alessandro Cucchiarelli, Danilo Luzi, Paola Velardi
Nat. Lang. Eng.3
1998 Using corpus evidence for automatic gazetteer extension
Alessandro Cucchiarelli, Danilo Luzi, Paola Velardi
LREC3
1998 Finding a domain-appropriate sense inventory for semantically tagging a corpus
Alessandro Cucchiarelli, Paola Velardi
Nat. Lang. Eng.2
1996 Unsupervised Learning of Syntactic Knowledge: Methods and Measures
Roberto Basili 0001, Alessandro Marziali, Maria Teresa Pazienza, Paola Velardi
EMNLP4
1996 An Empirical Symbolic Approach to Natural Language Processing
Roberto Basili 0001, Maria Teresa Pazienza, Paola Velardi
Artif. Intell.3
1996 Integrating General-purpose and Corpus-based Verb Classification
Roberto Basili 0001, Paola Velardi, Maria Teresa Pazienza
Comput. Linguistics2
1994 A "not-so-shallow" parser for collocational analysis
Roberto Basili 0001, Maria Teresa Pazienza, Paola Velardi
COLING3
1993 What can be learned from raw texts?
Roberto Basili 0001, Maria Teresa Pazienza, Paola Velardi
Mach. Transl.3
1993 Acquisition of selectional patterns in sublanguages
Roberto Basili 0001, Maria Teresa Pazienza, Paola Velardi
Mach. Transl.3
1991 How to Encode Semantic Knowledge: A Method for Meaning Representation and Computer-Aided Acquisition
Paola Velardi, Maria Teresa Pazienza, Michela Fasolo
Comput. Linguistics1
1990 Why Human Translators Still Sleep In Peace? (Four Engineering And Linguistic Gaps In NLP)
Paola Velardi
COLING1
1989 Computer Aided Interpretation of Lexical Coocurrences
abstract
This paper addresses the problem of developing a large semantic lexicon for natural language processing. The increasing availability of machine readable documents offers an opportunity to the field of lexical semantics, by providing experimental evidence of word uses (on-line texts) and word definitions (on-line dictionaries).The system presented hereafter, PETRARCA, detects word cooccurrences from a large sample of press agency releases on finance and economics, and uses these associations to build a case-based semantic lexicon. Syntáctically valid cooccurences including a new word W are detected by a high-coverage morphosyntactic analyzer. Syntactic relations are interpreted e.g. replaced by case relations, using a a catalogue of patterns/interpretation pairs, a concept type hierarchy, and a set of selectional restriction rules on semantic interpretation types.
Paola Velardi, Maria Teresa Pazienza
ACL1
1989 A System for Text Analysis and Lexical Knowledge Acquisition
Francesco Antonacci, M. Russo, Maria Teresa Pazienza, Paola Velardi
Data Knowl. Eng.4
1988 Interfaces for Advanced Data Base Systems: Iconic, Graphical, Natural Language Oriented - Panel
Stefano Spaccapietra, Yirki Nummenmaa, Rainer Melchert, Michael Schrefl, Paola Velardi
ER5
1987 A Structured Representation Of Word-Senses For Semantic Analysis
Maria Teresa Pazienza, Paola Velardi
EACL2
1986 Reliability analysis of multipath interconnection networks
Paola Velardi, Alessandro Forcina
Microprocessing and Microprogramming1
1985 Hardware-Related Software Errors: Measurement and Analysis
abstract
This paper describes an analysis of hardware-related software (HW/SW) errors on an MVS/SP operating system at Stanford University. The analysis procedure demonstrates a methodology for evaluating the interaction between hardware and software as it relates to system reliability. The paper examines the operating system's handling of HW/SW errors and also the effectiveness of recovery management. Nearly 35 percent of all observed software failures were found to be hareware-related. The analysis shows that the operating system is seldom able to diagnose that a software error may be hardware-related. The impact of HW/SW errors on the system is evaluated by measuring the effectiveness of system recovery in containing the propagation of HW/SW errors. The system failure probability for HW/SW errors is close to three times that for software errors in general. The observed HW/SW errors are seen to have a specific pattern, suggesting the possibility of the use of such error patterns for intelligent error prediction and recovery.
Ravishankar K. Iyer, Paola Velardi
IEEE Trans. Software Eng.2
1984 A Study of Software Failures and Recovery in the MVS Operating System
abstract
This paper describes an analysis of system detected software errors on the MVS operating system at the Center for Information Technology (CIT), Stanford University. The analysis procedure demonstrates a methodology by which systems with automatic recovery features can be evaluated. Most common error categories are determined and related to the program in execution at the time of the error. The severity of the error is measured by evaluating the criticality of the program for continued system operation. The system recovery and error correction features are then analyzed and an estimate of the system fault tolerance to errors of different levels of severity is made.
Paola Velardi, Ravishankar K. Iyer
IEEE Trans. Computers1
1983 A fully distributed arbiter for multiprocessor systems
Giacomo Cioffi, Paola Velardi
Microprocessing and Microprogramming2
1983 Recovery blocks for communicating systems
Paola Velardi, Bruno Ciciani
Microprocessing and Microprogramming1