Elio Masciari

dblp:28/168 · DBLP profile ↗
← Back
69ranked-venue papers in the field
16as first author
18since 2021 · last 2025
0000-0002-1778-5321ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 49 (11 first)Knowledge Engineering, Semantic Web & Information Systems · 12 (3 first)Data Mining & Knowledge Discovery · 4 (1 first)Business Process & Enterprise Data · 2Information Retrieval & Web Search · 1 (1 first)Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2025 A Hybrid Approach to Estimating AI Carbon Emissions
Salvatore Borraccia, Elio Masciari, Enea Vincenzo Napolitano
DEXA (1)2
2025 From Sound to Success: An AI Framework for Predicting Music Popularity and Sentiment Analysis
Simona Fioretto, Elio Masciari, Enea Vincenzo Napolitano
MEDI2
2025 A comparative analysis of predictive process monitoring: object-centric versus classical event logs
abstract
Abstract Predictive Process Monitoring (PPM) techniques are emerging as part of the more general research scenario of Process Mining (PM). They play a crucial role in the continuously evolving process of digital transformation by constantly supporting the organizational decision-making processes providing (accurate) predictions on the future behavior of processes. The state of the art of PPM application methodologies is mainly focused on Single ID Event Logs, commonly known as Traditional Event Logs or Classical Event Logs. As a matter of fact, the importance of Object-Centric Event Logs (OCEL) is being increasingly recognized as many emerging PPM approaches benefited of the usage of OCEL by obtaining a significative increase of the prediction accuracy. This survey aims to explore the current proposals in the context of OCEL-based PPM approaches. More in detail, we contribute to the state of the art by adding new classification features by differentiating between the approaches based on the input Event Log (Traditional or OCEL). We also analyzed the existing literature considering the prediction task addressed, the methodology used, the specific contribution area they addressed and the application domain.
Simona Fioretto, Elio Masciari
Knowl. Inf. Syst.2
2024 Integrating Flow and Structure in Diagrams for Data Science
abstract
In Data Science, data modeling (including relational databases and big data) has traditionally used Entity-Relationship (ER) diagrams to represent structural characteristics of data. However, ER diagrams lack the capability to capture the flow and transformation of data through analytic pipelines, which are essential to manage modern data science workflows. On the other hand, flowcharts have been used for decades to describe the processing. Based on this motivation, this paper provides a historical perspective identifying the limitations of employing ER diagrams and Data Flow Diagrams (DFDs), separately, emphasizing the need to integrate both solutions. We examine established diagram notations and design models, including traditional models such as UML, ER, DFD, BPMN, and FLOWER. Our literature analysis suggests that integrative diagram approaches can provide a more intuitive and comprehensive understanding of data collection, data integration, and data transformation for big data analytics in the future.
Enea Vincenzo Napolitano, Elio Masciari, Carlos Ordonez 0001
IEEE Big Data2
2024 Machine Learning for KPI Development in Public Administration
Simona Fioretto, Elio Masciari, Enea Vincenzo Napolitano
DATA2
2024 Sustainability and High Performance Computing
Elio Masciari, Enea Vincenzo Napolitano
iiWAS (2)1
2024 Special issue on intelligent systems
Michelangelo Ceci, Sergio Flesca, Giuseppe Manco 0001, Elio Masciari
J. Intell. Inf. Syst.4
2023 An Advanced BERT LayerSum Model for Sentiment Classification of COVID-19 Tweets
Areeba Umair, Elio Masciari
DATA2
2023 How Pandemic Affected the Adoption of e-Health Systems
abstract
The COVID-19 pandemic has dramatically transformed healthcare systems globally, therefore improving health information technology sector. From the moment that pandemic has broken out, the use of information and communication technologies (ICT) has become absolutely necessary for the continuation of healthcare services. Furthermore, data digitization has enabled the extraction of meaningful insights through big data analytics. E-Health, encompassing a wide range of ICTs used in healthcare, has become a critical component in addressing the challenges posed by the pandemic. The main goal of this paper is to examine the impact of COVID-19 on health information technology and explores the rapid growth of e-Health and big data-driven innovation in healthcare processes. Through the analysis of the major tools, techniques, and innovative processes that have emerged in response to the pandemic, this paper has the aim of highlight their potential to improve system efficiency and enhance citizen health. We explore the current state of e-Health and big data in healthcare and discuss the future implications of these technologies for the sector. Our analysis underscores the need for continued investments in health information technology and highlights the role of policymakers, healthcare providers, and researchers in fostering innovation and driving positive change in the sector.
Enea Vincenzo Napolitano, Simona Fioretto, Elio Masciari, Arianna Anniciello
IDEAS3
2023 Sentimental and spatial analysis of COVID-19 vaccines tweets
abstract
The world has to face health concerns due to huge spread of COVID. For this reason, the development of vaccine is the need of hour. The higher vaccine distribution, the higher the immunity against coronavirus. Therefore, there is a need to analyse the people's sentiment for the vaccine campaign. Today, social media is the rich source of data where people share their opinions and experiences by their posts, comments or tweets. In this study, we have used the twitter data of vaccines of COVID and analysed them using methods of artificial intelligence and geo-spatial methods. We found the polarity of the tweets using the TextBlob() function and categorized them. Then, we designed the word clouds and classified the sentiments using the BERT model. We then performed the geo-coding and visualized the feature points over the world map. We found the correlation between the feature points geographically and then applied hotspot analysis and kernel density estimation to highlight the regions of positive, negative or neutral sentiments. We used precision, recall and F score to evaluate our model and compare our results with the state-of-the-art methods. The results showed that our model achieved 55% & 54% precision, 69% & 85% recall and 58% & 64% F score for positive class and negative class respectively. Thus, these sentimental and spatial analysis helps in world-wide pandemics by identify the people's attitudes towards the vaccines.
Areeba Umair, Elio Masciari
J. Intell. Inf. Syst.2
2022 Clustered Majority Judgement
Emanuele d'Ajello, Davide Formica, Elio Masciari, Gaia Mattia, Arianna Anniciello, Cristina Moscariello, Stefano Quintarelli, Davide Zaccarella
DATA3
2022 Decision making with Clustered Majority Judgment
abstract
Making decisions quickly and efficiently is essential in all areas of existence; when the decision concerns a fair number of voters and options to vote, it is sometimes appropriate to choose a voting system that represents the voting population well, without excluding the preferences of minorities. In this paper we present a voting system that aims to be an enhancement of Majority Judgment through unsupervised machine learning techniques, in particular the cluster and, in addition, a criterion for obtaining a multiwinner result has also been added. After exposing its functioning, a case study is presented to test its applicability, which leads to multiple fields of interest and it is not limited exclusively to purely political occasions.
Arianna Anniciello, Emanuele d'Ajello, Davide Formica, Elio Masciari, Gaia Mattia, Stefano Quintarelli, Davide Zaccarella
IDEAS4
2022 Applications of Majority Judgement for Winner Selection in Eurovision Song Contest
abstract
The existence of big data, social media interactions, and digital globalization has changed the way people make decisions either in their life or those of collective importance. Computational Social Choice (COMSOC), as an emerged field, has tried to join various social fields (social choice theory) and technical fields (computer science, mathematics, economics and logic). In the last few decades, expert rating was used to select the winner in the contest or competition, that was, later, merged with crowd voting. However, the results of voting based on aggregation of crowd opinion was not considered satisfied. The majority judgement is a new method of election. It is the consequence of a new theory of social choice where voters judge candidates instead of ranking them. In this research, we used Eurovision song contest data of 2021 final round. Eurovision song contest is held annually, in which almost 40 countries participate. We applied majority judgement on the Eurovision song contest data and found that Italy got highest position, followed by Croatia and Australia acquiring second and third positions respectively in competition of 2021.
Areeba Umair, Elio Masciari, Giusi Madeo, Muhammad Habib Ullah
IDEAS2
2022 A comprehensive Benchmark for fake news detection
abstract
Nowadays, really huge volumes of fake news are continuously posted by malicious users with fraudulent goals thus leading to very negative social effects on individuals and society and causing continuous threats to democracy, justice, and public trust. This is particularly relevant in social media platforms (e.g., Facebook, Twitter, Snapchat), due to their intrinsic uncontrolled publishing mechanisms. This problem has significantly driven the effort of both academia and industries for developing more accurate fake news detection strategies: early detection of fake news is crucial. Unfortunately, the availability of information about news propagation is limited. In this paper, we provided a benchmark framework in order to analyze and discuss the most widely used and promising machine/deep learning techniques for fake news detection, also exploiting different features combinations w.r.t. the ones proposed in the literature. Experiments conducted on well-known and widely used real-world datasets show advantages and drawbacks in terms of accuracy and efficiency for the considered approaches, even in the case of limited content information.
Antonio Galli, Elio Masciari, Vincenzo Moscato, Giancarlo Sperlì
J. Intell. Inf. Syst.2
2021 Data Mining for Animal Health to Improve Human Quality of Life: Insights from a University Veterinary Hospital
Oscar Tamburis, Elio Masciari, Christian Esposito 0001, Gerardo Fatone
DATA2
2021 Exploratory analysis of methods for automated classification of clinical diagnoses in Veterinary Medicine
abstract
The present work describes the analysis conducted on the diagnoses made during the general physical examinations in the decade 2010–2020, starting from the DB of the EMR previously implemented in the University Veterinary Teaching Hospital at Federico II University of Naples. A decision tree algorithm was implemented to work out a predictive model for an effective recognition of neoplastic diseases and zoonoses for cats and dogs from Campania Region. The results achievable by data mining techniques for what concerns computer aided disease diagnosis and exploration of risk factors and their relations to diseases, show the increasing importance of Veterinary Informatics within the wider field of Biomedical and Health Informatics, and in particular its capacity to point out the existing connections between humans, animals, and surrounding environment, according to the One (Digital) Health perspective specifics.
Oscar Tamburis, Elio Masciari, Gerardo Fatone
IDEAS2
2021 Sentimental Analysis Applications and Approaches during COVID-19: A Survey
abstract
The social media and electronic media has a vast amount of user-generated data such as people’ comment and reviews about different product, diseases, government policies etc. Sentimental analysis is the emerging field in text mining where people’s feeling and emotions are extracted using different techniques. COVID-19 has declared as pandemic and effected people’s lives all over the globe. It caused the feelings of fear, anxiety, anger, depression and many other psychological issues. In this survey paper, the sentimental analysis applications and methods which are used for COVID-19 research are briefly presented. The comparison of thirty primary studies shows that Naive Bayes and SVM are the widely used algorithms of sentimental analysis for COVID-19 research. The applications of sentimental analysis during COVID includes the analysis of people’s sentiments specially students, reopening sentiments, analysis of restaurants reviews and analysis of vaccine sentiments.
Areeba Umair, Elio Masciari, Muhammad Habib Ullah
IDEAS2
2021 A survey of Big Data dimensions vs Social Networks analysis
abstract
The pervasive diffusion of Social Networks (SN) produced an unprecedented amount of heterogeneous data. Thus, traditional approaches quickly became unpractical for real life applications due their intrinsic properties: large amount of user-generated data (text, video, image and audio), data heterogeneity and high speed generation rate. More in detail, the analysis of user generated data by popular social networks (i.e Facebook (https://www.facebook.com/), Twitter (https://www.twitter.com/), Instagram (https://www.instagram.com/), LinkedIn (https://www.linkedin.com/)) poses quite intriguing challenges for both research and industry communities in the task of analyzing user behavior, user interactions, link evolution, opinion spreading and several other important aspects. This survey will focus on the analyses performed in last two decades on these kind of data w.r.t. the dimensions defined for Big Data paradigm (the so called Big Data 6 V's).
Michele Ianni, Elio Masciari, Giancarlo Sperlì
J. Intell. Inf. Syst.2
2020 Leveraging Machine Learning for Fake News Detection
Elio Masciari, Vincenzo Moscato, Antonio Picariello, Giancarlo Sperlì
DATA1
2020 Detecting fake news by image analysis
abstract
The uncontrolled growth of fake news creation and dissemination we observed in recent years causes continuous threats to democracy, justice, and public trust. This problem has significantly driven the effort of both academia and industries for developing more accurate fake news detection strategies. Early detection of fake news is crucial, however the availability of information about news propagation is limited. Moreover, it has been shown that people tend to believe more fake news due to their features [10]. In this paper, we present our framework for fake news detection and we discuss in detail an approach based on deep learning that we implemented by using Google Bert features. Our experiments conducted on two well-known and widely used real-world datasets suggest that our method can outperform the state-of-the-art approaches and allows fake news accurate detection, even in the case of limited content information.
Elio Masciari, Vincenzo Moscato, Antonio Picariello, Giancarlo Sperlì
IDEAS1
2019 An Overview of the Endless Battle between Virus Writers and Detectors: How Compilers Can Be Used as an Evasion Technique
Michele Ianni, Elio Masciari, Domenico Saccà
DATA2
2019 HIKE: A Step Beyond Data Exchange
Sergio Greco, Elio Masciari, Domenico Saccà, Irina Trubitsyna
ER2
2019 An Effective System for User Queries Assistance
Elio Masciari, Domenico Saccà, Irina Trubitsyna
FQAS1
2019 Simplified data posting in practice
abstract
The data posting framework introduced in [8] adapts the well-known Data Exchange techniques to the new Big Data management and analysis challenges that can be found in real world scenarios. Although it is expressive enough, it requires the ability of using count constraints and may be difficult for a non expert user. Moreover, the data posting problem is NP-complete under the data complexity in the general case, then the use of the non-deterministic variables is performed. Indeed, identifying the conditions that guarantee polynomial-time execution in the presence of non-deterministic choices is very important for practical purposes. In this paper, we present a simplified version of data posting framework, based on the use of the smart mapping rules, that integrate the simple mapping description with some parameters, avoiding the complex specifications with count constraints. We show that the data posting problem in the new setting is NP- complete and identify the conditions under which this problem becomes polynomial even in the presence of non-deterministic choices.
Elio Masciari, Irina Trubitsyna, Domenico Saccà
IDEAS1
2018 Clustering Big Data
abstract
The need to support advanced analytics on Big Data is driving data scientist' interest toward massively parallel distributed systems and software platforms, such as Map-Reduce and Spark, that make possible their scalable utilization.However, when complex data mining algorithms are required, their fully scalable deployment on such platforms faces a number of technical challenges that grow with the complexity of the algorithms involved.Thus algorithms, that were originally designed for a sequential nature, must often be redesigned in order to effectively use the distributed computational resources.In this paper, we explore these problems, and then propose a solution which has proven to be very effective on the complex hierarchical clustering algorithm CLUBS+.By using four stages of successive refinements, CLUBS+ delivers high-quality clusters of data grouped around their centroids, working in a totally unsupervised fashion.Experimental results confirm the accuracy and scalability of CLUBS+ on Map-Reduce platforms.
Michele Ianni, Elio Masciari, Giuseppe M. Mazzeo, Carlo Zaniolo
DATA2
2018 Efficient Big Data Clustering
abstract
The need to support advanced analytics on Big Data is driving data scientist' interest toward massively parallel distributed systems and software platforms, such as Map-Reduce and Spark, that make possible their scalable utilization. However, when complex data mining algorithms are required, their fully scalable deployment on such platforms faces a number of technical challenges that grow with the complexity of the algorithms involved. Thus algorithms, that were originally designed for a sequential nature, must often be redesigned in order to effectively use the distributed computational resources. In this paper, we explore these problems, and then propose a solution which has proven to be very effective on the complex hierarchical clustering algorithm CLUBS+. By using four stages of successive refinements, CLUBS+ delivers high-quality clusters of data grouped around their centroids, working in a totally unsupervised fashion. Experimental results confirm the accuracy and scalability of CLUBS+.
Michele Ianni, Elio Masciari, Giuseppe M. Mazzeo, Carlo Zaniolo
IDEAS2
2018 Efficiently interpreting traces of low level events in business process logs
Bettina Fazzinga, Sergio Flesca, Filippo Furfaro, Elio Masciari, Luigi Pontieri
Inf. Syst.4
2017 Choose The Best!: Ranking Group of Users In Collaborative Networks
abstract
Social Networks analysis is driving both research and industrial effort as the outcomes of this activity are relevant both from a merely theoretical point of view and for the potential market advantages they can provide to companies. Indeed, there is a growing number of applications that call for user (social) intervention with the aim of helping each other in solving complex tasks or rating other users work. The topic is even more intriguing when a reward is given to users that properly complete their tasks. In this paper, we focus on the analysis of user mutual rankings in a collaborative network where they contribute to the solution of complex tasks. We leverage Exponential Random Graph to model user interaction rankings and we evaluate our approach in a real life scenario.
Nunzio Cassavia, Sergio Flesca, Elio Masciari
ASONAM3
2017 An Open Source System for Big Data Warehousing
Nunzio Cassavia, Elio Masciari, Domenico Saccà
DATA2
2017 WFinger: a joint-decoder for very short Tardos fingerprinting codes
abstract
research-article Share on WFinger: a joint-decoder for very short Tardos fingerprinting codes Authors: Bettina Fazzinga ICAR-CNR, Rende (CS), Italy ICAR-CNR, Rende (CS), ItalyView Profile , Sergio Flesca DIMES, University of Calabria, Rende (CS), Italy DIMES, University of Calabria, Rende (CS), ItalyView Profile , Filippo Furfaro DIMES, University of Calabria, Rende (CS), Italy DIMES, University of Calabria, Rende (CS), ItalyView Profile , Elio Masciari ICAR-CNR, Rende (CS), Italy ICAR-CNR, Rende (CS), ItalyView Profile Authors Info & Claims IDEAS '17: Proceedings of the 21st International Database Engineering & Applications SymposiumJuly 2017 Pages 176–183https://doi.org/10.1145/3105831.3105860Published:12 July 2017Publication History 1citation36DownloadsMetricsTotal Citations1Total Downloads36Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Bettina Fazzinga, Sergio Flesca, Filippo Furfaro, Elio Masciari
IDEAS4
2017 A fast and accurate algorithm for unsupervised clustering around centroids
Giuseppe M. Mazzeo, Elio Masciari, Carlo Zaniolo
Inf. Sci.2
2016 A Big Data Approach For Querying Data in EHR Systems
abstract
Information management in healthcare is nowadays experiencing a great revolution. After the impressive progress in digitizing medical data by private organizations, also the federal government and other public stakeholders have also started to make use of healthcare data for data analysis purposes in order to extract actionable knowledge. In this paper, we propose an architecture for supporting interoperability in healthcare systems by exploiting Big Data techniques. In particular, we describe a proposal based on big data techniques to implement a nationwide system able to improve EHR data access efficiency and reduce costs.
Nunzio Cassavia, Mario Ciampi, Giuseppe De Pietro, Elio Masciari
IDEAS4
2016 How, Who and When: Enhancing Business Process Warehouses By Graph Based Queries
abstract
Log analysis and querying recently received a renewed interest from the research community, as the effective understanding of process behavior is crucial for improving business process management. Indeed, currently available log querying tools are not completely satisfactory, especially from the viewpoint of easiness of use. As a matter of fact, there is no framework which meets the requirements of easiness of use, flexibility and efficiency of query evaluation. In this paper, we propose a framework for graphical querying of (process) log data that makes the log analysis task quite easy and efficient, adopting a very general model of process log data which guarantees a high level of flexibility. We implemented our framework by using a flexible storage architecture and a user-friendly data analysis interface, based on an intuitive and yet expressive graph-based query language. Experiments performed on real data confirm the validity of the approach.
Bettina Fazzinga, Sergio Flesca, Filippo Furfaro, Elio Masciari, Luigi Pontieri, Chiara Pulice
IDEAS4
2016 Privacy or Security?: Take A Look And Then Decide
abstract
Big data paradigm is currently the leading paradigm for data production and management. As a matter of fact, new information are generated at high rates in specialized fields (e.g., cybersecurity scenario). This may cause that the events to be studied occur at rates that are too fast to be effectively analyzed in real time. For example, in order to detect possible security threats, millions of records in a high-speed flow stream must be screened. To ameliorate this problem, a viable solution is the use of data compression for reducing the amount of data to be analyzed. In this paper we propose the use of privacy-preserving histograms, that provide approximate answers to 'safe' queries, for analyzing data in the cybersecurity scenario without compromising individuals' privacy, and we describe our system that has been used in a real life scenario.
Bettina Fazzinga, Filippo Furfaro, Elio Masciari, Giuseppe M. Mazzeo
SSDBM3
2016 Recent advances in mining patterns from complex data
Annalisa Appice, Michelangelo Ceci, Corrado Loglisci, Giuseppe Manco 0001, Elio Masciari
J. Intell. Inf. Syst.5
2015 Surfing Big Data Warehouses for Effective Information Gathering
abstract
Due to the emerging Big Data paradigm traditional data management techniques result inadequate in many real life scenarios. In particular, OLAP techniques require substantial changes in order to offer useful analysis due to huge amount of data to be analyzed and their velocity and variety. In this paper, we describe an approach for dynamic Big Data searching that based on data collected by a suitable storage system, enriches data in order to guide users through data exploration in an efficient and effective way.
Nunzio Cassavia, Pietro Dicosta, Elio Masciari, Domenico Saccà
DATA3
2015 Big Data Techniques For Supporting Accurate Predictions of Energy Production From Renewable Sources
abstract
Predicting the output power of renewable energy production plants distributed on a wide territory is a really valuable goal, both for marketing and energy management purposes. Vi-POC (Virtual Power Operating Center) project aims at designing and implementing a prototype which is able to achieve this goal. Due to the heterogeneity and the high volume of data, it is necessary to exploit suitable Big Data analysis techniques in order to perform a quick and secure access to data that cannot be obtained with traditional approaches for data management. In this paper, we describe Vi-POC -- a distributed system for storing huge amounts of data, gathered from energy production plants and weather prediction services. We use HBase over Hadoop framework on a cluster of commodity servers in order to provide a system that can be used as a basis for running machine learning algorithms. Indeed, we perform one-day ahead forecast of PV energy production based on Artificial Neural Networks in two learning settings, that is, structured and non-structured output prediction. Preliminary experimental results confirm the validity of the approach, also when compared with a baseline approach.
Michelangelo Ceci, Roberto Corizzo, Fabio Fumarola, Michele Ianni, Donato Malerba, Gaspare Maria, Elio Masciari, Marco Oliverio, Aleksandra Rashkovska
IDEAS7
2015 A compression-based framework for the efficient analysis of business process logs
abstract
The increasing availability of large process log repositories calls for efficient solutions for their analysis. In this regard, a novel specialized compression technique for process logs is proposed, that builds a synopsis supporting a fast estimation of aggregate queries, which are of crucial importance in exploratory and high-level analysis tasks. The synopsis is constructed by progressively merging the original log-tuples, which represent single activity executions within the process instances, into aggregate tuples, summarizing sets of activity executions. The compression strategy is guided by a heuristic aiming at limiting the loss of information caused by summarization, while guaranteeing that no information is lost on the set of activities performed within the process instances and on the order among their executions. The selection conditions in an aggregate query are specified in terms of a graph pattern, that allows precedence relationships over activity executions to be expressed, along with conditions on their starting times, durations, and executors. The efficacy of the compression technique, in terms of capability of reducing the size of the log and of accuracy of the estimates retrieved from the synopsis, has been experimentally validated.
Bettina Fazzinga, Sergio Flesca, Filippo Furfaro, Elio Masciari, Luigi Pontieri
SSDBM4
2015 An end to end framework for building data cubes over trajectory data streams
Elio Masciari
J. Intell. Inf. Syst.1
2014 Data Preparation for Tourist Data Big Data Warehousing
abstract
The pervasive diffusion of new generation devices like smart phones and tablets along with the widespread use of social networks causes the generation of massive data flows containing heterogeneous information generated at different rates and having different formats. These data are referred as \emph{Big Data} and require new storage and analysis approaches to be investigated for managing them. In this paper we will describe a system for dealing with massive tourism flows that we exploited for the analysis of tourist behavior in Italy. We defined a framework that exploits a NoSQL approach for data management and map reduce for improving the analysis of the data gathered from different sources
Nunzio Cassavia, Pietro Dicosta, Elio Masciari, Domenico Saccà
DATA3
2014 Innovative power operating center management exploiting big data techniques
abstract
The problem of accurately predicting the energy production from renewable sources has recently received an increasing attention from both the industrial and the research communities. It presents several challenges, such as facing with the rate data are provided by sensors, the heterogeneity of the data collected, power plants efficiency, as well as uncontrollable factors, such as weather conditions and user consumption profiles. In this paper we describe Vi-POC (Virtual Power Operating Center), a project conceived to assist energy producers and decision makers in the energy market. In this paper we present the Vi-POC project and how we face with challenges posed by the specific application. The solutions we propose have roots both in big data management and in stream data mining.
Michelangelo Ceci, Nunzio Cassavia, Roberto Corizzo, Pietro Dicosta, Donato Malerba, Gaspare Maria, Elio Masciari, Camillo Pastura
IDEAS7
2014 Analysing microarray expression data through effective clustering
Elio Masciari, Giuseppe M. Mazzeo, Carlo Zaniolo
Inf. Sci.1
2014 Mining complex patterns
Annalisa Appice, Michelangelo Ceci, Corrado Loglisci, Elio Masciari, Giuseppe Manco 0001
J. Intell. Inf. Syst.4
2014 Dealing with trajectory streams by clustering and mathematical transforms
Gianni Costa, Giuseppe Manco 0001, Elio Masciari
J. Intell. Inf. Syst.3
2013 Sequential pattern mining from trajectory data
abstract
In this paper, we study the problem of mining for frequent trajectories, which is crucial in many application scenarios, such as vehicle traffic management, hand-off in cellular networks, supply chain management. We approach this problem as that of mining for frequent sequential patterns. Our approach consists of a partitioning strategy for incoming streams of trajectories in order to reduce the trajectory size and represent trajectories as strings. We mine frequent trajectories using a sliding windows approach combined with a counting algorithm that allows us to promptly update the frequency of patterns. In order to make counting really efficient, we represent frequent trajectories by prime numbers, whereby the Chinese reminder theorem can then be used to expedite the computation.
Elio Masciari, Shi Gao, Carlo Zaniolo
IDEAS1
2013 A New, Fast and Accurate Algorithm for Hierarchical Clustering on Euclidean Distances
Elio Masciari, Giuseppe M. Mazzeo, Carlo Zaniolo
PAKDD (2)1
2013 RFID-data compression for supporting aggregate queries
abstract
RFID-based systems for object tracking and supply chain management have been emerging since the RFID technology proved effective in monitoring movements of objects. The monitoring activity typically results in huge numbers of readings, thus making the problem of efficiently retrieving aggregate information from the collected data a challenging issue. In fact, tackling this problem is of crucial importance, as fast answers to aggregate queries are often mandatory to support the decision making process. In this regard, a compression technique for RFID data is proposed, and used as the core of a system supporting the efficient estimation of aggregate queries. Specifically, this technique aims at constructing a lossy synopsis of the data over which aggregate queries can be estimated, without accessing the original data. Owing to the lossy nature of the compression, query estimates are approximate, and are returned along with intervals that are guaranteed to contain the exact query answers. The effectiveness of the proposed approach has been experimentally validated, showing a remarkable trade-off between the efficiency and the accuracy of the query estimation.
Bettina Fazzinga, Sergio Flesca, Filippo Furfaro, Elio Masciari
ACM Trans. Database Syst.4
2012 Warehousing and querying trajectory data streams with error estimation
abstract
In this paper, we address the problem of trajectory data streams warehousing and querying, that revealed really challenging as we deal with data (trajectories) for which the order of elements is relevant. We propose an end to end framework in order to make the querying step quite effective. We performed several tests on real world datasets that confirmed the efficiency and effectiveness of the proposed techniques.
Elio Masciari
DOLAP1
2012 Efficient MD5 hash reversing using D.E.A. framework for sharing computational resources
abstract
The recent advances in computing technology lead to the availability of a huge number of computational resources that can be easily connected through network infrastructures. Indeed, a really small fraction of the available computing power is fully exploited for performing effective computation of user tasks. On the contrary, there are several research projects that require a lot of computing power to reach their goals, but they usually lack adequate resources thus making the project activities quite hard to be completed. In this paper we describe D.E.A. (Distributed Execution Agent), a framework for sharing computational resources. We will exploit D.E.A. framework to tame the high computational demanding problem of hash MD5 reversing. We performed several experiments that confirmed the validity of our approach.
Nunzio Cassavia, Elio Masciari
IDEAS2
2012 XML class outlier detection
abstract
XML (eXtensible Markup Language) became in recent years the new standard for data representation and exchange on the WWW. This has resulted in a great need for data cleaning techniques in order to identify outlying data. In this paper, we present a technique for outlier detection that singles out anomalies with respect to a relevant group of objects. We exploit a suitable encoding of XML documents that are encoded as signals of fixed frequency that can be transformed using Fourier Transforms. Outliers are identified by simply looking at the signal spectra. The results show the effectiveness of our approach.
Giuseppe Manco 0001, Elio Masciari
IDEAS2
2012 SMART: Stream Monitoring enterprise Activities by RFID Tags
Elio Masciari
Inf. Sci.1
2011 Efficient and Effective Query Answering for Trajectory Cuboids
Elio Masciari
FQAS1
2011 Query answering on trajectory cuboids using prime numbers encodings
abstract
Trajectory data streams are huge amounts of data pertaining to time and position of moving objects generated by different sources continuously using a wide variety of technologies (e.g., RFID tags, GPS, GSM networks). Mining such amounts of data is challenging, since the possibility to extract useful information from this peculiar kind of data is crucial in many application scenarios such as vehicle traffic management, hand-off in cellular networks, supply chain management. Moreover, spatial data streams poses interesting challenges both for their proper definition and acquisition, thus making the mining process harder than for classical point data. In this paper, we address the problem of trajectory data streams On Line Analytical Processing, that revealed really challenging as we deal with data (trajectories) for which the order of elements is relevant. We propose an end to end framework in order to make the querying step quite effective. We performed several tests on real world datasets that confirmed the efficiency and effectiveness of the proposed techniques.
Elio Masciari
IDEAS1
2011 Fast and Accurate Trajectory Streams Clustering
Elio Masciari
SSDBM1
2011 A Fuzzy Logic Approach to Wrapping PDF Documents
abstract
The PDF format represents the de facto standard for print-oriented documents. In this paper, we address the problem of wrapping PDF documents, which raises new challenges in several contexts of text data management. Our proposal is based on a novel bottom-up hierarchical wrapping approach that exploits fuzzy logic to handle the “uncertainty” which is intrinsic to the structure and presentation of PDF documents. A PDF wrapper is defined by specifying a set of group type definitions that impose a target structure to groups of tokens containing the required information. Constraints on token groupings are formulated as fuzzy conditions, which are defined on spatial and content predicates of tokens. We define a formal semantics for PDF wrappers and propose an algorithm for wrapper evaluation working in polynomial time with respect to the size of a PDF document. The proposed approach has been implemented in a wrapper generation system that offers visual capabilities to assist the designer in specifying and evaluating a PDF wrapper. Experimental results have shown good accuracy and applicability of our system to PDF documents of various domains.
Sergio Flesca, Elio Masciari, Andrea Tagarelli
IEEE Trans. Knowl. Data Eng.2
2010 Effectively Monitoring RFID Based Systems
Fabrizio Angiulli, Elio Masciari
ADBIS2
2009 Trajectory Clustering via Effective Partitioning
Elio Masciari
FQAS1
2009 Efficient and effective RFID data warehousing
abstract
Radio Frequency Identification (RFID) applications are emerging as key components in object tracking and supply chain management systems since in the next future almost every major retailer will use RFID systems to track the shipment of products from suppliers to warehouses. Due to the streaming nature of RFID readings, large amounts of data are generated by these devices at high production rates. This phenomenon is even more relevant since RFIDs are so cheap that every individual item can be tagged thus leaving a "trail" of data as it moves across different locations. This scenario raises new challenges in effectively and efficiently exploiting such large amounts of data. In this paper we address the problem of compressing RFID data in order to enable devices with limited amount of available memory (such as PDAs) to issue queries on RFID warehouses. In particular, we designed a lossy strategy for collapsing tuples carrying information about items being delivered at different location of the supply chain.
Bettina Fazzinga, Sergio Flesca, Elio Masciari, Filippo Furfaro
IDEAS3
2008 Mining categories for emails via clustering and pattern discovery
Giuseppe Manco 0001, Elio Masciari, Andrea Tagarelli
J. Intell. Inf. Syst.2
2007 A Framework for Outlier Mining in RFID data
abstract
Radio frequency identification (RFID) applications are emerging as key components in object tracking and supply chain management systems. In next future almost every major retailer will use RFID systems to track the shipment of products from suppliers to warehouses. Due to RFID readings features this will result in a huge amount of information generated by such systems when costs will be at a level such that each individual item could be tagged thus leaving a trail of data as it moves through different locations. We define a technique for efficiently detecting anomalous data in order to prevent problems related to inefficient shipment or fraudulent actions. Since items usually move together in large groups through distribution centers and only in stores do they move in smaller groups we exploit such a feature in order to design our technique. The preliminary experiments show the effectiveness of our approach.
Elio Masciari
IDEAS1
2007 Exploiting structural similarity for effective Web information extraction
Sergio Flesca, Giuseppe Manco 0001, Elio Masciari, Luigi Pontieri, Andrea Pugliese 0001
Data Knowl. Eng.3
2006 Wrapping PDF Documents Exploiting Uncertain Knowledge
Sergio Flesca, Salvatore Garruzzo, Elio Masciari, Andrea Tagarelli
CAiSE3
2005 Fast Detection of XML Structural Similarity
abstract
Because of the widespread diffusion of semistructured data in XML format, much research effort is currently devoted to support the storage and retrieval of large collections of such documents. XML documents can be compared as to their structural similarity, in order to group them into clusters so that different storage, retrieval, and processing techniques can be effectively exploited. In this scenario, an efficient and effective similarity function is the key of a successful data management process. We present an approach for detecting structural similarity between XML documents which significantly differs from standard methods based on graph-matching algorithms, and allows a significant reduction of the required computation costs. Our proposal roughly consists of linearizing the structure of each XML document, by representing it as a numerical sequence and, then, comparing such sequences through the analysis of their frequencies. First, some basic strategies for encoding a document are proposed, which can focus on diverse structural facets. Moreover, the theory of discrete Fourier transform is exploited to effectively and efficiently compare the encoded documents (i.e., signals) in the domain of frequencies. Experimental results reveal the effectiveness of the approach, also in comparison with standard methods.
Sergio Flesca, Giuseppe Manco 0001, Elio Masciari, Luigi Pontieri, Andrea Pugliese 0001
IEEE Trans. Knowl. Data Eng.3
2003 On the minimization of Xpath queries
Sergio Flesca, Filippo Furfaro, Elio Masciari
VLDB3
2003 Efficient and effective Web change detection
Sergio Flesca, Elio Masciari
Data Knowl. Eng.2
2002 Detecting Structural Similarities between XML Documents
Sergio Flesca, Giuseppe Manco 0001, Elio Masciari, Luigi Pontieri, Andrea Pugliese 0001
WebDB3
2001 Meaningful Change Detection on the Web
Sergio Flesca, Filippo Furfaro, Elio Masciari
DEXA3
2000 A Hybrid Technique for Data Mining on Balance-Sheet Data
Giuseppe Dattilo, Sergio Greco, Elio Masciari, Luigi Pontieri
DaWaK3
2000 Combining Different Data Mining Techniques to Improve Data Analysis
abstract
In this paper we propose the combined use of different methods to improve the data analysis process. This is obtained by combining inductive and deductive techniques. Inductive techniques are used for generating hypotheses from data whereas deductive techniques are used to derive knowledge and to verify hypotheses. In order to guide users in the the analysis process, we have developed a system which integrates deductive tools, data mining tools (such as classification algorithms and features selection algorithms), visualization tools and tools for the easy manipulation of data sets. The system developed is currently used in a large project whose aim is the integration of information sources containing data concerning the socio-economic aspects of Calabria and the analysis of the integrated data. Several experiments on socio-economic indicators of Calabrian cities have shown that the combined use of different techniques improves both the comprehensibility and the accuracy of models. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
Sergio Greco, Elio Masciari, Luigi Pontieri
FQAS2