Alfredo Cuzzocrea

dblp:c/AlfredoCuzzocrea · DBLP profile ↗
← Back
206ranked-venue papers in the field
111as first author
83since 2021 · last 2026
0000-0002-7104-6415ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 81 (19 first)Database Systems & Data Management · 77 (58 first)Data Mining & Knowledge Discovery · 18 (12 first)Information Retrieval & Web Search · 15 (9 first)Knowledge Engineering, Semantic Web & Information Systems · 11 (9 first)Other / Interdisciplinary · 4 (4 first)
YearPublicationVenuePosition
2026 Advancing m6Am Site Prediction Through Deep Learning on Diverse mRNA Sequence Landscapes: The m6Am-DLcat Approach
Alfredo Cuzzocrea, Tasmin Karim, Md. Shazzad Hossain Shaon, Md. Fahim Sultan, Mst. Shapna Akter
DATA (1)1
2026 SMART-DAI: A Cloud/Edge Big Data Computing Platform for Supporting Multidimensional Big Data Analytics in Advanced Artificial Intelligence Applications
Alfredo Cuzzocrea
DEXA (2)1
2026 Multidimensional risk analysis and prediction over digital twins via big data paradigms
abstract
Digital twins are increasingly adopted to support data-driven decision-making in complex and dynamic systems . However, existing digital twin–based approaches for risk analysis and prediction often rely on isolated predictive models and lack native support for multidimensional analytics , thus limiting their ability to provide comprehensive and interpretable insights. In order to fulfill this gap, the paper proposes a novel multidimensional big data analytics framework for risk analysis and prediction over digital twins . The framework integrates streaming data ingestion, distributed multidimensional modeling, analytics, and predictive modeling within a unified cloud-native architecture. Risk prediction is embedded within the multidimensional analytics pipeline, enabling context-aware and interpretable predictions across multiple analytical perspectives. The proposed approach is evaluated using real-world datasets through a combination of multidimensional exploratory analysis and quantitative baseline comparisons with standard Machine Learning (ML) models. Experimental results demonstrate that the framework achieves competitive predictive performance while supporting scalable, interpretable, and multidimensional risk analysis. These characteristics make the proposed solution particularly suitable for complex digital twin scenarios where both analytical flexibility and predictive reliability are required.
Alfredo Cuzzocrea, Ismail Benlaredj
Inf. Sci.1
2025 Machine Learning on Big Data for Crime Trend Analytics and Prediction
M. Akroma E. Ansah, Carson K. Leung, Omodesayo A. Owolabi, Alfredo Cuzzocrea
IEEE Big Data4
2025 Secure Database Sharing in Healthcare: An LLM Based HIPAA Compliant Solution for Data Privacy and Security
Md Abdul Barek, Md Bajlur Rashid, ABM Kamrul Islam Riad, Sharmin Yeasmin, Md. Jobair Hossain Faruk, Hakki Erhan Sevil, Guillermo A. Francia III, Hossain Shahriar, Alfredo Cuzzocrea, Sheikh Iqbal Ahamed, Coskun Cetinkaya
IEEE Big Data10
2025 Segment-Aware Learning for Adaptive Customer Churn Prediction
Soumia Benkrid, Alfredo Cuzzocrea
IEEE Big Data2
2025 Multidimensional Big Healthcare Data Generation for Alzheimer Disease Early Prevention and Analytics: The HCRS Algorithm
Alfredo Cuzzocrea
IEEE Big Data1
2025 ALCMAEON: A Multidimensional Big Data AnaLytiCs PlatforM for Supporting Predictive Analysis and Mining over Clinical and MEdical (Big) Data ON Alzheimer's Disease Patients
Alfredo Cuzzocrea
IEEE Big Data1
2025 Generation, Analysis and Experimental Validation of an Emotion-Enlightened Synthetic Dialogue-Dataset via Advanced LLM-Based Methodologies
Alfredo Cuzzocrea, Abderraouf Hafsaoui, Ismail Benlaredj
IEEE Big Data1
2025 A Self-Adaptation Framework for Supporting Distributed Computing Based on Industry-Scale Digital Twins
Alfredo Cuzzocrea, Cristina Cerschi Seceleanu, Tiberiu Seceleanu
IEEE Big Data1
2025 Exposing Privacy Vulnerabilities in Federated Learning: A GAN-Based Model Inversion Attack
Md Morshedul Islam, Suraj Neupane, Md. Jobair Hossain Faruk, Hossain Shahriar, Alfredo Cuzzocrea
IEEE Big Data5
2025 Machine Learning on Big Economic Data for Predictive Financial Analytics
Carson K. Leung, Thanh Trung Jack Nguyen, Jing Xiang Ong, Alfredo Cuzzocrea
IEEE Big Data5
2025 A Benchmark Dataset for Code-Level Vulnerability Detection and Analysis
Tasmin Karim, Mst. Shapna Akter, Alfredo Cuzzocrea
IEEE Big Data3
2025 Mind the Shift: A Study on Transfer Learning and Domain Adaptation in Vehicular Intrusion Detection
abstract
This paper addresses the need for adaptable intrusion detection systems (IDS) in intra-vehicle networks. We evaluated two scalable IDS strategies-combined training and transfer learning-on novel traffic data to balance predictive accuracy and computational efficiency under distributional shifts. Previous IDS models for intra-vehicle attacks achieved high accuracy but relied heavily on the simple CarHacking dataset. To address this limitation, we integrate the newer CIC-IoV-2024 dataset, which reflects realistic vehicular traffic. In Strategy 1, we retrain models from scratch on a combined dataset. These models achieve strong classification across all classes, with accuracies ranging from 90.57% to 100% and F1 scores of 88.91% to 100%. However, training takes longer (61-184 minutes) and inference per packet is slower (30-80 ms). In Strategy 2, we apply transfer learning by fine-tuning pre-trained models while freezing earlier layers. This approach reduces training time (4-16 minutes) and improves latency (21.29-126.76 ms), but predictive performance declines. The models primarily distinguish between benign and malicious traffic, with F1 scores ranging from 82.99% to 89.32%, and exhibit high uncertainty in classifying diverse attack types. Our findings highlight a trade-off between predictive power and computational efficiency. These insights can guide the deployment of IDS frameworks in real-time vehicular environments.
Jakob Richard Proos, Giuseppe Cascavilla, Cristoffer Leite, Alfredo Cuzzocrea
IEEE Big Data4
2025 A Survey of Large Language Models (LLMs) for Cybersecurity: Opportunities and Directions
Md Abdur Rahman, Guillermo A. Francia III, Hossain Shahriar, Alfredo Cuzzocrea, Atef Mohamed, Sheikh Iqbal Ahamed
IEEE Big Data4
2025 Explaining Network Intrusion Detection System with SHAP and LIME
Md Abdur Rahman, Guillermo A. Francia III, Hossain Shahriar, Eman El-Sheikh, Alfredo Cuzzocrea, Sheikh Iqbal Ahamed
IEEE Big Data5
2025 Healthcare Solutions for Noisy Clinical Text: A Federated Privacy-Preserving Approach
ABM Kamrul Islam Riad, Salma Akter, Md Abdul Barek, Maliha Zaman Nizum, Guillermo A. Francia III, Hossain Shahriar, Alfredo Cuzzocrea, Sheikh Iqbal Ahamed
IEEE Big Data8
2025 Phishing Defense: An ML-Based URL Detection System with Real-World Deployment
ABM Kamrul Islam Riad, Md Reazul Hassan Rizvi, Shakil Miah, Md Abdul Barek, Yasmeen Rawajfih, Hossain Shahriar, Alfredo Cuzzocrea
IEEE Big Data9
2025 A Parallel and Distributed Rust Library for Core Decomposition on Large Graphs
Davide Rucci, Sebastian Parfeniuc, Matteo Mordacchini, Emanuele Carlini 0001, Alfredo Cuzzocrea, Patrizio Dazzi
IEEE Big Data5
2025 ResVul-LLM: A Neurosymbolic Framework Combining Large Language Models and Symbolic Reasoning for C/C++ Vulnerability Analysis
Md. Shazzad Hossain Shaon, Mst. Shapna Akter, Alfredo Cuzzocrea
IEEE Big Data3
2025 P3R: Parallel Plugin-Based Parameter Efficient Fine-Tuning for Code Understanding Through Hierarchical Representation Refinement
Md. Fahim Sultan, Mst. Shapna Akter, Alfredo Cuzzocrea
IEEE Big Data3
2025 CodeVul+: A Structure-Aware Framework for Cross-Repository Vulnerability Detection
Md. Fahim Sultan, Mst. Shapna Akter, Alfredo Cuzzocrea
IEEE Big Data3
2025 Neuro-Symbolic Methods in Natural Language Processing: A Review
Mst. Shapna Akter, Md. Fahim Sultan, Alfredo Cuzzocrea
DATA3
2025 Improving Software Security Through a LLM-Based Vulnerability Detection Model
Syeda Sadia Alam, Mst. Shapna Akter, Alfredo Cuzzocrea
DEXA (1)3
2025 Vector Databases for Modelling, Managing and Querying Big Scientific Data: Models, Issues, Paradigms
Alfredo Cuzzocrea
SSDBM1
2025 Privacy-preserving OLAP against big query workloads: innovative theories and theorems
Alfredo Cuzzocrea
Distributed Parallel Databases1
2025 Innovative research aspects of big data modelling, management and analytics
Alfredo Cuzzocrea
Knowl. Inf. Syst.1
2024 An Innovative Big Data Framework for Supporting Multidimensional Risk Analysis and Prediction over Digital Twins
abstract
Focusing on the innovative context that predicates the integration between digital twin technologies and big data paradigms, including management and analytics, this paper proposes models, principles and implementations of a framework for supporting multidimensional risk analysis and prediction over digital twins. We complement our research contributions by means of a comprehensive experimental campaign over real-life datasets that focus on the emerging digital twin healthcare setting.
Alfredo Cuzzocrea, Ismail Benlaredj
IEEE Big Data1
2024 Privacy-Preserving Identity and Access Management in Multiple Cloud Environments: Models, Issues, and Solutions
abstract
This paper focuses the attention on privacy-preserving identity and access management in multiple Cloud environments, which is an annoying problem in the modern big data era. Within this conceptual context, the paper describes contemporaneous models and issues, and put the basis for future solid solutions. Finally, we provide a summary table where we embed an innovative taxonomy of state-of-the-art research proposals in the reference scientific field.
Alfredo Cuzzocrea, Islam Belmerabet
IEEE Big Data1
2024 Experimental Analysis and Assessment of a Real-Life Cloud Mobile Big OLAP System Enhanced with Compression and Approximation Paradigms
abstract
The increasing complexity and volume of data in Cloud-based environments have posed significant challenges for Online Analytical Processing (OLAP), especially on resource-constrained mobile devices. This paper introduces the Indexed Quad-Tree Summary (IQTS) algorithm, a novel technique for compressing and approximating multidimensional OLAP data cubes. IQTS efficiently handles large data cubes by flattening dimensions, employing Quad-Tree partitioning, and leveraging efficient Indexing methods. These processes significantly reduce storage requirements while ensuring high accuracy in query responses. Our experimental evaluation, utilizing real-world medical datasets, demonstrates that IQTS outperforms existing data compression techniques such as MinSkew, STHoles, and GenHist, delivering superior results in terms of space efficiency and query accuracy. The paper demonstrates the algorithm’s potential for mobile Cloud environments, showing its scalability and adaptability for big data applications.
Alfredo Cuzzocrea, Mojtaba Hajian
IEEE Big Data1
2024 MALAGA - MultidimensionAL Big DAta Analytics over Massive Graph DAta
abstract
Focusing on the main research context represented by the issue of supporting big data analytics over big graph data, this paper introduces and experimentally assesses the framework MALAGA (MultidimensionAL Big DAta Analytics over Massive Graph DAta). MALAGA incorporates several innovations, including OLAP analysis of big graph data, columnar-OLAP methodologies, and Apache Hive extensions. A comprehensive experimental assessment and analysis of the framework’s performance is finally presented and discussed, by significantly integrating the conceptual contributions of our research.
Alfredo Cuzzocrea, Mojtaba Hajian, Abderraouf Hafsaoui
IEEE Big Data1
2024 Creating, Using and Assessing a Generative-AI-Based Human-Chatbot-Dialogue Dataset with User-Interaction Learning Capabilities
abstract
The study illustrates a first step towards an ongoing work aimed at developing a dataset of dialogues potentially useful for customer service conversation management between humans and AI chatbots. The approach exploits ChatGPT 3.5 to generate dialogues. One of the requirements is that the dialogue is characterized by a specific language proficiency level of the user; the other one is that the user expresses a specific emotion during the interaction. The generated dialogues were then evaluated for overall quality. The complexity of the language used by both humans and AI agents, has been evaluated by using standard complexity measurements. Furthermore, the attitudes and interaction patterns exhibited by the chatbot at each turn have been stored for further detection of common conversation patterns in specific emotional contexts. The methodology could improve human-AI dialogue effectiveness and serve as a basis for systems that can learn from user interactions.
Alfredo Cuzzocrea, Giovanni Pilato, Pablo García Bringas
IEEE Big Data1
2024 A Systematic Literature Review of Decentralized Applications in Web3: Identifying Challenges and Opportunities for Blockchain Developers
abstract
The Internet has opened the floor to stakeholders by redefining the way of organizing, communicating, and collaborating that was initiated by the Web’s development. The advancement of the World Wide Web is an outright phenomenon and significant and we witnessed the evolution of the Web. As decentralized technologies continue to gain traction, Web3, or the decentralized internet, has emerged as a promising approach to enable a more secure, transparent, and privacy-preserved digital landscape. In this paper, we thoroughly conduct a systematic study to explore the challenges and opportunities encountered by blockchain developers in the context of decentralized applications (dApps) in Web3. We analyze a set of peer-reviewed research articles, whitepapers, and technical reports and present an in-depth understanding of the current state of Web3 development and its implications. Our finding indicates the opportunities that Web3 can facilitate, such as expanded use cases, enhanced security and privacy, decentralized infrastructure, and the potential for enabling inclusive development resources for blockchain developers. Additionally, we highlight various challenges that blockchain developers deal with including scalability, security, privacy, interoperability, and the need for standardized tools and frameworks along with various challenges in the software development lifecycle (SDLC). While there are significant challenges to overcome, the potential benefits of Web3 are substantial and could lead to a more inclusive, secure, and transparent digital ecosystem. Furthermore, we emphasize the importance of continued research, collaboration, and innovation among stakeholders to address the identified challenges and capitalize on Web3’s opportunities.
Md. Jobair Hossain Faruk, Pratusha Raya, Md Kamrul Siam, Jerry Q. Cheng, Hossain Shahriar, Alfredo Cuzzocrea, Pablo García Bringas
IEEE Big Data6
2024 Toxic language based echo chambers on the Incels.net community: A network analysis approach
abstract
This study examines interaction patterns on the web forum Incels.net through Social Network Analysis, focusing on the formation of echo chambers characterized by toxic language. We explore several hypotheses: (H2) users tend to engage in animated one-to-one interactions by quoting each other; (H4) frequent posters are less likely to be quoted by less active users, indicating a lack of clear leadership; and (H5) users sharing the same sentiment are more likely to participate in large discussions (echo chambers) but do not engage in one-to-one conversations. Our findings enhance the understanding of behaviors that may contribute to radicalization within fringe online communities and offer insights for future research into similar extreme hate forums.
Mathieu Janssen, Giuseppe Cascavilla, Claudia Zucca, Alfredo Cuzzocrea
IEEE Big Data4
2024 Privacy-Preserving Publishing with Generative Adversarial Network (GAN) for Supporting Contact Tracing of Infectious Diseases
abstract
Generative artificial intelligence (AI) has become popular. The combination of increasingly complex datasets beyond human comprehension and the widespread availability of advanced computing systems—such as graphics processing unit (GPU) and tensor processing unit (TPU)—has driven the rapid advancement of generative AI. This technology has found applications in areas such as voice recognition, recommendation systems and data privacy preservation, which foster more data sharing and reuse. While challenges related to bias, fairness and uncertainty in AI continue to evolve, emerging government regulations aim to ensure ethical use and maximize societal benefits. In this paper, we present a system that leverages generative adversarial network (GAN) to enable privacy-preserving data publishing. The system supports contact tracing for infectious diseases like coronavirus disease 2019 (COVID-19) and monkey-pox. Evaluation using COVID-19 data highlights the practicality and effectiveness of our system.
Anifat M. Olawoyin, Carson K. Leung, Hoang Hai Nguyen, Alfredo Cuzzocrea
IEEE Big Data4
2024 EnStack: An Ensemble Stacking Framework of Large Language Models for Enhanced Vulnerability Detection in Source Code
abstract
Automated detection of software vulnerabilities is critical for enhancing security, yet existing methods often struggle with the complexity and diversity of modern codebases. In this paper, we propose a novel ensemble stacking approach that synergizes multiple pre-trained large language models (LLMs)—CodeBERT, GraphCodeBERT, and UniXcoder—to improve vulnerability detection in source code. Our method uniquely combines the semantic understanding of CodeBERT, the structural code representations of GraphCodeBERT, and the cross-modal capabilities of UniXcoder. By fine-tuning these models on the Draper VDISC dataset and integrating their predictions using meta-classifiers such as Logistic Regression, Support Vector Machines (SVM), Random Forest, and XGBoost, we effectively capture complex code patterns that individual models may miss. The meta-classifiers aggregate the strengths of each model, enhancing overall predictive performance. Our ensemble demonstrates significant performance gains over existing methods, with notable improvements in accuracy, precision, recall, F1-score, and AUC-score. This advancement addresses the challenge of detecting subtle and complex vulnerabilities in diverse programming contexts. The results suggest that our ensemble stacking approach offers a more robust and comprehensive solution for automated vulnerability detection, potentially influencing future AI-driven security practices.
Shahriyar Zaman Ridoy, Md. Shazzad Hossain Shaon, Alfredo Cuzzocrea, Mst. Shapna Akter
IEEE Big Data3
2024 NeuroBooster: A Robust Classifier for the Discovery of Neuropeptide Sequences based on Meta-learning Approach
abstract
Neuropeptides (NPs) are fragile proteins that serve as essential signaling molecules in the neurological system, playing a key role in modulating various physiological processes. Identifying particular neuropeptide sequences relevant to specific disorders would be beneficial for accelerating the development of diagnostic tools. The study proposed another approach to detecting NPs with multi-layer perception (MLP) and a bagging classifier-based meta-learning method called NeuroBooster. This investigation initially focused on five feature extractions based on composition, such as AAC, PAAC, physicochemical properties, QSO, and transfer-learning, such as Bert, and F2V strategies. Subsequently, we used the XGB feature selection method in the Bert and F2V methods to obtain the most 100D crucial features. The predicted probabilistic outcomes of NPs from the 8 preliminary models merged and derived a two-stage dataset with 40 dimensions of features and transmitted them into three classic models and two meta-models, through rigorous criteria for evaluation. Compared with the existing predictor, our proposed model NeuroBooster achieved a higher accuracy of 91.91% in the independent test method. Consequently, we discovered important features in these five models underscoring that physicochemical properties are potential targets for identification, thereby revealing new avenues for therapies.
Md. Shazzad Hossain Shaon, Md. Fahim Sultan, Tasmin Karim, Md. Shoaib Hossain Alshan, Alfredo Cuzzocrea, Mst. Shapna Akter
IEEE Big Data5
2024 An Advanced Liver Disease Detection Tool with a Stacking-Ensemble-based Machine Learning Approach
abstract
Liver diseases (LD) encompass a variety of disorders associated with the liver, including infectious hepatitis, obesity, cirrhosis, and malignancy, which constitute significant health issues in the world. Due to the minimal symptoms, comprehending the disease becomes extremely difficult until its severe stages; earlier detection is advantageous for appropriate action, which could save lives. This study developed a StackLD framework based on the stacking-ensemble-based machine learning approach. In the preliminary phase, we collected the Indian Liver Patient Dataset (ILPD) dataset, which contains 11 features, and the dataset highlighted a significant discrepancy. To overcome this, we used the SMOTE to rebalance the dataset, facilitating the development of robust machine learning models. We applied 7 different models such as XGB, LGBM, DT, KNN, RF, KNN, and stacking approaches with several evaluation metrics on independent test methods The analysis presented that the stacking technique executed superior in accuracy, sensitivity, specificity, and area under the curve, with values of 0.8622, 0.8933, 0.8369, and 0.9275, respectively. These outcomes indicate that our approach effectively differentiates between positive and negative classes. This study illustrates that the Alkphos, Sgot, and Sgpt elements have a significant role in determining liver disease (LD) from various features. As a result, a web server was built using these attributes, demonstrating that our model accurately predicts liver disorders at an early stage which is useful insight into the medical field with the potential to improve diagnostic procedures and patient outcomes.
Md. Shazzad Hossain Shaon, Md. Fahim Sultan, Tasmin Karim, Alfredo Cuzzocrea, Mst. Shapna Akter
IEEE Big Data4
2024 Privacy-Preserving Big Hierarchical Data Analytics via Co-Occurrence Analysis
Alfredo Cuzzocrea, Selim Soufargi
DATA1
2024 From Theory to Practice of Multidimensional Big Data Analytics over Big Healthcare Data: A Real-Life Case Study
Alfredo Cuzzocrea, Abderraouf Hafsaoui, Carmine Gallo
IDEAS1
2024 A bayesian-neural-networks framework for scaling posterior distributions over different-curation datasets
Alfredo Cuzzocrea, Alessandro Baldo 0001, Edoardo Fadda
J. Intell. Inf. Syst.1
2023 Quantum Cryptography for Enhanced Network Security: A Comprehensive Survey of Research, Developments, and Future Directions
abstract
With the ever-growing concern for internet security, the field of quantum cryptography emerges as a promising solution for enhancing the security of networking systems. In this paper, 20 notable papers from leading conferences and journals are reviewed and categorized based on their focus on various aspects of quantum cryptography, including key distribution, quantum bit commitment, post-quantum cryptography, and counterfactual quantum key distribution. The paper explores the motivations and challenges of employing quantum cryptography, addressing security and privacy concerns along with existing solutions. Secure key distribution, a critical component in ensuring the confidentiality and integrity of transmitted information over a network, is emphasized in the discussion. The survey examines the potential of quantum cryptography to enable secure key exchange between parties, even when faced with eavesdropping, and other applications of quantum cryptography. Additionally, the paper analyzes the methodologies, findings, and limitations of each reviewed study, pinpointing trends such as the increasing focus on practical implementation of quantum cryptography protocols and the growing interest in post-quantum cryptography research. Furthermore, the survey identifies challenges and open research questions, including the need for more efficient quantum repeater networks, improved security proofs for continuous variable quantum key distribution, and the development of quantum-resistant cryptographic algorithms, showing future directions for the field of quantum cryptography.
Mst. Shapna Akter, Juanjose Rodriguez-Cardenas, Hossain Shahriar, Alfredo Cuzzocrea, Fan Wu 0013
IEEE Big Data4
2023 A Trustable LSTM-Autoencoder Network for Cyberbullying Detection on Social Media Using Synthetic Data
abstract
Social media cyberbullying has a detrimental effect on human life. As online social networking grows daily, the amount of hate speech also increases. Such terrible content can cause depression and actions related to suicide. This paper proposes a trustable LSTM-Autoencoder Network for cyberbullying detection on social media using synthetic data. We have demonstrated a cutting-edge method to address data availability difficulties by producing machine-translated data. However, several languages such as Hindi and Bangla still lack adequate investigations due to a lack of datasets. We carried out experimental identification of aggressive comments on Hindi, Bangla, and English datasets using the proposed model and traditional models, including Long Short-Term Memory (LSTM), Bidirectional Long ShortTerm Memory (BiLSTM), LSTM-Autoencoder, Word2vec, Bidirectional Encoder Representations from Transformers (BERT), and Generative Pre-trained Transformer 2 (GPT-2) models. We employed evaluation metrics such as f1-score, accuracy, precision, and recall to assess the models’ performance. Our proposed model outperformed all the models on all datasets, achieving the highest accuracy of 95%. Our model achieves state-of-the-art results among all the previous works on the dataset we used in this paper.
Mst. Shapna Akter, Hossain Shahriar, Alfredo Cuzzocrea, Fan Wu 0013, Juanjose Rodriguez-Cardenas
IEEE Big Data3
2023 Adversarial Data-Augmented Resilient Intrusion Detection System for Unmanned Aerial Vehicles
abstract
With the growing adoption of unmanned aerial vehicles (UAVs) across various domains, the security of their operations is paramount. UAVs, heavily dependent on GPS navigation, are at risk of jamming and spoofing cyberattacks, which can severely jeopardize their performance, safety, and mission integrity. Intrusion detection systems (IDSs) are typically employed as defense mechanisms, often leveraging traditional machine learning techniques. However, these IDSs are susceptible to adversarial attacks that exploit machine learning models by introducing input perturbations. In this work, we propose a novel IDS for UAVs to enhance resilience against such attacks using generative adversarial networks (GAN). We also comprehensively study several evasion-based adversarial attacks and utilize them to compare the performance of the proposed IDS with existing ones. The resilience is achieved by generating synthetic data based on the identified weak points in the IDS and incorporating these adversarial samples in the training process to regularize the learning. The evaluation results demonstrate that the proposed IDS is significantly robust against adversarial machine learning-based attacks compared to the state-of-the-art IDSs while maintaining a low false positive rate.
Muneeba Asif, Mohammad Ashiqur Rahman, Kemal Akkaya, Hossain Shahriar, Alfredo Cuzzocrea
IEEE Big Data5
2023 BigData Fusion for Trajectory Prediction of Multi-Sensor Surveillance Information Systems
abstract
Video surveillance information systems assist forensics to examine and analyze the evidence from crime scenes to develop objective findings in the investigation of crime. Often, the existing surveillance information systems exploit an array of security cameras and IoT devices monitoring the same crime scene from different points of view while the crime unfolds over a range of time. However, none can automatically and selectively merge big data streams connected to such systems to provide a holistic, end-to-end safety picture.This work proposes a trajectory prediction architecture framework within a multi-sensor surveillance system. We developed a novel position measurement technique using monocular depth perception networks with multi-camera setup using triangulation. We tested and compared our technique with a single camera sensor in our first experiment and as the multi-camera setup determines the position of our target more accurately, we used our measurement function in our second experiment. In our second experiment, we employed the Unscented Kalman Filter (UKF) for predicting the trajectory of the target, and proved that UKF has good potential for being used in surveillance systems. Lastly, we designed a general architecture framework for big data analysis in multi-sensor surveillance systems consisting the four layers: the Sensor Layer, the Single Sensor Computation Layer, the Data Fusion and Interpretation Layer, and the Human Interaction Layer.
Giuseppe Cascavilla, Alfredo Cuzzocrea, Daniel De Pascale, Mandana Omidbakhsh, Damian A. Tamburri
IEEE Big Data2
2023 F-TBDA: A Frequency-Based Temporal Big Data Analytics Technique for Mining and Analyzing Quality-Of-Life Indicators of Cancer Patients
abstract
In this paper, we introduce and experimentally assess an innovative big data analytics technique for mining and analyzing Quality-of-Life Indicators (QoL) over time among patients with lung cancer and treated with immunotherapy. In more details, given datasets of QoL indicators collected over time, at regular intervals, the F-TBDA technique (Frequency-based Temporal Big Data Analytics) computes temporal relative frequency tables over fixed-time intervals where data of subsequent observations (i.e., intermediate therapy) are compared with the baseline observation (i.e., starting therapy). Then, on the basis of these relative frequency tables, both simple and complex frequency-based big data analytics tools are developed, in order to unveil hidden patterns over cancer patient therapies. Experimental results on top of a real-life dataset nicely complete the theoretical contributions we provide in our research.
Alfredo Cuzzocrea, Geertruida H. de Bock, Willemijn J. Maas, Selim Soufargi
IEEE Big Data1
2023 Effective and Efficient Big OLAP Data Cube Compression in Mobile Cloud Environments: The IQTS Algorithm
abstract
This paper introduces the Indexed Quad-Tree Summary (IQTS) algorithm, along with its concepts, models, and “philosophy”. IQTS allows us to compress multidimensional OLAP views derived from big OLAP data cubes that populate Mobile Clouds. In addition to this, IQTS effectively and efficiently supports approximate query answering over compressed multidimensional OLAP views. These functionalities turn to be enabling functionalities for big data applications in a wide collection of modern scenarios, such as healthcare analytics. In light of these considerations, this paper provides motivations, anatomy and procedures of IQTS, enriched by several case studies that contribute to pinpoint the benefits coming from our proposed algorithm.
Alfredo Cuzzocrea, Mojtaba Hajian
IEEE Big Data1
2023 Machine-Learning-Based Multidimensional Big Data Analytics over Clouds via Multi-Columnar Big OLAP Data Cube Compression
abstract
This paper proposes a new theory on combining innovative Multidimensional Big Data Analytics with well-known Machine Learning (ML) in order to magnify the expressive power and the accuracy of knowledge insights discovery from massive big datasets. At the level of enabling technology, with the goal of fully supporting this novel paradigm, the issue of managing and mining big OLAP data cubes over Clouds arises. Due to computational complexity requirements, the latter challenge is addressed by proposing an innovative solution for (1) representing big OLAP data cubes over Clouds via a multi-column-based representation, and (2) compressing the deriving multi-column representations for achieving the desired effectiveness and efficiency. This paper introduces the fundamental model of Machine-Learning-Based Multidimensional Big Data Analytics, along with a reference architecture implementing it.
Alfredo Cuzzocrea, Abderraouf Hafsaoui, Carson K. Leung
IEEE Big Data1
2023 Towards Big Data Analytics over Mobile User Data using Machine Learning
abstract
Machine Learning (ML) is a science that forces computers to learn and behave like humans. As these systems interact with data, networks, and people, they automatically become smarter so that they can eventually solve or predict a practical issue in the world for us. The use of ML can be a giant leap for cannot simply be integrated as the top layer. This requires redefining workflow, architecture, data collection and storage, analytics, and other modules. This paper aims to discuss the issue of machine learning technique for analysis data of mobile user. First, we identified the machine learning benefits and drawbacks, challenges, advantages of using Machine Learning. Then, we propose a generic model of analytic mobile user data using ML, the model is centered on the machine learning component, which interacts with two other components, including mobile user data, and system. The interactions go in both directions. For instance, mobile user data serves as inputs to the learning component and the latter generates outputs; system architecture has impact on how learning algorithms should run and how efficient it is to run them, and simultaneously meeting. Mobile user data goes through several stages: prepossessing which includes the steps we need to follow to transform or encode the data so that it can be easily analyzed by the machine. Then, modelling in this step we will be clustering and classification the data obtained. Finally, evaluation, various measures of performance, accuracy, recall, precision, and F-measure were used to analyze the results of the naive Bayes, SVM, and K-nearest neighbor classification algorithms.
Sabrina Ichou, Slimane Hammoudi, Alfredo Cuzzocrea, Abdelkrim Meziane, Amel Benna
IEEE Big Data3
2023 QFLS: A Cloud-Based Framework for Supporting Big Healthcare Data Management and Analytics from Big Data Lakes: Definitions, Requirements, Models and Techniques
Alfredo Cuzzocrea, Selim Soufargi
DATA1
2023 Supporting Big Healthcare Data Management and Analytics: The Cloud-Based QFLS Framework
Alfredo Cuzzocrea, Selim Soufargi
DaWaK1
2023 Effective and Efficient Heuristic Algorithms for Supporting Optimal Location of Hubs over Networks with Demand Uncertainty
Alfredo Cuzzocrea, Luigi Canadè, Giulia Fornari, Vittorio Gatto, Abderraouf Hafsaoui
DEXA (1)1
2023 A Machine-Learning Framework for Supporting Content Recommendation via User Feedback Data and Content Profiles in Content Managements Systems
Debashish Roy, Chen Ding 0004, Alfredo Cuzzocrea, Islam Belmerabet
DEXA (2)3
2023 Privacy-Preserving OLAP via Modeling and Analysis of Query Workloads: Innovative Theories and Theorems
abstract
This paper proposes innovative theories and theorems in the context of a state-of-the-art paper that computes privacy-preserving OLAP cubes via modeling and analyzing query workloads. The work contributes to actual literature by devising a solid theoretical framework that can be used for future optimization opportunities.
Alfredo Cuzzocrea
SSDBM1
2022 Software Supply Chain Vulnerabilities Detection in Source Code: Performance Comparison between Traditional and Quantum Machine Learning Algorithms
abstract
The software supply chain (SSC) attack has become one of the crucial issues that are being increased rapidly with the advancement of the software development domain. In general, SSC attacks execute during the software development processes lead to vulnerabilities in software products targeting downstream customers and even involved stakeholders. Machine Learning approaches are proven in detecting and preventing software security vulnerabilities. Besides, emerging quantum machine learning can be promising in addressing SSC attacks. Considering the distinction between traditional and quantum machine learning, performance could be varies based on the proportions of the experimenting dataset. In this paper, we conduct a comparative analysis between quantum neural networks (QNN) and conventional neural networks (NN) with a software supply chain attack dataset known as ClaMP. Our goal is to distinguish the performance between QNN and NN and to conduct the experiment, we develop two different models for QNN and NN by utilizing Pennylane for quantum and TensorFlow and Keras for traditional respectively. We evaluated the performance of both models with different proportions of the ClaMP dataset to identify the f1 score, recall, precision, and accuracy. We also measure the execution time to check the efficiency of both models. The demonstration result indicates that execution time for QNN is slower than NN with a higher percentage of datasets. Due to recent advancements in QNN, a large level of experiments shall be carried out to understand both models accurately in our future research.
Mst. Shapna Akter, Md. Jobair Hossain Faruk, Nafisa Anjum, Mohammad Masum, Hossain Shahriar, Akond Ashfaque Ur Rahman, Fan Wu 0013, Alfredo Cuzzocrea
IEEE Big Data9
2022 Deep Learning Approach for Classifying the Aggressive Comments on Social Media: Machine Translated Data Vs Real Life Data
abstract
Aggressive comments on social media negatively impact human life. Such offensive contents are responsible for depression and suicidal-related activities. Since online social networking is increasing day by day, the hate content is also increasing. Several investigations have been done on the domain of cyberbullying, cyberaggression, hate speech, etc. The majority of the inquiry has been done in the English language. Some languages (Hindi and Bangla) still lack proper investigations due to the lack of a dataset. This paper particularly worked on the Hindi, Bangla, and English datasets to detect aggressive comments and have shown a novel way of generating machine-translated data to resolve data unavailability issues. A fully machine-translated English dataset has been analyzed with the models such as the Long Short term memory model (LSTM), Bidirectional Long-short term memory model (BiLSTM), LSTM-Autoencoder, word2vec, Bidirectional Encoder Representations from Transformers (BERT), and generative pre-trained transformer (GPT-2) to make an observation on how the models perform on a machine-translated noisy dataset. We have compared the performance of using the noisy data with two more datasets such as raw data, which does not contain any noises, and semi-noisy data, which contains a certain amount of noisy data. We have classified both the raw and semi-noisy data using the aforementioned models. To evaluate the performance of the models, we have used evaluation metrics such as F1-score, accuracy, precision, and recall. We have achieved the highest accuracy on raw data using the gpt2 model, semi-noisy data using the BERT model, and fully machine-translated data using the BERT model. Since many languages do not have proper data availability, our approach will help researchers create machine-translated datasets for several analysis purposes.
Mst. Shapna Akter, Hossain Shahriar, Nova Ahmed, Alfredo Cuzzocrea
IEEE Big Data4
2022 Handwritten Word Recognition using Deep Learning Approach: A Novel Way of Generating Handwritten Words
abstract
A handwritten word recognition system comes with issues such as-lack of large and diverse datasets. It is necessary to resolve such issues since millions of official documents can be digitized by training deep learning models using a large and diverse dataset. Due to the lack of data availability, the trained model does not give the expected result. Thus, it has a high chance of showing poor results. This paper proposes a novel way of generating diverse handwritten word images using handwritten characters. The idea of our project is to train the BiLSTM-CTC architecture with generated synthetic handwritten words. The whole approach shows the process of generating two types of large and diverse handwritten word datasets: overlapped and non-overlapped. Since handwritten words also have issues like overlapping between two characters, we have tried to put it into our experimental part. We have also demonstrated the process of recognizing handwritten documents using the deep learning model. For the experiments, we have targeted the Bangla language, which lacks the handwritten word dataset, and can be followed for any language. Our approach is less complex and less costly than traditional GAN models. Finally, we have evaluated our model using Word Error Rate (WER), accuracy, f1-score, precision, and recall metrics. The model gives 39% WER score, 92% percent accuracy, and 92% percent f1 scores using non-overlapped data and 63% percent WER score, 83% percent accuracy, and 85% percent f1 scores using overlapped data.
Mst. Shapna Akter, Hossain Shahriar, Alfredo Cuzzocrea, Nova Ahmed, Carson K. Leung
IEEE Big Data3
2022 Multi-class Skin Cancer Classification Architecture Based on Deep Convolutional Neural Network
abstract
Skin cancer is a deadly disease. Melanoma is a type of skin cancer responsible for the high mortality rate. Early detection of skin cancer can enable patients to treat the disease and minimize the death rate. Skin cancer detection is challenging since different types of skin lesions share high similarities. This paper proposes a computer-based deep learning approach that will accurately identify different kinds of skin lesions. Deep learning approaches can detect skin cancer very accurately since the models learn each pixel of an image. Sometimes humans can get confused by the similarities of the skin lesions, which we can minimize by involving the machine. However, not all deep learning approaches can give better predictions. Some deep learning models have limitations, leading the model to a false-positive result. We have introduced several deep learning models to classify skin lesions to distinguish skin cancer from different types of skin lesions. Before classifying the skin lesions, data preprocessing and data augmentation methods are used. Finally, a Convolutional Neural Network (CNN) model and six transfer learning models such as Resnet-50, VGG-16, Densenet, Mobilenet, Inceptionv3, and Xception are applied to the publically available benchmark HAM10000 dataset to classify seven classes of skin lesions and to conduct a comparative analysis. The models will detect skin cancer by differentiating the cancerous cell from the non-cancerous ones. The models’ performance is measured using performance metrics such as precision, recall, f1 score, and accuracy. We receive accuracy of 90, 88, 88, 87, 82, and 77 percent for inceptionv3, Xception, Densenet, Mobilenet, Resnet, CNN, and VGG16, respectively. Furthermore, we develop five different stacking models such as inceptionv3-inceptionv3, Densenet-mobilenet, inceptionv3-Xception, Resnet50-Vgg16, and stack-six for classifying the skin lesions and found that the stacking models perform poorly. We achieve the highest accuracy of 78 percent among all the stacking models.
Mst. Shapna Akter, Hossain Shahriar, Sweta Sneha, Alfredo Cuzzocrea
IEEE Big Data4
2022 Explaining IoT Attacks: An Effective and Efficient Semi-Supervised Learning Framework
abstract
Cyber-attacks targeting Internet-of-Things (IoT) devices are prevalent due to the limited security resources of the target devices and their often limited connectivity. Explaining such attacks is therefore greatly important to construct countermeasures. Current methods of automated IoT attack analysis require either large amounts of labelled data for classification, or use clustering methods which can be inaccurate. However, when a desired grouping of the data, as well as some prior knowledge about some observations in the data is available, approximate semi-supervised learning methods may be used to create accurate cluster arrangements. We therefore investigated the use of semi-supervised clustering approaches for creating accurate clusters of IoT attack sessions based on their goals and characteristic commonalities. We first manually created a ground-truth grouping of recent IoT attacks based on their goal. We differentiated the goal of each session according to the purpose of the used commands and the taken approach, resulting in a total of five classes. We then automatically constructed a feature set suitable for clustering similar IoT attack sessions using a method proposed in recent literature, and passed it to two different semi-supervised clustering algorithms using either labelled data (SeededKMeans) or pairwise constraints (PCKMeans) as prior knowledge. We found that both semi-supervised approaches were able to create accurate cluster arrangements using only small amounts of prior knowledge. Moreover, they outperformed an entirely unsupervised KMeans algorithm in terms of accuracy.
Giuseppe Cascavilla, Reinier Zwart, Damian A. Tamburri, Alfredo Cuzzocrea
IEEE Big Data4
2022 GAPS: Generality and Precision with Shapley Attribution
abstract
In an age of the growing use of Machine-learning, it has become an imperative task to be able to explain the processes behind the functions of many "black box" models. The explainability feature of artificial intelligence is key to building trust between humans and computers' algorithmic predictions. One of the main ways to generate this interpretability is through attribution methods, which produce importance values of each feature for a single instance in a dataset. There are many different ways of attribution for various Machine-learning models, including ones designed for specific models or "model agnostic" attribution methods—ones that do not require a specific model to achieve importance values. These attribution methods are valued because of their easily understood nature. While evaluation procedures exist such as generality and precision for rule-based explanation methods, these have not been used on attribution methods until recently. A recent experiment by Ratul et al. [1] proved that the two most popular local model-agnostic attribution methods, LIME and SHAP, have poor precision and generality. In this paper, we propose a new attribution method, the Generality and Precision Shapley Attributions (GAPS). To evaluate these models, we use the generality and precision equations used previously to evaluate the other models. We present our findings that GAPS produces higher generality and precision scores than the existing LIME and SHAP models.
Brian Daley, Qudrat E. Alahy Ratul, Edoardo Serra, Alfredo Cuzzocrea
IEEE Big Data4
2022 Authentic Learning of Machine Learning in Cybersecurity with Portable Hands-on Labware: Neural Network Algorithms for Network Denial of Service (DOS) Detection
abstract
The primary goal of the authentic learning approach is to engage and motivate students in a learning environment that encourages all students in learning. This approach provides students with hands-on experiences in solving real-world security problems. We designed and developed ten learning modules based on 10 cybersecurity cases with different ML solutions. Each learning module consists of pre-lab, lab, and post-lab (Pre/Lab/Post) activities. All portable labs are made available on Google CoLab for ML to cybersecurity so that students can access and practice these hands-on labs anywhere and anytime without time tedious installation and configuration which will engage students in learning concepts and getting more experience for hands-on problem-solving skills. In this paper, we adopt Neural Network Algorithms for Network Denial of Service (DOS) Detection where we apply the KDDCup 1999 datasets contain a standard set of data to be audited, which includes a wide variety of intrusions simulated in a military network environment. Our primary goal of this lab is to show whether a link is a malicious or safe connection. Our demonstration shows an achieved accuracy of 99.89%.
Md. Jobair Hossain Faruk, Hossain Shahriar, Dan Chia-Tien Lo, Michael E. Whitman, Alfredo Cuzzocrea, Fan Wu 0013, Victor Clincy
IEEE Big Data6
2022 Multi-criteria Rating and Review based Recommendation Model
abstract
These days, due to the advancement of information technology, recommendation system has become one of the key tools for e-commerce business. E-commerce platforms allow users to provide feedback in both written comments and numerical ratings. Recommendation systems are utilized to recommend users to new or unseen items based on these previously collected comments or ratings. In recent years, multi-aspect or multi-criteria based recommendation systems have been studied a lot by the recommendation research community. However, these research works are conducted either with reviews or ratings, not with both. In this project, we argue that integrating both textual reviews (with multiple aspects) and numerical multi-criteria ratings can further enhance the overall rating prediction accuracy. We propose a Multi-criteria Rating and Review based Recommendation model (MRRRec). We show that incorporating multi-criteria ratings into multi-aspect ratings from reviews has a great impact on performance. Our proposed model outperforms several state-of-the-art models such as ANR, DeepCoNN, and Deep Multi-criteria Recommendation System in terms of MSE, MAE, precision, recall, and F1. We show that our proposed model achieves an average of 19% and 23.0% lower MSE and MAE respectively and 7.0%, 1.0% and 3.8% higher precision, recall, and F1 score respectively. We further show that our model performs significantly better with Word2Vec word embedding than the GloVe word embedding method.
Emrul Hasan, Chen Ding 0004, Alfredo Cuzzocrea
IEEE Big Data3
2022 Privacy and Security of Mobile Users in Smart Cities: A Reference Architecture
abstract
Privacy and Security in big data and smart cities play a major role to ensure better quality of citizens life. Privacy emphasizes on the data being collected, shared, and used in the right manner, and security focuses on protecting the data from intruders’ attack, and exploitation of data for other purposes such as criminal behavior. This paper aims to discuss the issue of privacy and security of mobile users in big data and smart cities. First, based on big data privacy and security challenges, classification, and models, we identify the privacy and security requirements for a mobile user. Then, we propose a generic architecture and an algorithm for the management of privacy and security of a mobile user in a smart environment. This architecture groups the main components required for implementing the proposed algorithm that ensure the privacy and security of mobile users in smart cities.
Sabrina Ichou, Slimane Hammoudi, Amel Benna, Abdelkrim Meziane, Alfredo Cuzzocrea
IEEE Big Data5
2022 Anomaly Detection in Cybersecurity Events Through Graph Neural Network and Transformer Based Model: A Case Study with BETH Dataset
abstract
With the increasing prevalence of the internet, detecting malicious behavior is becoming a greater need. This problem can be formulated as an anomaly detection task on provenance data, where attacks are detectable as anomalies in the behavior of the system. While network data is quite prevalent, we focus on system logs and propose a novel approach with two main components. The first is to make use of the graph-like structure of the logs in which processes enact events and generate additional processes, using a graph neural network (GNN) to produce representations of each event which encode information about their neighboring events in an unsupervised manner. The second is to make use of the complex features such as command arguments which vary widely and cannot be used in the presented format as features in typical machine learning algorithms. If these features are instead encoded using transformer models, they can then be used in other algorithms such as a GNN or anomaly detector. These two approaches combined improve anomaly detection results for the BETH dataset by around 8 percent as compared to the manually engineered features alone.
Bishal Lakha, Sara Lilly Mount, Edoardo Serra, Alfredo Cuzzocrea
IEEE Big Data4
2022 A Novel Machine Learning Based Framework for Bridge Condition Analysis
abstract
Bridges play a vital part in the transportation system by ensuring the connectedness of transportation systems, which is critical for a country’s social and economic prosperity by offering daily mobility to the people. However, according to the American Society of Civil Engineers (ASCE 2017), many U.S. bridges are in critical condition, raising safety issues, with 9.1 and 13.6 percent of the country’s 614,387 bridges, respectively, structurally defective, and functionally obsolete. Every day, 178 million people traverse these structurally defective bridges. Furthermore, the average annual failure rate is expected to be between 87 and 222. Bridge breakdowns have disastrous repercussions, and in many cases, result in death. While bridge authorities strive to improve bridge conditions, budget limits make it difficult to make cost-effective maintenance decisions. Bridge authorities distribute limited repair resources based on projected future bridge conditions. As a result, building a data-driven, autonomous, and effective bridge condition prediction model is critical for improving maintenance decision-making. In this paper, we present a novel bridge condition prediction framework using advanced Machine Learning (ML) algorithms on the National Bridge Inventory (NBI) dataset. The framework consists of two stages, where the most informative features from the NBI dataset are selected using the Recursive Feature Elimination process and in the 2ndstep, ML classifiers are applied to the selected features for bridge condition prediction. The experimental results show that the proposed framework can effectively predict bridge conditions by producing highly accurate results in terms of accuracy, precision, recall, and f1-score.
Mohammad Masum, Nafisa Anjum, Md. Jobair Hossain Faruk, Hossain Shahriar, Maria Valero, Mohammed Karim, Akond Ashfaque Ur Rahman, Fan Wu 0013, Alfredo Cuzzocrea
IEEE Big Data10
2022 A Crowd Source System for YouTube Big Data Analytics: Unpacking Values from Data Sprawl
abstract
YouTube has emerged as the most popular video platform across the world. This paper proposes a system for stream based meta-data analytics, to gain insights and uncover hidden patterns, for YouTube videos. The system reports the number of videos uploaded for each category, the videos with the highest views and highest likes etc. It crowd-sources the calls to the YouTube’s search and data APIs, and feeds it to Kafka Stream to process using PySpark. Finally, the processed video meta-data in stored in Apache Cassandra for downstream consumption. By experimenting with different time windows from 5 to 30 minutes, it is observed that 15-minute window is optimum for getting adequate video data de-duplication. From the available 30 video categories, with the minimum video length of 30 minutes, our study observed that the top 10 video categories with the highest number of videos are Film & Animation, Autos & Vehicles, Music, Sports, Travel & Events, Entertainment, News & politics, Documentary, Science & Technology and Education. The highest videos are uploaded under Entertainment category, which clearly captures the content creator’s interest.
Jagan Mohan Reddy, Abhishek Attuluri, Abhinay Kolli, Hossain Shahriar, Alfredo Cuzzocrea
IEEE Big Data6
2022 Mahalanobis Distance Based K-Means Clustering
Paul O. Brown, Meng Ching Chiang, Shiqing Guo, Yingzi Jin, Carson K. Leung, Evan L. Murray, Adam G. M. Pazdor, Alfredo Cuzzocrea
DaWaK8
2022 An Innovative Risk Assessment Methodology for Medical Information Systems
abstract
Modern Medical Information Systems very often comprise Medical Devices and governed by regulations which require stringent Risk Management activities to be implemented to minimize the occurrence of safety risks. Currently, the reference standard adopted by manufacturers for Risk Management is ISO 14971, which, however, was devised for traditional (mostly hardware) Medical Devices and does not either take into account the peculiarities of modern Medical Information Systems, or define a formal methodology to conduct Risk Assessment. Moreover, the approaches currently implemented by manufacturers typically aims at obtaining qualitative Risk Assessment results. Within the so-delineated application scenario, this paper proposes a methodology for the Dynamic Probabilistic Risk Assessment of Medical Information Systems, by specifically looking at medical devices that are intended as one of the most relevant components in such systems. The methodology complies with ISO 14971 and improves current practices because it allows the analyst to conduct a quantitative analysis, also taking into account the temporal dimension. It relies on a Probabilistic Risk Model, defined as a set of Markov Models, which is model-checked to obtain quantitative information about the risks. The proposed methodology is also adopted to improve definitively the Medical Device post-market surveillance, which is currently implemented as a ”wait for an incident” activity. In other words, currently a manufacturer sets up a service that has to ”react” to an incident by starting an investigation activity. Instead, the methodology proposes the adoption of risk models defined during the development phase also to re-assess periodically the risks related to the product during the post-market surveillance. This may prevent some incidents because risks are assessed using data collected in the field (no longer guesstimated as during the development phase) and taking into account the temporal effects on probability distributions (such as the deterioration of hardware/software components over the time).
Antonio Coronato, Alfredo Cuzzocrea
IEEE Trans. Knowl. Data Eng.2
2021 Advances in Social Network Analysis and Mining in the Big Data Era: Overview of the IEEE/ACM ASONAM 2021 International Conference
Michele Coscia, Alfredo Cuzzocrea, Kai Shu
ASONAM2
2021 Detecting Botnet Nodes via Structural Node Representation Learning
abstract
Botnets are an ever-growing threat to private users, small companies, and even large corporations. They are known for spamming, mass downloads, and launching distributed denial-of-service (DDoS) attacks that have a destructive impact on large corporations. With the rise of internet-of-things (IoT) devices, they are also used to mine cryptocurrency, intercept data in transit and send logs containing sensitive information to the master botnet. Many approaches have been developed to detect botnet activities. A few approaches employ graph neural networks (GNN) to analyze the behavior of hosts using a directed graph to represent their communications. However, while designed to capture structural graph properties, GNN may overfit, and therefore fail to capture these properties when the network is unknown. In this work we hypothesize that structural graph patterns can be used to effectively detect Botnets. We then propose a structural iterative representation learning approach for graph nodes, which is designed to perform well on unseen data, called Inferential SIR-GN. Our model creates a vector representation for each node that epitomizes its structural information. We demonstrate that this set of node representation vectors can be used with a neural network classifier to identify bot nodes within an unknown network with better performance than the current state-of-the-art GNN based method.
Justin Carpenter, Janet Layne, Edoardo Serra, Alfredo Cuzzocrea
IEEE BigData4
2021 Unsupervised Risk for Privacy
abstract
This position paper deals with privacy for deep neural networks, more precisely with robustness to membership inference attacks. The current state-of-the-art methods, such as the ones based on differential privacy and training loss regularization, mainly propose approaches that try to improve the compromise between privacy guarantees and decrease in model accuracy. We propose a new research direction that challenges this view, and that is based on novel approximations of the training objective of deep learning models. The resulting loss offers several important advantages with respect to both privacy and model accuracy: it may exploit unlabeled corpora, it both regularizes the model and improves its generalization properties, and it encodes corpora into a latent low-dimensional parametric representation that complies with Federated Learning architectures. Arguments are detailed in the paper to support the proposed approach and its potential beneficial impact with regard to preserving both privacy and quality of deep learning.
Christophe Cerisara, Alfredo Cuzzocrea
IEEE BigData2
2021 Privacy-Preserving Big Data Exchange: Models, Issues, Future Research Directions
abstract
Big data exchange is an emerging problem in the context of big data management and analytics. In big data exchange, multiple entities exchange big datasets beyond the common data integration or data sharing paradigms, mostly in the context of data federation architectures. How to make big data exchange while ensuring privacy preservation constraintsƒ The latter is a critical research challenge that is gaining momentum on the research community, especially due to the wide family of application scenarios where it plays a critical role (e.g., social networks, bio-informatics tools, smart cities systems and applications, and so forth). Inspired by these considerations, in this paper we provide an overview of models and issues in the context of privacy-preserving big data exchange research, along with a selection of future research directions that will play a critical role in next-generation research.
Alfredo Cuzzocrea, Ernesto Damiani
IEEE BigData1
2021 Identifying Malicious Users in the Offshore Leaks Networks via Structural Node Representation Learning
abstract
Starting in 2013, the International Consortium of Investigative Journalists released a series of networks, known as the Offshore Leaks Networks, detailing the information of entities and transactions of offshore accounts. Through cross-referencing with known blacklists of entities, illicit individuals and transactions were able to be identified in the networks provided. In machine learning research, the Offshore Leaks Networks draws off of large databases of data to classify many nodes in high dimensional space. The chief problem with node classification is that the illicit entities are not always known, and techniques have been devised to tackle this problem, such as centrality and structural-based learning. In this paper, SparseStruct—the algorithm developed by Serra et al. [1]— is shown to achieve the best results. This is because it uses a structural node representational learning technique able to identify specific structural patterns in the graph. This technique achieved AUROC scores of between 0.61 and 0.81, with three of the four scores being the top score of all classifiers compared.
Brian Daley, Edoardo Serra, Alfredo Cuzzocrea
IEEE BigData3
2021 Malware Detection and Prevention using Artificial Intelligence Techniques
abstract
With the rapid technological advancement, security has become a major issue due to the increase in malware activity that poses a serious threat to the security and safety of both computer systems and stakeholders. To maintain stakeholder’s, particularly, end user’s security, protecting the data from fraudulent efforts is one of the most pressing concerns. A set of malicious programming code, scripts, active content, or intrusive software that is designed to destroy intended computer systems and programs or mobile and web applications is referred to as malware. According to a study, naive users are unable to distinguish between malicious and benign applications. Thus, computer systems and mobile applications should be designed to detect malicious activities towards protecting the stakeholders. A number of algorithms are available to detect malware activities by utilizing novel concepts including Artificial Intelligence, Machine Learning, and Deep Learning. In this study, we emphasize Artificial Intelligence (AI) based techniques for detecting and preventing malware activity. We present a detailed review of current malware detection technologies, their shortcomings, and ways to improve efficiency. Our study shows that adopting futuristic approaches for the development of malware detection applications shall provide significant advantages. The comprehension of this synthesis shall help researchers for further research on malware detection and prevention using AI.
Md. Jobair Hossain Faruk, Hossain Shahriar, Maria Valero, Farhat Lamia Barsha, Shahriar Sobhan, Md Abdullah Khan, Michael E. Whitman, Alfredo Cuzzocrea, Dan Chia-Tien Lo, Akond Ashfaque Ur Rahman, Fan Wu 0013
IEEE BigData8
2021 Bayesian Hyperparameter Optimization for Deep Neural Network-Based Network Intrusion Detection
abstract
Traditional network intrusion detection approaches encounter feasibility and sustainability issues to combat modern, sophisticated, and unpredictable security attacks. Deep neural networks (DNN) have been successfully applied for intrusion detection problems. The optimal use of DNN-based classifiers requires careful tuning of the hyper-parameters. Manually tuning the hyperparameters is tedious, time-consuming, and computationally expensive. Hence, there is a need for an automatic technique to find optimal hyperparameters for the best use of DNN in intrusion detection. This paper proposes a novel Bayesian optimization-based framework for the automatic optimization of hyperparameters, ensuring the best DNN architecture. We evaluated the performance of the proposed framework on NSL-KDD, a benchmark dataset for network intrusion detection. The experimental results show the framework’s effectiveness as the resultant DNN architecture demonstrates significantly higher intrusion detection performance than the random search optimization-based approach in terms of accuracy, precision, recall, and f1-score.
Mohammad Masum, Hossain Shahriar, Hisham M. Haddad, Md. Jobair Hossain Faruk, Maria Valero, Md Abdullah Khan, Mohammad Ashiqur Rahman, Muhaiminul I. Adnan, Alfredo Cuzzocrea, Fan Wu 0013
IEEE BigData9
2021 Open Data Lake to Support Machine Learning on Arctic Big Data
abstract
The era of big data is evolving with the introduction of the data lake concept. While a data warehouse provides a well-structured model to manage big data, a data lake accepts data of any types and formats with or without schema and provides access to the data for diverse communities of users. A data lake provides flexible, agile, and scalable solution to manage the ever-increasing volume of big data we are witnessing in the world today, including many siloed data collected over the years by researchers through Arctic expeditions. In this paper, we present our conceptual model of a data lake for integrating the diverse huge amount of data collected by researchers during Arctic expedition. We also design a baseline metadata using a data-driven approach to manage the disparately huge structured, semi-structured, and unstructured data collected from the Arctic region. The resulting open data lake not only effectively manages big Arctic data but also supports machine learning on these big data.
Anifat M. Olawoyin, Carson K. Leung, Alfredo Cuzzocrea
IEEE BigData3
2021 Privacy-Preserving Publishing and Visualization of Spatial-Temporal Information
abstract
Partially due to technological advancements as well as the availability of affordable global positioning system (GPS) and cellular devices, more spatio-temporal data can be generated and collected. The presence of spatial and temporal dimensions uniquely differentiate spatio-temporal data from classical data as spatio-temporal data points are structurally related in the context of space and time. In this paper, we present a solution for privacy-preserving publishing and visualization of spatiotemporal big data information. Specifically, it consists of a spatiotemporal hierarchy model (STHM) for some common big data management tasks such as visualization. Our data visualizer provides actionable insight to enhance data-driven decision making. It also enables the discovery of hidden patterns, clusters of events, and outliers. We design two different metrics to preprocess the spatio-temporal for data visualization. Although we demonstrate the usefulness of our solution in privacy-preserving publishing and visualization of spatio-temporal information by using big real-life parking data from two cities, our solution can be applicable for publishing and visualizing spatio-temporal information from many other big data.
Anifat M. Olawoyin, Carson K. Leung, Alfredo Cuzzocrea
IEEE BigData3
2021 Android Malware Identification and Polymorphic Evolution Via Graph Representation Learning
abstract
Developing techniques to identify malware is critical. The polymorphic nature of malware makes it difficult to detect, especially if the detection is done with Hash-based based techniques. Image-based binary representations have been shown to be more robust to popular polymorphic obfuscation techniques. In contrast to image-based techniques, in this paper, we employed a graph-based technique that extracts control flow graphs from Android APK binary. To process the resulting graph, we use a procedure combining a new graph representation learning method, called Inferential SIR-GN for Graph representation, that preserves graph structural similarities, with XGboost, which is a standard machine learning model. Then, we apply this procedure to MALNET, which is a publicly available cybersecurity database that provides image and graph-based Android APK binary representations for a total 1,262,024 million Android APK binary with 47 types and 696 families. Experimental results show that this graph-based procedure is even more accurate than the image-based approach. Moreover, this paper provides a procedure that, by leveraging Inferential SIR-GN is able to create malware polymorphic evolution representations to use during the train of the XGboost that strengthens the malware classification tasks when the train and test datasets are split temporally according to the binary creation date. This means that our procedure can predict malware polymorphic evolution.
Miguel Quebrado, Edoardo Serra, Alfredo Cuzzocrea
IEEE BigData3
2021 Evaluating Attribution Methods in Machine Learning Interpretability
abstract
Interpretability is a key feature to broaden a conscious adoption of machine learning models in domains involving safety, security, and fairness. To achieve the interpretability of complex machine learning models, one approach consists in explaining the outcome of machine learning models through input features attribution. Attribution consists in scoring the features of an input instance by establishing how important is each feature value in a fixed instance to obtain a specific classification outcome from the machine learning model. In literature, several attribution methods are defined for specific machine learning models (e.g., neural networks) or more general ones that are model agnostic (i.e., can interpret any machine learning models). Attribution is particularly appreciated for its easy understanding of the interpretation, which is the attribution. In domains involving safety, security, and fairness, properties of the explanation such as precision and generality are crucial to establish human trust in machine learning interpretability and then on the machine learning model itself. However, even if precision and generality are clearly defined in rule-based interpretation models, they are not defined or measure on attribution models. In this work, we propose a general methodology to estimate the degree of precision and generality in attribution methods. In addition, we propose a way to measured consistency in attribution between two attribution methods. Our experiments focus on the two most popular model agnostic attribution methods, SHAP and LIME, and we evaluate them to two real applications in the field of attack detection. Our proposed methodology shows in these experiments that both SHAP and LIME lack precision, generality, and consistency and that still more investigation in the attribution research field is required.
Qudrat E. Alahy Ratul, Edoardo Serra, Alfredo Cuzzocrea
IEEE BigData3
2021 Ride-Hailing for Autonomous Vehicles: Hyperledger Fabric-Based Secure and Decentralize Blockchain Platform
abstract
Ride-hailing and ride-sharing applications have recently gained popularity as a convenient alternative to traditional modes of travel. Current research into autonomous vehicles is accelerating rapidly and will soon become a critical component of a ride-hailing platform’s architecture. Implementing an autonomous vehicle ride-hailing platform proves a difficult challenge due to the centralized nature of traditional ride-hailing architectures. In a traditional ride-hailing environment the drivers operate their own personal vehicles so it follows that a fleet of autonomous vehicles would be required for a centralized ride-hailing platform to succeed. Decentralization of the ride-hailing platform would remove a roadblock along the way to an autonomous vehicle ride-hailing platform by allowing owners of autonomous vehicles to add their vehicle to a community-driven fleet when not in use. Blockchain technology is an attractive choice for this decentralized architecture due to its immutability and fault tolerance. This thesis proposes a framework for developing a decentralized ride-hailing architecture that is verifiably secure. This framework is implemented on the Hyperledger Fabric blockchain platform. The evaluation of the implementation is done by applying known security models, utilizing a static analysis tool, and performing a performance analysis under heavy network load.
Ryan Shivers, Mohammad Ashiqur Rahman, Md. Jobair Hossain Faruk, Hossain Shahriar, Alfredo Cuzzocrea, Victor Clincy
IEEE BigData5
2021 Enhancing Scan Matching Algorithms via Genetic Programming for Supporting Big Moving Objects Tracking and Analysis in Emerging Environments
Alfredo Cuzzocrea, Kristijan Lenac, Enzo Mumolo
DEXA (1)1
2021 A Novel Approach for Supporting Italian Satire Detection Through Deep Learning
Gabriella Casalino, Alfredo Cuzzocrea, Giosuè Lo Bosco, Mariano Maiorana, Giovanni Pilato, Daniele Schicchi
FQAS2
2021 Scalable Query Processing and Query Engines over Cloud Databases: Models, Paradigms, Techniques, Future Challenges
abstract
Scalable query processing and scalable query engines over Cloud databases is a vibrant area of research, which has recently emerged within both the academic and industrial research community. This area has been further stirred-up by the current explosion of big data management and analytics models and techniques that, usually executed within the internal layer of public as well as private Clouds, pose severe (and new!) challenges to the annoying distributed query processing optimization problem in (distributed) database systems. Among other, taming the complexity of query execution plays a leading role, especially considering the typical Cloud environment that includes tens and tens of different-in-granularity data processing tasks (also at a different scale) over large-scale clusters. Inspired by these considerations, this paper focuses on models, paradigms, techniques and future challenges of scalable query processing and query engines over Cloud databases, by reporting on state-of-the-art results as well as emerging trends, with also criticisms on future work that we should expect from the community.
Alfredo Cuzzocrea
SSDBM1
2020 An Innovative Framework for Supporting Remote Sensing in Image Processing Systems via Deep Transfer Learning
abstract
In this work, we propose a method based on Deep-Learning and Convolutional Neural Network (CNN) ensemble fine-tuning for the task of remote sensing imagery registration and processing. Our method is based on the CNN transfer learning technique that allows the use of large-scale models that are already pre-trained on big general datasets and fine-tunes them for a particular application area. This approach can significantly decrease the needed size of the training set, for cases where such big training datasets are not available, and improve the quality of classification using a larger CNN or an ensemble of CNNs. This paper addresses the challenges encountered at each stage of the proposed pipeline. For image registration, objects of predefined type are detected, such as roads with hardcover and buildings using a CNN ensemble. Also, a CNN ensemble is used to detect undesirable structures in the image, such as clouds or rocks on the agricultural fields. Our image segmentation method can be used for image matching and fusion. To test our approach, we use an annotated dataset from the Kaggle contest "Dstl Satellite Imagery Feature Detection," UC Merced Land Use Dataset, and a custom annotated dataset of remote sensing imagery of agricultural areas.
Oxana Korzh, Ashish Sharma 0012, Mikel Joaristi, Edoardo Serra, Alfredo Cuzzocrea
IEEE BigData5
2020 Machine Learning and OLAP on Big COVID-19 Data
abstract
In the current technological era, huge amounts of big data are generated and collected from a wide variety of rich data sources. These big data can be of different levels of veracity in the sense that some of them are precise while some others are imprecise and uncertain. Embedded in these big data are useful information and valuable knowledge to be discovered. An example of these big data is healthcare and epidemiological data such as data related to patients who suffered from epidemic diseases like the coronavirus disease 2019 (COVID-19). Knowledge discovered from these epidemiological data-via data science techniques such as machine learning, data mining, and online analytical processing (OLAP)-helps researchers, epidemiologists and policy makers to get a better understanding of the disease, which may inspire them to come up ways to detect, control and combat the disease. In this paper, we present a machine learning and big data analytic tool for processing and analyzing COVID-19 epidemiological data. Specifically, the tool makes good use of taxonomy and OLAP to generalize some specific attributes into some generalized attributes for effective big data analytics. Instead of ignoring unknown or unstated values of some attributes, the tool provides users with flexibility of including or excluding these values, depending on their preference and applications. Moreover, the tool discovers frequent patterns and their related patterns, which help reveal some useful knowledge such as absolute and relative frequency of the patterns. Furthermore, the tool learns from the patterns discovered from historical data and predicts useful information such as clinical outcomes for future data. As such, the tool helps users to get a better understanding of information about the confirmed cases of COVID-19. Although this tool is designed for machine learning and analytics of big epidemiological data, it would be applicable to machine learning and analytics of big data in many other real-life applications and services.
Carson K. Leung, Yubo Chen 0003, Calvin S. H. Hoi, Siyuan Shang, Alfredo Cuzzocrea
IEEE BigData5
2020 Actionable Knowledge Extraction Framework for COVID-19
abstract
In response to the COVID-19 pandemic, the White House and a coalition of leading research groups have prepared the COVID-19 Open Research Dataset (CORD-19) containing over 51,000 scholarly articles, including over 40,000 with full text, about COVID-19, SARS-CoV-2, and related coronaviruses. Medical professional including physicians frequently seek answers to specific questions to improve guidelines and decisions. The huge resource of medical literature is important sources to generate new insights that can help medical communities to provide relevant knowledge and overall fight against the infectious disease. There are ongoing attempts to develop intelligent systems to automatically extract relevant knowledge from many unstructured documents. In this paper, we propose an efficient question answering framework based on automatically analyzing thousands of articles to generate both long text answers (sections/ paragraphs) in response to the questions that are posed by medical communities. In the process of developing the framework, we explored natural language processing techniques like query expansion, data preprocessing, and vector space models early. We show the initial results of an example query answering for the incubation period.
Mohammad Masum, Hossain Shahriar, Hisham M. Haddad, Sheikh Iqbal Ahamed, Sweta Sneha, Mohammad Ashiqur Rahman, Alfredo Cuzzocrea
IEEE BigData7
2020 Preserving Privacy of Temporal Big Data
abstract
In the current technological era, huge amounts of big data are generated and collected from a wide variety of rich data sources. Embedded in these big data are useful information and valuable knowledge to be utilized. With the popularity of initiatives of open data, more big data have been published on open data platforms and made accessible to the public. To preserve privacy while maintaining the utility of data, research on privacy-preserving publishing has focused on preserving privacy of sensitive personal data such as patient data for health related applications. However, there are many other real-life situations, in which personal data of individual citizens and their daily routines need to be preserved when publishing. In this paper, we examine the problem of preserving privacy of temporal big data. Specifically, we present a temporal hierarchy privacy preserving model (THPPM) for some common daily routines-for example, parking. The model adapts and extends temporal hierarchy to generalize temporal data related to timestamp and spatial data related to check-in location. It also makes good use of generalized temporal representative points to preserve privacy of specific temporal data points. Evaluations on two real-life datasets on parking tickets for the US city of Buffalo and the Canadian city of Toronto shows that effectiveness and practicality of our THPPM in preserving privacy of temporal big data. Although this model is demonstrated and evaluated on parking ticket data, it would be applicable to preserving privacy of temporal big data for many other real-life applications and services.
Anifat M. Olawoyin, Carson K. Leung, Alfredo Cuzzocrea
IEEE BigData3
2020 Adaptive and Efficient Streaming Time Series Forecasting with Lambda Architecture and Spark
abstract
The rise of the Internet of Things (IoT) devices and the streaming platform has tremendously increased the data in motion or streaming data. It incorporates a wide variety of data, for example, social media posts, online gamers in-game activities, mobile or web application logs, online e-commerce transactions, financial trading, or geospatial services. Accurate and efficient forecasting based on real-time data is a critical part of the operation in areas like energy & utility consumption, healthcare, industrial production, supply chain, weather forecasting, financial trading, agriculture, etc. Statistical time series forecasting methods like Autoregression (AR), Autoregressive integrated moving average (ARIMA), and Vector Autoregression (VAR), face the challenge of concept drift in the streaming data, i.e., the properties of the stream may change over time. Another challenge is the efficiency of the system to update the Machine Learning (ML) models which are based on these algorithms to tackle the concept drift. In this paper, we propose a novel framework to tackle both of these challenges. The challenge of adaptability is addressed by applying the Lambda architecture to forecast future state based on three approaches simultaneously: batch (historic) data-based prediction, streaming (real-time) data-based prediction, and hybrid prediction by combining the first two. To address the challenge of efficiency, we implement a distributed VAR algorithm on top of the Apache Spark big data platform. To evaluate our framework, we conducted experiments on streaming time series forecasting with four types of data sets of experiments: data without drift (no drift), data with gradual drift, data with abrupt drift and data with mixed drift. The experiments show the differences of our three forecasting approaches in terms of accuracy and adaptability.
Arjun Pandya, Oluwatobiloba Odunsi, Chen Liu 0007, Alfredo Cuzzocrea, Jianwu Wang 0001
IEEE BigData4
2020 Large-scale Sparse Structural Node Representation
abstract
In the BigData era, large graph datasets are becoming increasingly popular due to their capability to integrate and interconnect large sources of data in many fields, e.g., social media, biology, communication networks, etc. Graph representation learning is a flexible tool that automatically extracts features from a graph node. These features can be directly used for machine learning tasks. Graph representation learning approaches producing features preserving the structural information of the graphs are still an open problem, especially in the context of largescale graphs. In this paper, we propose a new fast and scalable structural representation learning approach called SparseStruct. Our approach uses a sparse internal representation for each node, and we formally proved its ability to preserve structural information. Thanks to a light-weight algorithm where each iteration costs only linear time in the number of the edges, SparseStruct is able to easily process large graphs. In addition, it provides improvements in comparison with state of the art in terms of prediction and classification accuracy by also providing strong robustness to noise data.
Edoardo Serra, Mikel Joaristi, Alfredo Cuzzocrea
IEEE BigData3
2020 Data science for healthcare predictive analytics
abstract
Big data are everywhere nowadays. Many businesses possess big data for their success because big data are very useful and are considered as new oil. For instance, big data are very important in predicting the trends on what will happen in the future. Many researchers have generated or gathered data to further enhance their research and to apply them to numerous real-life applications. Examples of big data include healthcare patient data. To improve the detection of illnesses and diseases, researchers have gathered healthcare patient data, examined the diagnosis on healthcare patient data (e.g., cells, blood count, antibodies count), and compared with previous data to determine if a specific illness or disease exist. Having an automatic predictive method for healthcare and disease analytics would be desirable. In this paper, we focus on healthcare mining, which aims to computationally discover knowledge from healthcare data. In particular, we present a data science framework with two predictive analytic algorithms for accurate prediction on the trends of cancer cases. The algorithms predict cancerous cells based on the information of the cell data from some data samples. Evaluation results on several real-life datasets related to the breast cancer demosntrate the effectiveness of our data science framework and predictive algorithms in healthcare data analytics.
Carson K. Leung, Daryl L. X. Fung, Saad B. Mushtaq, Owen T. Leduchowski, Robert Luc Bouchard, Alfredo Cuzzocrea, Christine Y. Zhang
IDEAS7
2020 Analysis and Comparison of Deep Learning Networks for Supporting Sentiment Mining in Text Corpora
abstract
In this paper, we tackle the problem of the irony and sarcasm detection for the Italian language to contribute to the enrichment of the sentiment analysis field. We analyze and compare five deep-learning systems. Results show the high suitability of such systems to face the problem by achieving 93% of F1-Score in the best case. Furthermore, we briefly analyze the model architectures in order to choose the best compromise between performances and complexity.
Teresa Alcamo, Alfredo Cuzzocrea, Giosuè Lo Bosco, Giovanni Pilato, Daniele Schicchi
iiWAS2
2020 Multidimensional Clustering over Big Data: Models, Issues, Analysis, Emerging Trends
abstract
Clustering is an essential task of the whole pattern recognition process, and it can serve under several roles, for instance in terms of data pre-processing tool for better (i.e., more accurate) pattern recognition analysis and mining. In this vest, a critical applicative setting is represented by applying pattern recognition tools over emerging big data. Here, clustering specially plays a challenging role within the context of this conceptual mining framework, and, under a larger vision, it can act as pre-processing task for general big data clustering problems. In this paper, we first focus on state-of-the-art solutions for big data clustering in the specific pattern recognition context, by highlighting benefits and limitations. Then, we focus the attention on the problem of effectively and efficiently clustering big data via innovative multidimensional metaphors, thus achieving the definition of so-called multidimensional clustering over big data. In this so-delineated research setting, based on the well-known challenges of big data management (e.g., volume, velocity, variety, and veracity), we provide critical review and discussion, complemented by a rich set of research directions, development perspectives and emerging trends of the investigated topics, as contextualized in the reference big-data-analytics scientific area.
Alfredo Cuzzocrea
SSDBM1
2020 Interpretable Anomaly Prediction: Predicting anomalous behavior in industry 4.0 settings via regularized logistic regression tools
Rocco Langone, Alfredo Cuzzocrea, Nikolaos Skantzos
Data Knowl. Eng.2
2020 Extending Big Data Management via Semantics: Recent Innovations
Alfredo Cuzzocrea, Timos K. Sellis
Inf. Syst.1
2019 RIBS: Risky Blind-Spots for Attack Classification Models
abstract
Nowadays, there has been an increment in the use of machine learning methods for cyber-security applications. These methods can be prone to generalization, especially in a binary attack classification setting, where the objective is to differentiate between benign vs. malicious behavior. This generalization creates risky security blind-spot weaknesses that make the system vulnerable. Current attackers are well aware of these blind-spots and as a counter-strategy, they exploit such vulnerabilities to bypass security measures and achieve their nefarious objectives. In this work, we propose a methodology to mitigate the problem, RIsky Blind-Spot (RIBS), by making the classification more robust. Our proposed approach creates a generator model that can learn the real characteristics of the data, and consequently, sample real examples targeting the blind-spots of a classifier. We validate our methodology in the context of power grids, where we show how this framework can improve the detection of unknown malicious behavior. Our approach provides an increment of 10% in terms of accuracy and detected attacks when compared to the baseline method.
Mikel Joaristi, Arthur Putnam, Alfredo Cuzzocrea, Edoardo Serra
IEEE BigData3
2019 Personalized DeepInf: Enhanced Social Influence Prediction with Deep Learning and Transfer Learning
abstract
Social influence is referred to as the phenomenon that one's opinions or behaviors be affected by others. Nowadays, the potential impact of social influence analysis (SIA) is significant. For example, SIA applications can include viral marketing, online content recommendation. Convention social influence analysis uses hand-crafted features and requires domain expert knowledge. Such an approach is not scalable and introduces a high cost. To overcome these disadvantages, deep learning based approaches was introduced. One of the most recent approaches is DeepInf, which is an end-to-end framework for predicting social influence by learning user's latent features. We extended DeefInf in the current paper by integrating teleport probability $\alpha$ from the domain of page rank into the graph convolution network (GCN) model to enhance the performance. Furthermore, we also propose an algorithm called hybrid personalized propagation of neural predictions (HPPNP), which shows an impressive performance in terms of prediction accuracy compared to existing methods. We reused the datasets from DeepInf and performed extensive experiments on Open Academic Graph, Twitter, DIGG datasets. By optimally sampling the teleport probability $\alpha$, the experimental results show that our model performs the best when compared with existing methods on different datasets. These results demonstrates the effectiveness of our enhanced personalized DeepInf-namely, HPPNP-in social influence prediction via both deep and transfer learning.
Carson K. Leung, Alfredo Cuzzocrea, Jiaxing Jason Mai, Deyu Deng, Fan Jiang 0001
IEEE BigData2
2019 Experiential Learning: Case Study-Based Portable Hands-on Regression Labware for Cyber Fraud Prediction
abstract
Machine Learning (ML) analyzes, and processes data and discover patterns. In cybersecurity, it effectively analyzes big data from existing cybersecurity attacks and develop proactive strategies to detect current and future cybersecurity attacks. Both ML and cybersecurity are important subjects in computing curriculum, but using ML for cybersecurity is not commonly explored. This paper designs and presents a case study-based portable labware experience built on Google's CoLaboratory (CoLab) for a ML cybersecurity application to provide students with hands-on labs accessing from anywhere and anytime, reducing or eliminating tedious installations and configurations. This approach allows students to focus on learning essential concepts and gaining valuable experience through hands-on problem solving skills. Our preliminary results and student evaluations are reported for a case-based hands-on regression labware in cyber fraud prediction using credit card fraud as an example.
Hossain Shahriar, Michael E. Whitman, Dan Chia-Tien Lo, Fan Wu 0013, Cassandra Thomas, Alfredo Cuzzocrea
IEEE BigData6
2019 Fast Privacy-Preserving Keyword Search on Encrypted Outsourced Data
abstract
Cloud providers offer storage as a service to the data owners to store emails and files on the cloud server. However, sensitive data should be encrypted before storing on the cloud server to avoid privacy concerns. With the encryption of documents, it is not feasible for data owners to retrieve documents based on keyword search as they can do with plain text documents. Hence, it is desirable to perform a multi-keyword search on encrypted data. To achieve this goal, we present a fast privacy-preserving model for keyword search on encrypted outsourced data in this paper. Specifically, the model first performs a keyword search on encrypted data and checks its support for dynamic operations. Based on keyword search results, it then sorts all the relevant data documents using the number of keywords matched for a given query. To evaluate its performance of our model, we applied the standard metrics like precision and recall. The results show the effectiveness of our privacy-preserving keyword search on encrypted outsourced data.
Bryan H. Wodi, Carson K. Leung, Alfredo Cuzzocrea, S. Sourav
IEEE BigData3
2019 An Innovative Online Process Mining Framework for Supporting Incremental GDPR Compliance of Business Processes
abstract
GDPR (General Data Protection Regulation) is a new regulation of the European Union that superimposes strict privacy constraints on storing, accessing and processing user data, as a way to ensure that personal user data are not violated neither disclosed without an explicit consent. As a consequence, business processes that interact with large amounts of such data may easily cause GDPR violations, due to the typical complexity of such processes. Inspired by these considerations, this paper highlights the challenges and critical aspects associated with the GDPR compliance journey when opting for naïve straight-forward solutions. We propose a business-aware GDPR compliance journey using online process mining. Using several large log files generated based on a real scenario, we show that the proposed tool is both effective and efficient. As such, it proves to be a powerful concept for usage in incremental GDPR compliance environments.
Rashid Zaman, Alfredo Cuzzocrea, Marwan Hassani
IEEE BigData2
2019 Urban Analytics of Big Transportation Data for Supporting Smart Cities
Carson K. Leung, Peter Braun 0004, Calvin S. H. Hoi, Joglas Souza, Alfredo Cuzzocrea
DaWaK5
2019 Flexible Querying and Analytics for Smart Cities and Smart Societies in the Age of Big Data: Overview of the FQAS 2019 International Conference
Alfredo Cuzzocrea, Sergio Greco
FQAS1
2019 A machine-learning framework for supporting intelligent web-phishing detection and analysis
abstract
This paper proposes a machine-learning framework for supporting intelligent web phishing detection and analysis, and provides its experimental evaluation. In particular we make use of state-of-the-art decision tree algorithms for detecting whether a Web site is able to perform phishing activities. If this is the case, the Web site is classified as a Web-phishing site. Our experimental evaluation confirms the benefits of applying machine learning methods to the well-known web-phishing detection problem.
Alfredo Cuzzocrea, Fabio Martinelli, Francesco Mercaldo
IDEAS1
2019 Big Data Management and Analytics in Intelligent Smart Environments: State-of-the-Art Analysis and Future Research Directions
abstract
This paper focuses on big data management and analytics in intelligent smart environments, with particular regards to intelligent transportation and logistics systems, and provides relevant research directions that may represent a milestone for future years.
Alfredo Cuzzocrea
iiWAS1
2019 A Theoretical Approach to Discover Mutual Friendships from Social Graph Networks
abstract
Due to popularity of social networking in the current era of big data, many social networking sites (e.g., Facebook) has been generating huge volumes of social data. In this paper, we aim to discover interesting relationships in a undirected (social) graph via a theoretical approach. Specifically, we examine the graph theory and linear algebra approaches to find mutual friendships from social networks represented in the form of big graphs or big graph databases.
Sehaj Pal Singh, Carson K. Leung, Fan Jiang 0001, Alfredo Cuzzocrea
iiWAS4
2019 Predictive monitoring of temporally-aggregated performance indicators of business processes against low-level streaming events
Alfredo Cuzzocrea, Francesco Folino, Massimo Guarascio 0001, Luigi Pontieri
Inf. Syst.1
2019 A novel GPU-aware Histogram-based algorithm for supporting moving object segmentation in big-data-based IoT application scenarios
Alfredo Cuzzocrea, Enzo Mumolo
Inf. Sci.1
2018 Set Similarity Joins with Complex Expressions on Distributed Platforms
Diego Junior do Carmo Oliveira, Felipe Ferreira Borges, Leonardo Andrade Ribeiro, Alfredo Cuzzocrea
ADBIS4
2018 A General Overview of Privacy-Preserving Big Data Management and Analytics Models, Methods and Techniques in Specific Domains: Static and Dynamic Distributed Environments
abstract
Privacy-preserving big data management and analytics is gaining the momentum within the research community, and several current research efforts aim to provide solutions to the challenges that emerge when models, techniques and algorithms must be delivered on top of massive, distributed big data repositories, especially with regards to emerging distributed settings such as Clouds and social networks. In this paper, at the convergence of the contexts of static and dynamic distributed environments, we provide a general overview of models, issues and approaches, along with some reference frameworks. Indeed, both static and dynamic distributed environments are relevant cases of settings where the privacy of big data turns to be critical. Finally, we discuss emerging research directions.
Alfredo Cuzzocrea, Carlo Mastroianni
IEEE BigData1
2018 Improving Machine Learning Tools with Embeddings: Applications to Big Data Security
abstract
Considering the widespread diffusion of machine learning techniques to solve several issues, from network security to malware detection, in this paper we propose the adoption of word embeddings aimed to improve the machine learning-based classifiers. Three real-world experiment we perform in order to demonstrate that the proposed method overcomes in terms of performances the mostly used machine-learning algorithms.
Alfredo Cuzzocrea, Fabio Martinelli, Francesco Mercaldo
IEEE BigData1
2018 Privacy-Preserving Frequent Pattern Mining from Big Uncertain Data
abstract
As we are living in the era of big data, high volumes of wide varieties of data which may be of different veracity (e.g., precise data, imprecise and uncertain data) are easily generated or collected at a high velocity in many real-life applications. Embedded in these big data is valuable knowledge and useful information, which can be discovered by big data science solutions. As a popular data science task, frequent pattern mining aims to discover implicit, previously unknown and potentially useful information and valuable knowledge in terms of sets of frequently co-occurring merchandise items and/or events. Many of the existing frequent pattern mining algorithms use a transaction-centric mining approach to find frequent patterns from precise data. However, there are situations in which an item-centric mining approach is more appropriate, and there are also situations in which data are imprecise and uncertain. Hence, in this paper, we present an item-centric algorithm for mining frequent patterns from big uncertain data. In recent years, big data have been gaining the attention from the research community as driven by relevant technological innovations (e.g., clouds) and novel paradigms (e.g., social networks). As big data are typically published online to support knowledge management and fruition processes, these big data are usually handled by multiple owners with possible secure multi-part computation issues. Thus, privacy and security of big data has become a fundamental problem in this research context. In this paper, we present, not only an item-centric algorithm for mining frequent patterns from big uncertain data, but also a privacy-preserving algorithm. In other words, we present- in this paper-a privacy-preserving item-centric algorithm for mining frequent patterns from big uncertain data. Results of our analytical and empirical evaluation show the effectiveness of our algorithm in mining frequent patterns from big uncertain data in a privacy-preserving manner.
Carson K. Leung, Calvin S. H. Hoi, Adam G. M. Pazdor, Bryan H. Wodi, Alfredo Cuzzocrea
IEEE BigData5
2018 CIKM 2018 Co-Located Workshops Summary
abstract
This paper provides an overview of the workshops co-located with the 27th ACM International Conference on Information and Knowledge Management (CIKM 2018), held during October 22-26, 2018 in Turin, Italy.
Alfredo Cuzzocrea, Francesco Bonchi, Dimitrios Gunopulos
CIKM1
2018 Learning Ranking Functions by Genetic Programming Revisited
Ricardo Baeza-Yates, Alfredo Cuzzocrea, Domenico Crea, Giovanni Lo Bianco
DEXA (2)2
2018 A Predictive Learning Framework for Monitoring Aggregated Performance Indicators over Business Process Events
abstract
In many application contexts, a business process' executions are subject to performance constraints expressed in an aggregated form, usually over predefined time windows, and detecting a likely violation to such a constraint in advance could help undertake corrective measures for preventing it. This paper illustrates a prediction-aware event processing framework that addresses the problem of estimating whether the process instances of a given (unfinished) window w will violate an aggregate performance constraint, based on the continuous learning and application of an ensemble of models, capable each of making and integrating two kinds of predictions: single-instance predictions concerning the ongoing process instances of w, and time-series predictions concerning the "future" process instances of w (i.e. those that have not started yet, but will start by the end of w). Notably, the framework can continuously update the ensemble, fully exploiting the raw event data produced by the process under monitoring, suitably lifted to an adequate level of abstraction. The framework has been validated against historical event data coming from real-life business processes, showing promising results in terms of both accuracy and efficiency.
Alfredo Cuzzocrea, Francesco Folino, Massimo Guarascio 0001, Luigi Pontieri
IDEAS1
2018 The Inverse Tree-OLAP Problem: Definitions, Models, Complexity Analysis, and a Possible Solution
abstract
Count constraint is a data dependency that requires the results of given count operations on a relation to be within a certain range. By means of count constraints a new decisional problem, called the Inverse OLAP, has been recently introduced: given a flat fact table, does there exist an instance satisfying a set of given count constraints? This paper focuses on a special case of Inverse OLAP, called Inverse Tree-OLAP, for which the flat fact table key is modeled by a Dimensional Fact Model (DFM) with a tree structure. The count constraints define aggregation patterns to be respected by both the many-to-many relationship among the basic dimensions and the one-to-many relationships within dimension hierarchies. A count constraint is required to have a particular structure so that the problem of handling fact table projections with duplicates is avoided. The simplified structure enables the invention of an effective method for its solution that consists of three main steps: (1) using some of the count constraints to extract a subproblem that is formulated as a known data mining problem (inverse frequent itemset mining), (2) solving the subproblem using a recent method that has been shown to be effective in practical situations also for large size instances and (3) enforcing the remaining count constraints on the solution returned by step 2 using a system of linear equations. The overall proposed approach can be effectively used to generate OLAP cubes for benchmarking that reflect patterns of real datasets.
Domenico Saccà, Edoardo Serra, Alfredo Cuzzocrea
IDEAS3
2018 Applying Machine Learning Techniques to Detect and Analyze Web Phishing Attacks
abstract
Phishing is a technique aimed to imitate an official websites of any company such as banks, institutes, etc. The purpose of phishing is to theft private and sensitive credentials of users such as password, username or PIN. Phishing detection is a technique to deal with this kind of malicious activity. In this paper we propose a method able to discriminate between web pages aimed to perform phishing attacks and legitimate ones. We exploit state of the art machine learning algorithms in order to build models using indicators that are able to detect phishing activities.
Alfredo Cuzzocrea, Fabio Martinelli, Francesco Mercaldo
iiWAS1
2017 Fighting fake news spread in online social networks: Actual trends and future research directions
abstract
In this paper we present how fake news spread in the current online social networks. We discuss how existing social network technologies such as influence maximization, information diffusion, and epidemiological models contributes to fake news creation and spreading. Solutions to reducing the creation and spreading of fake news are also reviewed. We make recommendations regarding future areas of research in this field.
Alina Campan, Alfredo Cuzzocrea, Traian Marius Truta
IEEE BigData2
2017 Tor traffic analysis and detection via machine learning techniques
abstract
Tor is an anonymous Internet communication system based on the second generation of onion routing network protocol. Using Tor is really difficult to trace the users Internet activity: this is the reason why the usage of Tor is intended in order to protect the privacy of users, their freedom and the ability to conduct confidential communications without being monitored. Tor is even more used by cyber-criminals in order to cover their illegal activities: the Tor community has observed, for instance an alarming increase in the number of malware that abuse of the popular anonymizing network to hide their command and control infrastructures. In this paper we present a technique able to identify whether an host is generating Tor-related traffic. We resort to well-known machine learning algorithms in order to evaluate the effectiveness of the proposed feature set in a real world environment. In addition we demonstrate that the proposed method is able to recognize the kind of activity (e.g., email or P2P applications) the user under analysis is doing on the Tor network.
Alfredo Cuzzocrea, Fabio Martinelli, Francesco Mercaldo, Gianni Viardo Vercelli
IEEE BigData1
2017 Data masking techniques for NoSQL database security: A systematic review
abstract
This paper first presents an in-depth study of potential security vulnerabilities in MongoDB and Cassandra, two popular NoSQL databases. We provide examples of attacks. We then explore some popular data masking techniques as ways of mitigating security threats in these databases.
Alfredo Cuzzocrea, Hossain Shahriar
IEEE BigData1
2017 MapReduce-Based Complex Big Data Analytics over Uncertain and Imprecise Social Networks
Peter Braun 0004, Alfredo Cuzzocrea, Fan Jiang 0001, Carson K. Leung, Adam G. M. Pazdor
DaWaK2
2017 Big Data Management: New Frontiers, New Paradigms
Alfredo Cuzzocrea, Alkis Simitsis, Il-Yeol Song
Inf. Syst.1
2016 Private databases on the cloud: Models, issues and research perspectives
abstract
Privacy and security of big data is emerging as one among the most relevant research challenges of recent years, also stirred-up by a wide family of critical applications ranging from scientific computing to social network analysis and mining, from data stream management to smart cities, and so forth. Traditionally, the issue of making (even very-large) databases private and secure has a long history in the context of encrypted databases but, when specifically considered in the Cloud setting, it poses new requirements and challenges to deal with, with particular regard to the scalability of solutions. In line with this emerging research trend, this paper focuses the attention on state-of-the-art proposals in the area of private databases over Clouds, and proposes critical comments about pros and cons of actual research efforts along with future research directions to be considered in future years.
Alfredo Cuzzocrea, Carlo Mastroianni, Giorgio Mario Grasso
IEEE BigData1
2016 Incorporating Clustering into Set Similarity Join Algorithms: The SjClust Framework
Leonardo Andrade Ribeiro, Alfredo Cuzzocrea, Karen Aline Alves Bezerra, Ben Hur Bahia do Nascimento
DEXA (1)2
2016 An Innovative Framework for Effectively and Efficiently Supporting Big Data Analytics over Geo-Located Mobile Social Media
abstract
Mobile Social Media are gaining momentum in the broader context of Big Data Analytics, where the main issue is represented by the problem of extracting interesting and actionable knowledge from big data repositories. Mobile social media sources like Twitter and Instagram are indeed producing massive amounts of data (namely, posts) that represent a very rich source of knowledge for predictive analytics. In line with this emerging trend, this paper proposes an innovative approach for effectively and efficiently supporting big data analytics over geo-localized mobile social media, with particular emphasis with the context of modern tourist information systems. In this context, the innovative FollowMe suite, which implements the proposed methodology, is also described in details. We complement our analytical contribution with a real-life case study focusing on the EXPO 2015 event in Milan, Italy which clearly shows benefits and potentialities of our proposed big data analytics framework.
Alfredo Cuzzocrea, Giuseppe Psaila, Maurizio Toccu
IDEAS1
2016 Computing Theoretically-Sound Upper Bounds to Expected Support for Frequent Pattern Mining Problems over Uncertain Big Data
Alfredo Cuzzocrea, Carson K. Leung
IPMU (2)1
2016 Advanced Query Answering Techniques over Big Mobile Data
abstract
Mobile environments are classical settings where big data arise, mostly raised up by emerging (mobile) environmentssuch as social networks, sensor networks, IoT infrastructures, and so forth. This phenomenon introduces a novel class of big data, the so-called big mobile data. Big mobile data demand for novel models, techniques and algorithms devoted to the annoying problem of effectively and efficiently querying large-scale, enormous, highly-heterogeneous amounts of data, which is now living a renewed season precisely due to the advent of the big data era. Indeed, classical approaches developed during decades of database research activities demand for novel adaptations and optimizations explicitly tailored to deal with the (many) V-requirements of big data management in mobile environments. In line with this emerging research trends, this panel will focus the attention on state-of-the-art proposals in the area of advanced query answering techniques over big mobile data, and will propose critical comments about pros and cons of actual research efforts along with future research directions to be considered in future years.
Alfredo Cuzzocrea
MDM1
2016 A Robust and Versatile Multi-View Learning Framework for the Detection of Deviant Business Process Instances
abstract
Increasing attention has been paid to the detection and analysis of “deviant” instances of a business process that are connected with some kind of “hidden” undesired behavior (e.g. frauds and faults). In particular, several recent works faced the problem of inducing a binary classification model (here named deviance detection model ) that can discriminate between deviant traces and normal ones, based on a set of historical log traces (labeled as either deviant or normal). Current solutions rely on applying standard classifier-induction methods to a feature-based representation of the given traces, where the features include sequence-based patterns extracted from the corresponding sequences of activities. However, there is no consensus on which kinds of patterns are the most suitable for such a task. On the other hand, mixing multiple pattern families together may produce a heterogenous, redundant and sparse representation of the traces that likely leads to poor deviance detection models. In this paper, we propose an ensemble-learning method for solving this problem, where multiple base classifiers are trained on different feature-based views of the log (each obtained by mapping the traces onto a distinguished collection of patterns). A stacking procedure is used to combine the discovered base models into an overall probabilistic model that associates any new trace with an estimate of the probability that it reflects a deviant process instance. This helps the analyst prioritize the inspection of the cases that are more likely to be deviant. The method also takes advantage of all nonstructural data available in the log, and employs a resampling mechanism to deal with the rarity of deviances in the training log. It has been conceived as the core of a comprehensive framework for detecting and analyzing business process deviances. The framework supports the analyst to investigate suspect deviances, and provides some feedback to the learning method for improving the accuracy of the discovered deviance detection models. Tests on several real-life datasets proved the validity of the approach, as concerns its capability to discover an accurate deviance detection model, and to effectively exploit new (originally unlabeled) traces via active learning and self-training mechanisms.
Alfredo Cuzzocrea, Francesco Folino, Massimo Guarascio 0001, Luigi Pontieri
Int. J. Cooperative Inf. Syst.1
2015 A distributed framework for supporting adaptive ensemble-based intrusion detection
abstract
This paper proposes anatomy and main functionalities of a distributed framework for supporting adaptive ensemble-based intrusion detection. We start from open issues and limitations of actual state-of-the-art proposals, and we derive a suitable architecture that, based on actual, emerging research trends, finally defines an innovative ensemble-based network intrusion detection system that combines following requirements: distribution, cooperativeness, scalability, multi-scale network traffic analysis, feature selection and extraction. These requirements are recognized by our study as first-class research challenges for next-generation intrusion detection systems.
Alfredo Cuzzocrea, Gianluigi Folino, Pietro Sabatino
IEEE BigData1
2015 Heterogeneous k-anonymization with high utility
abstract
Among the privacy-preserving approaches that are known in the literature, h-anonymity remains the basis of more advanced models while still being useful as a stand-alone solution. Applying h-anonymity in practice, though, incurs severe loss of data utility, thus limiting its effectiveness and reliability in real-life applications and systems. However, such loss in utility does not necessarily arise from an inherent drawback of the model itself, but rather from the deficiencies of the algorithms used to implement the model. Conventional approaches rely on a methodology that publishes data in homogeneous generalized groups. An alternative modern data publishing scheme focuses on publishing the data in heterogeneous groups and achieves higher utility, while ensuring the same privacy guarantees. As conventional approaches cannot anonymize data following this heterogeneous scheme, innovative solutions are required for this purpose. Following this approach, in this paper we provide a set of algorithms that ensure high-utility h-anonymity, via solving an equivalent graph processing problem.
Katerina Doka, Mingqiang Xue, Dimitrios Tsoumakos, Panagiotis Karras, Alfredo Cuzzocrea, Nectarios Koziris
IEEE BigData5
2015 Distributed Classification of Data Streams: An Adaptive Technique
Alfredo Cuzzocrea, Mohamed Medhat Gaber, Ary Mazharuddin Shiddiqi
DaWaK1
2015 Towards OLAP Analysis of Multidimensional Tweet Streams
abstract
Social media and networks are used by millions of people to share with their friends across the world: tastes, opinions, ideas, etc. The volume and the speed at which these data are produced make it a challenging task to discover meaningful patterns in the data. Nevertheless, very interesting business goals could be achieved collecting these data and performing analytics on social media data streams, such as: addressing marketing strategies, targeting advertisements, and so forth. We emphasize that there is a need to investigate and define suitable knowledge mining approaches to go beyond explicitly available metadata by analyzing unstructured data to provide intelligent analytics services. Specifically, in this paper we provide first results on applying OLAP analysis to multidimensional Tweet streams.
Alfredo Cuzzocrea, Carmen De Maio, Giuseppe Fenza, Vincenzo Loia, Mimmo Parente
DOLAP1
2015 OLAP-enabled web search of complex objects
abstract
Inspired by the actual trend of empowering traditional Web search methodologies by means of novel computational paradigms, in this paper we propose and experimentally assess WebClustCube, a novel system that allows OLAP-enabled Web search of complex objects, thus adding new value to the potentialities of current Web search paradigms. In particular, WebClustCube supports the building and the interactive manipulation of OLAP-enabled Web views over complex objects extracted from distributed databases. The data management, OLAP-like support of WebClustCube is provided by ClustCube, a state-of-the-art framework for coupling OLAP methodologies and clustering algorithms with the goal of analyzing and mining of complex database objects. A case study that clearly shows the potentialities of WebClustCube in the context of next-generation Web search environments is provided. We complement of analytical contribution by means of an experimental assessment and analysis of WebClustCube according to several metric perspectives.
Alfredo Cuzzocrea, Guandong Xu, Giorgio Mario Grasso
iiWAS1
2015 Knowledge Discovery from Geo-Located Tweets for Supporting Advanced Big Data Analytics: A Real-Life Experience
Alfredo Cuzzocrea, Giuseppe Psaila, Maurizio Toccu
MEDI1
2015 Computing and Mining ClustCube Cubes Efficiently
Alfredo Cuzzocrea
PAKDD (2)1
2015 Aggregation and multidimensional analysis of big data for large-scale scientific applications: models, issues, analytics, and beyond
abstract
Aggregation and multidimensional analysis are well-known powerful tools for extracting useful knowledge, shaped in a summarized manner, which are being successfully applied to the annoying problem of managing and mining big data produced by large-scale scientific applications. Indeed, in the context of big data analytics, aggregation approaches allow us to provide meaningful descriptions of these data, otherwise impossible for alternative data-intensive analysis tools. On the other hand, multidimensional analysis methodologies introduce fortunate metaphors that significantly empathize the knowledge discovery phase from such huge amounts of data. Following this main trend, several big data aggregation and multidimensional analysis approaches have been proposed recently. The goal of this paper is to (i) provide a comprehensive overview of state-of-the-art techniques and (ii) depict open research challenges and future directions adhering to the reference scientific field.
Alfredo Cuzzocrea
SSDBM1
2015 Advances in data warehousing and OLAP in the big Data Era
Ladjel Bellatreche, Alfredo Cuzzocrea, Il-Yeol Song
Inf. Syst.2
2015 Effectively and efficiently supporting roll-up and drill-down OLAP operations over continuous dimensions via hierarchical clustering
Michelangelo Ceci, Alfredo Cuzzocrea, Donato Malerba
J. Intell. Inf. Syst.2
2015 Design, implementation and validation of AI-inspired information systems
abstract
While there is an emerging and always-growing interest for novel paradigms appeared recently (e.g., social networks, Cloud computing, NoSQL databases, Big Data, and so forth), Artificial Intelligence (AI) always plays a critical role in next-generation Information Systems.Indeed, as technology and paradigms pervade our life, there is a challenging need for smarter and more sophisticated Information Systems, for instance using innovative methodologies like crowdsourcing.As a consequence, it is natural to foresee the advancement of a novel class of Information Systems, which we call as AI-Inspired Information Systems.Basically, these are Information Systems which incorporate in their critical layers (i.e., design, implementation, validation) AI methodologies, yet extending their roots to classical foundations, with, indeed, exciting innovations.From this main evidence, there emerges a great interest in methodologies that can improve the variegate aspects of Information Systems life-cycle, with special emphasis on the modeling, representation, querying, retrieval, reasoning, mining and end-user phases.On the other hand, AI-inspired approaches tend to be computationally expensive, so that complexity analysis of proposed solutions must be rigorously evaluated and reasonable execution bounds should be derived.As a consequence, it naturally follows that designing, implementing and validating AI-Inspired Information Systems not only conveys in the need for elegant and scalable approaches, but also these approaches need to be theoretically-sound and exposing bounded complexities. Along this line, this special issue on "Design, Implementation and Validation of AI-Inspired Information Systems" of Journal of Intelligent Information Systems focuses on latest research resultsand open research challenges on the problem of designing, implementing and validating Information Systems, based on AI paradigms, according to principles and guidelines provided above.With the aim of adequately fulfilling both theoretical and practical issues of the investigated research context, this special issue contains five papers, which have gone through two rigorous review rounds before being accepted for the final inclu-
Alfredo Cuzzocrea
J. Intell. Inf. Syst.1
2014 Efficient Frequent Itemset Mining from Dense Data Streams
Alfredo Cuzzocrea, Fan Jiang 0001, Wookey Lee, Carson K. Leung
APWeb1
2014 PSBD 2014: Overview of the 1st International Workshop on Privacy and Security of Big Data
abstract
The ACM 1st International Workshop on Privacy and Security of Big Data (PSBD 2014), held in Shanghai, China on November 7, 2014, in conjunction with the ACM 23rd International Conference on Information and Knowledge Management (CIKM 2014), presents research on privacy and security of big data, an emerging challenge in actual database and data mining research. PSBD 2014 program has two interesting s/essions on (i) scalable privacy-preserving and security-control methods for big data processing, and (ii) user-oriented and data-oriented privacy methods for big data processing, plus a panel discussing current challenges and future research perspectives of privacy and security of big data.
Alfredo Cuzzocrea
CIKM1
2014 A Pattern-Oriented Approach for Supporting ETL Conceptual Modelling and Its YAWL-based Implementation
Bruno Oliveira 0001, Orlando Belo, Alfredo Cuzzocrea
DATA3
2014 Real-Time Data Warehousing: A Rewrite/Merge Approach
Alfredo Cuzzocrea, Nickerson Ferreira, Pedro Furtado 0001
DaWaK1
2014 Big Graph Analytics: The State of the Art and Future Research Agenda
abstract
Analytics over big graphs is becoming a first-class challenge in database research, with fast-growing interest from both the academia and the industrial community. This problem arises in several application scenarios, ranging from social networks to large-scale network systems, from knowledge discovery to cybersecurity, and so forth. Following this major trend, this paper explores actual state-of-the-art results in the area of analytics over big graphs and discusses open research issues and actual trends in such area.
Alfredo Cuzzocrea, Il-Yeol Song
DOLAP1
2014 A confidence-based entity resolution approach with incomplete information
abstract
Entity resolution identifies entities from different data sources that refer to the same real-world entity and it is an important prerequisite for integrating data from multiple sources. Entity resolution mainly relies on similarity measures on data records. Unfortunately, the data quality of data sources is not so good in practice. Especially web data sources often only provide incomplete information, which leads to the difficulties of direct applying similarity measures to identify the same entities. In order to address this problem, the concept of confidence is introduced to measure the trustworthy of the similarity calculation. An adaptive rule-based approach is used to calculate the similarity between records and its confidence is also derived. Then the similarity and confidence are propagated on the entity relational graph until fix point is reached. Finally, any pair of two records can be determined as matched or unmatched based on a threshold. We performed a series of experiments on real data sets and experiment results show that our approach has a better performance comparing with others.
Jian Cao 0001, Guandong Xu, Alfredo Cuzzocrea
DSAA5
2014 Adaptive data stream mining for wireless sensor networks
abstract
Data stream mining in wireless sensor networks has many important applications. Realizing these applications is faced by resource constraints of the sensor nodes that form the network. Adaptation to availability of resources is crucial to the success of these applications. In this paper, we propose a distributed data stream classification technique that has been tested on a real sensor network platform, namely, Sun SPOT. Experimental results evidenced the applicability of our technique to operate in such an environment of scarce resources.
Alfredo Cuzzocrea, Mohamed Medhat Gaber, Ary Mazharuddin Shiddiqi
IDEAS1
2014 Big Data Mining or Turning Data Mining into Predictive Analytics from Large-Scale 3Vs Data: The Future Challenge for Knowledge Discovery
Alfredo Cuzzocrea
MEDI1
2014 Optimization issues of querying and evolving sensor and stream databases
Alfredo Cuzzocrea
Inf. Syst.1
2014 Processing and mining complex data streams
Jerzy Stefanowski, Alfredo Cuzzocrea, Dominik Slezak
Inf. Sci.2
2013 Designing Parallel Relational Data Warehouses: A Global, Comprehensive Approach
Soumia Benkrid, Ladjel Bellatreche, Alfredo Cuzzocrea
ADBIS (2)3
2013 Energy Efficiency in W-Grid Data-Centric Sensor Networks via Workload Balancing
Alfredo Cuzzocrea, Gianluca Moro, Claudio Sartori 0001
APWeb1
2013 DOLAP 2013 workshop summary
abstract
The ACM DOLAP workshop presents research on data warehousing and On-Line Analytical Processing (OLAP). The DOLAP 2013 program has three interesting sessions on Design and Exploitation of Social Data Warehouses, ETL and modeling and new trends, as well as a keynote talk on OLAP query processing and a panel on OLAP and DataWarehousing Technology in Big Data era.
Ladjel Bellatreche, Alfredo Cuzzocrea, Il-Yeol Song
CIKM2
2013 Approximate OLAP Query Processing over Uncertain and Imprecise Multidimensional Data Streams
Alfredo Cuzzocrea
DEXA (2)1
2013 Data warehousing and OLAP over big data: current challenges and future research directions
abstract
In this paper, we highlight open problems and actual research trends in the field of Data Warehousing and OLAP over Big Data, an emerging term in Data Warehousing and OLAP research. We also derive several novel research directions arising in this field, and put emphasis on possible contributions to be achieved by future research efforts.
Alfredo Cuzzocrea, Ladjel Bellatreche, Il-Yeol Song
DOLAP1
2013 DynamicNet: an effective and efficient algorithm for supporting community evolution detection in time-evolving information networks
abstract
DynamicNet, an effective and efficient algorithm for supporting community evolution detection in time-evolving information networks is presented and experimentally evaluated in this paper. DynamicNet introduces a graph-based model-theoretic approach to represent time-evolving information networks, and to capture how they change over time. A central feature of DynamicNet is represented by the ability of supporting matching-based community evolution detection, by identifying several classes of community transitions. Experimental results clearly demonstrate the reliability and the efficiency of our proposal.
Alfredo Cuzzocrea, Francesco Folino, Clara Pizzuti
IDEAS1
2013 Big data: a research agenda
abstract
Recently, a great deal of interest for Big Data has risen, mainly driven from a widespread number of research problems strongly related to real-life applications and systems, such as representing, modeling, processing, querying and mining massive, distributed, large-scale repositories (mostly being of unstructured nature). Inspired by this main trend, in this paper we discuss three important aspects of Big Data research, namely OLAP over Big Data, Big Data Posting, and Privacy of Big Data. We also depict future research directions, hence implicitly defining a research agenda aiming at leading future challenges in this research field.
Alfredo Cuzzocrea, Domenico Saccà, Jeffrey D. Ullman
IDEAS1
2013 OLAP*: Effectively and Efficiently Supporting Parallel OLAP over Big Data
Alfredo Cuzzocrea, Rim Moussa, Guandong Xu
MEDI1
2013 Mining Frequent Itemsets from Sparse Data Streams in Limited Memory Environments
Juan J. Cameron, Alfredo Cuzzocrea, Fan Jiang 0001, Carson K. Leung
WAIM2
2013 Community Detection in Multi-relational Social Networks
Zhiang Wu 0001, Wenpeng Yin 0001, Jie Cao 0001, Guandong Xu, Alfredo Cuzzocrea
WISE (2)5
2013 Advances in Managing, Updating and Querying Uncertain and Imprecise Sensor and Stream Databases
Alfredo Cuzzocrea
Inf. Syst.1
2012 Enforcing Interaction and Cooperation in Content-Based Web3.0 Applications
Antonio Bevacqua, Marco Carnuccio, Alfredo Cuzzocrea, Riccardo Ortale, Ettore Ritacco
APWeb3
2012 Enhancing Coverage and Expressive Power of Spatial Data Warehousing Modeling: The SDWM Approach
Alfredo Cuzzocrea, Robson do Nascimento Fidalgo
DaWaK1
2012 Enhanced clustering of complex database objects in the clustcube framework
abstract
This paper significantly extends our previous research contribution [1], where we introduced the OLAP-based ClustCube framework for clustering and mining complex database objects extracted from distributed database settings. In particular, in this research we provide the following two novel contributions over [1]. First, we provide an innovative tree-based distance function over complex objects that takes into account the typical tree-like nature of these objects in distributed database settings. This novel distance is a relevant contribution over the simpler low-level-field-based distance presented in [1]. Second, we provide a comprehensive experimental campaign of ClustCube algorithms for computing ClustCube cubes, according to both performance metrics and accuracy metrics, against a well-known benchmark data set, and in comparison with a state-of-the-art subspace clustering algorithm for high-dimensional data. Retrieved results clearly demonstrate the superiority of our approach.
Alfredo Cuzzocrea, Paolo Serafino
DOLAP1
2012 Polynomial Asymptotic Complexity of Multiple-Objective OLAP Data Cube Compression
Alfredo Cuzzocrea, Marco Fisichella
IPMU (2)1
2012 Intelligent knowledge-based models and methodologies for complex information systems
Alfredo Cuzzocrea
Inf. Sci.1
2012 Effectively and Efficiently Designing and Querying Parallel Relational Data Warehouses on Heterogeneous Database Clusters: The F&A Approach
abstract
In this paper, a comprehensive methodology for designing and querying Parallel Rational Data Warehouses (PRDW) over database clusters, called Fragmentation & Allocation (F&A) is proposed. F&A assumes that cluster nodes are heterogeneous in processing power and storage capacity, contrary to traditional design approaches that assume that cluster nodes are instead homogeneous, and fragmentation and allocation phases are performed in a simultaneous manner. In classical approaches, two different cost models are used to perform fragmentation and allocation, separately, whereas F&A makes use of one cost model that considers fragmentation and allocation parameters simultaneously. Therefore, according to the F&A methodology proposed, the allocation phase/decision is done at fragmentation. At the fragmentation phase, F&A uses two well-known algorithms, namely Hill Climbing (HC) and Genetic Algorithm (GA), which the authors adapt to the main PRDW design problem over heterogeneous database clusters, as these algorithms are capable of taking into account the heterogeneous characteristics of the reference application scenario. At the allocation phase, F&A introduces an innovative matrix-based formalism capable of capturing the interactions among fragments, input queries, and cluster node characteristics, driving the data allocation task accordingly, and a related affinity-based algorithm, called F&A-ALLOC. Finally, their proposal is experimentally assessed and validated against the widely-known data warehouse benchmark APB-1 release II.
Ladjel Bellatreche, Alfredo Cuzzocrea, Soumia Benkrid
J. Database Manag.2
2011 A Constraint-Based Framework for Computing Privacy Preserving OLAP Aggregations on Data Cubes
Alfredo Cuzzocrea, Domenico Saccà
ADBIS (2)1
2011 DOLAP 2011: overview of the 14th international workshop on data warehousing and olap
abstract
The ACM 14th International Workshop on Data Warehousing and OLAP (DOLAP 2011), held in Glasgow, Scotland, UK on October 28, 2011, in conjunction with the ACM 20th International Conference on Information and Knowledge Management (CIKM 2011), presents research on data warehousing and On-Line Analytical Processing (OLAP). The DOLAP 2011 program has three interesting sessions on data warehouse modeling and maintenance, ETL and performance, and OLAP visualization and extensions, and a panel discussing analytics in data warehouses.
Alfredo Cuzzocrea, Karen C. Davis, Il-Yeol Song
CIKM1
2011 Analytics over large-scale multidimensional data: the big data revolution!
abstract
In this paper, we provide an overview of state-of-the-art research issues and achievements in the field of analytics over big data, and we extend the discussion to analytics over big multidimensional data as well, by highlighting open problems and actual research trends. Our analytical contribution is finally completed by several novel research directions arising in this field, which plays a leading role in next-generation Data Warehousing and OLAP research.
Alfredo Cuzzocrea, Il-Yeol Song, Karen C. Davis
DOLAP1
2011 A family of graph-theory-driven algorithms for managing complex probabilistic graph data efficiently
abstract
Traditionally, a great deal of attention has been devoted to the problem of effectively modeling and querying probabilistic graph data. State-of-the-art proposals are not prone to deal with complex probabilistic data, as they essentially introduce simple data models (e.g., based on confidence intervals) and straightforward query methodologies (e.g., based on the reachability property). According to our vision, these proposals need to be extended towards achieving the definition of innovative models and algorithms capable of dealing with the hardness of novel requirements posed by managing complex probabilistic graph data efficiently. Inspired by this main motivation, in this paper we propose and experimentally assess an innovative family of graph-theory-driven algorithms for managing complex probabilistic graph data, whose main double-fold goal consists in enhancing the expressive power of the underlying probabilistic graph data model and the expressive power of graph queries.
Alfredo Cuzzocrea, Paolo Serafino
IDEAS1
2011 Hand-OLAP: Semantics-Aware Compression of Data Cubes for Effective and Efficient OLAP in Mobile Enviroments
abstract
In this paper, we present a complete demonstration of Hand-OLAP, a Java-based distributed system that relies on intelligent data cube compression techniques for effectively and efficiently supporting OLAP in mobile environments. Hand-OLAP is based on an innovative systematic technique according to which first a two-dimensional OLAP view of interest is extracted from the target multidimensional data cube via the so-called OLAP dimension flattening process, and then this view is compressed by means of a meaningful semantics-based data cube compression approach. The compressed two-dimensional view is finally delivered to mobile devices, and used to support interactive OLAP exploration and querying tasks in an off-line manner.
Alfredo Cuzzocrea, Domenico Saccà
Mobile Data Management (1)1
2011 Detecting Health Events on the Social Web to Enable Epidemic Intelligence
Marco Fisichella, Avare Stewart, Alfredo Cuzzocrea, Kerstin Denecke
SPIRE3
2011 Retrieving Accurate Estimates to OLAP Queries over Uncertain and Imprecise Multidimensional Data Streams
Alfredo Cuzzocrea
SSDBM1
2011 Pushing artificial intelligence in database and data warehouse systems
Alfredo Cuzzocrea
Data Knowl. Eng.1
2011 Enhancing accuracy and expressive power of range query answers over incomplete spatial databases via a novel reasoning approach
Alfredo Cuzzocrea, Andrea Nucita
Data Knowl. Eng.1
2011 Data warehousing and knowledge discovery from sensors and streams
Alfredo Cuzzocrea
Knowl. Inf. Syst.1
2010 Efficiently Computing and Querying Multidimensional OLAP Data Cubes over Probabilistic Relational Data
Alfredo Cuzzocrea, Dimitrios Gunopulos
ADBIS1
2010 F&A: A Methodology for Effectively and Efficiently Designing Parallel Relational Data Warehouses on Heterogenous Database Clusters
Ladjel Bellatreche, Alfredo Cuzzocrea, Soumia Benkrid
DaWak2
2010 Balancing accuracy and privacy of OLAP aggregations on data cubes
abstract
In this paper we propose an innovative framework based on flexible sampling-based data cube compression techniques for computing privacy preserving OLAP aggregations on data cubes while allowing approximate answers to be efficiently evaluated over such aggregations. In our proposal, this scenario is accomplished by means of the so-called accuracy/privacy contract, which determines how OLAP aggregations must be accessed throughout balancing accuracy of approximate answers and privacy of sensitive ranges of multidimensional data.
Alfredo Cuzzocrea, Domenico Saccà
DOLAP1
2010 Effectively and efficiently selecting access control rules on materialized views over relational databases
abstract
A novel framework for effectively and efficiently selecting fine-grained access control rules from a target relational database to the set of materialized views defined on such a database is presented and experimentally assessed in this paper, along with the main algorithm implementing the focal selection task, called VSP-Bucket. The proposed security framework introduces a number of research innovations, ranging from a novel Datalog-based syntax, and related semantics, aimed at modeling and expressing access control rules over relational databases to algorithm VSP-Bucket itself, which is a meaningful adaptation of a well-know view-based query re-writing algorithm for query optimization purposes. Our framework exposes a high flexibility, due to the fact it allows several classes of access control rules to be expressed and handled on top of large relational databases, and, at the same, it introduces high effectiveness and efficiency, as demonstrated by our comprehensive experimental evaluation and analysis of performance and scalability of algorithm VSP-Bucket.
Alfredo Cuzzocrea, Mohand-Said Hacid, Nicola Grillo
IDEAS1
2010 Advanced knowledge-based systems
Alfredo Cuzzocrea
Data Knowl. Eng.1
2010 Event-based lossy compression for effective and efficient OLAP over data streams
Alfredo Cuzzocrea, Sharma Chakravarthy
Data Knowl. Eng.1
2010 A top-down approach for compressing data cubes under the simultaneous evaluation of multiple hierarchical range queries
Alfredo Cuzzocrea
J. Intell. Inf. Syst.1
2009 CAMS: OLAPing Multidimensional Data Streams Efficiently
Alfredo Cuzzocrea
DaWaK1
2009 LCS-Hist: taming massive high-dimensional data cube compression
abstract
The problem of efficiently compressing massive high-dimensional data cubes still waits for efficient solutions capable of overcoming well-recognized scalability limitations of state-of-the-art histogram-based techniques, which perform well on small-in-size low-dimensional data cubes, whereas their performance in both representing the input data domain and efficiently supporting approximate query answering against the generated compressed data structure decreases dramatically when data cubes grow in dimension number and size. To overcome this relevant research challenge, in this paper we propose LCS-Hist, an innovative multidimensional histogram devising a complex methodology that combines intelligent data modeling and processing techniques in order to tame the annoying problem of compressing massive high-dimensional data cubes. With respect to similar histogram-based proposals, our technique introduces (i) a surprising consumption of the storage space available to house the compressed representation of the input data cube, and (ii) a superior scalability on high-dimensional data cubes. Finally, several experimental results performed against various classes of data cubes confirm the advantages of LCS-Hist, even in comparison with those given by state-of-the-art similar techniques.
Alfredo Cuzzocrea, Paolo Serafino
EDBT1
2009 Reasoning on Incompleteness of Spatial Information for Effectively and Efficiently Answering Range Queries over Incomplete Spatial Databases
Alfredo Cuzzocrea, Andrea Nucita
FQAS1
2009 Enabling OLAP in mobile environments via intelligent data cube compression techniques
Alfredo Cuzzocrea, Filippo Furfaro, Domenico Saccà
J. Intell. Inf. Syst.1
2008 Multiple-Objective Compression of Data Cubes in Cooperative OLAP Environments
Alfredo Cuzzocrea
ADBIS1
2008 A Robust Sampling-Based Framework for Privacy Preserving OLAP
Alfredo Cuzzocrea, Vincenzo Russo, Domenico Saccà
DaWaK1
2008 A Probabilistic Approach for Computing Approximate Iceberg Cubes
Alfredo Cuzzocrea, Filippo Furfaro, Giuseppe M. Mazzeo
DEXA1
2008 H-IQTS: a semantics-aware histogram for compressing categorical OLAP data
abstract
This paper introduces a novel data cube compression technique for data cubes whose main idea consists in exploiting the knowledge kept in OLAP hierarchies to drive the compression process. This approach leads to the so-called knowledge-oriented data cube compression paradigm, which is a noticeable alternative to the traditional algorithmic-oriented paradigm that focuses the attention on the issue of compressing the data cube like the latter would be a simple multidimensional array without additional knowledge. This amenity allows us to achieve several benefits, among which a more meaningful exploration of the compressed data cube enriched by semantics-aware metaphors. Our analytical contribution is finally completed by a comprehensive experimental evaluation of our proposed technique on both benchmark and real-life data cubes, also in comparison with well-established histogram-based data cube compression techniques.
Alfredo Cuzzocrea, Domenico Saccà
IDEAS1
2007 An OLAM-Based Framework for Complex Knowledge Pattern Discovery in Distributed-and-Heterogeneous-Data-Sources and Cooperative Information Systems
Alfredo Cuzzocrea
DaWaK1
2007 Efficient Fragmentation of Large XML Documents
Angela Bonifati, Alfredo Cuzzocrea
DEXA2
2007 Approximate range-sum query answering on data cubes with probabilistic guarantees
Alfredo Cuzzocrea, Wei Wang 0011
J. Intell. Inf. Syst.1
2006 A Hierarchy-Driven Compression Technique for Advanced OLAP Visualization of Multidimensional Data Cubes
Alfredo Cuzzocrea, Domenico Saccà, Paolo Serafino
DaWaK1
2006 On Semantically-Augmented XML-Based P2P Information Systems
Alfredo Cuzzocrea
FQAS1
2006 Towards a Lightweight Framework for Privacy Preserving P2P XML Databases
abstract
The problem of securing XML databases is rapidly gaining interest for both academic and industrial research. It becomes even more challenging when XML data are managed and delivered according to the P2P paradigm, as malicious attacks could take advantage from the totally-decentralized and untrusted nature of P2P networks. Starting from these considerations, in this paper we propose the guidelines of a distributed framework for supporting (i) secure fragmentation of XML documents into P2P XML databases by means of lightweight XPath-based identifiers, and (it) the creation of trusted groups of peers by means of "self-certifying" XPath links that exploit the benefits of well-known fingerprinting techniques
Angela Bonifati, Alfredo Cuzzocrea
IDEAS2
2006 Accuracy Control in Compressed Multidimensional Data Cubes for Quality of Answer-based OLAP Tools
abstract
An innovative technique supporting accuracy control in compressed multidimensional data cubes is presented in this paper. The proposed technique can be efficiently used in QoA-based OLAP tools, where OLAP users/applications and DW servers are allowed to mediate on the accuracy of (approximate) answers, similarly to what happens in QoS-based systems for the quality of services. The compressed data structure KLSA, which implements the technique, is also extensively presented and discussed. We complement our analytical contributions with an experimental evaluation on several kinds of synthetic multidimensional data cubes, demonstrating the superiority of our approach in comparison with other similar techniques
Alfredo Cuzzocrea
SSDBM1
2006 Storing and retrieving XPath fragments in structured P2P networks
Angela Bonifati, Alfredo Cuzzocrea
Data Knowl. Eng.2
2006 Improving range-sum query evaluation on data cubes via polynomial approximation
Alfredo Cuzzocrea
Data Knowl. Eng.1
2005 Providing probabilistically-bounded approximate answers to non-holistic aggregate range queries in OLAP
abstract
A novel framework for providing probabilistically-bounded approximate answers to non-holistic aggregate range queries in OLAP is presented in this paper. Such a framework allows us to efficiently support OLAP applications, as answering queries is the main bottleneck for this kind of applications. To this end, scalability of the techniques and accuracy of the answers are recognized as important limitations of state-of-the-art approximate query answering proposals in OLAP. Specifically, this paper is focused on the latter limitation, whereas it refers to results presented in [9] for the first one. The KSyn synopsis data structure, which implements the guidelines of the proposed framework and overcomes the recognized limitations, is also presented and discussed in detail, along with a query-conscious error metrics-based storage space allocation scheme. Finally, encouraging preliminary experimental results stating the goodness of our proposal are presented and discussed.
Alfredo Cuzzocrea
DOLAP1
2005 Overcoming Limitations of Approximate Query Answering in OLAP
abstract
Two important limitations of approximate query answering in OLAP are recognized and investigated. These limitations are: (i) scalability of the techniques, i.e. their reliability on highly-dimensional data cubes; and (ii) need for guarantees on the degree of approximation of the answers. In this paper, we focus on the first limitation, and propose adopting the well-known Karhunen-Loeve transform (KLT) to obtain dimensionality reduction of data cubes, thus devising a transformation methodology that is independent by the number of dimensions of the data cubes. To tailor the KLT for the specific OLAP context, effective optimizations are also proposed, by taking into account the query-consciousness feature. Finally, some encouraging preliminary experimental results are presented.
Alfredo Cuzzocrea
IDEAS1
2005 Towards a Semantics-Based Framework for KD- and IR-style Resource Querying on XML-Based P2P Information Systems
abstract
A semantics-based framework for KD- and IR-style resource querying on XML-based P2P information systems is described in this paper. Particularly, we present and discuss in detail the XML data model and the knowledge representation model of such a framework, which are both based on the amenity of adding semantics to data. As main result, we obtain a collection of techniques that allows us to efficiently process XML-formatted knowledge in P2P IS.
Alfredo Cuzzocrea
Web Intelligence1
2004 Answering Approximate Range Aggregate Queries on OLAP Data Cubes with Probabilistic Guarantees
Alfredo Cuzzocrea, Wei Wang 0011, Ugo Matrangolo
DaWaK1
2004 Analytical Synopses for Approximate Query Answering in OLAP Environments
Alfredo Cuzzocrea, Ugo Matrangolo
DEXA1
2004 Pushing Knowledge Management in Web Information Systems Engineering
Alfredo Cuzzocrea, Carlo Mastroianni
IDEAS1
2004 Knowledge on the Web: Making Web Services Knowledge-Aware
abstract
Knowledge personalization is currently the most investigated issue in the context of service-oriented systems on the Web. Knowledge representation and management are the critical issues for knowledge personalization, and actually are currently being widely investigated, mainly due to the explosion of data modeling technologies such as XML and XML Schema. Despite some progress, a widely approved standard for delivering knowledge is still missing. In this paper we propose a new approach for representing, managing, and delivering knowledge on the Web and the correspondent framework, called Distributed Knowledge Networks (DKN), that implements it. We also provide a reference architecture for DKN and some experimental results about knowledge personalization.
Alfredo Cuzzocrea
Web Intelligence1
2003 A Reference Architecture for Knowledge Management-Based Web Systems
abstract
Knowledge management-based Web systems (KMbWS) are a novel class of Web information systems whose main goal is to adapt contents and presentations with respect to user needs and backgrounds through the execution of knowledge management processes. KMbWS involve complex issues such as knowledge representation, classification and clustering, reasoning, and, more recently, ontologies and semantic Web. In this paper we present a methodology for the designing and developing of KM-bWS, starting from the application domain analysis. We also provide a reference multi-layer architecture for KM-bWS. In our opinion, a KM-bWS can be considered as an "intelligent knowledge hub" because it makes distributed Web resources available by means of knowledge management techniques such as classification and clustering. Finally, we present a reference model for developing a KM-bMS authoring tool able to support our methodology.
Alfredo Cuzzocrea, Carlo Mastroianni
WISE1