Michele Ianni

dblp:165/6510 · DBLP profile ↗
← Back
20ranked-venue papers
10as first author
11since 2021 · last 2026
0000-0003-0562-7462ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 4 first-author · 2 since 2021Systems, architecture and hardware · 5 · 1 first-author · 2 since 2021Security and privacy · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author
YearPublicationVenuePosition
2026 A formal framework for LLM-assisted automated generation of Zeek signatures from binary artifacts
abstract
Designing semantically meaningful and operationally effective intrusion detection signatures remains a labor-intensive and expertise-driven task, particularly within the Zeek network monitoring framework. In this paper, we introduce a formalized and modular system for automating Zeek signature generation using Large Language Models (LLMs). Our pipeline begins with static analysis of binary artifacts, extracts salient behavioral features, and transforms them into structured prompts for an LLM tasked with synthesizing Zeek scripts. We provide a rigorous formal framework that defines each stage of this transformation, along with theoretical models for prompt distortion, injection resilience, and sanitization. Furthermore, we explore the adversarial surface exposed by LLMs—introducing a taxonomy of injection attacks, prompt inversion risks, and behavioral feedback loops—and propose mitigations grounded in filtering and robust prompt engineering. Our approach not only accelerates signature creation but also enhances interpretability and adaptability in evolving threat environments. The framework lays the groundwork for future extensions involving dynamic analysis and automated post-validation of generated signatures.
Claudia Greco, Michele Ianni
Future Gener. Comput. Syst.2
2025 CheckMATE '25: Research on Offensive and Defensive Techniques in the Context of Man At The End (MATE) Attacks
abstract
Man-At-The-End (MATE) attackers operate with full access to software or hardware targets and can observe, analyze, and modify running systems to extract secrets or alter behavior. CheckMATE explores both offensive and defensive research in this space: measurement studies and tooling that expose realistic attack techniques, alongside defenses such as obfuscation, tamper-resistance, watermarking, white-box cryptography, and hardware-assisted protections. This workshop collects rigorous, reproducible research aimed at bridging academic advances and industry practice. The CheckMATE '25 complete workshop proceedings can be found at: https://dl.acm.org/citation.cfm?id=3733817
Sébastien Bardin, Michele Ianni, Hyungon Moon
CCS2
2025 Light sensor based covert channels on mobile devices
Mila Dalla Preda, Claudia Greco, Michele Ianni, Francesco Lupia, Andrea Pugliese 0001
Inf. Sci.3
2025 A cross-architecture malware detection approach based on intermediate representation
abstract
Detecting malware across diverse architectures and evasion techniques has become a critical challenge as modern malware increasingly targets non-traditional platforms such as IoT devices. Traditional signature-based approaches, which rely on architecture-specific bytecode patterns, often fail when malware is recompiled for different platforms or obfuscated to evade detection. In this paper, we propose a novel framework for cross-architecture, signature-based malware detection. Our approach leverages Intermediate Representation (IR) to identify malicious behaviors in a platform-independent manner. By matching higher-level patterns in the IR, our framework generates signatures capable of detecting malware across multiple architectures and resisting common obfuscation techniques. The proposed framework adopts the YARA syntax, a widely used tool for malware detection, while introducing custom high-level primitives that abstract complex IR constructs. These primitives simplify the rule-writing process, enabling more efficient and precise signature creation. Additionally, we discuss the limitations of current approaches and demonstrate how our framework advances the state of the art in signature-based malware detection.
Claudia Greco, Michele Ianni
J. Inf. Secur. Appl.2
2024 CheckMATE '24 - Research on Offensive and Defensive Techniques in the context of Man At The End (MATE) Attacks
abstract
MATE (Man-At-The-End) is an attacker model where an adversary has access to the target software and/or hardware environment of his victim and the ability to observe and modify it in order to extract secrets such as cryptographic keys or sensitive information, possibly with the subsequent goal of compromising code integrity or inserting backdoors, among others. A typical example of such a scenario is the case of an attack on a stolen smartphone or against software leveraging protection to offer premium content and/or features such as paid TV channels.
Sebastian Schrittwieser, Michele Ianni
CCS2
2024 Model aggregation techniques in federated learning: A comprehensive survey
abstract
Federated learning (FL) is a distributed machine learning (ML) approach that enables models to be trained on client devices while ensuring the privacy of user data. Model aggregation, also known as model fusion, plays a vital role in FL. It involves combining locally generated models from client devices into a single global model while maintaining user data privacy. However, the accuracy and reliability of the resulting global model depend on the aggregation method chosen, making the selection of an appropriate method crucial. Initially, the simple averaging of model weights was the most commonly used method. However, due to its limitations in handling low-quality or malicious models, alternative techniques have been explored. As FL gains popularity in various domains, it is crucial to have a comprehensive understanding of the available model aggregation techniques and their respective strengths and limitations. However, there is currently a significant gap in the literature when it comes to systematic and comprehensive reviews of these techniques. To address this gap, this paper presents a systematic literature review encompassing 201 studies on model aggregation in FL. The focus is on summarizing the proposed techniques and the ones currently applied for model fusion. This survey serves as a valuable resource for researchers to enhance and develop new aggregation techniques, as well as for practitioners to select the most appropriate method for their FL applications.
Pian Qi, Diletta Chiaro, Antonella Guzzo, Michele Ianni, Giancarlo Fortino, Francesco Piccialli
Future Gener. Comput. Syst.4
2024 Editorial: Special issue on software protection and attacks
Michele Ianni, Mila Dalla Preda, Kim-Kwang Raymond Choo, Miguel Correia 0001
J. Inf. Secur. Appl.1
2023 SCOUT: Security by computing OUTliers on activity logs
abstract
The current increase in cybercrime is demanding more effective and efficient data exploration and analysis solutions that can help analysts to detect cyberattacks. However, the huge amount of data generated continuously confronts a number of technological difficulties and classical algorithms must be often redesigned to be able to deal with this seemingly endless stream of information coming from past activity logs and real-time data. In this paper, we propose a novel methodology able to identify security threats in activity logs. The contribution of the paper is twofold: we propose an encoding technique, based on prime numbers, that can be used to represent in a compact way a set of activities, we then describe an outlier detection algorithm which, based on the encoded activities, is able to detect malicious behavior. The extensive experimental analysis proved the effectiveness of the proposed methodology.
Michele Ianni, Elio Masciari
Comput. Secur.1
2022 Some Experiments on High Performance Anomaly Detection
abstract
The rise of cyber crime observed in recent years calls for more efficient and effective data exploration and analysis tools. In this respect, the need to support advanced analytics on activity logs and real time data is driving data scientist’ interest in designing and implementing scalable cyber security solutions. However, when data science algorithms are leveraged for huge amounts of data, their fully scalable deployment faces a number of technical challenges that grow with the complexity of the algorithms involved and the task to be tackled. Thus algorithms, that were originally designed for classical scenarios, need to be redesigned in order to be effectively used for cyber security purposes. In this paper, we explore these problems and then propose a solution which has proven to be very effective in identifying malicious activities.
Michele Ianni, Elio Masciari
PDP1
2021 A compact encoding of security logs for high performance activity detection
abstract
The rise of cyber crime observed in recent years calls for more efficient and effective Data Exploration and analysis tools. In this respect, the need to support advanced analytics on activity logs and real time data is driving data scientist' interest in order to design and implement scalable cyber security solutions. However, when data science algorithms are leveraged for huge amount of data, their fully scalable deployment faces a number of technical challenges that grow with the complexity of the algorithms involved and the task to be tackled. Thus algorithms, that were originally designed for classical scenarios, must often be redesigned in order to be effectively used for cyber security purposes. In this paper, we explore these problems and then propose a solution which has proven to be very effective for compressing suspicious activities in smaller and larger cyber environment in order to make the intelligent analysis and simulation of threats more efficient.
Michele Ianni, Elio Masciari
PDP1
2021 A survey of Big Data dimensions vs Social Networks analysis
abstract
The pervasive diffusion of Social Networks (SN) produced an unprecedented amount of heterogeneous data. Thus, traditional approaches quickly became unpractical for real life applications due their intrinsic properties: large amount of user-generated data (text, video, image and audio), data heterogeneity and high speed generation rate. More in detail, the analysis of user generated data by popular social networks (i.e Facebook (https://www.facebook.com/), Twitter (https://www.twitter.com/), Instagram (https://www.instagram.com/), LinkedIn (https://www.linkedin.com/)) poses quite intriguing challenges for both research and industry communities in the task of analyzing user behavior, user interactions, link evolution, opinion spreading and several other important aspects. This survey will focus on the analyses performed in last two decades on these kind of data w.r.t. the dimensions defined for Big Data paradigm (the so called Big Data 6 V's).
Michele Ianni, Elio Masciari, Giancarlo Sperlì
J. Intell. Inf. Syst.1
2020 Modeling and efficiently detecting security-critical sequences of actions
Antonella Guzzo, Michele Ianni, Andrea Pugliese 0001, Domenico Saccà
Future Gener. Comput. Syst.2
2020 Fast and effective Big Data exploration by clustering
Michele Ianni, Elio Masciari, Giuseppe M. Mazzeo, Mario Mezzanzanica, Carlo Zaniolo
Future Gener. Comput. Syst.1
2019 An Overview of the Endless Battle between Virus Writers and Detectors: How Compilers Can Be Used as an Evasion Technique
Michele Ianni, Elio Masciari, Domenico Saccà
DATA1
2018 Clustering Big Data
abstract
The need to support advanced analytics on Big Data is driving data scientist' interest toward massively parallel distributed systems and software platforms, such as Map-Reduce and Spark, that make possible their scalable utilization.However, when complex data mining algorithms are required, their fully scalable deployment on such platforms faces a number of technical challenges that grow with the complexity of the algorithms involved.Thus algorithms, that were originally designed for a sequential nature, must often be redesigned in order to effectively use the distributed computational resources.In this paper, we explore these problems, and then propose a solution which has proven to be very effective on the complex hierarchical clustering algorithm CLUBS+.By using four stages of successive refinements, CLUBS+ delivers high-quality clusters of data grouped around their centroids, working in a totally unsupervised fashion.Experimental results confirm the accuracy and scalability of CLUBS+ on Map-Reduce platforms.
Michele Ianni, Elio Masciari, Giuseppe M. Mazzeo, Carlo Zaniolo
DATA1
2018 Efficient Big Data Clustering
abstract
The need to support advanced analytics on Big Data is driving data scientist' interest toward massively parallel distributed systems and software platforms, such as Map-Reduce and Spark, that make possible their scalable utilization. However, when complex data mining algorithms are required, their fully scalable deployment on such platforms faces a number of technical challenges that grow with the complexity of the algorithms involved. Thus algorithms, that were originally designed for a sequential nature, must often be redesigned in order to effectively use the distributed computational resources. In this paper, we explore these problems, and then propose a solution which has proven to be very effective on the complex hierarchical clustering algorithm CLUBS+. By using four stages of successive refinements, CLUBS+ delivers high-quality clusters of data grouped around their centroids, working in a totally unsupervised fashion. Experimental results confirm the accuracy and scalability of CLUBS+.
Michele Ianni, Elio Masciari, Giuseppe M. Mazzeo, Carlo Zaniolo
IDEAS1
2018 Clustering Goes Big: CLUBS-P, an Algorithm for Unsupervised Clustering Around Centroids Tailored For Big Data Applications
abstract
The need to support advanced analytics on Big Data is driving data scientist' interest toward massively parallel distributed systems and software platforms, such as Map- Reduce and Spark, that make possible their scalable utilization. However, when complex data mining algorithms are required, their fully scalable deployment on such platforms faces a number of technical challenges that grow with the complexity of the algorithms involved. Thus algorithms, that were originally designed for a sequential nature, must often be redesigned in order to effectively use the distributed computational resources. In this paper, we explore these problems, and then propose a solution which has proven to be very effective on the complex hierarchical clustering algorithm CLUBS+. We present a parallel version of CLUBS+ named CLUBS-P with an ad-hoc implementation based on message passing: CLUBS-MP.
Michele Ianni, Elio Masciari, Giuseppe M. Mazzeo, Carlo Zaniolo
PDP1
2018 Distributed computing by leveraging and rewarding idling user resources from P2P networks
Nunzio Cassavia, Sergio Flesca, Michele Ianni, Elio Masciari, Chiara Pulice
J. Parallel Distributed Comput.3
2017 A Peer to Peer Approach to Efficient High Performance Computing
abstract
Nowadays, many applications call for collaborative solutions in order to accomplish complex projects requiring huge amounts of computing resources, e.g., physical science simulation. Many approaches have been proposed in order to design a task partitioning strategy able to assign pieces of execution to the appropriate workers. In this paper, we describe our peer to peer solution for solving complex works by using the idling computational resources of users connected to our network. More in detail, we designed a framework that allows users to share their CPU and memory in a secure and efficient way. By doing this, users help each others by asking the network available computational resources when they face high computing demanding tasks. Differently from many proposal available for volunteer computing, users providing their resources are rewarded with tangible credits, i.e., they can redeem their credits by asking computation power to solve their own task or/and they can redeem them earning coins. As we do not require to power additional resources for solving tasks (we better exploit unused resources already powered instead).
Nunzio Cassavia, Sergio Flesca, Michele Ianni, Elio Masciari, Giuseppe Papuzzo, Chiara Pulice
PDP3
2015 Big Data Techniques For Supporting Accurate Predictions of Energy Production From Renewable Sources
abstract
Predicting the output power of renewable energy production plants distributed on a wide territory is a really valuable goal, both for marketing and energy management purposes. Vi-POC (Virtual Power Operating Center) project aims at designing and implementing a prototype which is able to achieve this goal. Due to the heterogeneity and the high volume of data, it is necessary to exploit suitable Big Data analysis techniques in order to perform a quick and secure access to data that cannot be obtained with traditional approaches for data management. In this paper, we describe Vi-POC -- a distributed system for storing huge amounts of data, gathered from energy production plants and weather prediction services. We use HBase over Hadoop framework on a cluster of commodity servers in order to provide a system that can be used as a basis for running machine learning algorithms. Indeed, we perform one-day ahead forecast of PV energy production based on Artificial Neural Networks in two learning settings, that is, structured and non-structured output prediction. Preliminary experimental results confirm the validity of the approach, also when compared with a baseline approach.
Michelangelo Ceci, Roberto Corizzo, Fabio Fumarola, Michele Ianni, Donato Malerba, Gaspare Maria, Elio Masciari, Marco Oliverio, Aleksandra Rashkovska
IDEAS4