Martin Molan

dblp:304/1331 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-6805-2232ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Distilling Graph Reasoning into Lightweight CNNs for Near-Sensor Point Cloud Corruption Detection
abstract
Real-world point cloud corruption on automotive LiDAR lenses can significantly degrade the reliability of down-stream perception, particularly object detection models trained on clean data, which may yield overconfident false positives. To address this, we propose a near-sensor gating module that classifies incoming point clouds as either clean or contaminated using a teacher–student knowledge distillation pipeline. A Graph Attention Network (GAT), trained directly on raw point clouds, serves as the teacher. On real-world contaminated LiDAR data, the distilled student achieves an average F1-score of 0.83, closely matching the GAT teacher’s 0.88, and consistently outperforming other supervised baselines across diverse test environments. Importantly, the student’s 2D-CNN architecture reduces preprocessing complexity from O(n log n) of graph construction to O(n), enabling faster and more efficient point cloud handling. The student model is quantized to 16-bit fixed-point and deployed on a GAP8 (RISC-V) platform. It achieves an inference latency of 26 milliseconds, consumes only 210µJ per inference, and fits within 84KB of L2 memory. This makes the proposed solution a practical and resource-efficient near-sensor gating module for robust, contaminant-aware perception in embedded automotive systems. The implementation will be available at https://gitlab.com/ecs-lab/distilling-pointcloud-corruption
Grafika Jati, Martin Molan, Francesco Barchi, Andrea Bartolini, Giuseppe Mercurio, Andrea Acquaviva
DATE2
2026 EXASAGE: The first data center operational data analysis assistant
abstract
• We propose EXASAGE, the first ODA operational data analysis assistant for data centers. To the best of our knowledge, this is the first prototype of a Large Language Model (LLM)-based tool that provides an AI-driven interoperable layer designed to interact with data collected at data center facilities, serving as an on-demand data access assistant that generates graph database query codes for timely, non-critical operational analysis. • The proposed framework leverages a Knowledge Graph (KG) approach instead of a standard NoSQL database at a data center. To achieve this, we provide a formal representation of the data collected at the data center using a Resource Description Framework (RDF) ontology. • We evaluated the framework in a real-world setting using 1,000 complex queries representative of the daily tasks performed by facility managers and engineers. The framework achieved a 93.6% accuracy for correctly generated and executed graph queries, compared to only 25% accuracy for standard NoSQL query generation, demonstrating the benefits of combining LLMs and KG. • We address the significant storage challenges caused by time-series data conversion into a KG, which results in a storage size increase of than 745x compared to NoSQL database storage, using virtualization of KGs. This results in a max storage overhead of just 52.62 MiB over all the 1000 user input queries. Data centers increasingly depend on Operational Data Analytics (ODA) for real-time insights from vast streams of telemetry data. They typically utilize NoSQL databases for scalability and data diversity, which leads to unstructured data representation and presents significant challenges for the data interoperability. Indeed, the lack of standardization, combined with schema flexibility and complex data structures, makes it difficult for system administrators to write and execute queries, ultimately complicating the automation of data retrieval tasks. Pre-trained Large Language Models (LLMs), with their latent knowledge, promise a ready-to-use AI-driven data interoperability layer, enabling data retrieval through natural language input. However, they often generate inaccurate or hallucinated query code when handling heterogeneous data sources and complex data structures. In this paper we present EXASAGE, the first ODA operational data analysis assistant that leverages a Knowledge Graph (KG)-based approach, addressing these LLM limitations and simplifying data retrieval tasks in data center facilities through a prototype implementation. EXASAGE employs an LLM based query generator as an interoperable layer to convert natural language into SPARQL queries (native to KGs), executed at a graph database endpoint, along with a virtual KG approach that retrieves only the data relevant to the user input query. In evaluations on 1,000 user input queries, EXASAGE achieved a 93.6% accuracy in generating correct SPARQL code and retrieving correct answers, significantly outperforming the 25% accuracy of NoSQL/SQLite queries, which frequently exhibited hallucinations. Furthermore, SPARQL queries are generally more concise and demonstrate shorter inference and execution times compared to compared to NoSQL/SQLite queries. For EXASAGE, the average end-to-end time for a single execution cycle is 12.77 seconds, which is suitable for interactive, non-critical operational data analysis tasks. The maximum observed storage overhead across all generated virtual KGs is just 52.62 MiB.
Junaid Ahmed Khan, Martin Molan, Andrea Bartolini
Future Gener. Comput. Syst.2
2025 ANZIL: Attention-based network for zero-risk inspection of LiDAR point cloud in self-driving cars
abstract
LiDAR is widely used in autonomous vehicle perception, and its performance relies on algorithmic confidence. However, sensor contamination can lead to catastrophic mistakes in downstream tasks, such as incorrect object detections occurring with high confidence. This underscores the need for point cloud contaminant detection to identify data reliability before downstream processing. To overcome this, we propose a model-agnostic approach that integrates contaminant detection with downstream tasks, using graph representations and attention networks. Trained on real contaminated LiDAR, the contaminant detector complements any existing clean-trained downstream model. Point clouds passing through the contaminant detector are discarded if contamination is detected. We propose a novel cost-benefit methodology, evaluating contaminant detectors in downstream processing. The benefit measures the proportion of discarded frames that would have caused high-confidence errors in the downstream task. The cost includes the model misalignment cost, representing frames wrongly discarded despite being correctly processed, and the true cost, which reflects uncontaminated frames falsely identified as contaminated. The contaminant detector was tested on 78,000 frames with water, dust, mud, salt, oil, and foam contamination from tunnel and outdoor locations. It achieved an F1-score of 0.884 and a recall of 0.957 on a static dataset, and a 0.975 F1-score in unseen real-world environments, demonstrating high sensitivity. In model-agnostic evaluation, the method is applicable to any object detector and reduces catastrophic mistakes by at least 97%, successfully identifying 4,046 failure cases. Deployed on the NVIDIA Jetson AGX Xavier, it achieved inference in under 5 ms and point cloud transformation in 200 ms, making it suitable for edge processing. This enhances sensor reliability, enables automatic cleaning, and improves vehicle safety.
Grafika Jati, Martin Molan, Francesco Barchi, Andrea Bartolini, Giuseppe Mercurio, Andrea Acquaviva
Expert Syst. Appl.2
2024 TinyLid: a RISC-V accelerated Neural Network For LiDAR Contaminant Classification in Autonomous Vehicle
abstract
LiDAR plays a critical role in autonomous car perception. Hence, the robustness of LiDAR data is imperative. However, malfunctions resulting from sensor cover contaminants are unavoidable and can lead to erroneous data that slowly degrade performance.
Grafika Jati, Martin Molan, Francesco Barchi, Andrea Bartolini, Andrea Acquaviva
CF2
2024 GRAAFE: GRaph Anomaly Anticipation Framework for Exascale HPC systems
Martin Molan, Mohsen Seyedkazemi Ardebili, Junaid Ahmed Khan, Francesco Beneventi, Daniele Cesarini, Andrea Borghesi, Andrea Bartolini
Future Gener. Comput. Syst.1
2023 The Graph-Massivizer Approach Toward a European Sustainable Data Center Digital Twin
abstract
Modeling and understanding an expensive next-generation data center operating at a sustainable exascale performance remains a challenge yet to solve. The paper presents the approach taken by the Graph-Massivizer project, funded by the European Union, towards a sustainable data center, targeting a massive graph representation and analysis of its digital twin. We introduce five interoperable open-source tools that support this undertaking, creating an automated, sustainable loop of graph creation, analytics, optimization, sustainable resource management, and operation, emphasizing state-of-the-art progress. We plan to employ the tools for designing a massive data center graph, representing a digital twin describing spatial, semantic, and temporal relationships between the monitoring metrics, hardware nodes, cooling equipment, and jobs. The project aims to strengthen Bologna Technopole as a leading European supercomputing and big data hub offering sustainable green computing for improved societally relevant science throughput.
Martin Molan, Junaid Ahmed Khan, Andrea Bartolini, Roberta Turra, Giorgio Pedrazzi, Michael Cochez, Alexandru Iosup, Dumitru Roman, Joze M. Rozanec, Ana Lucia Varbanescu, Radu Prodan
COMPSAC1
2023 RUAD: Unsupervised anomaly detection in HPC systems
Martin Molan, Andrea Borghesi, Daniele Cesarini, Luca Benini, Andrea Bartolini
Future Gener. Comput. Syst.1
2022 Semi-supervised anomaly detection on a Tier-0 HPC system
abstract
Automated and data-driven methodologies are being introduced to assist system administrators in managing increasingly complex modern HPC systems. Anomaly detection (AD) is an integral part of improving the overall availability as it eases the system administrators' burden and reduces the time between an anomaly and its resolution. This work improves upon the current state-of-the-art (SoA) AD model by considering temporal dependencies in the data and including long-short term memory cells in the architecture of the AD model. The proposed model is evaluated on a complete ten-month history of a Tier-0 system (Marconi100 from CINECA consisting of 985 nodes). The proposed model achieves an area under the curve (AUC) of 0.758, improving upon the state-of-the-art approach that achieves an AUC of 0.747.
Martin Molan, Andrea Borghesi, Luca Benini, Andrea Bartolini
CF1
2022 Analysing Supercomputer Nodes Behaviour with the Latent Representation of Deep Learning Models
Martin Molan, Andrea Borghesi, Luca Benini, Andrea Bartolini
Euro-Par1
2022 Anomaly Detection and Anticipation in High Performance Computing Systems
abstract
In their quest toward Exascale, High Performance Computing (HPC) systems are rapidly becoming larger and more complex, together with the issues concerning their maintenance. Luckily, many current HPC systems are endowed with data monitoring infrastructures that characterize the system state, and whose data can be used to train Deep Learning (DL) anomaly detection models, a very popular research area. However, the lack of labels describing the state of the system is a wide-spread issue, as annotating data is a costly task, generally falling on human system administrators and thus does not scale toward exascale. In this article we investigate the possibility to extract labels from a service monitoring tool (Nagios) currently used by HPC system administrators to flag the nodes which undergo maintenance operations. This allows to automatically annotate data collected by a fine-grained monitoring infrastructure; this labelled data is then used to train and validate a DL model for anomaly detection. We conduct the experimental evaluation on a tier-0 production supercomputer hosted at CINECA, Bologna, Italy. The results reveal that the DL model can accurately detect the real failures, and, moreover, it canpredictthe insurgency of anomalies, by systematically anticipating the actual labels (i.e., the moment when system administrators realize when an anomalous event happened); the average advance time computed on historical traces is around 45 minutes. The proposed technology can be easily scaled toward exascale systems to easy their maintenance.
Andrea Borghesi, Martin Molan, Michela Milano, Andrea Bartolini
IEEE Trans. Parallel Distributed Syst.2