Giulia De Masi

dblp:147/8719 · DBLP profile ↗
← Back
17ranked-venue papers
0as first author
17since 2021 · last 2026
0000-0003-3284-880XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Leveraging model explainability and fine-grained cutmix augmentation for robust detection of apricot diseases in UAV images
Jamil Ahmad 0003, Wail Gueaieb, Abdulmotaleb El Saddik, Giulia De Masi, Fakhri Karray
Expert Syst. Appl.4
2025 Enhancing collaboration in uncertain environment: Multi-Agent Reinforcement Learning for underwater monitoring
abstract
Underwater monitoring is extremely complex due to the lack of a global localization system, limited communication and environmental factors such as turbidity and darkness that limit visibility, affecting control and situational awareness. Typically, monitoring relies on a single autonomous underwater vehicle (AUV) or a set of independent AUVs; techniques which are prone to failure as they rely only on onboard odometry and sensors, making missions vulnerable to malfunctions, damage, and noise. To address these challenges, we propose a Multi-Agent Reinforcement Learning (MARL) framework to enable cooperation among multiple AUVs, mitigating the limitations of the underwater environment. Our in-silico solution focuses on a group of robots learning a strategy to follow a partially hidden underwater pipe without global localization, while dealing with environmental disturbances affecting sensors and actuators. The numerosity of the agents, and most importantly their collaboration, helps overcome underwater visibility constraints. By sharing relative position information of neighboring agents with respect to the pipe, navigation is improved. By introducing quantitative measures for pipe exploration, we show that cooperation significantly enhances system performance compared to independent agents. Emerging collaboration among robots allows the swarm to complete pipe inspections faster and more efficiently than non-cooperative baseline models of non-interacting agents, even under extremely reduced visibility scenarios. Moreover, single agents also benefit from cooperation, learning effective policies more quickly and covering a longer portion of the pipe. Finally, our model guarantees explainability. We analyze learned strategies and provide a visualization method that allows the interpretation of the learned policies. • Reinforcement learning is applied to pipeline following by underwater robotic agent. • In conditions of poor visibility, a single agent is not able to complete the mission. • In contrast, multi-agent team completes this task using reinforcement learning. • Swarm collaboration enables faster and more efficient task completion. • Collaboration enables agents develop more efficient individual strategies.
Alberto Luvisutto, Antonio Celani, Federico Renda, Cesare Stefanini, Giulia De Masi
Expert Syst. Appl.5
2024 Enhancing Training of Spiking Neural Network with Stochastic Latency
abstract
Spiking neural networks (SNNs) have garnered significant attention for their low power consumption when deployed on neuromorphic hardware that operates in orders of magnitude lower power than general-purpose hardware. Direct training methods for SNNs come with an inherent latency for which the SNNs are optimized, and in general, the higher the latency, the better the predictive powers of the models, but at the same time, the higher the energy consumption during training and inference. Furthermore, an SNN model optimized for one particular latency does not necessarily perform well in lower latencies, which becomes relevant in scenarios where it is necessary to switch to a lower latency because of the depletion of onboard energy or other operational requirements. In this work, we propose Stochastic Latency Training (SLT), a direct training method for SNNs that optimizes the model for the given latency but simultaneously offers a minimum reduction of predictive accuracy when shifted to lower inference latencies. We provide heuristics for our approach with partial theoretical justification and experimental evidence showing the state-of-the-art performance of our models on datasets such as CIFAR-10, DVS-CIFAR-10, CIFAR-100, and DVS-Gesture. Our code is available at https://github.com/srinuvaasu/SLT
Srinivas Anumasa, Bhaskar Mukhoty, Velibor Bojkovic, Giulia De Masi, Huan Xiong, Bin Gu 0001
AAAI4
2024 Dynamic Spiking Graph Neural Networks
abstract
The integration of Spiking Neural Networks (SNNs) and Graph Neural Networks (GNNs) is gradually attracting attention due to the low power consumption and high efficiency in processing the non-Euclidean data represented by graphs. However, as a common problem, dynamic graph representation learning faces challenges such as high complexity and large memory overheads. Current work often uses SNNs instead of Recurrent Neural Networks (RNNs) by using binary features instead of continuous ones for efficient training, which overlooks graph structure information and leads to the loss of details during propagation. Additionally, optimizing dynamic spiking models typically requires the propagation of information across time steps, which increases memory requirements. To address these challenges, we present a framework named Dynamic Spiking Graph Neural Networks (Dy-SIGN). To mitigate the information loss problem, Dy-SIGN propagates early-layer information directly to the last layer for information compensation. To accommodate the memory requirements, we apply the implicit differentiation on the equilibrium state, which does not rely on the exact reverse of the forward computation. While traditional implicit differentiation methods are usually used for static situations, Dy-SIGN extends it to the dynamic graph setting. Extensive experiments on three large-scale real-world dynamic graph datasets validate the effectiveness of Dy-SIGN on dynamic node classification tasks with lower computational costs.
Mengzhu Wang, Zhenghan Chen, Giulia De Masi, Huan Xiong, Bin Gu 0001
AAAI4
2024 Data Driven Threshold and Potential Initialization for Spiking Neural Networks
abstract
Spiking neural networks (SNNs) present an increasingly popular alternative to artificial neural networks (ANNs), due to their energy and time efficiency when deployed on neuromorphic hardware. However, due to their discrete and highly non-differentiable nature, training SNNs is a challenging task and remains an active area of research. Some of the most prominent ways to train SNNs are based on ANN-to-SNN conversion where an SNN model is initialized with parameters from the corresponding, pre-trained ANN model. SNN models trained through ANN-to-SNN conversion or hybrid training show state of the art performance among SNNs on many machine learning tasks, comparable to those of ANNs. However, the top performing models need high latency or tailored ANNs to perform well, and in general are not using the full information available from ANNs. In this work, we propose novel method to initialize SNN’s thresholds and initial membrane potential after ANN-to-SNN conversion, using distributions of ANN’s activation values. We provide a theoretical framework for feature distribution-based conversion error, providing theoretical results on optimal membrane initialization and thresholds which minimize this error, as well as a practical algorithm for finding these optimal values. We test our method, both as a stand-alone ANN-to-SNN conversion and in combination with other methods, and show state of the art results on high-dimensional datasets such as CIFAR10, CIFAR100 and ImageNet and various architectures. Our code is available at \url{https://github.com/srinuvaasu/data_driven_init}
Velibor Bojkovic, Srinivas Anumasa, Giulia De Masi, Bin Gu 0001, Huan Xiong
AISTATS3
2024 Knowledge-Infused Learning for Fine-Grained Plant Disease Recognition
abstract
Domain knowledge exists in various forms, including text, ontologies, graphs, images, audio, and videos. In plant disease detection, most works solely utilize images with disease labels, neglecting textual descriptions of visual disease symptoms used by human experts for diagnosis. These text descriptions and sample images aid expert identification of visual symptoms. We propose a novel method that leverages text descriptions and image data by modeling domain-specific knowledge about visual symptoms in leaf images as separate feature channels. Each channel corresponds to specific features whose absence or presence in the image influences model predictions. We introduce a channel attention-guided fusion module for weighting each channel based on the input and corresponding output. The combined feature channels are transformed into a standardized 3-channel input format, which can then be processed by any pre-trained convolutional neural network (CNN) as input for feature extraction and subsequent classification. Furthermore, intermediate activations of the channel attention layer combined with the weights from the fusion layer make model predictions explainable. Experimental results on three publicly available datasets of apple and cucumber leaf diseases demonstrate improvements of up to 5% utilizing various state-of-the-art CNN architectures, indicating the efficacy of incorporating textual disease descriptions using the proposed approach.
Jamil Ahmad 0003, Wail Gueaieb, Abdulmotaleb El Saddik, Giulia De Masi, Fakhri Karray
ICIP4
2024 TAB: Temporal Accumulated Batch Normalization in Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) are attracting growing interest for their energy-efficient computing when implemented on neuromorphic hardware. However, directly training SNNs, even adopting batch normalization (BN), is highly challenging due to their non-differentiable activation function and the temporally delayed accumulation of outputs over time. For SNN training, this temporal accumulation gives rise to Temporal Covariate Shifts (TCS) along the temporal dimension, a phenomenon that would become increasingly pronounced with layer-wise computations across multiple layers and multiple time-steps. In this paper, we introduce TAB (Temporal Accumulated Batch Normalization), a novel SNN batch normalization method that addresses the temporal covariate shift issue by aligning with neuron dynamics (specifically the accumulated membrane potential) and utilizing temporal accumulated statistics for data normalization. Within its framework, TAB effectively encapsulates the historical temporal dependencies that underlie the membrane potential accumulation process, thereby establishing a natural connection between neuron dynamics and TAB batch normalization. Experimental results on CIFAR-10, CIFAR-100, and DVS-CIFAR10 show that our TAB method outperforms other state-of-the-art methods.
Vincent Zoonekynd, Giulia De Masi, Bin Gu 0001, Huan Xiong
ICLR3
2024 Certified Adversarial Robustness for Rate Encoded Spiking Neural Networks
abstract
The spiking neural networks are inspired by the biological neurons that employ binary spikes to propagate information in the neural network. It has garnered considerable attention as the next-generation neural network, as the spiking activity simplifies the computation burden of the network to a large extent and is known for its low energy deployment enabled by specialized neuromorphic hardware. One popular technique to feed a static image to such a network is rate encoding, where each pixel is encoded into random binary spikes, following a Bernoulli distribution that uses the pixel intensity as bias. By establishing a novel connection between rate-encoding and randomized smoothing, we give the first provable robustness guarantee for spiking neural networks against adversarial perturbation of inputs bounded under $l_1$-norm. We introduce novel adversarial training algorithms for rate-encoded models that significantly improve the state-of-the-art empirical robust accuracy result. Experimental validation of the method is performed across various static image datasets, including CIFAR-10, CIFAR-100 and ImageNet-100. The code is available at \url{https://github.com/BhaskarMukhoty/CertifiedSNN}.
Bhaskar Mukhoty, Hilal AlQuabeh, Giulia De Masi, Huan Xiong, Bin Gu 0001
ICLR3
2024 NDOT: Neuronal Dynamics-based Online Training for Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) are attracting great attention for their energy-efficient and fast-inference properties in neuromorphic computing. However, the efficient training of deep SNNs poses challenges in gradient calculation due to the non-differentiability of their binary spike-generating activation functions. The widely used surrogate gradient (SG) method, combined with the back-propagation through time (BPTT), has shown considerable effectiveness. Yet, BPTT’s process of unfolding and back-propagating along the computation graph requires storing intermediate information at all time-steps, resulting in huge memory consumption and failing to meet online requirements. In this work, we propose Neuronal Dynamics-based Online Training (NDOT) for SNNs, which uses the neuronal dynamics-based temporal dependency/sensitivity in gradient computation. NDOT enables forward-in-time learning by decomposing the full gradient into temporal and spatial gradients. To illustrate the intuition behind NDOT, we employ the Follow-the-Regularized-Leader (FTRL) algorithm. FTRL explicitly utilizes historical information and addresses limitations in instantaneous loss. Our proposed NDOT method accurately captures temporal dependencies through neuronal dynamics, functioning similarly to FTRL’s explicit utilizing historical information. Experiments on CIFAR-10, CIFAR-100, and CIFAR10-DVS demonstrate the superior performance of our NDOT method on large-scale static and neuromorphic datasets within a small number of time steps. The codes are available at https://github.com/HaiyanJiang/SNN-NDOT.
Giulia De Masi, Huan Xiong, Bin Gu 0001
ICML2
2024 CamoFocus: Enhancing Camouflage Object Detection with Split-Feature Focal Modulation and Context Refinement
abstract
Camouflage Object Detection (COD) involves the challenge of isolating a target object from a visually similar background, presenting a formidable challenge for learning algorithms. Drawing inspiration from state-of-the-art (SOTA) Focal Modulation Networks, our objective is to proficiently modulate the foreground and background components, thereby capturing the distinct features of each. We introduce a Feature Split and Modulation (FSM) module to attain this goal. This module efficiently separates the object from the background by utilizing foreground and background modulators guided by a supervisory mask. For enhanced feature refinement, we propose a Context Refinement Module (CRM), which considers features acquired from FSM across various spatial scales, leading to comprehensive enrichment and highly accurate prediction maps. Through extensive experimentation, we showcase the superiority of CamoFocus over recent SOTA COD methods. Our evaluations encompass diverse benchmark datasets, including CAMO, COD10K, CHAMELEON, and NC4K. The findings underscore the potential and significance of the proposed CamoFocus model and establish its efficacy in addressing the critical challenges of camouflage object detection.
Mustaqeem Khan 0001, Wail Gueaieb, Abdulmotaleb El Saddik, Giulia De Masi, Fakhri Karray
WACV5
2024 Yield estimation and health assessment of temperate fruits: A modular framework
Jamil Ahmad 0003, Wail Gueaieb, Abdulmotaleb El Saddik, Giulia De Masi, Fakhri Karray
Eng. Appl. Artif. Intell.4
2023 A Unified Optimization Framework of ANN-SNN Conversion: Towards Optimal Mapping from Activation Values to Firing Rates
abstract
Spiking Neural Networks (SNNs) have gained significant attention for their energy-efficient and fast-inference capabilities, but training SNNs from scratch can be challenging due to the discrete nature of spikes. One alternative method is to convert an Artificial Neural Network (ANN) into an SNN, known as ANN-SNN conversion. Currently, existing ANN-SNN conversion methods often involve redesigning the ANN with a new activation function, rather than utilizing the traditional ReLU, and converting it to an SNN. However, these methods do not take into account the potential performance loss between the regular ANN with ReLU and the tailored ANN. In this work, we propose a unified optimization framework for ANN-SNN conversion that considers both performance loss and conversion error. To achieve this, we introduce the SlipReLU activation function, which is a weighted sum of the threshold-ReLU and the step function. Theoretical analysis demonstrates that conversion error can be zero on a range of shift values $\delta \in [-0.5,0.5]$ rather than a fixed shift term 0.5. We evaluate our SlipReLU method on CIFAR datasets, which shows that SlipReLU outperforms current ANN-SNN conversion methods and supervised training methods in terms of accuracy and latency. To the best of our knowledge, this is the first ANN-SNN conversion method that enables SNN inference using only 1 time step. Code is available at https://github.com/HaiyanJiang/SNN_Conversion_unified.
Srinivas Anumasa, Giulia De Masi, Huan Xiong, Bin Gu 0001
ICML3
2023 CEAFFOD: Cross-Ensemble Attention-based Feature Fusion Architecture Towards a Robust and Real-time UAV-based Object Detection in Complex Scenarios
abstract
Deploying object detectors in embedded devices such as unmanned aerial vehicles (UAVs) comes with many challenges. This is due to both the UAV itself having low embedded resources in terms of computation and memory, and also due to the nature of the captured visual data with the variations in objects' scale, orientation, density, viewpoint, distribution, shape, context and others. It is crucial for the object detector to be robust with high accuracy, real-time with fast inference and light-weight to be applicable. Inspired by YOLO architecture, we propose a novel single-stage detection architecture. Our contributions are, first, feature fusion spatial pyramid pooling (FFSPP) block that applies attention-based feature fusion across both time and space utilizing the information of subsequent frames and scales in an efficient manner. Secondly, we introduce a multi-dilated attention-based cross-stage partial connection (MDACSP) block that helps in increasing the receptive field and producing per-channel modulation weights after aggregating the feature maps across their spatial domain. Third, scaled feature fusion head (SFFH) fuses both the FFSPP block features and the connected MDACSP block features specific for this head. For a more robust result across different scenarios, we perform cross-ensembling with three of the top UAV/traffic surveillance datasets: UAVDT, UA-DETRAC and VisDrone. Our ablation study shows how every contribution improves over the baseline. Our approach yielded the state-of-the-art results in all the aforementioned datasets achieving 89.3% mAP, 93.5% mAP, and 42.9% mAP respectively. Testing the model performance on NVIDIA Jetson Xavier NX board shows a desirable balance between the inference time and the memory cost. We also show qualitatively the model robustness and efficiency across the diverse complex scenarios of these datasets. We hope this work facilitates the advancement of the UAV-based perception in such crucial industrial applications.
Ahmed Elhagry, Hang Dai, Abdulmotaleb El Saddik, Wail Gueaieb, Giulia De Masi
ICRA5
2023 Direct Training of SNN using Local Zeroth Order Method
abstract
Spiking neural networks are becoming increasingly popular for their low energy requirement in real-world tasks with accuracy comparable to traditional ANNs. SNN training algorithms face the loss of gradient information and non-differentiability due to the Heaviside function in minimizing the model loss over model parameters. To circumvent this problem, the surrogate method employs a differentiable approximation of the Heaviside function in the backward pass, while the forward pass continues to use the Heaviside as the spiking function. We propose to use the zeroth-order technique at the local or neuron level in training SNNs, motivated by its regularizing and potential energy-efficient effects and establish a theoretical connection between it and the existing surrogate methods. We perform experimental validation of the technique on standard static datasets (CIFAR-10, CIFAR-100, ImageNet-100) and neuromorphic datasets (DVS-CIFAR-10, DVS-Gesture, N-Caltech-101, NCARS) and obtain results that offer improvement over the state-of-the-art results. The proposed method also lends itself to efficient implementations of the back-propagation method, which could provide 3-4 times overall speedup in training time. The code is available at \url{https://github.com/BhaskarMukhoty/LocalZO}.
Bhaskar Mukhoty, Velibor Bojkovic, William de Vazelhes, Xiaohan Zhao, Giulia De Masi, Huan Xiong, Bin Gu 0001
NeurIPS5
2023 A Universal Multimode (Acoustic, Magnetic Induction, Optical, RF) Software Defined Modem Architecture for Underwater Communication
abstract
In this paper, a Universal Underwater Software Defined Modem (UniSDM) architecture is proposed that may operate in different modes (acoustic, magnetic induction, optical and RF), in order to utilize the advantages of each mode and accordingly satisfy the requirements of many latest use cases in underwater communication systems. A detailed description of the novel UniSDM architecture is presented first. The novelty of this architecture is its flexibility, i.e., allowing the designers to produce a device that may include any type of modes operating seamlessly and jointly by exchanging data, control and synchronization. Many challenges, including high system costs and coordination between different modes, are addressed in the paper. Moreover, numerical evaluation is conducted to assess the performance of the proposed UniSDM architecture. Finally, the performance evaluation shows that the utilization of the UniSDM allows to decrease the transmission latency and improve the energy efficiency, while maintaining high reliability and robustness in underwater communication systems.
Igor V. Zhilin, Osama M. Bushnaq, Giulia De Masi, Enrico Natalizio, Ian F. Akyildiz
IEEE Trans. Wirel. Commun.3
2022 Automatic Network Slicing for Multi-Mode Internet of Underwater Things (MM-IoUT)
abstract
In recent years, many underwater communication applications have been proposed and tested, leading to the Internet of underwater things (IoUT) concept. In the IoUT, sensors may be deployed individually on sea surface and seabed as well as in water at different depths. They also may be integrated into underwater items such as fish, plants, autonomous underwater vehicles (AUVs), remotely operated underwater vehicles (ROVs), and divers. Depending on the application, different connectivity requirements must be satisfied such as data rate, latency, and reliability. However, the underwater communication channels pose a significant challenge to meet these requirements. To accommodate various applications with different service level agreements (SLAs), a multimode (acoustic, optical, and magnetic induction (MI)) communication system is proposed to take advantage of the modes' complementary features. Furthermore, an automatic network slicing (ANS) solution is proposed to provide globally optimized resource management, enhanced quality of service (QoS), simplified network operation, reduced deployment cost, and functional isolation of services. Taking the channel characteristics of the acoustic, optical, and MI communication modes into consideration, an optimization problem is formulated and solved for multimode IoUT (MM-IoUT) to enable admission control, routing, and dynamic resource allocation based on the different SLAs. Numerical analysis is conducted to verify the proposed solution and evaluate its performance.
Osama M. Bushnaq, Igor V. Zhilin, Giulia De Masi, Enrico Natalizio, Ian F. Akyildiz
GLOBECOM3
2021 A bio-inspired spatial defence strategy for collective decision making in self-organized swarms
abstract
In collective decision-making, individuals in a swarm reach consensus on a decision using only local interactions without any centralized control. In the context of the best-of-n problem - characterized by n discrete alternatives - it has been shown that consensus to the best option can be reached if individuals disseminate that option more than the other options. Besides being used as a mechanism to modulate positive feedback, long dissemination times could potentially also be used in an adversarial way, whereby adversarial swarms could infiltrate the system and propagate bad decisions using aggressive dissemination strategies. Motivated by the above scenario, in this paper we propose a bio-inspired defence strategy that allows the swarm to be resilient against options that can be disseminated for longer times. This strategy mainly consists in reducing the mobility of the agents that are associated to options disseminated for a shorter amount of time, allowing the swarm to converge to this option. We study the effectiveness of this strategy using two classical decision mechanisms, the voter model and the majority rule, showing that the majority rule is necessary in our setting for this strategy to work. The strategy has also been validated on a real Kilobots experiment.
Judhi Prasetyo, Giulia De Masi, Raina Zakir, Muhanad H. Mohammed Alkilabi, Elio Tuci, Eliseo Ferrante
GECCO2