Alessandro Bruno

dblp:97/8841 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-0707-6131ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Closing the Foveal Gap: Perceptually Grounded Scanpath Comparison with Disc IoU
abstract
Scanpath comparison metrics universally represent fixations as dimensionless (x, y) points, despite the human fovea subtending approximately 1° ∼ 2° of visual angle. This mismatch penalises spatially proximate fixations that process identical visual content. We propose FDISS (Foveal Disc IoU Scanpath Score), a metric that models each fixation as a foveal disc and computes similarity via bidirectional nearest-neighbour IoU matching, yielding interpretable Precision, Recall, and F1 scores. FDISS introduces a biologically grounded perceptual tolerance zone absent from all existing metrics and requires no free parameters at runtime.
Mohamed Amine Kerkouri, Marouane Tliba, Abdellah Zakaria Sellam, Cosimo Distante, Alessandro Bruno, Aladine Chetouani
ETRA5
2026 What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metrics
abstract
Scanpath similarity metrics are central to eye-movement research, yet existing methods predominantly evaluate spatial and temporal alignment while neglecting semantic equivalence between attended image regions. We present a semantic scanpath similarity framework that integrates vision-language models (VLMs) into eye-tracking analysis. Each fixation is encoded under controlled visual context (patch-based and marker-based strategies) and transformed into concise textual descriptions, which are aggregated into scanpath-level representations. Semantic similarity is then computed using embedding-based and lexical NLP metrics and compared against established spatial measures, including MultiMatch and DTW. Experiments on free-viewing eye-tracking data demonstrate that semantic similarity captures partially independent variance from geometric alignment, revealing cases of high content agreement despite spatial divergence. We further analyze the impact of contextual encoding on description fidelity and metric stability. Our findings suggest that multimodal foundation models enable interpretable, content-aware extensions of classical scanpath analysis, providing a complementary dimension for gaze research within the ETRA community.
Mohamed Amine Kerkouri, Marouane Tliba, Bin Wang 0068, Aladine Chetouani, Ulas Bagci, Alessandro Bruno
ETRA6
2025 Lung sound disease detection using attention over pre-trained efficientnet architecture
Anuja Nair, Himanshu Vadher, Pal Patel, Tarjni Vyas, Chintan M. Bhatt, Alessandro Bruno
Multim. Tools Appl.6
2024 AVAtt : Art Visual Attention dataset for diverse painting styles
abstract
Preserving cultural heritage is paramount for societal and historical identity. Paintings, spanning ancient to modern eras, are pivotal subjects under constant scrutiny. As art reflects human creativity, studying human visual behavior toward paintings becomes increasingly vital. Thus, we introduce the AVAtt dataset, providing eye movement data for a diverse collection of painting styles across various ages and geographical origins. This dataset aims to facilitate the development and evaluation of computational saliency and scanpath prediction methods in the unique domain of painting (Dataset available at Github).
Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Alessandro Bruno
ETRA4
2024 Perceptual Evaluation of Masked AutoEncoder Emergent Properties Through Eye-Tracking-Based Policy
abstract
The advancement of image restoration, especially in reconstructing missing or damaged image areas, has benefited significantly from self-supervised learning techniques, notably through the recent Masked Auto Encoder (MAE) strategy. In this project, we leverage eye-tracking data to enhance image reconstruction quality, and more specifically, with fixation-based saliency combined with the MAE strategy. By examining the emergent properties of representation learning and drawing parallels to human perceptual observation, we focus on how eye-tracking data informs the selection of image patches for reconstruction, aligning computational methods with human visual perception. Our findings reveal the potential of integrating eye-tracking insights to improve the accuracy and perceptual relevance of self-supervised learning models in computer vision. This study thus underscores the synergy between computational image restoration methods and human perception, facilitated by eye-tracking technology, opening new directions and insights for both fields. Our experiments are available for reproducibility in this GitHub Repository.
Marouane Tliba, Mohamed Amine Kerkouri, Aladine Chetouani, Alessandro Bruno, Mohammed El Hassouni, Arzu Çöltekin
ETRA4
2024 Special issue: Multimedia data analysis for smart city environment safety
Alessandro Bruno, Aladine Chetouani, Zoheir A. Sabeur, Marouane Tliba, Evangelos Maltezos, Miguel Gonzalez San Emeterio
Multim. Tools Appl.1
2024 Object Counting via Group and Graph Attention Network
abstract
Object counting, defined as the task of accurately predicting the number of objects in static images or videos, has recently attracted considerable interest. However, the unavoidable presence of background noise prevents counting performance from advancing further. To address this issue, we created a group and graph attention network (GGANet) for dense object counting. GGANet is an encoder-decoder architecture incorporating a group channel attention (GCA) module and a learnable graph attention (LGA) module. The GCA module groups the feature map into several subfeatures, each of which is assigned an attention factor through the identical channel attention. The LGA module views the feature map as a graph structure in which the different channels represent diverse feature vertices, and the responses between channels represent edges. The GCA and LGA modules jointly avoid the interference of irrelevant pixels and suppress the background noise. Experiments are conducted on four crowd-counting datasets, two vehicle-counting datasets, one remote-sensing counting dataset, and one few-shot object-counting dataset. Comparative results prove that the proposed GGANet achieves superior counting performance.
Mingliang Gao 0001, Guofeng Zou, Alessandro Bruno, Abdellah Chehri, Gwanggil Jeon
IEEE Trans. Neural Networks Learn. Syst.4
2023 Automatic diagnosis of knee osteoarthritis severity using Swin transformer
abstract
Knee osteoarthritis (KOA) is a widespread condition that can cause chronic pain and stiffness in the knee joint. Early detection and diagnosis are crucial for successful clinical intervention and management to prevent severe complications, such as loss of mobility. In this paper, we propose an automated approach that employs the Swin Transformer to predict the severity of KOA. Our model uses publicly available radiographic datasets with Kellgren and Lawrence scores to enable early detection and severity assessment. To improve the accuracy of our model, we employ a multi-prediction head architecture that utilizes multi-layer perceptron classifiers. Additionally, we introduce a novel training approach that reduces the data drift between multiple datasets to ensure the generalization ability of the model. The results of our experiments demonstrate the effectiveness and feasibility of our approach in predicting KOA severity accurately.
Aymen Sekhri, Mohamed Amine Kerkouri, Aladine Chetouani, Marouane Tliba, Yassine Nasser, Rachid Jennane, Alessandro Bruno
CBMI7
2023 Detecting colour vision deficiencies via Webcam-based Eye-tracking: A case study
abstract
Webcam-based eye-tracking platforms have recently re-emerged due to improvements in machine learning-supported calibration processes and offer a scalable option for conducting eye movement studies. Although not yet comparable to the infrared-based ones regarding accuracy and frequency, some compelling performances have been observed, especially in those scenarios with medium-sized AOI (Areas of Interest) in images. In this study, we test the reliability of webcam-based eye-tracking on a specific task: Eye movement distribution analysis for CVD (Colour Vision Deficiency) detection. We introduce a new publicly available eye movement dataset based on a pilot study (n=12) on images with dominant red colour (previously shown to be difficult with dichromatic AOI to investigate CVD by comparing attention patterns obtained in webcam eye-tracking sessions). We hypothesized that webcam eye tracking without infrared support could detect differing attention patterns between CVD and non-CVD participants and observed statistically significant differences, allowing the retention of our hypothesis.
Alessandro Bruno, Marouane Tliba, Mohamed Amine Kerkouri, Aladine Chetouani, Carlo Calogero Giunta, Arzu Çöltekin
ETRA1
2023 An Inter-Observer Consistent Deep Adversarial Training for Visual Scanpath Prediction
abstract
The visual scanpath represents the fundamental concept upon which visual attention research is based. As a result, the ability to predict them has emerged as a crucial task in recent years. It is represented as a sequence of points through which the human gaze moves while exploring a scene. In this paper, we propose an inter-observer consistent adversarial training approach for scanpath prediction through a lightweight deep neural network. The proposed method employs a discriminative neural network as a dynamic loss that better models the natural stochastic phenomenon while maintaining consistency between the distributions related to the subjective nature of scanpaths traversed by different observers. The competitiveness of our approach against state-of-the-art methods is shown through a testing phase.
Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Alessandro Bruno
ICIP4
2022 A domain adaptive deep learning solution for scanpath prediction of paintings
abstract
Cultural heritage understanding and preservation is an important issue for society as it represents a fundamental aspect of its identity. Paintings represent a significant part of cultural heritage, and are the subject of study continuously. However, the way viewers perceive paintings is strictly related to the so-called HVS (Human Vision System) behaviour. This paper focuses on the eye-movement analysis of viewers during the visual experience of a certain number of paintings. In further details, we introduce a new approach to predicting human visual attention, which impacts several cognitive functions for humans, including the fundamental understanding of a scene, and then extend it to painting images. The proposed new architecture ingests images and returns scanpaths, a sequence of points featuring a high likelihood of catching viewers’ attention. We use an FCNN (Fully Convolutional Neural Network), in which we exploit a differentiable channel-wise selection and Soft-Argmax modules. We also incorporate learnable Gaussian distributions onto the network bottleneck to simulate visual attention process bias in natural scene images. Furthermore, to reduce the effect of shifts between different domains (i.e. natural images, painting), we urge the model to learn unsupervised general features from other domains using a gradient reversal classifier. The results obtained by our model outperform existing state-of-the-art ones in terms of accuracy and efficiency.
Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Alessandro Bruno
CBMI4
2021 Toward a head movement-based system for multilayer digital content exploration
abstract
Abstract In this article, we propose a novel technique based on Head Movement tracking to explore multilayer digital content. We extend an existing method by Kazemi et al. dealing with the extraction of facial landmarks to define the “head‐gaze” of the user. We use the “head‐gaze” to calculate the users' on‐screen coordinates. Hovering the cursor over an interactive area for a given time threshold allows users to explore the next layer contents. Our experimental sessions allowed us to measure the technique's level of control and usability. Our results were promising, and users were able to interact with considerably small regions. Furthermore, our lightweight method can be used with a low‐cost camera or webcam and a wide range of screen sizes and distances.
Alessandro Bruno, Morgan Moore, Jinglu Zhang, Stéphane Lancette, Ville P. Ward, Jian Chang 0001
Comput. Animat. Virtual Worlds1
2015 An Unsupervised Method for Suspicious Regions Detection in Mammogram Images
Marco Insalaco, Alessandro Bruno, Alfonso Farruggia, Salvatore Vitabile, Edoardo Ardizzone
ICPRAM (2)2
2015 Copy-Move Forgery Detection by Matching Triangles of Keypoints
abstract
Copy-move forgery is one of the most common types of tampering for digital images. Detection methods generally use block-matching approaches, which first divide the image into overlapping blocks and then extract and compare features to find similar ones, or point-based approaches, in which relevant keypoints are extracted and matched to each other to find similar areas. In this paper, we present a very novel hybrid approach, which compares triangles rather than blocks, or single points. Interest points are extracted from the image, and objects are modeled as a set of connected triangles built onto these points. Triangles are matched according to their shapes (inner angles), their content (color information), and the local feature vectors extracted onto the vertices of the triangles. Our methods are designed to be robust to geometric transformations. Results are compared with a state-of-the-art block matching method and a point-based method. Furthermore, our data set is available for use by academic researchers.
Edoardo Ardizzone, Alessandro Bruno, Giuseppe Mazzola
IEEE Trans. Inf. Forensics Secur.2
2014 Video Object Recognition and Modeling by SIFT Matching Optimization
abstract
In this paper we present a novel technique for object modeling and object recognition in video. Given a set of videos containing 360 degrees views of objects we compute a model for each object, then we analyze short videos to determine if the object depicted in the video is one of the modeled objects. The object model is built from a video spanning a 360 degree view of the object taken against a uniform background. In order to create the object model, the proposed techniques selects a few representative frames from each video and local features of such frames. The object recognition is performed selecting a few frames from the query video, extracting local features from each frame and looking for matches in all the representative frames constituting the models of all the objects. If the number of matches exceed a fixed threshold the corresponding object is considered the recognized objects .To evaluate our approach we acquired a dataset of 25 videos representing 25 different objects and used these videos to build the objects model. Then we took 25 test videos containing only one of the known objects and 5 videos containing only unknown objects. Experiments showed that, despite a significant compression in the model, recognition results are satisfactory.
Alessandro Bruno, Luca Greco 0002, Marco La Cascia
ICPRAM1
2013 Object Recognition and Modeling Using SIFT Features
Alessandro Bruno, Luca Greco 0002, Marco La Cascia
ACIVS1
2013 Scale detection via keypoint density maps in regular or near-regular textures
Edoardo Ardizzone, Alessandro Bruno, Giuseppe Mazzola
Pattern Recognit. Lett.2
2010 Detecting multiple copies in tampered images
abstract
Copy-move forgeries are parts of the image that are duplicated elsewhere into the same image, often after being modified by geometrical transformations. In this paper we present a method to detect these image alterations, using a SIFT-based approach. First we describe a state of the art SIFT-point matching method, which inspired our algorithm, then we compare it with our SIFT-based approach, which consists of three parts: keypoint clustering, cluster matching, and texture analysis. The goal is to find copies of the same object, i.e. clusters of points, rather than points that match. Cluster matching proves to give better results than single point matching, since it returns a complete and coherent comparison between copied objects. At last, textures of matching areas are analyzed and compared to validate results and to eliminate false positives.
Edoardo Ardizzone, Alessandro Bruno, Giuseppe Mazzola
ICIP2