EDBT 2026 Demo / reviewers in the wild / expert
Alessandro Bruno
dblp:97/8841
· DBLP profile ↗
18ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-0707-6131ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Closing the Foveal Gap: Perceptually Grounded Scanpath Comparison with Disc IoUabstractScanpath comparison metrics universally represent fixations as dimensionless (x, y) points, despite the human fovea subtending approximately 1° ∼ 2° of visual angle. This mismatch penalises spatially proximate fixations that process identical visual content. We propose FDISS (Foveal Disc IoU Scanpath Score), a metric that models each fixation as a foveal disc and computes similarity via bidirectional nearest-neighbour IoU matching, yielding interpretable Precision, Recall, and F1 scores. FDISS introduces a biologically grounded perceptual tolerance zone absent from all existing metrics and requires no free parameters at runtime. Mohamed Amine Kerkouri, Marouane Tliba, Abdellah Zakaria Sellam, Cosimo Distante, Alessandro Bruno, Aladine Chetouani |
ETRA | 5 |
| 2026 | What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metricsabstractScanpath similarity metrics are central to eye-movement research, yet existing methods predominantly evaluate spatial and temporal alignment while neglecting semantic equivalence between attended image regions. We present a semantic scanpath similarity framework that integrates vision-language models (VLMs) into eye-tracking analysis. Each fixation is encoded under controlled visual context (patch-based and marker-based strategies) and transformed into concise textual descriptions, which are aggregated into scanpath-level representations. Semantic similarity is then computed using embedding-based and lexical NLP metrics and compared against established spatial measures, including MultiMatch and DTW. Experiments on free-viewing eye-tracking data demonstrate that semantic similarity captures partially independent variance from geometric alignment, revealing cases of high content agreement despite spatial divergence. We further analyze the impact of contextual encoding on description fidelity and metric stability. Our findings suggest that multimodal foundation models enable interpretable, content-aware extensions of classical scanpath analysis, providing a complementary dimension for gaze research within the ETRA community. Mohamed Amine Kerkouri, Marouane Tliba, Bin Wang 0068, Aladine Chetouani, Ulas Bagci, Alessandro Bruno |
ETRA | 6 |
| 2025 | Lung sound disease detection using attention over pre-trained efficientnet architecture
Anuja Nair, Himanshu Vadher, Pal Patel, Tarjni Vyas, Chintan M. Bhatt, Alessandro Bruno |
Multim. Tools Appl. | 6 |
| 2024 | AVAtt : Art Visual Attention dataset for diverse painting stylesabstractPreserving cultural heritage is paramount for societal and historical identity. Paintings, spanning ancient to modern eras, are pivotal subjects under constant scrutiny. As art reflects human creativity, studying human visual behavior toward paintings becomes increasingly vital. Thus, we introduce the AVAtt dataset, providing eye movement data for a diverse collection of painting styles across various ages and geographical origins. This dataset aims to facilitate the development and evaluation of computational saliency and scanpath prediction methods in the unique domain of painting (Dataset available at Github). Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Alessandro Bruno |
ETRA | 4 |
| 2024 | Perceptual Evaluation of Masked AutoEncoder Emergent Properties Through Eye-Tracking-Based PolicyabstractThe advancement of image restoration, especially in reconstructing missing or damaged image areas, has benefited significantly from self-supervised learning techniques, notably through the recent Masked Auto Encoder (MAE) strategy. In this project, we leverage eye-tracking data to enhance image reconstruction quality, and more specifically, with fixation-based saliency combined with the MAE strategy. By examining the emergent properties of representation learning and drawing parallels to human perceptual observation, we focus on how eye-tracking data informs the selection of image patches for reconstruction, aligning computational methods with human visual perception. Our findings reveal the potential of integrating eye-tracking insights to improve the accuracy and perceptual relevance of self-supervised learning models in computer vision. This study thus underscores the synergy between computational image restoration methods and human perception, facilitated by eye-tracking technology, opening new directions and insights for both fields. Our experiments are available for reproducibility in this GitHub Repository. Marouane Tliba, Mohamed Amine Kerkouri, Aladine Chetouani, Alessandro Bruno, Mohammed El Hassouni, Arzu Çöltekin |
ETRA | 4 |
| 2024 | Special issue: Multimedia data analysis for smart city environment safety
Alessandro Bruno, Aladine Chetouani, Zoheir A. Sabeur, Marouane Tliba, Evangelos Maltezos, Miguel Gonzalez San Emeterio |
Multim. Tools Appl. | 1 |
| 2024 | Object Counting via Group and Graph Attention NetworkabstractObject counting, defined as the task of accurately predicting the number of objects in static images or videos, has recently attracted considerable interest. However, the unavoidable presence of background noise prevents counting performance from advancing further. To address this issue, we created a group and graph attention network (GGANet) for dense object counting. GGANet is an encoder-decoder architecture incorporating a group channel attention (GCA) module and a learnable graph attention (LGA) module. The GCA module groups the feature map into several subfeatures, each of which is assigned an attention factor through the identical channel attention. The LGA module views the feature map as a graph structure in which the different channels represent diverse feature vertices, and the responses between channels represent edges. The GCA and LGA modules jointly avoid the interference of irrelevant pixels and suppress the background noise. Experiments are conducted on four crowd-counting datasets, two vehicle-counting datasets, one remote-sensing counting dataset, and one few-shot object-counting dataset. Comparative results prove that the proposed GGANet achieves superior counting performance. Mingliang Gao 0001, Guofeng Zou, Alessandro Bruno, Abdellah Chehri, Gwanggil Jeon |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Automatic diagnosis of knee osteoarthritis severity using Swin transformerabstractKnee osteoarthritis (KOA) is a widespread condition that can cause chronic pain and stiffness in the knee joint. Early detection and diagnosis are crucial for successful clinical intervention and management to prevent severe complications, such as loss of mobility. In this paper, we propose an automated approach that employs the Swin Transformer to predict the severity of KOA. Our model uses publicly available radiographic datasets with Kellgren and Lawrence scores to enable early detection and severity assessment. To improve the accuracy of our model, we employ a multi-prediction head architecture that utilizes multi-layer perceptron classifiers. Additionally, we introduce a novel training approach that reduces the data drift between multiple datasets to ensure the generalization ability of the model. The results of our experiments demonstrate the effectiveness and feasibility of our approach in predicting KOA severity accurately. Aymen Sekhri, Mohamed Amine Kerkouri, Aladine Chetouani, Marouane Tliba, Yassine Nasser, Rachid Jennane, Alessandro Bruno |
CBMI | 7 |
| 2023 | Detecting colour vision deficiencies via Webcam-based Eye-tracking: A case studyabstractWebcam-based eye-tracking platforms have recently re-emerged due to improvements in machine learning-supported calibration processes and offer a scalable option for conducting eye movement studies. Although not yet comparable to the infrared-based ones regarding accuracy and frequency, some compelling performances have been observed, especially in those scenarios with medium-sized AOI (Areas of Interest) in images. In this study, we test the reliability of webcam-based eye-tracking on a specific task: Eye movement distribution analysis for CVD (Colour Vision Deficiency) detection. We introduce a new publicly available eye movement dataset based on a pilot study (n=12) on images with dominant red colour (previously shown to be difficult with dichromatic AOI to investigate CVD by comparing attention patterns obtained in webcam eye-tracking sessions). We hypothesized that webcam eye tracking without infrared support could detect differing attention patterns between CVD and non-CVD participants and observed statistically significant differences, allowing the retention of our hypothesis. Alessandro Bruno, Marouane Tliba, Mohamed Amine Kerkouri, Aladine Chetouani, Carlo Calogero Giunta, Arzu Çöltekin |
ETRA | 1 |
| 2023 | An Inter-Observer Consistent Deep Adversarial Training for Visual Scanpath PredictionabstractThe visual scanpath represents the fundamental concept upon which visual attention research is based. As a result, the ability to predict them has emerged as a crucial task in recent years. It is represented as a sequence of points through which the human gaze moves while exploring a scene. In this paper, we propose an inter-observer consistent adversarial training approach for scanpath prediction through a lightweight deep neural network. The proposed method employs a discriminative neural network as a dynamic loss that better models the natural stochastic phenomenon while maintaining consistency between the distributions related to the subjective nature of scanpaths traversed by different observers. The competitiveness of our approach against state-of-the-art methods is shown through a testing phase. Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Alessandro Bruno |
ICIP | 4 |
| 2022 | A domain adaptive deep learning solution for scanpath prediction of paintingsabstractCultural heritage understanding and preservation is an important issue for society as it represents a fundamental aspect of its identity. Paintings represent a significant part of cultural heritage, and are the subject of study continuously. However, the way viewers perceive paintings is strictly related to the so-called HVS (Human Vision System) behaviour. This paper focuses on the eye-movement analysis of viewers during the visual experience of a certain number of paintings. In further details, we introduce a new approach to predicting human visual attention, which impacts several cognitive functions for humans, including the fundamental understanding of a scene, and then extend it to painting images. The proposed new architecture ingests images and returns scanpaths, a sequence of points featuring a high likelihood of catching viewers’ attention. We use an FCNN (Fully Convolutional Neural Network), in which we exploit a differentiable channel-wise selection and Soft-Argmax modules. We also incorporate learnable Gaussian distributions onto the network bottleneck to simulate visual attention process bias in natural scene images. Furthermore, to reduce the effect of shifts between different domains (i.e. natural images, painting), we urge the model to learn unsupervised general features from other domains using a gradient reversal classifier. The results obtained by our model outperform existing state-of-the-art ones in terms of accuracy and efficiency. Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Alessandro Bruno |
CBMI | 4 |
| 2021 | Toward a head movement-based system for multilayer digital content explorationabstractAbstract In this article, we propose a novel technique based on Head Movement tracking to explore multilayer digital content. We extend an existing method by Kazemi et al. dealing with the extraction of facial landmarks to define the “head‐gaze” of the user. We use the “head‐gaze” to calculate the users' on‐screen coordinates. Hovering the cursor over an interactive area for a given time threshold allows users to explore the next layer contents. Our experimental sessions allowed us to measure the technique's level of control and usability. Our results were promising, and users were able to interact with considerably small regions. Furthermore, our lightweight method can be used with a low‐cost camera or webcam and a wide range of screen sizes and distances. Alessandro Bruno, Morgan Moore, Jinglu Zhang, Stéphane Lancette, Ville P. Ward, Jian Chang 0001 |
Comput. Animat. Virtual Worlds | 1 |
| 2015 | An Unsupervised Method for Suspicious Regions Detection in Mammogram Images
Marco Insalaco, Alessandro Bruno, Alfonso Farruggia, Salvatore Vitabile, Edoardo Ardizzone |
ICPRAM (2) | 2 |
| 2015 | Copy-Move Forgery Detection by Matching Triangles of KeypointsabstractCopy-move forgery is one of the most common types of tampering for digital images. Detection methods generally use block-matching approaches, which first divide the image into overlapping blocks and then extract and compare features to find similar ones, or point-based approaches, in which relevant keypoints are extracted and matched to each other to find similar areas. In this paper, we present a very novel hybrid approach, which compares triangles rather than blocks, or single points. Interest points are extracted from the image, and objects are modeled as a set of connected triangles built onto these points. Triangles are matched according to their shapes (inner angles), their content (color information), and the local feature vectors extracted onto the vertices of the triangles. Our methods are designed to be robust to geometric transformations. Results are compared with a state-of-the-art block matching method and a point-based method. Furthermore, our data set is available for use by academic researchers. Edoardo Ardizzone, Alessandro Bruno, Giuseppe Mazzola |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2014 | Video Object Recognition and Modeling by SIFT Matching OptimizationabstractIn this paper we present a novel technique for object modeling and object recognition in video. Given a set of videos containing 360 degrees views of objects we compute a model for each object, then we analyze short videos to determine if the object depicted in the video is one of the modeled objects. The object model is built from a video spanning a 360 degree view of the object taken against a uniform background. In order to create the object model, the proposed techniques selects a few representative frames from each video and local features of such frames. The object recognition is performed selecting a few frames from the query video, extracting local features from each frame and looking for matches in all the representative frames constituting the models of all the objects. If the number of matches exceed a fixed threshold the corresponding object is considered the recognized objects .To evaluate our approach we acquired a dataset of 25 videos representing 25 different objects and used these videos to build the objects model. Then we took 25 test videos containing only one of the known objects and 5 videos containing only unknown objects. Experiments showed that, despite a significant compression in the model, recognition results are satisfactory. Alessandro Bruno, Luca Greco 0002, Marco La Cascia |
ICPRAM | 1 |
| 2013 | Object Recognition and Modeling Using SIFT Features
Alessandro Bruno, Luca Greco 0002, Marco La Cascia |
ACIVS | 1 |
| 2013 | Scale detection via keypoint density maps in regular or near-regular textures
Edoardo Ardizzone, Alessandro Bruno, Giuseppe Mazzola |
Pattern Recognit. Lett. | 2 |
| 2010 | Detecting multiple copies in tampered imagesabstractCopy-move forgeries are parts of the image that are duplicated elsewhere into the same image, often after being modified by geometrical transformations. In this paper we present a method to detect these image alterations, using a SIFT-based approach. First we describe a state of the art SIFT-point matching method, which inspired our algorithm, then we compare it with our SIFT-based approach, which consists of three parts: keypoint clustering, cluster matching, and texture analysis. The goal is to find copies of the same object, i.e. clusters of points, rather than points that match. Cluster matching proves to give better results than single point matching, since it returns a complete and coherent comparison between copied objects. At last, textures of matching areas are analyzed and compared to validate results and to eliminate false positives. Edoardo Ardizzone, Alessandro Bruno, Giuseppe Mazzola |
ICIP | 2 |