VLDB 2026 Research / reviewers in the wild / expert
Simone Mentasti
dblp:242/5430
· DBLP profile ↗
18ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0001-7059-413XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hypernetwork-driven Weight Adaptation for Personalized PSOG Eye-Trackers
Flavia Nicotri, Marco Paracchini, Simone Mentasti, Giulio Marano, Luca Merigo, Marco Marcon |
ETRA | 3 |
| 2026 | Lightweight Neural Networks for Event-Based 3D Gaze Estimation on Wearable Devices
Aaron Tognoli, Chiara Fossà, Andrea Simpsi, Andrea Aspesi, Luca Merigo, Matteo Matteucci, Simone Mentasti, Marco Cannici |
ETRA | 7 |
| 2026 | DGCMam: Fusing Distance Graph Convolution and Mamba for Skeleton-based Action Recognition
Juncen Long, Gianluca Bardaro, Simone Mentasti, Matteo Matteucci |
FG | 3 |
| 2026 | EgoAfford: Affordance-Aware Zero-Shot Open-Vocabulary Egocentric Action Recognition
Davide Gesualdi, Riccardo Santambrogio, Francesca Palermo, Chiara Plizzari, Simone Mentasti, Matteo Matteucci |
ICPR (16) | 5 |
| 2026 | Continuous Online Action Detection from Egocentric Videos
Riccardo Santambrogio, Chiara Plizzari, Francesca Palermo, Simone Mentasti, Matteo Matteucci |
ICPR (16) | 4 |
| 2026 | Overcoming Data Scarcity for Event-Based Pupil Tracking with Synthetic and Unlabeled Data ETRA018abstractSmart eyewear is emerging as an always-on platform capable of perceiving the environment and inferring intent through eye movements, making pupil tracking essential for personalized interaction. However, reliable tracking on wearable hardware remains challenging due to strict power limits and scarce annotated data. Event-based cameras offer a low-power, microsecond-latency solution, but labeled recordings still remain limited. We address this issue with a training framework that combines limited annotated real data with synthetic events and unlabeled real recordings, learning event-based pupil trackers with strong real-world generalization. We pair the U2Eyes tool with the v2e event camera simulator to generate realistic event streams, showing that networks trained on these events exhibit smaller sim-to-real gaps than networks trained on synthetic images. Moreover, our training procedure further bridges this gap, enabling our models to outperform networks trained exclusively on real data across all benchmarks, advancing toward more robust event-based eye tracking on wearable platforms. Aaron Tognoli, Andrea Simpsi, Andrea Aspesi, Marco Cannici, Luca Merigo, Matteo Matteucci, Simone Mentasti |
Proc. ACM Hum. Comput. Interact. | 7 |
| 2025 | EETnet: a CNN for Gaze Detection and Tracking for Smart-EyewearabstractEvent-based cameras are becoming a popular solution for efficient, low-power eye tracking. Due to the sparse and asynchronous nature of event data, they require less processing power and offer latencies in the microsecond range. However, many existing solutions are limited to validation on powerful GPUs, with no deployment on real embedded devices. In this paper, we present EETnet, a convolutional neural network designed for eye tracking using purely event-based data, capable of running on microcontrollers with limited resources. Additionally, we outline a methodology to train, evaluate, and quantize the network using a public dataset. Finally, we propose two versions of the architecture: a classification model that detects the pupil on a grid superimposed on the original image, and a regression model that operates at the pixel level. Andrea Aspesi, Andrea Simpsi, Aaron Tognoli, Simone Mentasti, Luca Merigo, Matteo Matteucci |
IJCNN | 4 |
| 2025 | Neuromorphic eye tracking: a surveyabstractEvent-based eye tracking represents an innovative approach to analyzing eye movements, leveraging event cameras’ high temporal resolution and asynchronous nature. This paper provides a comprehensive field review, focusing on datasets, algorithms, and challenges. It examines key datasets, highlighting their diversity and limitations, and categorizes algorithms into two core approaches: frame-based and spiking neural network (SNN)-based approaches. Frame-based methods leverage traditional techniques and deep learning, while SNNs represent an emerging field offering biologically inspired, energy-efficient solutions. However, the computational demands of emerging deep-learning methods raise questions about their feasibility for deployment on consumer-grade devices, particularly embedded systems like smart eyewear. This review provides a structured analysis of current advancements, datasets, and challenges, offering insights to guide the development of efficient and deployable event-based eye-tracking systems. Simone Mentasti, Andrea Simpsi, Andrea Aspesi, Aaron Tognoli, Luca Merigo, Matteo Matteucci |
IJCNN | 1 |
| 2025 | A Spatio-temporal Graph Network Allowing Incomplete Trajectory Input for Pedestrian Trajectory Prediction
Juncen Long, Gianluca Bardaro, Simone Mentasti, Matteo Matteucci |
SMC | 3 |
| 2025 | A neural approach to the Turing Test: The role of emotionsabstractAs is well known, the Turing Test proposes the possibility of distinguishing the behavior of a machine from that of a human being through an experimental session. The Turing Test assesses whether a person asking questions to two different entities, can tell from their answers which of them is the human being and which is the machine. With the progress of Artificial Intelligence, the number of contexts in which the capacities of response of a machine will be indistinguishable from those of a human being is expected to increase rapidly. In order to configure a Turing Test in which it is possible to distinguish human behavior from machine behavior independently from the advances of Artificial Intelligence, at least in the short-medium term, it would be important to base it not on the differences between man and machine in terms of performance and dialogue capacity, but on some specific characteristic of the human mind that cannot be reproduced by the machine even in principle. We studied a new kind of test based on the hypothesis that such characteristic of the human mind exists and can be made experimentally evident. This peculiar characteristic is the emotional content of human cognition and, more specifically, its link with memory enhancement. To validate this hypothesis we recorded the EEG signals of 39 subjects that underwent a specific test and analyzed their signals with a neural network able to label similar signal patterns with similar binary codes. The results showed that, with a statistically significant difference, the test participants more easily recognized images associated in the past with an emotional reaction than those not associated with such a reaction. This distinction in our view is not accessible to a software system, even AI-based, and a Turing Test based on this feature of the mind may make distinguishable human versus machine responses. Rita Pizzi, Hao Quan 0002, Matteo Matteucci, Simone Mentasti, Roberto Sassi |
Neural Networks | 4 |
| 2024 | Event-based eye tracking for smart eyewearabstractThis paper presents an innovative approach to gaze tracking in the context of smart eyewear, utilizing a fully event-based algorithm. Traditional gaze-tracking methods often rely on grayscale or infrared imaging, which can be computationally intensive and raise privacy concerns. Our research addresses these issues by developing an algorithm that exclusively uses data from event-based sensors, optimizing for the limited computational capabilities of smart eyewear. The system uses simple geometrical operations, enabling efficient real-time processing. Experimental results demonstrate the feasibility of this approach, offering a promising solution for gaze tracking in compact, computationally constrained devices. Despite certain limitations in accuracy due to optimization for efficiency, the research underscores the practicality of this approach for practical, privacy-conscious applications in smart eyewear technology. Simone Mentasti, Francesco Lattari, Riccardo Santambrogio, Gianmario Careddu, Matteo Matteucci |
ETRA | 1 |
| 2024 | LiDAR-Aided Cooperative Localization and Environmental Perception for CAVsabstractThis paper explores the potentialities of deploying vehicular Cooperative Positioning (CP) systems in urban scenarios utilizing real-world data collected via experimental campaigns. We examine the case of two prototype vehicles equipped with LiDAR sensors for perceiving their surrounding environment and with Global Navigation Satellite System (GNSS) receivers for positioning. The considered use case focuses on the cooperative detection of static landmarks, to be used for improving the vehicles’ GNSS positioning. The experimental campaign points out a severe degradation in ego vehicle localization performances due to complex multipath propagation experienced in the urban scenario. To cope with such a problem, we integrate into the CP system a compensation method able to mitigate the position bias originating from the adverse propagating conditions. Experimental results show that integrating the developed compensation into the CP solution enables an accurate detection of the landmark positions, leading to an enhancement of the vehicle localization accuracy. Akif Adas, Luca Barbieri, Satyesh Awasthi, Pietro Morri, Simone Mentasti, Stefano Arrigoni, Edoardo Sabbioni, Monica Nicoli |
FUSION | 5 |
| 2024 | RadarLCD: Learnable Radar-based Loop Closure Detection PipelineabstractLoop Closure Detection (LCD) is an essential task in robotics and computer vision, serving as a fundamental component for various applications across diverse domains. These applications encompass object recognition, image retrieval, and video analysis. LCD consists in identifying whether a robot has returned to a previously visited location, referred to as a loop, and then estimating the related roto-translation with respect to the analyzed location. Despite the numerous advantages of radar sensors, such as their ability to operate under diverse weather conditions and provide a wider range of view compared to other commonly used sensors (e.g., cameras or LiDARs), integrating radar data remains an arduous task due to intrinsic noise and distortion. To address this challenge, this research introduces RadarLCD, a novel supervised deep learning pipeline specifically designed for Loop Closure Detection using the FMCW Radar (Frequency Modulated Continuous Wave) sensor. RadarLCD, a learning-based LCD methodology explicitly designed for radar systems, makes a significant contribution by leveraging the pre-trained HERO (Hybrid Estimation Radar Odometry) model. Being originally developed for radar odometry, HERO’s features are used to select key points crucial for LCD tasks. The methodology undergoes evaluation across a variety of FMCW Radar dataset scenes, and it is compared to state-of-the-art systems such as Scan Context for Place Recognition and ICP for Loop Closure. The results demonstrate that RadarLCD surpasses the alternatives in multiple aspects of Loop Closure Detection. Mirko Usuelli, Matteo Frosi, Paolo Cudrano, Simone Mentasti, Matteo Matteucci |
IJCNN | 4 |
| 2024 | OptimusLine: Consistent Road Line Detection Through TimeabstractIn the field of autonomous vehicles, the detection of road line markings is a crucial yet versatile component. It provides real-time guidance for navigation and low-level vehicle control, while it also enables the generation of lane-level HD maps. These maps require high precision to provide low-level details to all future map users. At the same time, control-oriented detection pipelines require increased inference frequency and high robustness to be deployed on a safety-critical system. With this work, we present OptimusLine, a versatile line detection pipeline tackling with ease both scenarios. Built around a frame-by-frame transformer-based neural model operating in image segmentation, we show that OptimusLine achieves state-of-the-art performance and analyze its computational impact. To provide robustness to perturbations when deployed on an actual vehicle, OptimusLine introduces a scheme exploiting temporal links between consecutive frames. Enforcing temporal consistency on each new line prediction, OptimusLine can generate more robust line descriptions and produce an estimate of its prediction uncertainty. Paolo Cudrano, Simone Mentasti, Riccardo Erminio Filippo Cortelazzo, Matteo Matteucci |
IV | 2 |
| 2024 | Heterogeneous Data Fusion for Accurate Road User Tracking: A Distributed Multi-Sensor Collaborative ApproachabstractThis work presents the design and validation of a distributed multi-sensor object tracking algorithm designed to integrate heterogeneous sensory data from multiple static acquisition stations. The primary challenge addressed is the accurate tracking of targets in complex urban environments, where occlusions and the dynamic nature of traffic frequently hinder detection and tracking efforts. This challenge is particularly relevant in multimodal exchange areas, where vehicular traffic merges with heavy pedestrian and bicycle flow. We also address the scenario of delayed detection, which can easily occur when data from multiple stations are combined or when intensive data processing is performed. Our algorithm ensures high coverage and accuracy by maintaining dual Extended Kalman Filter states for each object, thus allowing for the assimilation of delayed detections and preserving optimal filter estimates at all times. The results of the proposed pipeline, tested using a digital twin of the Milano Bovisa Campus, demonstrate its efficacy, achieving high tracking precision across various scenarios and sensor combinations. Moreover, the results highlight the advantages of a distributed multi-sensor acquisition system compared to a single central station. Simone Mentasti, Alessandro Barbiero, Matteo Matteucci |
IV | 1 |
| 2023 | Fault Resistant Odometry Estimation using Message Passing Neural NetworkabstractMulti-modal sensor fusion constitutes an essential ingredient for safe autonomous navigation. In the last years, many works have improved the accuracy of Deep-Learning-based odometry estimators. However, the robustness of these algorithms to sensor failure or measurement degradation, which are very likely to happen during navigation, has been studied less extensively. Furthermore, works studying the robustness of the fusion modules are developed without modeling the correlation between sensor features, which is crucial to filter out features derived from noisy measurements and in sensor faults scenarios. To bridge this gap, in this paper, we propose a fault-resistant odometry estimator, which produces robust estimates even when the sensors completely fail, or measurements progressively degrade. Our framework models the correlation between the sensor embedding using Message Passing Neural Network (MPNN), a particular type of Graph Neural Network (GNN). A mask is then computed from the updated node features of the graph to weigh the multi-modal features computed from different sensors. We evaluate the proposed fusion strategy on the modified raw KITTI dataset with sensor degradation scenarios. Finally, we compare against state-of-the-art baselines based on trivial features concatenation and soft-fusion to demonstrate our method’s superiority in terms of accuracy and robustness to sensor degradation and failures. Pragyan Dahal, Simone Mentasti, Luca Paparusso, Stefano Arrigoni, Francesco Braghin |
IV | 2 |
| 2022 | Clothoidal Mapping of Road Line Markings for Autonomous Driving High-Definition MapsabstractLane-level HD maps are crucial for trajectory planning and control in current autonomous vehicles. For this reason, appropriate line models should be adopted to define them. Whereas mapping algorithms often rely on inaccurate representations, clothoid curves possess peculiar smoothness properties that make them desirable representations of road lines in control algorithms. We propose a multi-stage pipeline for the generation of lane-level HD maps from monocular vision relying on clothoidal spline models. We obtain measurements of the line positions using a line detection algorithm, and we exploit a graph-based optimization framework to reach an optimal fitting. An iterative greedy procedure reduces the model complexity removing unnecessary clothoids. We validate our system on a real-world dataset, which we make publicly available for further research at https://airlab.deib.polimi.it/datasets-and-tools/. Barbara Gallazzi, Paolo Cudrano, Matteo Frosi, Simone Mentasti, Matteo Matteucci |
IV | 4 |
| 2020 | Advances in centerline estimation for autonomous lateral controlabstractThe ability of autonomous vehicles to maintain an accurate trajectory within their road lane is crucial for safe operation. This requires detecting the road lines and estimating the car relative pose within its lane. Lateral lines are usually retrieved from camera images. Still, most of the works on line detection are limited to image mask retrieval and do not provide a usable representation in world coordinates. What we propose in this paper is a complete perception pipeline based on monocular vision and able to retrieve all the information required by a vehicle lateral control system: road lines equation, centerline, vehicle heading and lateral displacement. We evaluate our system by acquiring data with accurate geometric ground truth. To act as a benchmark for further research, we make this new dataset publicly available at http://airlab.deib.polimi.it/datasets/. Paolo Cudrano, Simone Mentasti, Matteo Matteucci, Mattia Bersani, Stefano Arrigoni, Federico Cheli |
IV | 2 |