VLDB 2026 Research / reviewers in the wild / expert
Matteo Fabbri
dblp:153/8126
· DBLP profile ↗
13ranked-venue papers
9as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Feeding the (short-video) feed: a design proposal for user control of social media recommender systems under the Digital Services ActabstractMost online platforms use recommender systems to reduce information overload and enhance user experience but do not usually address the implications of their influence on users from an ethical and legal perspective. The Digital Services Act (DSA) is the first supranational regulation that requires online platforms to explain the criteria underlying their recommender systems and to allow users to control them by modifying their parameters. However, the effectiveness of user control features is mainly dependent on the design of the interface: in fact, users may not be interested in using tools that, despite supporting their empowerment, increase the cognitive load of their experience. This paper introduces a controllable and transparent RS for short videos integrated into an interactive user interface, through which a preliminary user study was conducted to provide insights into how the DSA-informed control features of RSs can enhance users’ understanding and willingness to intervene on the recommendation process. After interacting with this platform, users participated in an evaluative survey featuring six dimensions: perceived control, perceived transparency, engagement with control features, perceived psychological workload, satisfaction, and impact on digital wellbeing. The findings indicate that: 1) users show the tendency not to use the control features when not prompted to do so; 2) users with different domain knowledge have similar preferences for different levels of control; 3) providing transparent and user-friendly control features can encourage their usage; 4) the availability of control features is associated with users’ feeling of empowerment and perceived ability to recognize how recommendations steer their attention. These results open future research directions on the design of DSA-informed transparent and controllable RSs and provide practical guidelines for the implementation of the DSA requirements for RSs. Matteo Fabbri, Jingyi Jia, Pablo Jerez Arnau, Wolfgang Wörndl |
Int. J. Hum. Comput. Stud. | 1 |
| 2025 | Auditing Recommender Systems for User Empowerment in Very Large Online Platforms under the Digital Services ActabstractThe governance of recommender systems (RSs) in very large online platforms (VLOPs) is expected to change significantly under the Digital Services Act (DSA), which imposes new obligations on transparency and user control.However, beyond legal compliance, a critical question remains: How can recommender systems be redesigned to genuinely empower users and foster meaningful personalization?This paper addresses this question by analyzing how three major short-video platforms-Instagram, TikTok, and YouTube-have implemented the DSA requirements for RSs.By reviewing their audit reports, systemic risk assessments, and compliance strategies, we evaluate the extent to which current approaches enhance user autonomy and control over content exposure.Building on this analysis, we outline a perspective for the future of VLOPs' RSs grounded in speculative design.We argue that meaningful personalization should integrate algorithmic choice, balancing proportionality and granularity in RS customization, and content curation, ensuring diversity and authoritativeness to mitigate systemic risks.By bridging legal analysis, platform governance, and user-centered design, this paper outlines actionable pathways for aligning technical developments with regulatory objectives.Our findings contribute to interdisciplinary research on RSs by highlighting how platforms can move beyond minimal compliance toward a model that prioritizes user empowerment and content pluralism. Matteo Fabbri, Ludovico Boratto |
RecSys | 1 |
| 2023 | Self-determination through explanation: an ethical perspective on the implementation of the transparency requirements for recommender systems set by the Digital Services Act of the European UnionabstractIn the contemporary information age, recommender systems (RSs) play a critical role in influencing online behaviour: from social media to e-commerce, from music streaming to news aggregators, individuals are constantly targeted by personalized recommendations suggesting contents that may interest them. Despite such diffusion, the extent to which recommendations influence users’ decisions is still underexplored, given that independent audits on the structure and functioning of RSs deployed on online platforms are usually prevented by proprietary constraints. The nudging potential of RSs can represent a risk for vulnerable people: indeed, judicial cases involving platforms’ responsibility for displaying recommendations that may lead to political radicalization or endangerment of minors have recently caught public attention. The Digital Services Act of the European Union (DSA) is the first supranational regulation that sets specific transparency and auditing requirements for RSs implemented by online platforms with the aim of enhancing users’ self-determination: in particular, it allows users to modify the parameters on which recommendations rely so to let them choose autonomously which kind of content they want to see. This research focuses on whether and how the enforcement of this regulation can mitigate the unfair consequences of the power imbalance between online platforms and users. To this aim, I discuss the harms arising from digital nudging based on RSs and propose explanations as a tool that can reduce the impact of those harms by increasing users’ awareness. Through a comparative analysis of relevant articles of the DSA, the General Data Protection Regulation (GDPR) and the AI Act, I outline how the provisions of the DSA fill some of the gaps left by other relevant European regulations, while leaving the so-called right to explanation substantially unaddressed. As a result of this analysis, I argue that, in order for the implementation of the DSA provisions on recommender systems to be effective, policy-makers should: 1) enhance users’ awareness through clear and easily accessible explanations on how the recommendation process works and how they can be influenced by it; 2) grant users the possibility of intervening directly on the strategies through which RSs target them on the platform’s interface. Matteo Fabbri |
AIES | 1 |
| 2023 | How and to which extent will the provisions of the Digital Services Act of the European Union impact on the relationship between users and platforms as information providers?abstractIn the contemporary information age, recommender systems (RSs) play a crucial role in determining the way in which people interact and obtain information online: in fact, from social media feeds to news aggregators and e-commerce websites, users are constantly targeted by personalized recommendations about what they may like. The Digital Services Act (DSA) of the European Union1 [3], which is the first supranational regulation addressing automated recommendations specifically, defines a RS as “a fully or partially automated system used by an online platform to suggest in its online interface specific information to recipients of the service or prioritize that information, including as a result of a search initiated by the recipient of the service or otherwise determining the relative order or prominence of information displayed” (DSA, art. 3 (s)). This definition highlights the method (“fully or partially automated”), aim (“to suggest”), content (“specific information”), target (“recipients of the service”), input (“as a result of a search initiated by the recipient”) and output (“determining the relative order or prominence of information displayed”) of a recommendation process. As it can be observed, RSs are involved in the main aspects of online interactions, and this is why their influencing potential should not be underestimated. In fact, whilst RSs are aimed to improve user’s experience by reducing the information overload, they can give rise to a variety of ethical concerns related to privacy, autonomy and fairness [5], to name but a few. However, independent research and users’ access to the design and functioning of the RSs implemented on mainstream platforms is usually prevented by their proprietary status. Matteo Fabbri |
AIES | 1 |
| 2023 | TrackFlow: Multi-Object Tracking with Normalizing FlowsabstractThe field of multi-object tracking has recently seen a renewed interest in the good old schema of tracking-by-detection, as its simplicity and strong priors spare it from the complex design and painful babysitting of tracking-by-attention approaches. In view of this, we aim at extending tracking-by-detection to multi-modal settings, where a comprehensive cost has to be computed from heterogeneous information e.g., 2D motion cues, visual appearance, and pose estimates. More precisely, we follow a case study where a rough estimate of 3D information is also available and must be merged with other traditional metrics (e.g., the IoU). To achieve that, recent approaches resort to either simple rules or complex heuristics to balance the contribution of each cost. However, i) they require careful tuning of tailored hyperparameters on a hold-out set, and ii) they imply these costs to be independent, which does not hold in reality. We address these issues by building upon an elegant probabilistic formulation, which considers the cost of a candidate association as the negative log-likelihood yielded by a deep density estimator, trained to model the conditional joint probability distribution of correct associations. Our experiments, conducted on both simulated and real benchmarks, show that our approach consistently enhances the performance of several tracking-by-detection algorithms. Gianluca Mancusi, Aniello Panariello, Angelo Porrello, Matteo Fabbri, Simone Calderara, Rita Cucchiara |
ICCV | 4 |
| 2022 | Fine-grained Human Analysis under Occlusions and Perspective Constraints in Multimedia SurveillanceabstractHuman detection in the wild is a research topic of paramount importance in computer vision, and it is the starting step for designing intelligent systems oriented to human interaction that work in complete autonomy. To achieve this goal, computer vision and machine learning should aim at superhuman capabilities. In this work, we address the problem of fine-grained human analysis under occlusions and perspective constraints. More specifically, we discuss some issues and some possible solutions to effectively detect people using pose estimation methods and to detect humans under occlusions both in the two-dimensional (2D) image plane and in the 3D space exploiting single monocular cameras. Dealing with occlusion can be done at the joint level or pixel level: We discuss two different solutions, the former based on a supervised neural network architecture for detecting occluded joints and the latter based on a semi-supervised specialized GAN that exploits both appearance and human shape attributes to determine the missing parts of the visible shape. To deal with perspective constraints, we further discuss a neural approach based on a double architecture that learns to create an optimal neural representation, which is useful to reconstruct the 3D position of human keypoints starting with simple RGB images. All these approaches have a critical point in common: the need for large annotated datasets. To have large, fair, consistent, transparent, and ethical datasets, we propose the adoption of synthetic datasets as, for example, JTA and MOTSynth. In this article, we discuss the pros and cons of using synthetic datasets while tackling several human-centered AI issues with respect to European GDPR rules for privacy. We further explore and discuss an application in the field of risk assessment by space occupancy estimation during the COVID-19 pandemic called Inter-Homines. Rita Cucchiara, Matteo Fabbri |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | MOTSynth: How Can Synthetic Data Help Pedestrian Detection and Tracking?abstractDeep learning-based methods for video pedestrian detection and tracking require large volumes of training data to achieve good performance. However, data acquisition in crowded public environments raises data privacy concerns – we are not allowed to simply record and store data without the explicit consent of all participants. Furthermore, the annotation of such data for computer vision applications usually requires a substantial amount of manual effort, especially in the video domain. Labeling instances of pedestrians in highly crowded scenarios can be challenging even for human annotators and may introduce errors in the training data. In this paper, we study how we can advance different aspects of multi-person tracking using solely synthetic data. To this end, we generate MOTSynth, a large, highly diverse synthetic dataset for object detection and tracking using a rendering game engine. Our experiments show that MOTSynth can be used as a replacement for real data on tasks such as pedestrian detection, re-identification, segmentation, and tracking. Matteo Fabbri, Guillem Brasó, Gianluca Maugeri, Orcun Cetintas, Riccardo Gasparini, Aljosa Osep, Simone Calderara, Laura Leal-Taixé, Rita Cucchiara |
ICCV | 1 |
| 2020 | Compressed Volumetric Heatmaps for Multi-Person 3D Pose EstimationabstractIn this paper we present a novel approach for bottom-up multi-person 3D human pose estimation from monocular RGB images. We propose to use high resolution volumetric heatmaps to model joint locations, devising a simple and effective compression method to drastically reduce the size of this representation. At the core of the proposed method lies our Volumetric Heatmap Autoencoder, a fully-convolutional network tasked with the compression of ground-truth heatmaps into a dense intermediate representation. A second model, the Code Predictor, is then trained to predict these codes, which can be decompressed at test time to re-obtain the original representation. Our experimental evaluation shows that our method performs favorably when compared to state of the art on both multi-person and single-person 3D human pose estimation datasets and, thanks to our novel compression strategy, can process full-HD images at the constant runtime of 8 fps regardless of the number of subjects in the scene. Code and models are publicly available. Matteo Fabbri, Fabio Lanzi, Simone Calderara, Stefano Alletto, Rita Cucchiara |
CVPR | 1 |
| 2020 | Face-from-Depth for Head Pose Estimation on Depth ImagesabstractDepth cameras allow to set up reliable solutions for people monitoring and behavior understanding, especially when unstable or poor illumination conditions make unusable common RGB sensors. Therefore, we propose a complete framework for the estimation of the head and shoulder pose based on depth images only. A head detection and localization module is also included, in order to develop a complete end-to-end system. The core element of the framework is a Convolutional Neural Network, called POSEidon+, that receives as input three types of images and provides the 3D angles of the pose as output. Moreover, a Face-from-Depth component based on a Deterministic Conditional GAN model is able to hallucinate a face from the corresponding depth image. We empirically demonstrate that this positively impacts the system performances. We test the proposed framework on two public datasets, namely Biwi Kinect Head Pose and ICT-3DHP, and on Pandora, a new challenging dataset mainly inspired by the automotive setup. Experimental results show that our method overcomes several recent state-of-art works based on both intensity and depth input data, running in real-time at more than 30 frames per second. Guido Borghi, Matteo Fabbri, Roberto Vezzani, Simone Calderara, Rita Cucchiara |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Can adversarial networks hallucinate occluded people with a plausible aspect?abstractWhen you see a person in a crowd, occluded by other persons, you miss visual information that can be used to recognize, re-identify or simply classify him or her. You can imagine its appearance given your experience, nothing more. Similarly, AI solutions can try to hallucinate missing information with specific deep learning architectures, suitably trained with people with and without occlusions. The goal of this work is to generate a complete image of a person, given an occluded version in input, that should be a) without occlusion b) similar at pixel level to a completely visible people shape c) capable to conserve similar visual attributes (e.g. male/female) of the original one. For the purpose, we propose a new approach by integrating the state-of-the-art of neural network architectures, namely U-nets and GANs, as well as discriminative attribute classification nets, with an architecture specifically designed to de-occlude people shapes. The network is trained to optimize a Loss function which could take into account the aforementioned objectives. As well we propose two datasets for testing our solution: the first one, occluded RAP, created automatically by occluding real shapes of the RAP dataset created by Li et al. (2016) (which collects also attributes of the people aspect); the second is a large synthetic dataset, AiC, generated in computer graphics with data extracted from the GTA video game, that contains 3D data of occluded objects by construction. Results are impressive and outperform any other previous proposal. This result could be an initial step to many further researches to recognize people and their behavior in an open crowded world. Federico Fulgeri, Matteo Fabbri, Stefano Alletto, Simone Calderara, Rita Cucchiara |
Comput. Vis. Image Underst. | 2 |
| 2018 | Learning to Detect and Track Visible and Occluded Body Joints in a Virtual World
Matteo Fabbri, Fabio Lanzi, Simone Calderara, Andrea Palazzi, Roberto Vezzani, Rita Cucchiara |
ECCV (4) | 1 |
| 2018 | Domain Translation with Conditional GANs: from Depth to RGB Face-to-FaceabstractCan faces acquired by low-cost depth sensors be useful to catch some characteristic details of the face? Typically the answer is no. However, new deep architectures can generate RGB images from data acquired in a different modality, such as depth data. In this paper, we propose a new Deterministic Conditional GAN, trained on annotated RGB-D face datasets, effective for a face-to-face translation from depth to RGB. Although the network cannot reconstruct the exact somatic features for unknown individual faces, it is capable to reconstruct plausible faces; their appearance is accurate enough to be used in many pattern recognition tasks. In fact, we test the network capability to hallucinate with some Perceptual Probes, as for instance face aspect classification or landmark detection. Depth face can be used in spite of the correspondent RGB images, that often are not available due to difficult luminance conditions. Experimental results are very promising and are as far as better than previously proposed approaches: this domain translation can constitute a new way to exploit depth data in new future applications. Matteo Fabbri, Guido Borghi, Fabio Lanzi, Roberto Vezzani, Simone Calderara, Rita Cucchiara |
ICPR | 1 |
| 2017 | Generative adversarial models for people attribute recognition in surveillanceabstractIn this paper we propose a deep architecture for detecting people attributes (e.g. gender, race, clothing ...) in surveillance contexts. Our proposal explicitly deal with poor resolution and occlusion issues that often occur in surveillance footages by enhancing the images by means of Deep Convolutional Generative Adversarial Networks (DCGAN). Experiments show that by combining both our Generative Reconstruction and Deep Attribute Classification Network we can effectively extract attributes even when resolution is poor and in presence of strong occlusions up to 80% of the whole person figure. Matteo Fabbri, Simone Calderara, Rita Cucchiara |
AVSS | 1 |