VLDB 2026 Research / reviewers in the wild / expert
Nicola Garau
dblp:182/7148
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0001-7147-9109ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Drone Agents: learning to fly to learn how to seeabstractRecent advances in reinforcement learning (RL) have opened promising opportunities for autonomous drone navigation. However, bridging RL-based methods with high-fidelity simulation environments remains an open challenge. In this paper, we introduce a novel multi-task approach for real-time RL-based autonomous drone navigation inside Unreal Engine 5. The idea is simple but effective: by using a physically accurate drone model and systematically increasing the complexity of simulated scenarios — from predefined path following to dynamic visual tracking with advanced sensing modalities — we achieve impressive generalization to multiple dynamically changing tasks. Our method addresses critical gaps in drone autonomy research, such as obstacle avoidance in swarm coordination and robust visual target tracking under varying environmental conditions. Through comprehensive experimental evaluation on pedestrian detection data in an urban scenario, we highlight the essential factors influencing drone learning performance. Our fine-tuned YOLOv8 model improves pedestrian detection recall from 72.41% to 84.71% on synthetic aerial images, demonstrating the effectiveness of drone-collected data for domain adaptation. https://mmlab-cv.github.io/DroneAgents/ Daniele Della Pietra, Karthik Govindarajan, Fabrizio Granelli, Andrea Rosani, Nicola Garau |
AVSS | 5 |
| 2025 | Render, Encode, Plan: A simple pipeline for hybrid RL-DL learning inside Unreal EngineabstractLearning is an iterative process that requires multiple forms of interaction with the environment. During learning, we experience the world through the repetition of observations and actions, gaining an insight into which combination of these leads to the best results, according to our goals. The same paradigm has been applied to traditional reinforcement learning (RL) over the years, with impressive results in 3D navigation and planning. On the other hand, the computer vision community has been focusing mostly on vision-related tasks (e.g. classification, segmentation, depth estimation) using deep learning (DL). We present REP: Render, Encode, Plan , a unified framework to train embodied agents of different kinds (humanoids, vehicles, and drones) inside Unreal Engine, showing how a combination of RL and DL can help to shape intelligent agents that can better sense the surrounding environment. The main advantage of our method is the combination of different sensory modalities, including game state observations and vision features, that allow the agents to share a similar structure in their observations and rewards, while defining separate rewards based on their goals. We demonstrate impressive generalization capabilities on large-scale realistic 3D environments and on multiple dynamically changing scenarios, with different goals and rewards. All code, complete experiments, and environments will be available at https://mmlab-cv.github.io/REP/ . • Hybrid RL-DL : REP trains agents in UE5 using visual and physical observations. • Multi-agent, multi-modal : Supports drones, cars, humans, and multiple tasks. • UE5-integrated : Uses NNE and shared memory for real-time training and inference. • Robust generalization : Solves diverse tasks in large, dynamic 3D environments. Daniele Della Pietra, Nicola Garau |
Comput. Graph. | 2 |
| 2024 | Unicrowd Simulator: Visual and Behavioral Fidelity For The Generation of Crowd DatasetsabstractWe introduce UniCrowd1, a human crowd simulator for the modeling of human-related dynamics. The simulator is accompanied by a meticulously collected dataset within its synthetic environment, along with a comprehensive validation pipeline. Leveraging simulation as a powerful tool for generating annotated data, UniCrowd addresses the increasing demand for large training datasets, mimicking both the behavioral and visual aspect of crowds. Recent advancements in rendering and virtualization engines have enhanced the simulators capabilities to represent complex scenes, encompassing environmental factors such as weather conditions, surface reflectance, and human-related events like actions and behaviors. The adaptability and the non-deterministic nature of the human behavioral module of UniCrowd, coupled with its 3D rendering represents an improvement over available crowd simulators. We demonstrate the suitability of our simulator and its associated dataset for various computer vision tasks. We highlight applications such as detection and segmentation, as well as specialized tasks including crowd counting, human pose estimation, trajectory analysis and prediction.1The simulator and the dataset can be accessed at github.com/mmlabcvUniCrowd Niccolò Bisagno, Antonio Luigi Stefani, Nicola Garau, Francesco G. B. De Natale, Nicola Conci |
ICIP | 3 |
| 2024 | All Skeletons are Created Equal! A Domain Adaptation Transformer to Handle Multiple TopologiesabstractDesigning an effective Human Pose Estimation (HPE) pipeline necessitates handling diverse datasets, a task often deemed necessary but burdensome by researchers. Existing datasets are annotated using different conventions, with the human body represented as a parametric 3D model or as a collection of 2D/3D joints and bones, known as skeleton-based annotation. Despite its widespread use in training both 2D and 3D HPE networks, the lack of standardization in the topologies of joint-based pose annotations requires considerable difficulties when evaluating the algorithms performances across different datasets. To solve this issue, we introduce a novel self-supervised human pose domain adaptation approach to map a given skeletal model into a target model of choice. We design a transformer-based architecture trained to reconstruct missing joints within a given topology, aligning them with the target model. During testing, the network seamlessly reconstructs the missing joints, treating them as if masked, based on the common joints shared by the original topology and the target one. Unlike previous approaches, our method works with arbitrary body poses, joints number, and skeleton topologies, across multiple datasets.11The code is available at https://github.com/mmlab-cv/DAT.git Giulia Martinelli, Nicola Garau, Niccolò Bisagno, Nicola Conci |
ICIP | 2 |
| 2024 | Dynamic Crowd Routing: RL-Driven Crowd DynamicsabstractThe simulation of crowds is complex and challenging. Every individual in a crowd exhibits a different behaviour, targets a different goal, and undergoes different types of interactions. Within crowds, groups can be identified in both static and dynamic configurations, with varying levels of responsiveness, leading to the emergence of complex avoidance mechanisms. In the past, rule-based models have been proposed to simulate crowds, unlocking the potential for large-scale simulations. Over the years, learning-based solutions have been presented, achieving acceptable results despite the lack of high-quality ground truth data for training. While both rule-based and learning-based methods have recently been integrated into 3D simulation engines, they usually rely on navigation meshes or B-spline functions, hindering their generalization to open-world scenarios. In this work, we propose a reinforcement learning-based solution to learn meaningful crowd dynamics inside the Unreal Engine 3D engine, enabling massive and highly dynamic crowd simulations. We show how our proposed method makes it possible to simulate crowd setups that require complex dynamic routing mechanisms, which are otherwise hard to achieve using rule-based approaches or even deep learning-based methods. Our approach also al-lows us to easily collect large synthetic datasets that are both photorealistic and provide accurate ground truth data without the need for any manual annotation. Some demonstration videos are available at mmlab-cv.github.io/DynamicCrowdRouting; Code, complete experiments and analysis will be made available upon acceptance. Daniele Della Pietra, Nicola Garau, Nicola Conci, Fabrizio Granelli |
MMSP | 2 |
| 2024 | MoMa: Skinned motion retargeting using masked pose modeling
Giulia Martinelli, Nicola Garau, Niccolò Bisagno, Nicola Conci |
Comput. Vis. Image Underst. | 2 |
| 2024 | Agglomerator++: Interpretable part-whole hierarchies and latent space representations in neural networksabstractDeep neural networks achieve outstanding results in a large variety of tasks, often outperforming human experts. However, a known limitation of current neural architectures is the poor accessibility in understanding and interpreting the network’s response to a given input. This is directly related to the huge number of variables and the associated non-linearities of neural models, which are often used as black boxes. This lack of transparency, particularly in crucial areas like autonomous driving, security, and healthcare, can trigger skepticism and limit trust, despite the networks’ high performance. In this work, we want to advance the interpretability in neural networks. We present Agglomerator++, a framework capable of providing a representation of part-whole hierarchies from visual cues and organizing the input distribution to match the conceptual-semantic hierarchical structure between classes. We evaluate our method on common datasets, such as SmallNORB, MNIST, FashionMNIST, CIFAR-10, and CIFAR-100, showing that our solution delivers a more interpretable model compared to other state-of-the-art approaches. Our code is available at https://mmlab-cv.github.io/Agglomeratorplusplus/ . • We introduce a novel model, called Agglomerator++, mimicking the functioning of the cortical columns in the human brain. • Our solution provides interpretability of relationships in data, namely the hierarchical organization of the feature space. • We introduce positional encoding and input masking during pre-training for self- supervised reconstruction. • This neural representation is more efficient and closely resembles human lexical similarities. Zeno Sambugaro, Nicola Garau, Niccolò Bisagno, Nicola Conci |
Comput. Vis. Image Underst. | 2 |
| 2023 | CapsulePose: A variational CapsNet for real-time end-to-end 3D human pose estimation
Nicola Garau, Nicola Conci |
Neurocomputing | 1 |
| 2022 | Interpretable part-whole hierarchies and conceptual-semantic relationships in neural networksabstractDeep neural networks achieve outstanding results in a large variety of tasks, often outperforming human experts. However, a known limitation of current neural architectures is the poor accessibility to understand and interpret the network response to a given input. This is directly related to the huge number of variables and the associated non-linearities of neural models, which are often used as black boxes. When it comes to critical applications as autonomous driving, security and safety, medicine and health, the lack of interpretability of the network behavior tends to induce skepticism and limited trustworthiness, despite the accurate performance of such systems in the given task. Furthermore, a single metric, such as the classification accuracy, provides a non-exhaustive evaluation of most realworld scenarios. In this paper, we want to make a step forward towards interpretability in neural networks, providing new tools to interpret their behavior. We present Agglomerator, a framework capable of providing a representation of part-whole hierarchies from visual cues and organizing the input distribution matching the conceptual-semantic hierarchical structure between classes. We evaluate our method on common datasets, such as SmallNORB, MNIST, FashionMNIST, CIFAR-10, and CIFAR-100, providing a more interpretable model than other state-of-the-art approaches. Nicola Garau, Niccolò Bisagno, Zeno Sambugaro, Nicola Conci |
CVPR | 1 |
| 2022 | A multimodal framework for the evaluation of patients' weaknesses, supporting the design of customised AAL solutions
Nicola Garau, Damiano Fruet, Alessandro Luchetti, Francesco G. B. De Natale, Nicola Conci |
Expert Syst. Appl. | 1 |
| 2021 | DECA: Deep viewpoint-Equivariant human pose estimation using Capsule AutoencodersabstractHuman Pose Estimation (HPE) aims at retrieving the 3D position of human joints from images or videos. We show that current 3D HPE methods suffer a lack of viewpoint equivariance, namely they tend to fail or perform poorly when dealing with viewpoints unseen at training time. Deep learning methods often rely on either scale-invariant, translation-invariant, or rotation-invariant operations, such as max-pooling. However, the adoption of such procedures does not necessarily improve viewpoint generalization, rather leading to more data-dependent methods. To tackle this issue, we propose a novel capsule autoencoder network with fast Variational Bayes capsule routing, named DECA. By modeling each joint as a capsule entity, combined with the routing algorithm, our approach can preserve the joints’ hierarchical and geometrical structure in the feature space, independently from the viewpoint. By achieving viewpoint equivariance, we drastically reduce the network data dependency at training time, resulting in an improved ability to generalize for unseen viewpoints. In the experimental validation, we outperform other methods on depth images from both seen and unseen viewpoints, both top-view, and front-view. In the RGB domain, the same network gives state-of-the-art results on the challenging viewpoint transfer task, also establishing a new framework for top-view HPE. The code can be found at https://github.com/mmlab-cv/DECA. Nicola Garau, Niccolò Bisagno, Piotr Bródka, Nicola Conci |
ICCV | 1 |