EDBT 2026 Demo / reviewers in the wild / expert
Niccolò Bisagno
dblp:207/6251
· DBLP profile ↗
13ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0001-5704-5785ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Signal processing for haptic surface modeling: A reviewabstractHaptic feedback has been integrated into Virtual and Augmented Reality, complementing acoustic and visual information and contributing to an all-round immersive experience in multiple fields, spanning from the medical domain to entertainment and gaming. Haptic technologies involve complex cross-disciplinary research that encompasses sensing, data representation, interactive rendering, perception, and quality of experience. The standard processing pipeline, consists of (I) sensing physical features in the real world using a transducer, (II) modeling and storing the collected information in some digital format, (III) communicating the information, and finally, (IV) rendering the haptic information through appropriate devices, thus producing a user experience (V) perceptually close to the original physical world. Among these areas, sensing, rendering and perception have been deeply investigated and are the subject of different comprehensive surveys available in the literature. Differently, research dealing with haptic surface modeling and data representation still lacks a comprehensive dissection. In this work, we aim at providing an overview on modeling and representation of haptic surfaces from a signal processing perspective, covering the aspects that lie in between haptic information acquisition on one side and rendering and perception on the other side. We analyze, categorize, and compare research papers that address the haptic surface modeling and data representation, pointing out existing gaps and possible research directions. Antonio Luigi Stefani, Niccolò Bisagno, Andrea Rosani, Nicola Conci, Francesco G. B. De Natale |
Signal Process. Image Commun. | 2 |
| 2024 | MapFlow: Multi-Agent Pedestrian Trajectory Prediction Using Normalizing FlowabstractIn the task of pedestrian trajectory prediction, multi-modal prediction has recently emerged, demonstrating how a good model should predict multiple socially acceptable futures. With this respect, Normalizing Flows (NFs) have shown remarkable generative capabilities that make them particularly suitable for multi-modal trajectory prediction. By sampling from the learned distribution, NFs can produce multiple socially acceptable trajectories, each one paired with its corresponding likelihood score. Taking advantage of the multi-modal prediction coupled with the likelihood score, with MapFlow we introduce a solution based on NFs that improves the accuracy in prediction by incorporating in the model the social influence of neighboring pedestrians.1 Antonio Luigi Stefani, Niccolò Bisagno, Nicola Conci |
ICASSP | 2 |
| 2024 | Unicrowd Simulator: Visual and Behavioral Fidelity For The Generation of Crowd DatasetsabstractWe introduce UniCrowd1, a human crowd simulator for the modeling of human-related dynamics. The simulator is accompanied by a meticulously collected dataset within its synthetic environment, along with a comprehensive validation pipeline. Leveraging simulation as a powerful tool for generating annotated data, UniCrowd addresses the increasing demand for large training datasets, mimicking both the behavioral and visual aspect of crowds. Recent advancements in rendering and virtualization engines have enhanced the simulators capabilities to represent complex scenes, encompassing environmental factors such as weather conditions, surface reflectance, and human-related events like actions and behaviors. The adaptability and the non-deterministic nature of the human behavioral module of UniCrowd, coupled with its 3D rendering represents an improvement over available crowd simulators. We demonstrate the suitability of our simulator and its associated dataset for various computer vision tasks. We highlight applications such as detection and segmentation, as well as specialized tasks including crowd counting, human pose estimation, trajectory analysis and prediction.1The simulator and the dataset can be accessed at github.com/mmlabcvUniCrowd Niccolò Bisagno, Antonio Luigi Stefani, Nicola Garau, Francesco G. B. De Natale, Nicola Conci |
ICIP | 1 |
| 2024 | All Skeletons are Created Equal! A Domain Adaptation Transformer to Handle Multiple TopologiesabstractDesigning an effective Human Pose Estimation (HPE) pipeline necessitates handling diverse datasets, a task often deemed necessary but burdensome by researchers. Existing datasets are annotated using different conventions, with the human body represented as a parametric 3D model or as a collection of 2D/3D joints and bones, known as skeleton-based annotation. Despite its widespread use in training both 2D and 3D HPE networks, the lack of standardization in the topologies of joint-based pose annotations requires considerable difficulties when evaluating the algorithms performances across different datasets. To solve this issue, we introduce a novel self-supervised human pose domain adaptation approach to map a given skeletal model into a target model of choice. We design a transformer-based architecture trained to reconstruct missing joints within a given topology, aligning them with the target model. During testing, the network seamlessly reconstructs the missing joints, treating them as if masked, based on the common joints shared by the original topology and the target one. Unlike previous approaches, our method works with arbitrary body poses, joints number, and skeleton topologies, across multiple datasets.11The code is available at https://github.com/mmlab-cv/DAT.git Giulia Martinelli, Nicola Garau, Niccolò Bisagno, Nicola Conci |
ICIP | 3 |
| 2024 | MoMa: Skinned motion retargeting using masked pose modeling
Giulia Martinelli, Nicola Garau, Niccolò Bisagno, Nicola Conci |
Comput. Vis. Image Underst. | 3 |
| 2024 | Agglomerator++: Interpretable part-whole hierarchies and latent space representations in neural networksabstractDeep neural networks achieve outstanding results in a large variety of tasks, often outperforming human experts. However, a known limitation of current neural architectures is the poor accessibility in understanding and interpreting the network’s response to a given input. This is directly related to the huge number of variables and the associated non-linearities of neural models, which are often used as black boxes. This lack of transparency, particularly in crucial areas like autonomous driving, security, and healthcare, can trigger skepticism and limit trust, despite the networks’ high performance. In this work, we want to advance the interpretability in neural networks. We present Agglomerator++, a framework capable of providing a representation of part-whole hierarchies from visual cues and organizing the input distribution to match the conceptual-semantic hierarchical structure between classes. We evaluate our method on common datasets, such as SmallNORB, MNIST, FashionMNIST, CIFAR-10, and CIFAR-100, showing that our solution delivers a more interpretable model compared to other state-of-the-art approaches. Our code is available at https://mmlab-cv.github.io/Agglomeratorplusplus/ . • We introduce a novel model, called Agglomerator++, mimicking the functioning of the cortical columns in the human brain. • Our solution provides interpretability of relationships in data, namely the hierarchical organization of the feature space. • We introduce positional encoding and input masking during pre-training for self- supervised reconstruction. • This neural representation is more efficient and closely resembles human lexical similarities. Zeno Sambugaro, Nicola Garau, Niccolò Bisagno, Nicola Conci |
Comput. Vis. Image Underst. | 3 |
| 2022 | Interpretable part-whole hierarchies and conceptual-semantic relationships in neural networksabstractDeep neural networks achieve outstanding results in a large variety of tasks, often outperforming human experts. However, a known limitation of current neural architectures is the poor accessibility to understand and interpret the network response to a given input. This is directly related to the huge number of variables and the associated non-linearities of neural models, which are often used as black boxes. When it comes to critical applications as autonomous driving, security and safety, medicine and health, the lack of interpretability of the network behavior tends to induce skepticism and limited trustworthiness, despite the accurate performance of such systems in the given task. Furthermore, a single metric, such as the classification accuracy, provides a non-exhaustive evaluation of most realworld scenarios. In this paper, we want to make a step forward towards interpretability in neural networks, providing new tools to interpret their behavior. We present Agglomerator, a framework capable of providing a representation of part-whole hierarchies from visual cues and organizing the input distribution matching the conceptual-semantic hierarchical structure between classes. We evaluate our method on common datasets, such as SmallNORB, MNIST, FashionMNIST, CIFAR-10, and CIFAR-100, providing a more interpretable model than other state-of-the-art approaches. Nicola Garau, Niccolò Bisagno, Zeno Sambugaro, Nicola Conci |
CVPR | 2 |
| 2021 | Out-of-Distribution Detection Using Union of 1-Dimensional SubspacesabstractThe goal of out-of-distribution (OOD) detection is to handle the situations where the test samples are drawn from a different distribution than the training data. In this paper, we argue that OOD samples can be detected more easily if the training data is embedded into a low-dimensional space, such that the embedded training samples lie on a union of 1-dimensional subspaces. We show that such embedding of the in-distribution (ID) samples provides us with two main advantages. First, due to compact representation in the feature space, OOD samples are less likely to occupy the same region as the known classes. Second, the first singular vector of ID samples belonging to a 1-dimensional subspace can be used as their robust representative. Motivated by these observations, we train a deep neural network such that the ID samples are embedded onto a union of 1-dimensional subspaces. At the test time, employing sampling techniques used for approximate Bayesian inference in deep learning, input samples are detected as OOD if they occupy the region corresponding to the ID samples with probability 0. Spectral components of the ID samples are used as robust representative of this region. Our method does not have any hyperparameter to be tuned using extra information and it can be applied on different modalities with minimal change. The effectiveness of the proposed method is demonstrated on different benchmark datasets, both in the image and video classification domains. Alireza Zaeemzadeh, Niccolò Bisagno, Zeno Sambugaro, Nicola Conci, Nazanin Rahnavard, Mubarak Shah |
CVPR | 2 |
| 2021 | DECA: Deep viewpoint-Equivariant human pose estimation using Capsule AutoencodersabstractHuman Pose Estimation (HPE) aims at retrieving the 3D position of human joints from images or videos. We show that current 3D HPE methods suffer a lack of viewpoint equivariance, namely they tend to fail or perform poorly when dealing with viewpoints unseen at training time. Deep learning methods often rely on either scale-invariant, translation-invariant, or rotation-invariant operations, such as max-pooling. However, the adoption of such procedures does not necessarily improve viewpoint generalization, rather leading to more data-dependent methods. To tackle this issue, we propose a novel capsule autoencoder network with fast Variational Bayes capsule routing, named DECA. By modeling each joint as a capsule entity, combined with the routing algorithm, our approach can preserve the joints’ hierarchical and geometrical structure in the feature space, independently from the viewpoint. By achieving viewpoint equivariance, we drastically reduce the network data dependency at training time, resulting in an improved ability to generalize for unseen viewpoints. In the experimental validation, we outperform other methods on depth images from both seen and unseen viewpoints, both top-view, and front-view. In the RGB domain, the same network gives state-of-the-art results on the challenging viewpoint transfer task, also establishing a new framework for top-view HPE. The code can be found at https://github.com/mmlab-cv/DECA. Nicola Garau, Niccolò Bisagno, Piotr Bródka, Nicola Conci |
ICCV | 2 |
| 2021 | Embedding group and obstacle information in LSTM networks for human trajectory prediction in crowded scenes
Niccolò Bisagno, Cristiano Saltori, Bo Zhang 0045, Francesco G. B. De Natale, Nicola Conci |
Comput. Vis. Image Underst. | 1 |
| 2021 | Where Are They Going? Predicting Human Behaviors in Crowded ScenesabstractIn this article, we propose a framework for crowd behavior prediction in complicated scenarios. The fundamental framework is designed using the standard encoder-decoder scheme, which is built upon the long short-term memory module to capture the temporal evolution of crowd behaviors. To model interactions among humans and environments, we embed both the social and the physical attention mechanisms into the long short-term memory. The social attention component can model the interactions among different pedestrians, whereas the physical attention component helps to understand the spatial configurations of the scene. Since pedestrians’ behaviors demonstrate multi-modal properties, we use the generative model to produce multiple acceptable future paths. The proposed framework not only predicts an individual’s trajectory accurately but also forecasts the ongoing group behaviors by leveraging on the coherent filtering approach. Experiments are carried out on the standard crowd benchmarks (namely, the ETH, the UCY, the CUHK crowd, and the CrowdFlow datasets), which demonstrate that the proposed framework is effective in forecasting crowd behaviors in complex scenarios. Bo Zhang 0045, Niccolò Bisagno, Nicola Conci, Francesco G. B. De Natale, Hongbo Liu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2020 | AG-GAN: An Attentive Group-Aware GAN for pedestrian trajectory predictionabstractUnderstanding human behaviors in crowded scenarios requires analyzing not only the position of the subjects in space, but also the scene context. Existing approaches mostly rely on the motion history of each pedestrian and model the interactions among people by considering the entire surrounding neighborhood. In our approach, we address the problem of motion prediction by applying coherent group clustering and a global attention mechanism on the LSTM-based Generative Adversarial Networks (GANs). The proposed model consists of an attentive group-aware GAN that observes the agents' past motion and predicts future paths, using (i) a group pooling module to model neighborhood interaction, and (ii) an attention module to specifically focus on hidden states. The experimental results demonstrate that our proposal outperforms state-of-the-art models on common benchmark datasets, and is able to generate socially-acceptable trajectories. Niccolò Bisagno, Syed Zohaib Hassan, Nicola Conci |
ICPR | 2 |
| 2017 | Data-Driven crowd simulationabstractIn this work we propose a framework for data-driven crowd simulation starting from a small set of trajectories. To model the macroscopic behavior, our method extracts pedestrian trajectories from real videos, clusters all trajectories of pedestrians who intend to reach the same goal, and computes the velocity field associated with each exit region in the scene to guide virtual agents toward their destinations, namely, the goal-dependent path selection. While at the microscopic level, the simulation is performed using on the one hand the Social Force Model to handle the collision-avoidance among agents and on the other hand the computed velocity fields to model the macroscopic behavior. The experimental results demonstrate that the velocity field can be exploited to effectively reproduce crowd behaviors. Niccolò Bisagno, Nicola Conci, Bo Zhang 0045 |
AVSS | 1 |