VLDB 2026 Research / reviewers in the wild / expert
Fabio Poiesi
dblp:87/8843
· DBLP profile ↗
45ranked-venue papers
8as first author
33since 2021 · last 2026
0000-0002-9769-1279ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 6 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 4 first-author · 21 since 2021Systems, architecture and hardware · 6 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Masked Clustering Prediction for Unsupervised Point Cloud Pre-trainingabstractVision transformers (ViTs) have recently been widely applied to 3D point cloud understanding, with masked autoencoding as the predominant pre-training paradigm. However, the challenge of learning dense and informative semantic features from point clouds via standard ViTs remains underexplored. We propose MaskClu, a novel unsupervised pre-training method for ViTs on 3D point clouds that integrates masked point modeling with clustering-based learning. MaskClu is designed to reconstruct both cluster assignments and cluster centers from masked point clouds, thus encouraging the model to capture dense semantic information. Additionally, we introduce a global contrastive learning mechanism that enhances instance-level feature learning by contrasting different masked views of the same point cloud. By jointly optimizing these complementary objectives, i.e., dense semantic reconstruction, and instance-level contrastive learning. MaskClu enables ViTs to learn richer and more semantically meaningful representations from 3D point clouds. We validate the effectiveness of MaskClu via multiple 3D tasks, including part segmentation, semantic segmentation, object detection, and classification, setting new competitive results. Bin Ren 0005, Xiaoshui Huang, Mengyuan Liu 0001, Hong Liu 0008, Fabio Poiesi, Nicu Sebe, Guofeng Mei |
AAAI | 5 |
| 2026 | High-Resolution Open-Vocabulary Object 6D Pose EstimationabstractThe generalisation to unseen objects in the 6D pose estimation task is very challenging. While Vision-Language Models (VLMs) enable using natural language descriptions to support 6D pose estimation of unseen objects, these solutions underperform compared to model-based methods. In this work we present Horyon, an open-vocabulary VLM-based architecture that addresses relative pose estimation between two scenes of an unseen object, described by a textual prompt only. We use the textual prompt to identify the unseen object in the scenes and then obtain high-resolution multi-scale features. These features are used to extract cross-scene matches for registration. We evaluate our model on a benchmark with a large variety of unseen objects across four datasets, namely REAL275, Toyota-Light, Linemod, and YCB-Video. Our method achieves state-of-the-art performance on all datasets, outperforming by 12.6 in Average Recall the previous best-performing approach. Jaime Corsetti, Davide Boscaini, Francesco Giuliari, Changjae Oh, Andrea Cavallaro, Fabio Poiesi |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Vocabulary-Free 3D Instance Segmentation with Vision-Language AssistantabstractMost recent 3D instance segmentation methods are open vocabulary, offering a greater flexibility than closedvocabulary methods. Yet, they are limited to reasoning within a specific set of concepts, i.e., the vocabulary, prompted by the user at test time. In essence, these models cannot reason in an open-ended fashion, i.e., answering “List the objects in the scene.” We introduce the first method to address 3D instance segmentation in a setting that is void of any vocabulary prior, namely a vocabularyfree setting. We leverage a large vision-language assistant and an open-vocabulary$2 D$instance segmenter to discover and ground semantic categories on the posed images. To form 3D instance masks, we first partition the input point cloud into dense superpoints, which are then merged into 3D instance masks. We propose a novel superpoint merging strategy via spectral clustering, accounting for both mask coherence and semantic coherence that are estimated from the 2D object instance masks. We evaluate our method using ScanNet200 and Replica, outperforming existing methods in both vocabulary-free and open-vocabulary settings. Guofeng Mei, Luigi Riz, Yiming Wang 0002, Fabio Poiesi |
3DV | 4 |
| 2025 | Fully-Geometric Cross-Attention for Point Cloud RegistrationabstractPoint cloud registration approaches often fail when the overlap between point clouds is low due to noisy point correspondences. This work introduces a novel cross-attention mechanism tailored for Transformer-based architectures that tackles this problem, by fusing information from coordinates and features at the super-point level between point clouds. This formulation has remained unexplored primarily because it must guarantee rotation and translation invariance since point clouds reside in different and independent reference frames. We integrate the Gromov-Wasserstein distance into the cross-attention formulation to jointly compute distances between points across different point clouds and account for their geometric structure. By doing so, points from two distinct point clouds can attend to each other under arbitrary rigid transformations. At the point level, we also devise a self-attention mechanism that aggregates the local geometric structure information into point features for fine matching. Our formulation boosts the number of inlier correspondences, thereby yielding more precise registration results compared to state-of-the-art approaches. We have conducted an extensive evaluation on 3DMatch, 3DLoMatch, KITTI, and 3DCSR datasets. Project page: https://github.com/twowwj/FLAT. Weijie Wang 0002, Guofeng Mei, Jian Zhang 0002, Nicu Sebe, Bruno Lepri, Fabio Poiesi |
3DV | 6 |
| 2025 | Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene UnderstandingabstractThe lack of a large-scale 3D-text corpus has led recent works to distill open-vocabulary knowledge from vision-language models (VLMs). However, these methods typically rely on a single VLM to align the feature spaces of 3D models within a common language space, which limits the potential of 3D models to leverage the diverse spatial and semantic capabilities encapsulated in various foundation models. In this paper, we propose Cross-modal and Uncertainty-aware Agglomeration for Open-vocabulary 3D Scene Understanding dubbed CUA-O3D, the first model to integrate multiple foundation models—such as CLIP, DINOv2, and Stable Diffusion—into 3D scene understanding. We further introduce a deterministic uncertainty estimation to adaptively distill and harmonize the heterogeneous 2D feature embeddings from these models. Our method addresses two key challenges: (1) incorporating semantic priors from VLMs alongside the geometric knowledge of spatially-aware vision foundation models, and (2) using a novel deterministic uncertainty estimation to capture model-specific uncertainties across diverse semantic and geometric sensitivities, helping to reconcile heterogeneous representations during training. Extensive experiments on ScanNetV2 and Matterport3D demonstrate that our method not only advances open-vocabulary segmentation but also achieves robust cross-domain alignment and competitive spatial perception capabilities. Project webpage: CUA-O3D. Jinlong Li 0003, Cristiano Saltori, Fabio Poiesi, Nicu Sebe |
CVPR | 3 |
| 2025 | Functionality Understanding and Segmentation in 3D ScenesabstractUnderstanding functionalities in 3D scenes involves interpreting natural language descriptions to locate functional interactive objects, such as handles and buttons, in a 3D environment. Functionality understanding is highly challenging, as it requires both world knowledge to interpret language and spatial perception to identify fine-grained objects. For example, given a task like ‘turn on the ceiling light,’ an embodied AI agent must infer that it needs to locate the light switch, even though the switch is not explicitly mentioned in the task description. To date, no dedicated methods have been developed for this problem. In this paper, we introduce Fun3DU, the first approach designed for functionality understanding in 3D scenes. Fun3DU uses a language model to parse the task description through Chain-of-Thought reasoning in order to identify the object of interest. The identified object is segmented across multiple views of the captured scene by using a vision and language model. The segmentation results from each view are lifted in 3D and aggregated into the point cloud using geometric information. Fun3DU is training-free, relying entirely on pre-trained models. We evaluate Fun3DU on SceneFun3D, the most recent and only dataset to benchmark this task, which comprises over 3000 task descriptions on 230 scenes. Our method significantly outperforms state-of-the-art open-vocabulary 3D segmentation approaches. Project page: https://tev-fbk.github.io/fun3du/ Jaime Corsetti, Francesco Giuliari, Alice Fasoli, Davide Boscaini, Fabio Poiesi |
CVPR | 5 |
| 2025 | PerLA: Perceptive 3D Language AssistantabstractEnabling Large Language Models (LLMs) to understand the 3D physical world is an emerging yet challenging research direction. Current strategies for processing point clouds typically downsample the scene or divide it into smaller parts for separate analysis. However, both approaches risk losing key local details or global contextual information. In this paper, we introduce PerLA, a 3D language assistant designed to be more perceptive to both details and context, making visual representations more informative for the LLM. PerLA captures high-resolution (local) details in parallel from different point cloud areas and integrates them with (global) context obtained from a lower-resolution whole point cloud. We present a novel algorithm that preserves point cloud locality through the Hilbert curve and effectively aggregates local-to-global information via cross-attention and a graph neural network. Lastly, we introduce a novel loss for local representation consensus to promote training stability. PerLA outperforms state-of-the-art 3D language assistants, with gains of up to +1.34 CiDEr on ScanQA for question answering, and +4.22 on ScanRefer and +3.88 on Nr3D for dense captioning. Project page: https://gfmei.github.io/PerLA Guofeng Mei, Luigi Riz, Yujiao Wu, Fabio Poiesi, Yiming Wang 0002 |
CVPR | 5 |
| 2025 | GRASPLAT: Enabling dexterous grasping through novel view synthesisabstractAchieving dexterous robotic grasping with multi-fingered hands remains a significant challenge. While existing methods rely on complete 3D scans to predict grasp poses, these approaches face limitations due to the difficulty of acquiring high-quality 3D data in real-world scenarios. In this paper, we introduce GRASPLAT, a novel grasping framework that leverages consistent 3D information while being trained solely on RGB images. Our key insight is that by synthesizing physically plausible images of a hand grasping an object, we can regress the corresponding hand joints for a successful grasp. To achieve this, we utilize 3D Gaussian Splatting to generate high-fidelity novel views of real hand-object interactions, enabling end-to-end training with RGB data. Unlike prior methods, our approach incorporates a photometric loss that refines grasp predictions by minimizing discrepancies between rendered and real images. We conduct extensive experiments on both synthetic and real-world grasping datasets, demonstrating that GRASPLAT improves grasp success rates up to 36.9% over existing image-based methods. Project page: https://mbortolon97.github.io/grasplat/ Matteo Bortolon, Nuno Ferreira Duarte, Plinio Moreno, Fabio Poiesi, José Santos-Victor, Alessio Del Bue |
IROS | 4 |
| 2025 | Distilling 3D distinctive local descriptors for 6D pose estimationabstractThree-dimensional local descriptors are crucial for encoding geometric surface properties, making them essential for various point cloud understanding tasks. Among these descriptors, GeDi has demonstrated strong zero-shot 6D pose estimation capabilities but remains computationally impractical for real-world applications due to its expensive inference process. Can we retain GeDi’s effectiveness, while significantly improving its efficiency? In this paper, we explore this question by introducing a knowledge distillation framework that trains an efficient student model to regress local descriptors from a GeDi teacher. Our key contributions include: an efficient large-scale training procedure that ensures robustness to occlusions and partial observations while operating under compute and storage constraints, and a novel loss formulation that handles weak supervision from non-distinctive teacher descriptors. We validate our approach on five BOP Benchmark datasets and demonstrate a significant reduction in inference time while maintaining competitive performance with existing methods, bringing zero-shot 6D pose estimation closer to real-time feasibility. Project website: https://tev-fbk.github.io/dGeDi. Amir Hamza, Andrea Caraffa, Davide Boscaini, Fabio Poiesi |
IROS | 4 |
| 2025 | Free-form language-based robotic reasoning and graspingabstractPerforming robotic grasping from a cluttered bin based on human instructions is a challenging task, as it requires understanding both the nuances of free-form language and the spatial relationships between objects. Vision-Language Models (VLMs) trained on web-scale data, such as GPT-4o, have demonstrated remarkable reasoning capabilities across both text and images. But can they truly be used for this task in a zero-shot setting? And what are their limitations? In this paper, we explore these research questions via the free-form language-based robotic grasping task and propose a novel method, FreeGrasp, leveraging the pre-trained VLMs’ world knowledge to reason about human instructions and object spatial arrangements. Our method detects all objects as keypoints and uses these keypoints to annotate marks on images, aiming to facilitate GPT-4o’s zero-shot spatial reasoning. This allows our method to determine whether a requested object is directly graspable or if other objects must be grasped and removed first. Since no existing dataset is specifically designed for this task, we introduce a synthetic dataset FreeGraspData by extending the MetaGraspNetV2 dataset with human-annotated instructions and ground-truth grasping sequences. We conduct extensive analyses with both FreeGraspData and real-world validation with a gripper-equipped robotic arm, demonstrating state-of-the-art performance in grasp reasoning and execution. Project website: https://tev-fbk.github.io/FreeGrasp/. Runyu Jiao, Alice Fasoli, Francesco Giuliari, Matteo Bortolon, Sergio Povoli, Guofeng Mei, Yiming Wang 0002, Fabio Poiesi |
IROS | 8 |
| 2025 | OpenHype: Hyperbolic Embeddings for Hierarchical Open-Vocabulary Radiance FieldsabstractModeling the inherent hierarchical structure of 3D objects and 3D scenes is highly desirable, as it enables a more holistic understanding of environments for autonomous agents. Accomplishing this with implicit representations, such as Neural Radiance Fields, remains an unexplored challenge. Existing methods that explicitly model hierarchical structures often face significant limitations: they either require multiple rendering passes to capture embeddings at different levels of granularity, significantly increasing inference time, or rely on predefined, closed-set discrete hierarchies that generalize poorly to the diverse and nuanced structures encountered by agents in the real world. To address these challenges, we propose OpenHype, a novel approach that represents scene hierarchies using a continuous hyperbolic latent space. By leveraging the properties of hyperbolic geometry, OpenHype naturally encodes multi-scale relationships and enables smooth traversal of hierarchies through geodesic paths in latent space. Our method outperforms state-of-the-art approaches on standard benchmarks, demonstrating superior efficiency and adaptability in 3D scene understanding. Lisa Weijler, Fabio Poiesi, Timo Ropinski, Pedro Hermosilla |
NeurIPS | 3 |
| 2025 | 3D Part Segmentation via Geometric Aggregation of 2D Visual FeaturesabstractSupervised 3D part segmentation models are tailored for a fixed set of objects and parts, limiting their transferability to open-set, real-world scenarios. Recent works have explored vision-language models (VLMs) as a promising alternative, using multi-view rendering and textual prompting to identify object parts. However, naively applying VLMs in this context introduces several drawbacks, such as the need for meticulous prompt engineering, and fails to leverage the 3D geometric structure of objects. To address these limitations, we propose COPS, a COmprehensive model for Parts Segmentation that blends the semantics extracted from visual concepts and 3D geometry to effectively identify object parts. COPS renders a point cloud from multiple viewpoints, extracts 2D features, projects them back to 3D, and uses a novel geometric-aware feature aggregation procedure to ensure spatial and semantic consistency. Finally, it clusters points into parts and labels them. We demonstrate that COPS is efficient, scalable, and achieves zero-shot state-of-the-art performance across five datasets, covering synthetic and real-world data, texture-less and coloured objects, as well as rigid and non-rigid shapes. The code is available at https://3d-cops.github.io. Marco Garosi, Riccardo Tedoldi, Davide Boscaini, Massimiliano Mancini, Nicu Sebe, Fabio Poiesi |
WACV | 6 |
| 2025 | Novel Class Discovery Meets Foundation Models for 3D Semantic Segmentation
Luigi Riz, Cristiano Saltori, Yiming Wang 0002, Elisa Ricci 0001, Fabio Poiesi |
Int. J. Comput. Vis. | 5 |
| 2024 | Open-vocabulary object 6D pose estimationabstractWe introduce the new setting of open-vocabulary object 6D pose estimation, in which a textual prompt is used to specify the object of interest. In contrast to existing approaches, in our setting (i) the object of interest is speci-fied solely through the textual prompt, (ii) no object model (e.g., CAD or video sequence) is required at inference, and (iii) the object is imaged from two RGBD viewpoints of dif-ferent scenes. To operate in this setting, we introduce a novel approach that leverages a Vision-Language Model to segment the object of interest from the scenes and to esti-mate its relative 6D pose. The key of our approach is a carefully devised strategy to fuse object-level information provided by the prompt with local image features, resulting in a feature space that can generalize to novel concepts. We validate our approach on a new benchmark based on two popular datasets, REAL275 and Toyota-Light, which collectively encompass 34 object instances appearing in four thousand image pairs. The results demonstrate that our approach outperforms both a well-established hand-crafted method and a recent deep learning-based base-line in estimating the relative 6D pose of objects in dif-ferent scenes. Code and dataset are available at https://jcorsetti.github.io/oryon. Jaime Corsetti, Davide Boscaini, Changjae Oh, Andrea Cavallaro, Fabio Poiesi |
CVPR | 5 |
| 2024 | Geometrically-Driven Aggregation for Zero-Shot 3D Point Cloud UnderstandingabstractZero-shot 3D point cloud understanding can be achieved via 2D Vision-Language Models (VLMs). Existing strategies directly map VLM representations from 2D pixels of rendered or captured views to 3D points, overlooking the inherent and expressible point cloud geometric structure. Geometrically similar or close regions can be exploited for bolstering point cloud understanding as they are likely to share semantic information. To this end, we introduce the first training-free aggregation technique that leverages the point cloud's 3D geometric structure to improve the quality of the transferred VLM representations. Our approach operates iteratively, performing local-to-global aggregation based on geometric and semantic point-level reasoning. We benchmark our approach on three downstream tasks, including classification, part segmentation, and semantic segmentation, with a variety of datasets representing both synthetic/real-world, and indoor/outdoor scenarios. Our approach achieves new state-of-the-art results in all benchmarks. Code and dataset are available at https://luigiriz.github.io/geoze-website/ Guofeng Mei, Luigi Riz, Yiming Wang 0002, Fabio Poiesi |
CVPR | 4 |
| 2024 | 6DGS: 6D Pose Estimation from a Single Image and a 3D Gaussian Splatting Model
Matteo Bortolon, Theodore Tsesmelis, Stuart James, Fabio Poiesi, Alessio Del Bue |
ECCV (52) | 4 |
| 2024 | FreeZe: Training-Free Zero-Shot 6D Pose Estimation with Geometric and Vision Foundation Models
Andrea Caraffa, Davide Boscaini, Amir Hamza, Fabio Poiesi |
ECCV (75) | 4 |
| 2024 | IFFNeRF: Initialisation Free and Fast 6DoF pose estimation from a single image and a NeRF modelabstractWe introduce IFFNeRF to estimate the six degrees-of-freedom (6DoF) camera pose of a given image, building on the Neural Radiance Fields (NeRF) formulation. IFFNeRF is specifically designed to operate in real-time and eliminates the need for an initial pose guess that is proximate to the sought solution. IFFNeRF utilizes the Metropolis-Hasting algorithm to sample surface points from within the NeRF model. From these sampled points, we cast rays and deduce the color for each ray through pixel-level view synthesis. The camera pose can then be estimated as the solution to a Least Squares problem by selecting correspondences between the query image and the resulting bundle. We facilitate this process through a learned attention mechanism, bridging the query image embedding with the embedding of parameterized rays, thereby matching rays pertinent to the image. Through synthetic and real evaluation settings, we show that our method can improve the angular and translation error accuracy by 80.1% and 67.3%, respectively, compared to iNeRF while performing at 34fps on consumer hardware and not requiring the initial pose guess. Project page: https://mbortolon97.github.io/frenerf/ Matteo Bortolon, Theodore Tsesmelis, Stuart James, Fabio Poiesi, Alessio Del Bue |
ICRA | 4 |
| 2024 | Delving into CLIP latent space for Video Anomaly Recognition
Luca Zanella, Benedetta Liberatori, Willi Menapace, Fabio Poiesi, Yiming Wang 0002, Elisa Ricci 0001 |
Comput. Vis. Image Underst. | 4 |
| 2024 | Unsupervised Point Cloud Representation Learning by Clustering and Neural RenderingabstractAbstract Data augmentation has contributed to the rapid advancement of unsupervised learning on 3D point clouds. However, we argue that data augmentation is not ideal, as it requires a careful application-dependent selection of the types of augmentations to be performed, thus potentially biasing the information learned by the network during self-training. Moreover, several unsupervised methods only focus on uni-modal information, thus potentially introducing challenges in the case of sparse and textureless point clouds. To address these issues, we propose an augmentation-free unsupervised approach for point clouds, named CluRender, to learn transferable point-level features by leveraging uni-modal information for soft clustering and cross-modal information for neural rendering. Soft clustering enables self-training through a pseudo-label prediction task, where the affiliation of points to their clusters is used as a proxy under the constraint that these pseudo-labels divide the point cloud into approximate equal partitions. This allows us to formulate a clustering loss to minimize the standard cross-entropy between pseudo and predicted labels. Neural rendering generates photorealistic renderings from various viewpoints to transfer photometric cues from 2D images to the features. The consistency between rendered and real images is then measured to form a fitting loss, combined with the cross-entropy loss to self-train networks. Experiments on downstream applications, including 3D object detection, semantic segmentation, classification, part segmentation, and few-shot learning, demonstrate the effectiveness of our framework in outperforming state-of-the-art techniques. Guofeng Mei, Cristiano Saltori, Elisa Ricci 0001, Nicu Sebe, Qiang Wu 0001, Jian Zhang 0002, Fabio Poiesi |
Int. J. Comput. Vis. | 7 |
| 2024 | Special Issue on Deep Learning for Intelligent Human Computer InteractionabstractSpecial Issue on Deep Learning for Intelligent Human Computer InteractionDeep Learning ( DL ) is growing at a fast pace with a plethora of related research being conducted and industrial applications being developed.Robotics, computer graphics, computer vision including areas such as feature extraction/matching, 3D reconstruction, manufacturing, medicine, knowledge acquisition, control theory, planning and scheduling, among others, have uncovered the potential of DL.The reason for such success mainly stems from two aspects.On the one hand, the theoretical advances in related disciplines such as optimization methods, pattern recognition, and hardware (GPUs) have undergone great breakthroughs.On the other hand, the technical applications in industry have addressed many actual problems that have in turn accelerated DL development.Leading technology companies like Microsoft, Apple, Google, Facebook, Nvidia, and Amazon all launched cutting-edge industrial products offering advanced functionalities that were only made possible thanks to the employment of DL algorithms.The effort in establishing new media and network technologies led to the identification of a niche for ML in the field of humancomputer interaction.Human-computer interaction ( HCI ) is a multidisciplinary field of study focusing on the design of computer technology through the analysis and understanding of how humans interact with computers.HCI addresses various areas, such as user interface design or computer-supported cooperative work for data modeling, system adaptation and optimization to fit user needs, complex systems, the development of smart equipment to ease utilization, and the design of smart environments to promote user comfort and safety.Nowadays, DL is largely employed in HCI to model human behavior as they have shown to outperform traditional modeling techniques.Although their potential is undoubtedly broad, there are several important research challenges that remain unaddressed.The explainability and interpretability of the decision-making process of DL approaches is key to assess models' behaviors, especially when unexpected or incorrect outputs occur.The complexity and the dimensionality of the underlying mathematical models are among the principal causes that hinder understanding.Privacy must be guaranteed as sensible private data is typically continually collected, processed, and stored over the Cloud in order to provide improved user-specific experiences.Such continual learning for DL approaches is itself a challenge as models may suffer from catastrophic forgetting, but it is a key ingredient that needs to be taken into consideration to enable scalability.Data annotation done at large scale is costly and hardly sustainable, therefore particular attention must be put on self-supervised or unsupervised learning mechanisms.DL-based user interfaces, such as gesture recognition, must be supported Zhihan Lyu, Fabio Poiesi, Qi Dong 0004, Jaime Lloret Mauri, Houbing Song |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Detect, Augment, Compose, and Adapt: Four Steps for Unsupervised Domain Adaptation in Object Detection
Mohamed Lamine Mekhalfi, Davide Boscaini, Fabio Poiesi |
BMVC | 3 |
| 2023 | Novel Class Discovery for 3D Point Cloud Semantic SegmentationabstractNovel class discovery (NCD) for semantic segmentation is the task of learning a model that can segment unlabelled (novel) classes using only the supervision from labelled (base) classes. This problem has recently been pioneered for 2D image data, but no work exists for 3D point cloud data. In fact, the assumptions made for 2D are loosely applicable to 3D in this case. This paper is presented to advance the state of the art on point cloud data analysis in four directions. Firstly, we address the new problem of NCD for point cloud semantic segmentation. Secondly, we show that the transposition of the only existing NCD method for 2D semantic segmentation to 3D data is sub-optimal. Thirdly, we present a new method for NCD based on online clustering that exploits uncertainty quantification to produce prototypes for pseudo-labelling the points of the novel classes. Lastly, we introduce a new evaluation protocol to assess the performance of NCD for point cloud semantic segmentation. We thoroughly evaluate our method on SemanticKITTI and SemanticPOSS datasets, showing that it can significantly outperform the baseline. Project page: https://github.com/LuigiRiz/NOPS. Luigi Riz, Cristiano Saltori, Elisa Ricci 0001, Fabio Poiesi |
CVPR | 4 |
| 2023 | Overlap-guided Gaussian Mixture Models for Point Cloud RegistrationabstractProbabilistic 3D point cloud registration methods have shown competitive performance in overcoming noise, outliers, and density variations. However, registering point cloud pairs in the case of partial overlap is still a challenge. This paper proposes a novel overlap-guided probabilistic registration approach that computes the optimal transformation from matched Gaussian Mixture Model (GMM) parameters. We reformulate the registration problem as the problem of aligning two Gaussian mixtures such that a statistical discrepancy measure between the two corresponding mixtures is minimized. We introduce a Transformer-based detection module to detect overlapping regions, and represent the input point clouds using GMMs by guiding their alignment through overlap scores computed by this detection module. Experiments show that our method achieves superior registration accuracy and efficiency than state-of-the-art methods when handling point clouds with partial overlap and different densities on synthetic and real-world datasets. https://github.com/gfmei/ogmm Guofeng Mei, Fabio Poiesi, Cristiano Saltori, Jian Zhang 0002, Elisa Ricci 0001, Nicu Sebe |
WACV | 2 |
| 2023 | PatchMixer: Rethinking network design to boost generalization for 3D point cloud understanding
Davide Boscaini, Fabio Poiesi |
Image Vis. Comput. | 2 |
| 2023 | Learning General and Distinctive 3D Local Deep Descriptors for Point Cloud RegistrationabstractAn effective 3D descriptor should be invariant to different geometric transformations, such as scale and rotation, robust to occlusions and clutter, and capable of generalising to different application domains. We present a simple yet effective method to learn general and distinctive 3D local descriptors that can be used to register point clouds that are captured in different domains. Point cloud patches are extracted, canonicalised with respect to their local reference frame, and encoded into scale and rotation-invariant compact descriptors by a deep neural network that is invariant to permutations of the input points. This design is what enables our descriptors to generalise across domains. We evaluate and compare our descriptors with alternative handcrafted and deep learning-based descriptors on several indoor and outdoor datasets that are reconstructed by using both RGBD sensors and laser scanners. Our descriptors outperform most recent descriptors by a large margin in terms of generalisation, and also become the state of the art in benchmarks where training and testing are performed in the same domain. Fabio Poiesi, Davide Boscaini |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Compositional Semantic Mix for Domain Adaptation in Point Cloud SegmentationabstractDeep-learning models for 3D point cloud semantic segmentation exhibit limited generalization capabilities when trained and tested on data captured with different sensors or in varying environments due to domain shift. Domain adaptation methods can be employed to mitigate this domain shift, for instance, by simulating sensor noise, developing domain-agnostic generators, or training point cloud completion networks. Often, these methods are tailored for range view maps or necessitate multi-modal input. In contrast, domain adaptation in the image domain can be executed through sample mixing, which emphasizes input data manipulation rather than employing distinct adaptation modules. In this study, we introduce compositional semantic mixing for point cloud domain adaptation, representing the first unsupervised domain adaptation technique for point cloud segmentation based on semantic and geometric sample mixing. We present a two-branch symmetric network architecture capable of concurrently processing point clouds from a source domain (e.g. synthetic) and point clouds from a target domain (e.g. real-world). Each branch operates within one domain by integrating selected data fragments from the other domain and utilizing semantic information derived from source labels and target (pseudo) labels. Additionally, our method can leverage a limited number of human point-level annotations (semi-supervised) to further enhance performance. We assess our approach in both synthetic-to-real and real-to-real scenarios using LiDAR datasets and demonstrate that it significantly outperforms state-of-the-art methods in both unsupervised and semi-supervised settings. Cristiano Saltori, Fabio Galasso, Giuseppe Fiameni, Nicu Sebe, Fabio Poiesi, Elisa Ricci 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Detection-aware multi-object tracking evaluationabstractHow would you fairly evaluate two multi-object tracking algorithms (i.e. trackers), each one employing a different object detector? Detectors keep improving, thus trackers can make less effort to estimate object states over time. Is it then fair to compare a new tracker employing a new detector with another tracker using an old detector? In this paper, we propose a novel performance measure, named Tracking Effort Measure (TEM), to evaluate trackers that use different detectors. TEM estimates the improvement that the tracker does with respect to its input data (i.e. detections) at frame level (intra-frame complexity) and sequence level (inter-frame complexity). We evaluate TEM over well-known datasets, four trackers and eight detection sets. Results show that, unlike conventional tracking evaluation measures, TEM can quantify the effort done by the tracker with a reduced correlation on the input detections. Its implementation will be made publicly available online.1.1https://github.com/vpulab/MOT-evaluation Juan C. SanMiguel, Jorge Muñoz, Fabio Poiesi |
AVSS | 3 |
| 2022 | Data Augmentation-free Unsupervised Learning for 3D Point Cloud Understanding
Guofeng Mei, Cristiano Saltori, Fabio Poiesi, Jian Zhang 0002, Elisa Ricci 0001, Nicu Sebe, Qiang Wu 0001 |
BMVC | 3 |
| 2022 | CoSMix: Compositional Semantic Mix for Domain Adaptation in 3D LiDAR Segmentation
Cristiano Saltori, Fabio Galasso, Giuseppe Fiameni, Nicu Sebe, Elisa Ricci 0001, Fabio Poiesi |
ECCV (33) | 6 |
| 2022 | GIPSO: Geometrically Informed Propagation for Online Adaptation in 3D LiDAR Segmentation
Cristiano Saltori, Evgeny Krivosheev, Stéphane Lathuilière, Nicu Sebe, Fabio Galasso, Giuseppe Fiameni, Elisa Ricci 0001, Fabio Poiesi |
ECCV (33) | 8 |
| 2022 | Virtual-reality and intelligent hardware in digital twinsabstractSeveral new models and formats for the digital transformation of the manufacturing industry appear because of the rapid integration of information technology and the real economy, as well as the increasingly obvious evolution trend of industrial digitalization, networking, and intelligence.Among them, digital twins have increasingly become a research hotspot in all sectors of the industry and have broad prospects.It maps physical objects in virtual space in a digital way and simulates their behavioral characteristics in real environments.It makes the gap between virtuality and reality disappear based on their closed-loop interaction.Digital twins are undoubtedly an important and strategic technology in response to familiar products, production, and services.It can also speculate some indicators that cannot be directly measured by machine learning through collecting the direct data of limited physical sensor indicators.This can realize an assessment of the current state, a diagnosis of past problems, and a prediction of future trends, and simulate possibilities to provide more comprehensive decision support.Driven by "Industry 4.0/5.0", the concept of "Digital Twins" is setting off a new wave of industrial simulation boom.Entity industrial objects or systems are fully digitized, and models including virtual objects, virtual processes, and virtual plant areas are established.These digital twins can be placed in virtual environments to analyze, simulate, verify, test, and adjust various situations such as product or process optimization, optimizing the operation of real objects or systems, and increasing their added value.Also, the expression of "virtual reality" for scene visualization and the new scene interaction mode will be more conducive to better visual effects and interactive operations in the digital world.The combination of digital twins and virtual reality is currently widely used as guidance instructions for operation and maintenance, early warning of error prevention, and information tips to improve the efficiency and accuracy of operators.Virtual reality brings users better vision and senses through smart hardware like glasses and helmets.However, there is still room for improvement in the technical level of smart hardware.And rational hardware configuration and cost-reduction can be realized gradually with the improvement of R & D. Now the focus in this field should be how to enhance user experience and enrich products.In "Integrating digital twins and deep learning for medical image analysis in the era of COVID-19", the authors introduce a digital-twin-based smart healthcare system integrated with medical devices to collect information regarding the current health condition, configuration, and maintenance history of the device/ machine/system.Furthermore, medical images, that is, X-rays, are analyzed by using a deep-learning model to detect the infection of COVID-19.The designed system is based on the cascade recurrent convolution neural network (RCNN) architecture.In this architecture, the detector stages are deeper and more sequentially selective against small and close false positives.This architecture is a multi-stage extension of the RCNN model and sequentially trained using the output of one stage for training the other.At each stage, the bounding Zhihan Lyu, Gustavo Marfia, Fabio Poiesi, Neil Vaughan, Jun Shen 0001 |
Virtual Real. Intell. Hardw. | 3 |
| 2022 | Virtual-reality and intelligent hardware in Digital Twins
Zhihan Lyu, Gustavo Marfia, Fabio Poiesi, Neil Vaughan, Jun Shen 0001 |
Virtual Real. Intell. Hardw. | 3 |
| 2020 | Novel-View Human Action Synthesis
Mohamed Ilyes Lakhal, Davide Boscaini, Fabio Poiesi, Oswald Lanz, Andrea Cavallaro |
ACCV (4) | 3 |
| 2020 | Distinctive 3D local deep descriptorsabstractWe present a simple but yet effective method for learning distinctive 3D local deep descriptors (DIPs) that can be used to register point clouds without requiring an initial alignment. Point cloud patches are extracted, canonicalised with respect to their estimated local reference frame and encoded into rotation-invariant compact descriptors by a PointNet-based deep neural network. DIPs can effectively generalise across different sensor modalities because they are learnt end-to-end from locally and randomly sampled points. Because DIPs encode only local geometric information, they are robust to clutter, occlusions and missing regions. We evaluate and compare DIPs against alternative hand-crafted and deep descriptors on several indoor and outdoor datasets consiting of point clouds reconstructed using different sensors. Results show that DIPs (i) achieve comparable results to the state-of-the-art on RGB-D indoor scenes (3DMatch dataset), (ii) outperform state-of-the-art by a large margin on laser-scanner outdoor scenes (ETH dataset), and (iii) generalise to indoor scenes reconstructed with the Visual-SLAM system of Android ARCore. Source code: https://github.com/fabiopoiesi/dip. Fabio Poiesi, Davide Boscaini |
ICPR | 1 |
| 2018 | A distributed vision-based consensus model for aerial-robotic teamsabstractWe present a distributed model for a team of autonomous aerial robots to collaboratively track a target without external control. The model uses distributed consensus to coordinate actions and to maintain formation via geometric constraints. Each robot uses its ego-centric view of a target and the relative distance from its two closest neighbors to infer its steering commands. To account for noisy and missing target detections, the robots exchange their estimated target position and formation configuration through shared PID-controlled steering responses. We show that the proposed model enables the team to maintain the view of a maneuvering target with varying acceleration under noisy detections and failures up to situations when all robots but one lose the target from their field of view. Fabio Poiesi, Andrea Cavallaro |
IROS | 1 |
| 2017 | Support Vector Motion ClusteringabstractWe present a closed-loop unsupervised clustering method for motion vectors extracted from highly dynamic video scenes. Motion vectors are assigned to nonconvex homogeneous clusters characterizing direction, size and shape of regions with multiple independent activities. The proposed method is based on support vector clustering. Cluster labels are propagated over time via incremental learning. The proposed method uses a kernel function that maps the input motion vectors into a high-dimensional space to produce nonconvex clusters. We improve the mapping effectiveness by quantifying feature similarities via a blend of position and orientation affinities. We use the Quasiconformal Kernel Transformation to boost the discrimination of outliers. The temporal propagation of the clusters’ identities is achieved via incremental learning based on the concept of feature obsolescence to deal with appearing and disappearing features. Moreover, we design an online clustering performance prediction algorithm used as a feedback that refines the cluster model at each frame in an unsupervised manner. We evaluate the proposed method on synthetic data sets and real-world crowded videos and show that our solution outperforms state-of-the-art approaches. Isah Abdullahi Lawal, Fabio Poiesi, Davide Anguita, Andrea Cavallaro |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Detection of fast incoming objects with a moving camera
Fabio Poiesi, Andrea Cavallaro |
BMVC | 1 |
| 2015 | Distributed vision-based flying cameras to film a moving targetabstractFormations of camera-equipped quadrotors (flying cameras) have the actuation agility to track moving targets from multiple viewing angles. In this paper we propose an infrastructure-free distributed control method for multiple flying cameras tracking a moving object. The proposed vision-based servoing can deal with noisy and missing target observations, accounts for quadrotor oscillations and does not require an external positioning system. The flight direction of each camera is inferred via geometric derivation, and the formation is maintained by employing a distributed algorithm that uses the target position information on the camera plane and the position of neighboring flying cameras. Simulations show that the proposed solution enables the tracking of a moving target by the cameras flying in formation despite noisy target detections and when the target is outside some of the fields of view. Fabio Poiesi, Andrea Cavallaro |
IROS | 1 |
| 2015 | Tracking Multiple High-Density Homogeneous TargetsabstractWe present a framework for multitarget detection and tracking that infers candidate target locations in videos containing a high density of homogeneous targets. We propose a gradient-climbing technique and an isocontor slicing approach for intensity maps to localize targets. The former uses Markov chain Monte Carlo to iteratively fit a shape model onto the target locations, whereas the latter uses the intensity values at different levels to find consistent object shapes. We generate trajectories by recursively associating detections with a hierarchical graph-based tracker on temporal windows. The solution to the graph is obtained with a greedy algorithm that accounts for false-positive associations. The edges of the graph are weighted with a likelihood function based on location information. We evaluate the performance of the proposed framework on challenging datasets containing videos with high density of targets and compare it with six alternative trackers. Fabio Poiesi, Andrea Cavallaro |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Assessing tracking assessment measuresabstractWe propose a methodology to quantitatively compare the relative performance of tracking evaluation measures. The proposed methodology is based on determining the probabilistic agreement between tracking result decisions made by measures and those made by humans. We use tracking results on publicly available datasets with different target types and varying challenges, and collect the judgments of 90 skilled, semi-skilled and unskilled human subjects using a web-based performance assessment test. The analysis of the agreements allows us to highlight the variation in performance of the different measures and the most appropriate ones for the various stages of tracking performance evaluation. Tahir Nawaz 0001, Fabio Poiesi, Andrea Cavallaro |
ICIP | 2 |
| 2014 | Measures of Effective Video TrackingabstractTo evaluate multitarget video tracking results, one needs to quantify the accuracy of the estimated target-size and the cardinality error as well as measure the frequency of occurrence of ID changes. In this paper, we survey existing multitarget tracking performance scores and, after discussing their limitations, we propose three parameter-independent measures for evaluating multitarget video tracking. The measures consider target-size variations, combine accuracy and cardinality errors, quantify long-term tracking accuracy at different accuracy levels, and evaluate ID changes relative to the duration of the track in which they occur. We conduct an extensive experimental validation of the proposed measures by comparing them with existing ones and by evaluating four state-of-the-art trackers on challenging real-world publicly-available data sets. The software implementing the proposed measures is made available online to facilitate their use by the research community. Tahir Nawaz 0001, Fabio Poiesi, Andrea Cavallaro |
IEEE Trans. Image Process. | 2 |
| 2013 | Detection and tracking of groups in crowdabstractWe propose a method to detect and track interacting people by employing a framework based on a Social Force Model (SFM). The method embeds plausible human behaviors to predict interactions in a crowd by iteratively minimizing the error between predictions and measurements. We model people approaching a group and restrict the group formation based on the relative velocity of candidate group members. The detected groups are then tracked by linking their interaction centers over time using a buffered graph-based tracker. We show how the proposed framework outperforms existing group localization techniques on three publicly available datasets, with improvements of up to 13% on group detection. Riccardo Mazzon, Fabio Poiesi, Andrea Cavallaro |
AVSS | 2 |
| 2013 | Multi-target tracking on confidence maps: An application to people tracking
Fabio Poiesi, Riccardo Mazzon, Andrea Cavallaro |
Comput. Vis. Image Underst. | 1 |
| 2010 | Detector-less ball localization using context and motion flow analysisabstractWe present a technique for estimating the location of the ball during a basketball game without using a detector. The technique is based on the analysis of the dynamics in the scene and allows us to overcome the challenges due to frequent occlusions of the ball and its similarity in appearance with the background. Based on the assumption that the ball is the point of focus of the game and that the motion flow of the players is dependent on its position during attack actions, the most probable candidates for the ball location are extracted from each frame. These candidates are then validated over time using a Kalman filter. Experimental results on a real basketball dataset show that the location of the ball can be estimated with an average accuracy of 82%. Fabio Poiesi, Fahad Daniyal, Andrea Cavallaro |
ICIP | 1 |