EDBT 2026 Demo / reviewers in the wild / expert
Riccardo Spezialetti
dblp:176/1487
· DBLP profile ↗
13ranked-venue papers
3as first author
7since 2021 · last 2024
0000-0001-9748-2824ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
3D vision · 80% Representation and self-supervised learning · 14% Deep learning architectures and training · 5% | |
| Computer graphics and multimedia
2 papers |
Rendering · 54% Geometric modeling and processing · 46% |
Topics — the 22 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d shape representation |
0.8 | 1 | 2024 | Deep Learning on Object-Centric 3D Neural Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Computer vision › 3D vision › implicit neural representation
neural field |
0.8 | 1 | 2024 | Deep Learning on Object-Centric 3D Neural Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Computer vision › 3D vision
neural radiance field |
0.8 | 1 | 2024 | Deep Learning on Object-Centric 3D Neural Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Computer vision › 3D vision
shape matching |
0.8 | 2 | 2019 | Learning an Effective Equivariant 3D Descriptor Without Supervision · ICCV 2019 GFrames: Gradient-Based Local Reference Frame for 3D Shape Matching · CVPR 2019 |
Computer vision › 3D vision › 3d shape analysis
3d shape understanding |
0.7 | 1 | 2023 | Deep Learning on Implicit Neural Representations of Shapes · ICLR 2023 |
Geometric modeling and processing
implicit neural representation |
0.7 | 1 | 2023 | Deep Learning on Implicit Neural Representations of Shapes · ICLR 2023 |
Rendering
neural radiance fields |
0.7 | 1 | 2023 | ReLight My NeRF: A Dataset for Novel View Synthesis and Relighting of Real World Objects · CVPR 2023 |
Rendering
relighting |
0.7 | 1 | 2023 | ReLight My NeRF: A Dataset for Novel View Synthesis and Relighting of Real World Objects · CVPR 2023 |
Geometric modeling and processing
shape representation |
0.7 | 1 | 2023 | Deep Learning on Implicit Neural Representations of Shapes · ICLR 2023 |
Computer vision › 3D vision
point cloud processing |
0.6 | 1 | 2022 | Unsupervised Learning of Local Equivariant Descriptors for Point Clouds · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Computer vision › 3D vision › geometric estimation › 3d registration
surface registration |
0.6 | 1 | 2022 | Unsupervised Learning of Local Equivariant Descriptors for Point Clouds · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Computer vision › 3D vision › 3d shape analysis
3d keypoint detection |
0.5 | 2 | 2018 | Learning to Detect Good 3D Keypoints · Int. J. Comput. Vis. 2018 Learning a Descriptor-Specific 3D Keypoint Detector · ICCV 2015 |
Computer vision › 3D vision
3d shape analysis |
0.4 | 1 | 2020 | Learning to Orient Surfaces by Self-supervised Spherical CNNs · NeurIPS 2020 |
Machine learning › Representation and self-supervised learning › equivariance
equivariant representation |
0.4 | 1 | 2020 | Learning to Orient Surfaces by Self-supervised Spherical CNNs · NeurIPS 2020 |
Machine learning › Deep learning architectures and training › equivariant neural network
spherical CNN |
0.4 | 1 | 2020 | Learning to Orient Surfaces by Self-supervised Spherical CNNs · NeurIPS 2020 |
Computer vision › 3D vision › local feature descriptor
3d local descriptors |
0.4 | 1 | 2019 | Learning an Effective Equivariant 3D Descriptor Without Supervision · ICCV 2019 |
Machine learning › Representation and self-supervised learning › equivariance
equivariant representation learning |
0.4 | 1 | 2019 | Learning an Effective Equivariant 3D Descriptor Without Supervision · ICCV 2019 |
Computer vision › 3D vision
local reference frame |
0.4 | 1 | 2019 | GFrames: Gradient-Based Local Reference Frame for 3D Shape Matching · CVPR 2019 |
Machine learning › Representation and self-supervised learning › representation learning
latent representation learning |
0.2 | 1 | 2024 | Deep Learning on Object-Centric 3D Neural Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Computer vision › 3D vision › feature matching
local feature matching |
0.2 | 1 | 2015 | Learning a Descriptor-Specific 3D Keypoint Detector · ICCV 2015 |
Rendering
novel view synthesis |
0.2 | 1 | 2023 | ReLight My NeRF: A Dataset for Novel View Synthesis and Relighting of Real World Objects · CVPR 2023 |
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning |
0.2 | 1 | 2022 | Unsupervised Learning of Local Equivariant Descriptors for Point Clouds · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Methods — techniques the papers use, named apart from their topics
deep learning · 1.6implicit neural representation · 1.3spherical CNN · 1.0plane folding decoder · 1.0single inference pass · 0.8neural field embedding · 0.8one-light-at-time acquisition · 0.7neural radiance field · 0.7self-supervised learning · 0.4tangent vector field · 0.4spherical convolutional neural network · 0.4intrinsic gradient · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Deep Learning on Object-Centric 3D Neural FieldsabstractIn recent years, Neural Fields (NFs) have emerged as an effective tool for encoding diverse continuous signals such as images, videos, audio, and 3D shapes. When applied to 3D data,NFs offer a solution to the fragmentation and limitations associated with prevalent discrete representations. However, given thatNFs are essentially neural networks, it remains unclear whether and how they can be seamlessly integrated into deep learning pipelines for solving downstream tasks. This paper addresses this research problem and introducesnf2vec, a framework capable of generating a compact latent representation for an inputNFin a single inference pass. We demonstrate thatnf2veceffectively embeds 3D objects represented by the inputNFs and showcase how the resulting embeddings can be employed in deep learning pipelines to successfully address various tasks, all while processing exclusivelyNFs. We test this framework on severalNFs used to represent 3D surfaces, such as unsigned/signed distance and occupancy fields. Moreover, we demonstrate the effectiveness of our approach with more complexNFs that encompass both geometry and appearance of 3D objects such as neural radiance fields. Pierluigi Zama Ramirez, Luca De Luigi, Daniele Sirocchi, Adriano Cardace, Riccardo Spezialetti, Francesco Ballerini, Samuele Salti, Luigi Di Stefano |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | ReLight My NeRF: A Dataset for Novel View Synthesis and Relighting of Real World ObjectsabstractIn this paper, we focus on the problem of rendering novel views from a Neural Radiance Field (NeRF) under unobserved light conditions. To this end, we introduce a novel dataset, dubbed ReNe (Relighting NeRF), framing real world objects under one-light-at-time (OLAT) conditions, annotated with accurate ground-truth camera and light poses. Our acquisition pipeline leverages two robotic arms holding, respectively, a camera and an omni-directional point-wise light source. We release a total of 20 scenes depicting a variety of objects with complex geometry and challenging materials. Each scene includes 2000 images, acquired from 50 different points of views under 40 different OLAT conditions. By leveraging the dataset, we perform an ablation study on the relighting capability of variants of the vanilla NeRF architecture and identify a lightweight architecture that can render novel views of an object under novel light conditions, which we use to establish a non-trivial baseline for the dataset. Dataset and benchmark are available at https://eyecan-ai. Marco Toschi, Riccardo De Matteo, Riccardo Spezialetti, Daniele De Gregorio, Luigi Di Stefano, Samuele Salti |
CVPR | 3 |
| 2023 | Deep Learning on Implicit Neural Representations of Shapes
Luca De Luigi, Adriano Cardace, Riccardo Spezialetti, Pierluigi Zama Ramirez, Samuele Salti, Luigi Di Stefano |
ICLR | 3 |
| 2023 | Self-Distillation for Unsupervised 3D Domain AdaptationabstractPoint cloud classification is a popular task in 3D vision. However, previous works, usually assume that point clouds at test time are obtained with the same procedure or sensor as those at training time. Unsupervised Domain Adaptation (UDA) instead, breaks this assumption and tries to solve the task on an unlabeled target domain, leveraging only on a supervised source domain. For point cloud classification, recent UDA methods try to align features across domains via auxiliary tasks such as point cloud reconstruction, which however do not optimize the discriminative power in the target domain in feature space. In contrast, in this work, we focus on obtaining a discriminative feature space for the target domain enforcing consistency between a point cloud and its augmented version. We then propose a novel iterative self-training methodology that exploits Graph Neural Networks in the UDA context to refine pseudo-labels. We perform extensive experiments and set the new state-of-the art in standard UDA benchmarks for point cloud classification. Finally, we show how our approach can be extended to more complex tasks such as part segmentation. Adriano Cardace, Riccardo Spezialetti, Pierluigi Zama Ramirez, Samuele Salti, Luigi Di Stefano |
WACV | 2 |
| 2022 | Unsupervised Learning of Local Equivariant Descriptors for Point CloudsabstractCorrespondences between 3D keypoints generated by matching local descriptors are a key step in 3D computer vision and graphic applications. Learned descriptors are rapidly evolving and outperforming the classical handcrafted approaches in the field. Yet, to learn effective representations they require supervision through labeled data, which are cumbersome and time-consuming to obtain. Unsupervised alternatives exist, but they lag in performance. Moreover, invariance to viewpoint changes is attained either by relying on data augmentation, which is prone to degrading upon generalization on unseen datasets, or by learning from handcrafted representations of the input which are already rotation invariant but whose effectiveness at training time may significantly affect the learned descriptor. We show how learning an equivariant 3D local descriptor instead of an invariant one can overcome both issues. LEAD (Local EquivAriant Descriptor) combines Spherical CNNs to learn an equivariant representation together with plane-folding decoders to learn without supervision. Through extensive experiments on standard surface registration datasets, we show how our proposal outperforms existing unsupervised methods by a large margin and achieves competitive results against the supervised approaches, especially in the practically very relevant scenario of transfer learning. Marlon Marcon, Riccardo Spezialetti, Samuele Salti, Luciano Silva, Luigi Di Stefano |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | RefRec: Pseudo-labels Refinement via Shape Reconstruction for Unsupervised 3D Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) for point cloud classification is an emerging research problem with relevant practical motivations. Reliance on multi-task learning to align features across domains has been the standard way to tackle it. In this paper, we take a different path and propose RefRec, the first approach to investigate pseudo-labels and self-training in UDA for point clouds. We present two main innovations to make self-training effective on 3D data: i) refinement of noisy pseudo-labels by matching shape descriptors that are learned by the unsupervised task of shape reconstruction on both domains; ii) a novel self-training protocol that learns domain-specific decision boundaries and reduces the negative impact of mislabelled target samples and in-domain intra-class variability. RefRec sets the new state of the art in both standard benchmarks used to test UDA for point cloud classification, showcasing the effectiveness of self-training for this important problem. Adriano Cardace, Riccardo Spezialetti, Pierluigi Zama Ramirez, Samuele Salti, Luigi Di Stefano |
3DV | 2 |
| 2021 | Go with the Flows: Mixtures of Normalizing Flows for Point Cloud Generation and ReconstructionabstractRecently Normalizing Flows (NFs) have demonstrated state-of-the-art performance on modeling 3D point clouds while allowing sampling with arbitrary resolution at inference time. However, these flow-based models still have fundamental limitations on complicated geometries. This work generalizes prior work by introducing additional discrete latent variable, i.e. mixture model. This circumvents limitations of prior approaches, leads to more parameter efficient models and reduces the inference runtime. Moreover, in this more general framework each component learns to specialize in a particular subregion of an object in a completely unsupervised fashion yielding promising clustering properties. We further demonstrate that by adding data augmentation, individual mixture components can learn to specialize in a semantically meaningful manner. We evaluate mixtures of NFs on generation, autoencoding and single-view reconstruction based on the ShapeNet dataset. Janis Postels, Riccardo Spezialetti, Luc Van Gool, Federico Tombari |
3DV | 3 |
| 2020 | A Divide et Impera Approach for 3D Shape Reconstruction from Multiple ViewsabstractEstimating the 3D shape of an object from a single or multiple images has gained popularity thanks to the recent breakthroughs powered by deep learning. Most approaches regress the full object shape in a canonical pose, possibly extrapolating the occluded parts based on the learned priors. However, their viewpoint invariant technique often discards the unique structures visible from the input images. In contrast, this paper proposes to rely on viewpoint variant reconstructions by merging the visible information from the given views. Our approach is divided into three steps. Starting from the sparse views of the object, we first align them into a common coordinate system by estimating the relative pose between all the pairs. Then, inspired by the traditional voxel carving, we generate an occupancy grid of the object taken from the silhouette on the images and their relative poses. Finally, we refine the initial reconstruction to build a clean 3D model which preserves the details from each viewpoint. To validate the proposed method, we perform a comprehensive evaluation on the ShapeNet reference benchmark in terms of relative pose estimation and 3D shape reconstruction. Riccardo Spezialetti, David Joseph Tan, Alessio Tonioni, Keisuke Tateno, Federico Tombari |
3DV | 1 |
| 2020 | Learning to Orient Surfaces by Self-supervised Spherical CNNsabstractDefining and reliably finding a canonical orientation for 3D surfaces is key to many Computer Vision and Robotics applications. This task is commonly addressed by handcrafted algorithms exploiting geometric cues deemed as distinctive and robust by the designer. Yet, one might conjecture that humans learn the notion of the inherent orientation of 3D objects from experience and that machines may do so alike. In this work, we show the feasibility of learning a robust canonical orientation for surfaces represented as point clouds. Based on the observation that the quintessential property of a canonical orientation is equivariance to 3D rotations, we propose to employ Spherical CNNs, a recently introduced machinery that can learn equivariant representations defined on the Special Ortoghonal group SO(3). Specifically, spherical correlations compute feature maps whose elements define 3D rotations. Our method learns such feature maps from raw data by a self-supervised training procedure and robustly selects a rotation to transform the input point cloud into a learned canonical orientation. Thereby, we realize the first end-to-end learning approach to define and extract the canonical orientation of 3D shapes, which we aptly dub Compass. Experiments on several public datasets prove its effectiveness at orienting local surface patches as well as whole objects. Riccardo Spezialetti, Federico Stella, Marlon Marcon, Luciano Silva, Samuele Salti, Luigi Di Stefano |
NeurIPS | 1 |
| 2019 | GFrames: Gradient-Based Local Reference Frame for 3D Shape MatchingabstractWe introduce GFrames, a novel local reference frame (LRF) construction for 3D meshes and point clouds. GFrames are based on the computation of the intrinsic gradient of a scalar field defined on top of the input shape. The resulting tangent vector field defines a repeatable tangent direction of the local frame at each point; importantly, it directly inherits the properties and invariance classes of the underlying scalar function, making it remarkably robust under strong sampling artifacts, vertex noise, as well as non-rigid deformations. Existing local descriptors can directly benefit from our repeatable frames, as we showcase in a selection of 3D vision and shape analysis applications where we demonstrate state-of-the-art performance in a variety of challenging settings. Simone Melzi, Riccardo Spezialetti, Federico Tombari, Michael M. Bronstein, Luigi Di Stefano, Emanuele Rodolà |
CVPR | 2 |
| 2019 | Learning an Effective Equivariant 3D Descriptor Without SupervisionabstractEstablishing correspondences between 3D shapes is a fundamental task in 3D Computer Vision, typically ad- dressed by matching local descriptors. Recently, a few at- tempts at applying the deep learning paradigm to the task have shown promising results. Yet, the only explored way to learn rotation invariant descriptors has been to feed neural networks with highly engineered and invariant representations provided by existing hand-crafted descriptors, a path that goes in the opposite direction of end-to-end learning from raw data so successfully deployed for 2D images. In this paper, we explore the benefits of taking a step back in the direction of end-to-end learning of 3D descriptors by disentangling the creation of a robust and distinctive rotation equivariant representation, which can be learned from unoriented input data, and the definition of a good canonical orientation, required only at test time to obtain an invariant descriptor. To this end, we leverage two re- cent innovations: spherical convolutional neural networks to learn an equivariant descriptor and plane folding de- coders to learn without supervision. The effectiveness of the proposed approach is experimentally validated by out- performing hand-crafted and learned descriptors on a standard benchmark. Riccardo Spezialetti, Samuele Salti, Luigi Di Stefano |
ICCV | 1 |
| 2018 | Learning to Detect Good 3D Keypoints
Alessio Tonioni, Samuele Salti, Federico Tombari, Riccardo Spezialetti, Luigi Di Stefano |
Int. J. Comput. Vis. | 4 |
| 2015 | Learning a Descriptor-Specific 3D Keypoint DetectorabstractKeypoint detection represents the first stage in the majority of modern computer vision pipelines based on automatically established correspondences between local descriptors. However, no standard solution has emerged yet in the case of 3D data such as point clouds or meshes, which exhibit high variability in level of detail and noise. More importantly, existing proposals for 3D keypoint detection rely on geometric saliency functions that attempt to maximize repeatability rather than distinctiveness of the selected regions, which may lead to sub-optimal performance of the overall pipeline. To overcome these shortcomings, we cast 3D keypoint detection as a binary classification between points whose support can be correctly matched by a predefined 3D descriptor or not, thereby learning a descriptor-specific detector that adapts seamlessly to different scenarios. Through experiments on several public datasets, we show that this novel approach to the design of a keypoint detector represents a flexible solution that, nonetheless, can provide state-of-the-art descriptor matching performance. Samuele Salti, Federico Tombari, Riccardo Spezialetti, Luigi Di Stefano |
ICCV | 3 |