EDBT 2026 Demo / reviewers in the wild / expert
Emanuele Rodolà
dblp:54/8401
· DBLP profile ↗
101ranked-venue papers
12as first author
43since 2021 · last 2026
0000-0003-0091-7241ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 69 · 10 first-author · 18 since 2021Artificial intelligence and machine learning · 59 · 7 first-author · 28 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Model-based Metric 3D Shape and Motion Reconstruction of Wild Bottlenose Dolphins in Drone-Shot Videos
Daniele Baieri, Riccardo Cicciarella, Michael Krützen, Emanuele Rodolà, Silvia Zuffi |
Int. J. Comput. Vis. | 4 |
| 2025 | Task Singular Vectors: Reducing Task Interference in Model MergingabstractTask Arithmetic has emerged as a simple yet effective method to merge models without additional training. However, by treating entire networks as flat parameter vectors, it overlooks key structural information and is susceptible to task interference. In this paper, we study task vectors at the layer level, focusing on task layer matrices and their singular value decomposition. In particular, we concentrate on the resulting singular vectors, which we refer to as Task Singular Vectors (TSV). Recognizing that layer task matrices are often low-rank, we propose TSV-Compress (TSV-C), a simple procedure that compresses them to 10% of their original size while retaining 99% of accuracy. We further leverage this low-rank space to define a new measure of task interference based on the interaction of singular vectors from different tasks. Building on these findings, we introduce TSV-Merge (TSV-M), a novel model merging approach that combines compression with interference reduction, significantly outperforming existing methods. Antonio Andrea Gargiulo, Donato Crisostomi, Maria Sofia Bucarelli, Simone Scardapane, Fabrizio Silvestri, Emanuele Rodolà |
CVPR | 6 |
| 2025 | Escaping Plato's Cave: Towards the Alignment of 3D and Text Latent SpacesabstractRecent works have shown that, when trained at scale, unimodal 2D vision and text encoders converge to learned features that share remarkable structural properties, despite arising from different representations. However, the role of 3D encoders with respect to other modalities remains unexplored. Furthermore, existing 3D foundation models that leverage large datasets are typically trained with explicit alignment objectives with respect to frozen encoders from other representations. In this work, we investigate the possibility of a posteriori alignment of representations obtained from uni-modal 3D encoders compared to text-based feature spaces. We show that naive post-training feature alignment of uni-modal text and 3D encoders results in limited performance. We then focus on extracting subspaces of the corresponding feature spaces and discover that by projecting learned representations onto well-chosen lower-dimensional subspaces the quality of alignment becomes significantly higher, leading to improved accuracy on matching and retrieval tasks. Our analysis further sheds light on the nature of these shared subspaces, which roughly separate between semantic and geometric data representations. Overall, ours is the first work that helps to establish a baseline for post-training alignment of 3D uni-modal and text feature spaces, and helps to highlight both the shared and unique properties of 3D data compared to other representations. Souhail Hadgi, Luca Moschella, Andrea Santilli, Diego Gomez 0003, Qixing Huang, Emanuele Rodolà, Simone Melzi, Maks Ovsjanikov |
CVPR | 6 |
| 2025 | COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio RepresentationsabstractWe present COCOLA (Coherence-Oriented Contrastive Learning for Audio), a contrastive learning method for musical audio representations that captures the harmonic and rhythmic coherence between samples. Our method operates at the level of the individual stems composing music tracks and can input features obtained via Harmonic-Percussive Separation (HPS). COCOLA allows an objective evaluation of generative models for music accompaniment generation, which are difficult to benchmark with established metrics. In this regard, we evaluate recent music accompaniment generation models, demonstrating the effectiveness of our proposed method. We release the model checkpoints trained on public datasets containing separate stems (MUSDB18-HQ, MoisesDB, Slakh2100, and CocoChorales). Ruben Ciranni, Giorgio Mariani, Michele Mancusi, Emilian Postolache, Giorgio Fabbro, Emanuele Rodolà, Luca Cosmo |
ICASSP | 6 |
| 2025 | Naturalistic Music Decoding from EEG Data via Latent Diffusion ModelsabstractIn this article, we explore the potential of using latent diffusion models, a family of powerful generative models, for the task of reconstructing naturalistic music from electroencephalogram (EEG) recordings. Unlike simpler music with limited timbres, such as MIDI-generated tunes or monophonic pieces, the focus here is on intricate music featuring a diverse array of instruments, voices, and effects, rich in harmonics and timbre. This study represents an initial foray into achieving general music reconstruction of high-quality using non-invasive EEG data, employing an end-to-end training approach directly on raw data without the need for manual pre-processing and channel selection. We train our models on the public NMED- T dataset and perform quantitative evaluation proposing neural embedding-based metrics. Our work contributes to the ongoing research in neural decoding and brain-computer interfaces, offering insights into the feasibility of using EEG data for complex auditory information reconstruction. Emilian Postolache, Natalia Polouliakh, Hiroaki Kitano, Akima Connelly, Emanuele Rodolà, Luca Cosmo, Taketo Akama |
ICASSP | 5 |
| 2025 | MERGE3: Efficient Evolutionary Merging on Consumer-grade GPUsabstractEvolutionary model merging enables the creation of high-performing multi-task models but remains computationally prohibitive for consumer hardware. We introduce MERGE$^3$, an efficient framework that makes evolutionary merging of Large Language Models (LLMs) feasible on a single GPU by reducing fitness computation costs 50$\times$ while retaining a large fraction of the original performance. MERGE$^3$ achieves this by Extracting a reduced dataset for evaluation, Estimating model abilities using Item Response Theory (IRT), and Evolving optimal merges via IRT-based performance estimators. Our method enables state-of-the-art multilingual and cross-lingual merging, transferring knowledge across languages with significantly lower computational overhead. We provide theoretical guarantees and an open-source library, democratizing high-quality model merging. Tommaso Mencattini, Adrian Robert Minut, Donato Crisostomi, Andrea Santilli, Emanuele Rodolà |
ICML | 5 |
| 2025 | Update Your Transformer to the Latest Release: Re-Basin of Task VectorsabstractFoundation models serve as the backbone for numerous specialized models developed through fine-tuning. However, when the underlying pretrained model is updated or retrained (e.g., on larger and more curated datasets), the fine-tuned model becomes obsolete, losing its utility and requiring retraining. This raises the question: is it possible to transfer fine-tuning to a new release of the model? In this work, we investigate how to transfer fine-tuning to a new checkpoint without having to re-train, in a data-free manner. To do so, we draw principles from model re-basin and provide a recipe based on weight permutations to re-base the modifications made to the original base model, often called task vector. In particular, our approach tailors model re-basin for Transformer models, taking into account the challenges of residual connections and multi-head attention layers. Specifically, we propose a two-level method rooted in spectral theory, initially permuting the attention heads and subsequently adjusting parameters within select pairs of heads. Through extensive experiments on visual and textual tasks, we achieve the seamless transfer of fine-tuned knowledge to new pre-trained backbones without relying on a single training step or datapoint. Code is available at https://github.com/aimagelab/TransFusion. Filippo Rinaldi, Giacomo Capitani, Lorenzo Bonicelli, Donato Crisostomi, Federico Bolelli, Elisa Ficarra, Emanuele Rodolà, Simone Calderara, Angelo Porrello |
ICML | 7 |
| 2025 | Implicit-ARAP: Efficient Handle-Guided Neural Field Deformation via Local Patch MeshingabstractNeural fields have emerged as a powerful representation for 3D geometry, enabling compact and continuous modeling of complex shapes. Despite their expressive power, manipulating neural fields in a controlled and accurate manner -- particularly under spatial constraints -- remains an open challenge, as existing approaches struggle to balance surface quality, robustness, and efficiency. We address this by introducing a novel method for handle-guided neural field deformation, which leverages discrete local surface representations to optimize the As-Rigid-As-Possible deformation energy. To this end, we propose the local patch mesh representation, which discretizes level sets of a neural signed distance field by projecting and deforming flat mesh patches guided solely by the SDF and its gradient. We conduct a comprehensive evaluation showing that our method consistently outperforms baselines in deformation quality, robustness, and computational efficiency. We also present experiments that motivate our choice of discretization over marching cubes. By bridging classical geometry processing and neural representations through local patch meshing, our work enables scalable, high-quality deformation of neural fields and paves the way for extending other geometric tasks to neural domains. Daniele Baieri, Filippo Maggioli, Emanuele Rodolà, Simone Melzi, Zorah Lähner |
NeurIPS | 3 |
| 2025 | Graph Kernel Neural NetworksabstractThe convolution operator at the core of many modern neural architectures can effectively be seen as performing a dot product between an input matrix and a filter. While this is readily applicable to data such as images, which can be represented as regular grids in the Euclidean space, extending the convolution operator to work on graphs proves more challenging, due to their irregular structure. In this article, we propose to use graph kernels, i.e., kernel functions that compute an inner product on graphs, to extend the standard convolution operator to the graph domain. This allows us to define an entirely structural model that does not require computing the embedding of the input graph. Our architecture allows to plug-in any type of graph kernels and has the added benefit of providing some interpretability in terms of the structural masks that are learned during the training process, similar to what happens for convolutional masks in traditional convolutional neural networks (CNNs). We perform an extensive ablation study to investigate the model hyperparameters' impact and show that our model achieves competitive performance on standard graph classification and regression datasets. Luca Cosmo, Giorgia Minello, Alessandro Bicciato, Michael M. Bronstein, Emanuele Rodolà, Luca Rossi 0004, Andrea Torsello |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Vector Quantile Regression on ManifoldsabstractQuantile regression (QR) is a statistical tool for distribution-free estimation of conditional quantiles of a target variable given explanatory features. QR is limited by the assumption that the target distribution is univariate and defined on an Euclidean domain. Although the notion of quantiles was recently extended to multi-variate distributions, QR for multi-variate distributions on manifolds remains underexplored, even though many important applications inherently involve data distributed on, e.g., spheres (climate and geological phenomena), and tori (dihedral angles in proteins). By leveraging optimal transport theory and c-concave functions, we meaningfully define conditional vector quantile functions of high-dimensional variables on manifolds (M-CVQFs). Our approach allows for quantile estimation, regression, and computation of conditional confidence sets and likelihoods. We demonstrate the approach’s efficacy and provide insights regarding the meaning of non-Euclidean quantiles through synthetic and real data experiments. Marco Pegoraro 0002, Sanketh Vedula, Aviv Rosenberg 0001, Irene Tallini, Emanuele Rodolà, Alexander M. Bronstein |
AISTATS | 5 |
| 2024 | ReMatching: Low-Resolution Representations for Scalable Shape Correspondence
Filippo Maggioli, Daniele Baieri, Emanuele Rodolà, Simone Melzi |
ECCV (37) | 3 |
| 2024 | Syncfusion: Multimodal Onset-Synchronized Video-to-Audio Foley SynthesisabstractSound design involves creatively selecting, recording, and editing sound effects for various media like cinema, video games, and virtual/augmented reality. One of the most time-consuming steps when designing sound is synchronizing audio with video. In some cases, environmental recordings from video shoots are available, which can aid in the process. However, in video games and animations, no reference audio exists, requiring manual annotation of event timings from the video. We propose a system to extract repetitive actions onsets from a video, which are then used - in conjunction with audio or textual embeddings - to condition a diffusion model trained to generate a new synchronized sound effects audio track. In this way, we leave complete creative control to the sound designer while removing the burden of synchronization with video. Furthermore, editing the onset track or changing the conditioning embedding requires much less effort than editing the audio track itself, simplifying the sonification process. We provide sound examples, source code, and pretrained models to faciliate reproducibility1. Marco Comunità, Riccardo F. Gramaccioni, Emilian Postolache, Emanuele Rodolà, Danilo Comminiello, Joshua D. Reiss |
ICASSP | 4 |
| 2024 | Generalized Multi-Source Inference for Text Conditioned Music Diffusion ModelsabstractMulti-Source Diffusion Models (MSDM) allow for compositional musical generation tasks: generating a set of coherent sources, creating accompaniments, and performing source separation. Despite their versatility, they require estimating the joint distribution over the sources, necessitating pre-separated musical data, which is rarely available, and fixing the number and type of sources at training time. This paper generalizes MSDM to arbitrary time-domain diffusion models conditioned on text embeddings. These models do not require separated data as they are trained on mixtures, can parameterize an arbitrary number of sources, and allow for rich semantic control. We propose an inference procedure enabling the coherent generation of sources and accompaniments. Additionally, we adapt the Dirac separator of MSDM to perform source separation. We experiment with diffusion models trained on Slakh2100 and MTG-Jamendo, showcasing competitive generation and separation results in a relaxed data setting. Emilian Postolache, Giorgio Mariani, Luca Cosmo, Emmanouil Benetos, Emanuele Rodolà |
ICASSP | 5 |
| 2024 | From Bricks to Bridges: Product of Invariances to Enhance Latent Space CommunicationabstractIt has been observed that representations learned by distinct neural networks conceal structural similarities when the models are trained under similar inductive biases. From a geometric perspective, identifying the classes of transformations and the related invariances that connect these representations is fundamental to unlocking applications, such as merging, stitching, and reusing different neural modules. However, estimating task-specific transformations a priori can be challenging and expensive due to several factors (e.g., weights initialization, training hyperparameters, or data modality). To this end, we introduce a versatile method to directly incorporate a set of invariances into the representations, constructing a product space of invariant components on top of the latent representations without requiring prior knowledge about the optimal invariance to infuse. We validate our solution on classification and reconstruction tasks, observing consistent latent similarity and downstream performance improvements in a zero-shot stitching setting. The experimental analysis comprises three modalities (vision, text, and graphs), twelve pretrained foundational models, nine benchmarks, and several architectures trained from scratch. Irene Cannistraci, Luca Moschella, Marco Fumero, Valentino Maiorca, Emanuele Rodolà |
ICLR | 5 |
| 2024 | Multi-Source Diffusion Models for Simultaneous Music Generation and SeparationabstractIn this work, we define a diffusion-based generative model capable of both music generation and source separation by learning the score of the joint probability density of sources sharing a context. Alongside the classic total inference tasks (i.e., generating a mixture, separating the sources), we also introduce and experiment on the partial generation task of source imputation, where we generate a subset of the sources given the others (e.g., play a piano track that goes well with the drums). Additionally, we introduce a novel inference method for the separation task based on Dirac likelihood functions. We train our model on Slakh2100, a standard dataset for musical source separation, provide qualitative results in the generation settings, and showcase competitive quantitative results in the source separation setting. Our method is the first example of a single model that can handle both generation and separation tasks, thus representing a step toward general audio models. Giorgio Mariani, Irene Tallini, Emilian Postolache, Michele Mancusi, Luca Cosmo, Emanuele Rodolà |
ICLR | 6 |
| 2024 | $C^2M^3$: Cycle-Consistent Multi-Model MergingabstractIn this paper, we present a novel data-free method for merging neural networks in weight space. Our method optimizes for the permutations of network neurons while ensuring global coherence across all layers, and it outperforms recent layer-local approaches in a set of challenging scenarios. We then generalize the formulation to the $N$-models scenario to enforce cycle consistency of the permutations with guarantees, allowing circular compositions of permutations to be computed without accumulating error along the path.
We qualitatively and quantitatively motivate the need for such a constraint, showing its benefits when merging homogeneous sets of models in scenarios spanning varying architectures and datasets. We finally show that, when coupled with activation renormalization, the approach yields the best results in the task. Donato Crisostomi, Marco Fumero, Daniele Baieri, Florian Bernard 0001, Emanuele Rodolà |
NeurIPS | 5 |
| 2024 | Latent Functional Maps: a spectral framework for representation alignmentabstractNeural models learn data representations that lie on low-dimensional manifolds, yet modeling the relation between these representational spaces is an ongoing challenge.
By integrating spectral geometry principles into neural modeling, we show that this problem can be better addressed in the functional domain, mitigating complexity, while enhancing interpretability and performances on downstream tasks.
To this end, we introduce a multi-purpose framework to the representation learning community, which allows to: (i) compare different spaces in an interpretable way and measure their intrinsic similarity; (ii) find correspondences between them, both in unsupervised and weakly supervised settings, and (iii) to effectively transfer representations between distinct spaces.
We validate our framework on various applications, ranging from stitching to retrieval tasks, and on multiple modalities, demonstrating that Latent Functional Maps can serve as a swiss-army knife for representation alignment. Marco Fumero, Marco Pegoraro 0002, Valentino Maiorca, Francesco Locatello, Emanuele Rodolà |
NeurIPS | 5 |
| 2024 | Geometric epitope and paratope predictionabstractMOTIVATION: Identifying the binding sites of antibodies is essential for developing vaccines and synthetic antibodies. In this article, we investigate the optimal representation for predicting the binding sites in the two molecules and emphasize the importance of geometric information. RESULTS: Specifically, we compare different geometric deep learning methods applied to proteins' inner (I-GEP) and outer (O-GEP) structures. We incorporate 3D coordinates and spectral geometric descriptors as input features to fully leverage the geometric information. Our research suggests that different geometrical representation information is useful for different tasks. Surface-based models are more efficient in predicting the binding of the epitope, while graph models are better in paratope prediction, both achieving significant performance improvements. Moreover, we analyze the impact of structural changes in antibodies and antigens resulting from conformational rearrangements or reconstruction errors. Through this investigation, we showcase the robustness of geometric deep learning methods and spectral geometric descriptors to such perturbations. AVAILABILITY AND IMPLEMENTATION: The python code for the models, together with the data and the processing pipeline, is open-source and available at https://github.com/Marco-Peg/GEP. Marco Pegoraro 0002, Clémentine C. J. Dominé, Emanuele Rodolà, Petar Velickovic, Andreea Deac |
Bioinform. | 3 |
| 2024 | Latent spectral regularization for continual learningabstractWhile biological intelligence grows organically as new knowledge is gathered throughout life, Artificial Neural Networks forget catastrophically whenever they face a changing training data distribution. Rehearsal-based Continual Learning (CL) approaches have been established as a versatile and reliable solution to overcome this limitation; however, sudden input disruptions and memory constraints are known to alter the consistency of their predictions. We study this phenomenon by investigating the geometric characteristics of the learner’s latent space and find that replayed data points of different classes increasingly mix up, interfering with classification. Hence, we propose a geometric regularizer that enforces weak requirements on the Laplacian spectrum of the latent space, promoting a partitioning behavior. Our proposal, called Continual Spectral Regularizer for Incremental Learning (CaSpeR-IL), can be easily combined with any rehearsal-based CL approach and improves the performance of SOTA methods on standard benchmarks. Emanuele Frascaroli, Riccardo Benaglia, Matteo Boschini, Luca Moschella, Cosimo Fiorini, Emanuele Rodolà, Simone Calderara |
Pattern Recognit. Lett. | 6 |
| 2023 | Latent Autoregressive Source SeparationabstractAutoregressive models have achieved impressive results over a wide range of domains in terms of generation quality and downstream task performance. In the continuous domain, a key factor behind this success is the usage of quantized latent spaces (e.g., obtained via VQ-VAE autoencoders), which allow for dimensionality reduction and faster inference times. However, using existing pre-trained models to perform new non-trivial tasks is difficult since it requires additional fine-tuning or extensive training to elicit prompting. This paper introduces LASS as a way to perform vector-quantized Latent Autoregressive Source Separation (i.e., de-mixing an input signal into its constituent sources) without requiring additional gradient-based optimization or modifications of existing models. Our separation method relies on the Bayesian formulation in which the autoregressive models are the priors, and a discrete (non-parametric) likelihood function is constructed by performing frequency counts over latent sums of addend tokens. We test our method on images and audio with several sampling strategies (e.g., ancestral, beam search) showing competitive results with existing approaches in terms of separation quality while offering at the same time significant speedups in terms of inference time and scalability to higher dimensional data. Emilian Postolache, Giorgio Mariani, Michele Mancusi, Andrea Santilli, Luca Cosmo, Emanuele Rodolà |
AAAI | 6 |
| 2023 | Accelerating Transformer Inference for Translation via Parallel DecodingabstractAndrea Santilli, Silvio Severino, Emilian Postolache, Valentino Maiorca, Michele Mancusi, Riccardo Marin, Emanuele Rodola. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Andrea Santilli, Silvio Severino, Emilian Postolache, Valentino Maiorca, Michele Mancusi, Riccardo Marin, Emanuele Rodolà |
ACL (1) | 7 |
| 2023 | Latent Code Disentanglement Using Orthogonal Latent Codes and Inter-Domain Signal Transformation
Babak Solhjoo, Emanuele Rodolà |
ICAART (2) | 2 |
| 2023 | Relative representations enable zero-shot latent space communication
Luca Moschella, Valentino Maiorca, Marco Fumero, Antonio Norelli, Francesco Locatello, Emanuele Rodolà |
ICLR | 6 |
| 2023 | Leveraging sparse and shared feature activations for disentangled representation learningabstractRecovering the latent factors of variation of high dimensional data has so far focused on simple synthetic settings. Mostly building on unsupervised and weakly-supervised objectives, prior work missed out on the positive implications for representation learning on real world data. In this work, we propose to leverage knowledge extracted from a diversified set of supervised tasks to learn a common disentangled representation. Assuming each supervised task only depends on an unknown subset of the factors of variation, we disentangle the feature space of a supervised multi-task model, with features activating sparsely across different tasks and information being shared as appropriate. Importantly, we never directly observe the factors of variations but establish that access to multiple tasks is sufficient for identifiability under sufficiency and minimality assumptions.
We validate our approach on six real world distribution shift benchmarks, and different data modalities (images, text), demonstrating how disentangled representations can be transferred to real settings. Marco Fumero, Florian Wenzel, Luca Zancato, Alessandro Achille, Emanuele Rodolà, Stefano Soatto, Bernhard Schölkopf, Francesco Locatello |
NeurIPS | 5 |
| 2023 | Latent Space Translation via Semantic AlignmentabstractWhile different neural models often exhibit latent spaces that are alike when exposed to semantically related data, this intrinsic similarity is not always immediately discernible. Towards a better understanding of this phenomenon, our work shows how representations learned from these neural modules can be translated between different pre-trained networks via simpler transformations than previously thought. An advantage of this approach is the ability to estimate these transformations using standard, well-understood algebraic procedures that have closed-form solutions. Our method directly estimates a transformation between two given latent spaces, thereby enabling effective stitching of encoders and decoders without additional training. We extensively validate the adaptability of this translation procedure in different experimental settings: across various trainings, domains, architectures (e.g., ResNet, CNN, ViT), and in multiple downstream tasks (classification, reconstruction). Notably, we show how it is possible to zero-shot stitch text encoders and vision decoders, or vice-versa, yielding surprisingly good classification performance in this multimodal setting. Valentino Maiorca, Luca Moschella, Antonio Norelli, Marco Fumero, Francesco Locatello, Emanuele Rodolà |
NeurIPS | 6 |
| 2023 | ASIF: Coupled Data Turns Unimodal Models to Multimodal without TrainingabstractCLIP proved that aligning visual and language spaces is key to solving many vision tasks without explicit training, but required to train image and text encoders from scratch on a huge dataset. LiT improved this by only training the text encoder and using a pre-trained vision network. In this paper, we show that a common space can be created without any training at all, using single-domain encoders (trained with or without supervision) and a much smaller amount of image-text pairs. Furthermore, our model has unique properties. Most notably, deploying a new version with updated training samples can be done in a matter of seconds. Additionally, the representations in the common space are easily interpretable as every dimension corresponds to the similarity of the input to a unique entry in the multimodal dataset. Experiments on standard zero-shot visual benchmarks demonstrate the typical transfer ability of image-text models. Overall, our method represents a simple yet surprisingly strong baseline for foundation multi-modal models, raising important questions on their data efficiency and on the role of retrieval in machine learning. Antonio Norelli, Marco Fumero, Valentino Maiorca, Luca Moschella, Emanuele Rodolà, Francesco Locatello |
NeurIPS | 5 |
| 2023 | A Physically-inspired Approach to the Simulation of Plant WiltingabstractPlants are among the most complex objects to be modeled in computer graphics. While a large body of work is concerned with structural modeling and the dynamic reaction to external forces, our work focuses on the dynamic deformation caused by plant internal wilting processes. To this end, we motivate the simulation of water transport inside the plant which is a key driver of the wilting process. We then map the change of water content in individual plant parts to branch stiffness values and obtain the wilted plant shape through a position based dynamics simulation. We show, that our approach can recreate measured wilting processes and does so with a higher fidelity than approaches ignoring the internal water flow. Realistic plant wilting is not only important in a computer graphics context but can also aid the development of machine learning algorithms in agricultural applications through the generation of synthetic training data. Filippo Maggioli, Jonathan Klein, Torsten Hädrich, Emanuele Rodolà, Wojtek Palubicki, Sören Pirk, Dominik L. Michels |
SIGGRAPH Asia | 4 |
| 2023 | Multimodal Neural DatabasesabstractThe rise in loosely-structured data available through text, images, and other modalities has called for new ways of querying them. Multimedia Information Retrieval has filled this gap and has witnessed exciting progress in recent years. Tasks such as search and retrieval of extensive multimedia archives have undergone massive performance improvements, driven to a large extent by recent developments in multimodal deep learning. However, methods in this field remain limited in the kinds of queries they support and, in particular, their inability to answer database-like queries. For this reason, inspired by recent work on neural databases, we propose a new framework, which we name Multimodal Neural Databases (MMNDBs). MMNDBs can answer complex database-like queries that involve reasoning over different input modalities, such as text and images, at scale. In this paper, we present the first architecture able to fulfill this set of requirements and test it with several baselines, showing the limitations of currently available models. The results show the potential of these new techniques to process unstructured data coming from different modalities, paving the way for future research in the area. Giovanni Trappolini, Andrea Santilli, Emanuele Rodolà, Alon Y. Halevy, Fabrizio Silvestri |
SIGIR | 3 |
| 2022 | 3D Human Pose Estimation Using Möbius Graph Convolutional Networks
Niloofar Azizi, Horst Possegger, Emanuele Rodolà, Horst Bischof |
ECCV (1) | 3 |
| 2022 | Reduced Representation of Deformation Fields for Effective Non-rigid Shape MatchingabstractIn this work we present a novel approach for computing correspondences between non-rigid objects, by exploiting a reduced representation of deformation fields. Different from existing works that represent deformation fields by training a general-purpose neural network, we advocate for an approximation based on mesh-free methods. By letting the network learn deformation parameters at a sparse set of positions in space (nodes), we reconstruct the continuous deformation field in a closed-form with guaranteed smoothness. With this reduction in degrees of freedom, we show significant improvement in terms of data-efficiency thus enabling limited supervision. Furthermore, our approximation provides direct access to first-order derivatives of deformation fields, which facilitates enforcing desirable regularization effectively. Our resulting model has high expressive power and is able to capture complex deformations. We illustrate its effectiveness through state-of-the-art results across multiple deformable shape matching benchmarks. Our code and data are publicly available at: https://github.com/Sentient07/DeformationBasis. Ramana Sundararaman, Riccardo Marin, Emanuele Rodolà, Maks Ovsjanikov |
NeurIPS | 3 |
| 2022 | Foreword to the Special Section on Smart Tools and Applications in Graphics (STAG 2021)
Patrizio Frosini, Daniela Giorgi, Simone Melzi, Emanuele Rodolà |
Comput. Graph. | 4 |
| 2022 | MoMaS: Mold Manifold Simulation for real-time procedural texturingabstractAbstract The slime mold algorithm has recently been under the spotlight thanks to its compelling properties studied across many disciplines like biology, computation theory, and artificial intelligence. However, existing implementations act only on planar surfaces, and no adaptation to arbitrary surfaces is available. Inspired by this gap, we propose a novel characterization of the mold algorithm to work on arbitrary curved surfaces. Our algorithm is easily parallelizable on GPUs and allows to model the evolution of millions of agents in real‐time over surface meshes with several thousand triangles, while keeping the simplicity proper of the slime paradigm. We perform a comprehensive set of experiments, providing insights on stability, behavior, and sensibility to various design choices. We characterize a broad collection of behaviors with a limited set of controllable and interpretable parameters, enabling a novel family of heterogeneous and high‐quality procedural textures. The appearance and complexity of these patterns are well‐suited to diverse materials and scopes, and we add another layer of generalization by allowing different mold species to compete and interact in parallel. Filippo Maggioli, Riccardo Marin, Simone Melzi, Emanuele Rodolà |
Comput. Graph. Forum | 4 |
| 2022 | Learning Spectral Unions of Partial Deformable 3D ShapesabstractAbstract Spectral geometric methods have brought revolutionary changes to the field of geometry processing. Of particular interest is the study of the Laplacian spectrum as a compact, isometry and permutation‐invariant representation of a shape. Some recent works show how the intrinsic geometry of a full shape can be recovered from its spectrum, but there are approaches that consider the more challenging problem of recovering the geometry from the spectral information of partial shapes. In this paper, we propose a possible way to fill this gap. We introduce a learning‐based method to estimate the Laplacian spectrum of the union of partial non‐rigid 3D shapes, without actually computing the 3D geometry of the union or any correspondence between those partial shapes. We do so by operating purely in the spectral domain and by defining the union operation between short sequences of eigenvalues. We show that the approximated union spectrum can be used as‐is to reconstruct the complete geometry [MRC*19], perform region localization on a template [RTO*19] and retrieve shapes from a database, generalizing ShapeDNA [RWP06] to work with partialities. Working with eigenvalues allows us to deal with unknown correspondence, different sampling, and different discretizations (point clouds and meshes alike), making this operation especially robust and general. Our approach is data‐driven and can generalize to isometric and non‐isometric deformations of the surface, as long as these stay within the same semantic class (e.g., human bodies or horses), as well as to partiality artifacts not seen at training time. Luca Moschella, Simone Melzi, Luca Cosmo, Filippo Maggioli, Or Litany, Maks Ovsjanikov, Leonidas J. Guibas, Emanuele Rodolà |
Comput. Graph. Forum | 8 |
| 2022 | Localized Shape Modelling with Global Coherence: An Inverse Spectral ApproachabstractAbstract Many natural shapes have most of their characterizing features concentrated over a few regions in space. For example, humans and animals have distinctive head shapes, while inorganic objects like chairs and airplanes are made of well‐localized functional parts with specific geometric features. Often, these features are strongly correlated – a modification of facial traits in a quadruped should induce changes to the body structure. However, in shape modelling applications, these types of edits are among the hardest ones; they require high precision, but also a global awareness of the entire shape. Even in the deep learning era, obtaining manipulable representations that satisfy such requirements is an open problem posing significant constraints. In this work, we address this problem by defining a data‐driven model upon a family of linear operators (variants of the mesh Laplacian), whose spectra capture global and local geometric properties of the shape at hand. Modifications to these spectra are translated to semantically valid deformations of the corresponding surface. By explicitly decoupling the global from the local surface features, our pipeline allows to perform local edits while simultaneously maintaining a global stylistic coherence. We empirically demonstrate how our learning‐based model generalizes to shape representations not seen at training time, and we systematically analyze different choices of local operators over diverse shape categories. Marco Pegoraro 0002, Simone Melzi, Umberto Castellani, Riccardo Marin, Emanuele Rodolà |
Comput. Graph. Forum | 5 |
| 2022 | 3D Shape Analysis Through a Quantum Lens: the Average Mixing Kernel SignatureabstractAbstract The Average Mixing Kernel Signature is a novel spectral signature for points on non-rigid three-dimensional shapes. It is based on a quantum exploration process of the shape surface, where the average transition probabilities between the points of the shape are summarised in the finite-time average mixing kernel. A band-filtered spectral analysis of this kernel then yields the AMKS. Crucially, we show that opting for a finite time-evolution allows the signature to account for a mixing of the Laplacian eigenspaces, similar to what is observed in the presence of noise, explaining the increased noise robustness of this signature when compared to alternative signatures. We perform an extensive experimental analysis of the AMKS under a wide range of problem scenarios, evaluating the performance of our descriptor under different sources of noise (vertex jitter and topological), shape representations (mesh and point clouds), as well as when only a partial view of the shape is available. Our experiments show that the AMKS consistently outperforms two of the most widely used spectral signatures, the Heat Kernel Signature and the Wave Kernel Signature, and suggest that the AMKS should be the signature of choice for various compute vision problems, including as input of deep convolutional architectures for shape analysis. Luca Cosmo, Giorgia Minello, Michael M. Bronstein, Emanuele Rodolà, Luca Rossi 0004, Andrea Torsello |
Int. J. Comput. Vis. | 4 |
| 2022 | Multimodal Feature Fusion and Knowledge-Driven Learning via Experts Consult for Thyroid Nodule ClassificationabstractComputer-aided diagnosis (CAD) is becoming a prominent approach to assist clinicians spanning across multiple fields. These automated systems take advantage of various computer vision (CV) procedures, as well as artificial intelligence (AI) techniques, to formulate a diagnosis of a given image, e.g., computed tomography and ultrasound. Advances in both areas (CV and AI) are enabling ever increasing performances of CAD systems, which can ultimately avoid performing invasive procedures such as fine-needle aspiration. In this study, a novel end-to-end knowledge-driven classification framework is presented. The system focuses on multimodal data generated by thyroid ultrasonography, and acts as a CAD system by providing a thyroid nodule classification into the benign and malignant categories. Specifically, the proposed system leverages cues provided by an ensemble of experts to guide the learning phase of a densely connected convolutional network (DenseNet). The ensemble is composed by various networks pretrained on ImageNet, including AlexNet, ResNet, VGG, and others. The previously computed multimodal feature parameters are used to create ultrasonography domain experts via transfer learning, decreasing, moreover, the number of samples required for training. To validate the proposed method, extensive experiments were performed, providing detailed performances for both the experts ensemble and the knowledge-driven DenseNet. As demonstrated by the results, the proposed system achieves relevant performances in terms of qualitative metrics for the thyroid nodule classification task, thus resulting in a great asset when formulating a diagnosis. Danilo Avola, Luigi Cinque, Alessio Fagioli 0001, Sebastiano Filetti, Giorgio Grani, Emanuele Rodolà |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Universal Spectral Adversarial Attacks for Deformable ShapesabstractMachine learning models are known to be vulnerable to adversarial attacks, namely perturbations of the data that lead to wrong predictions despite being imperceptible. However, the existence of "universal" attacks (i.e., unique perturbations that transfer across different data points) has only been demonstrated for images to date. Part of the reason lies in the lack of a common domain, for geometric data such as graphs, meshes, and point clouds, where a universal perturbation can be defined. In this paper, we offer a change in perspective and demonstrate the existence of universal attacks for geometric data (shapes). We introduce a computational procedure that operates entirely in the spectral domain, where the attacks take the form of small perturbations to short eigenvalue sequences; the resulting geometry is then synthesized via shape-from-spectrum recovery. Our attacks are universal, in that they transfer across different shapes, different representations (meshes and point clouds), and generalize to previously unseen data. Arianna Rampini, Franco Pestarini, Luca Cosmo, Simone Melzi, Emanuele Rodolà |
CVPR | 5 |
| 2021 | Learning disentangled representations via product manifold projectionabstractWe propose a novel approach to disentangle the generative factors of variation underlying a given set of observations. Our method builds upon the idea that the (unknown) low-dimensional manifold underlying the data space can be explicitly modeled as a product of submanifolds. This definition of disentanglement gives rise to a novel weakly-supervised algorithm for recovering the unknown explanatory factors behind the data. At training time, our algorithm only requires pairs of non i.i.d. data samples whose elements share at least one, possibly multidimensional, generative factor of variation. We require no knowledge on the nature of these transformations, and do not make any limiting assumption on the properties of each subspace. Our approach is easy to implement, and can be successfully applied to different kinds of data (from images to 3D surfaces) undergoing arbitrary transformations. In addition to standard synthetic benchmarks, we showcase our method in challenging real-world applications, where we compare favorably with the state of the art. Marco Fumero, Luca Cosmo, Simone Melzi, Emanuele Rodolà |
ICML | 4 |
| 2021 | Efficiently Parallelizable Strassen-Based Multiplication of a Matrix by its TransposeabstractThe multiplication of a matrix by its transpose, ATA, appears as an intermediate operation in the solution of a wide set of problems. In this paper, we propose a new cache-oblivious algorithm (AtA) for computing this product, based upon the classical Strassen algorithm as a sub-routine. In particular, we decrease the computational cost to the time required by Strassen’s algorithm, amounting to floating point operations. AtA works for generic rectangular matrices, and exploits the peculiar symmetry of the resulting product matrix for saving memory. In addition, we provide an extensive implementation study of AtA in a shared memory system, and extend its applicability to a distributed environment. To support our findings, we compare our algorithm with state-of-the-art solutions specialized in the computation of ATA. Our experiments highlight good scalability with respect to both the matrix size and the number of involved processes, as well as favorable performance for both the parallel paradigms and the sequential implementation, when compared with other methods in the literature. Viviana Arrigoni, Filippo Maggioli, Annalisa Massini, Emanuele Rodolà |
ICPP | 4 |
| 2021 | Shape Registration in the Time of TransformersabstractIn this paper, we propose a transformer-based procedure for the efficient registration of non-rigid 3D point clouds. The proposed approach is data-driven and adopts for the first time the transformers architecture in the registration task. Our method is general and applies to different settings. Given a fixed template with some desired properties (e.g. skinning weights or other animation cues), we can register raw acquired data to it, thereby transferring all the template properties to the input geometry. Alternatively, given a pair of shapes, our method can register the first onto the second (or vice-versa), obtaining a high-quality dense correspondence between the two.In both contexts, the quality of our results enables us to target real applications such as texture transfer and shape interpolation.Furthermore, we also show that including an estimation of the underlying density of the surface eases the learning process. By exploiting the potential of this architecture, we can train our model requiring only a sparse set of ground truth correspondences ($10\sim20\%$ of the total points). The proposed model and the analysis that we perform pave the way for future exploration of transformer-based architectures for registration and matching applications. Qualitative and quantitative evaluations demonstrate that our pipeline outperforms state-of-the-art methods for deformable and unordered 3D data registration on different datasets and scenarios. Giovanni Trappolini, Luca Cosmo, Luca Moschella, Riccardo Marin, Simone Melzi, Emanuele Rodolà |
NeurIPS | 6 |
| 2021 | Wavelet-based Heat Kernel Derivatives: Towards Informative Localized Shape AnalysisabstractAbstract In this paper, we propose a new construction for the Mexican hat wavelets on shapes with applications to partial shape matching. Our approach takes its main inspiration from the well‐established methodology of diffusion wavelets. This novel construction allows us to rapidly compute a multi‐scale family of Mexican hat wavelet functions, by approximating the derivative of the heat kernel. We demonstrate that this leads to a family of functions that inherit many attractive properties of the heat kernel (e.g. local support, ability to recover isometries from a single point, efficient computation). Due to its natural ability to encode high‐frequency details on a shape, the proposed method reconstructs and transfers ‐functions more accurately than the Laplace‐Beltrami eigenfunction basis and other related bases. Finally, we apply our method to the challenging problems of partial and large‐scale shape matching. An extensive comparison to the state‐of‐the‐art shows that it is comparable in performance, while both simpler and much faster than competing approaches. Maxime Kirgo, Simone Melzi, Giuseppe Patanè 0001, Emanuele Rodolà, Maks Ovsjanikov |
Comput. Graph. Forum | 4 |
| 2021 | Orthogonalized Fourier Polynomials for Signal Approximation and TransferabstractAbstract We propose a novel approach for the approximation and transfer of signals across 3D shapes. The proposed solution is based on taking pointwise polynomials of the Fourier‐like Laplacian eigenbasis, which provides a compact and expressive representation for general signals defined on the surface. Key to our approach is the construction of a new orthonormal basis upon the set of these linearly dependent polynomials. We analyze the properties of this representation, and further provide a complete analysis of the involved parameters. Our technique results in accurate approximation and transfer of various families of signals between near‐isometric and non‐isometric shapes, even under poor initialization. Our experiments, showcased on a selection of downstream tasks such as filtering and detail transfer, show that our method is more robust to discretization artifacts, deformation and noise as compared to alternative approaches. Filippo Maggioli, Simone Melzi, Maks Ovsjanikov, Michael M. Bronstein, Emanuele Rodolà |
Comput. Graph. Forum | 5 |
| 2021 | Spectral Shape Recovery and Analysis Via Data-driven ConnectionsabstractWe introduce a novel learning-based method to recover shapes from their Laplacian spectra, based on establishing and exploring connections in a learned latent space. The core of our approach consists in a cycle-consistent module that maps between a learned latent space and sequences of eigenvalues. This module provides an efficient and effective link between the shape geometry, encoded in a latent vector, and its Laplacian spectrum. Our proposed data-driven approach replaces the need for ad-hoc regularizers required by prior methods, while providing more accurate results at a fraction of the computational cost. Moreover, these latent space connections enable novel applications for both analyzing and controlling the spectral properties of deformable shapes, especially in the context of a shape collection. Our learning model and the associated analysis apply without modifications across different dimensions (2D and 3D shapes alike), representations (meshes, contours and point clouds), nature of the latent space (generated by an auto-encoder or a parametric model), as well as across different shape classes, and admits arbitrary resolution of the input spectrum without affecting complexity. The increased flexibility allows us to address notoriously difficult tasks in 3D vision and geometry processing within a unified framework, including shape generation from spectrum, latent space exploration and analysis, mesh super-resolution, shape exploration, style transfer, spectrum estimation for point clouds, segmentation transfer and non-rigid shape matching. SUPPLEMENTARY INFORMATION: The online version supplementary material available at 10.1007/s11263-021-01492-6. Riccardo Marin, Arianna Rampini, Umberto Castellani, Emanuele Rodolà, Maks Ovsjanikov, Simone Melzi |
Int. J. Comput. Vis. | 4 |
| 2020 | Instant recovery of shape from spectrum via latent space connectionsabstractWe introduce the first learning-based method for recovering shapes from Laplacian spectra. Our model consists of a cycle-consistent module that maps between learned latent vectors of an auto-encoder and sequences of eigenvalues. This module provides an efficient and effective linkage between Laplacian spectrum and geometry. Our data-driven approach replaces the need for ad-hoc regularizers required by prior methods, while providing more accurate results at a fraction of the computational cost. Our learning model applies without modifications across different dimensions (2D and 3D shapes alike), representations (meshes, contours and point clouds), as well as across different shape classes, and admits arbitrary resolution of the input spectrum without affecting complexity. The increased flexibility allows us to address notoriously difficult tasks in 3D vision and geometry processing within a unified framework, including shape generation from spectrum, mesh super-resolution, shape exploration, style transfer, spectrum estimation from point clouds, segmentation transfer and point-to-point matching. Riccardo Marin, Arianna Rampini, Umberto Castellani, Emanuele Rodolà, Maks Ovsjanikov, Simone Melzi |
3DV | 4 |
| 2020 | LIMP: Learning Latent Shape Representations with Metric Preservation Priors
Luca Cosmo, Antonio Norelli, Oshri Halimi, Ron Kimmel, Emanuele Rodolà |
ECCV (3) | 5 |
| 2020 | Towards Precise Completion of Deformable Shapes
Oshri Halimi, Ido Imanuel, Or Litany, Giovanni Trappolini, Emanuele Rodolà, Leonidas J. Guibas, Ron Kimmel |
ECCV (24) | 5 |
| 2020 | Generating Adversarial Surfaces via Band-Limited PerturbationsabstractAbstract Adversarial attacks have demonstrated remarkable efficacy in altering the output of a learning model by applying a minimal perturbation to the input data. While increasing attention has been placed on the image domain, however, the study of adversarial perturbations for geometric data has been notably lagging behind. In this paper, we show that effective adversarial attacks can be concocted for surfaces embedded in 3D, under weak smoothness assumptions on the perceptibility of the attack. We address the case of deformable 3D shapes in particular, and introduce a general model that is not tailored to any specific surface representation, nor does it assume access to a parametric description of the 3D object. In this context, we consider targeted and untargeted variants of the attack, demonstrating compelling results in either case. We further show how discovering adversarial examples, and then using them for adversarial training, leads to an increase in both robustness and accuracy. Our findings are confirmed empirically over multiple datasets spanning different semantic classes and deformations. Giorgio Mariani, Luca Cosmo, Alexander M. Bronstein, Emanuele Rodolà |
Comput. Graph. Forum | 4 |
| 2020 | FARM: Functional Automatic Registration Method for 3D Human BodiesabstractAbstract We introduce a new method for non‐rigid registration of 3D human shapes. Our proposed pipeline builds upon a given parametric model of the human, and makes use of the functional map representation for encoding and inferring shape maps throughout the registration process. This combination endows our method with robustness to a large variety of nuisances observed in practical settings, including non‐isometric transformations, downsampling, topological noise and occlusions; further, the pipeline can be applied invariably across different shape representations (e.g. meshes and point clouds), and in the presence of (even dramatic) missing parts such as those arising in real‐world depth sensing applications. We showcase our method on a selection of challenging tasks, demonstrating results in line with, or even surpassing, state‐of‐the‐art methods in the respective areas. Riccardo Marin, Simone Melzi, Emanuele Rodolà, Umberto Castellani |
Comput. Graph. Forum | 3 |
| 2020 | A parametric analysis of discrete Hamiltonian functional mapsabstractAbstract In this paper we develop an in‐depth theoretical investigation of the discrete Hamiltonian eigenbasis, which remains quite unexplored in the geometry processing community. This choice is supported by the fact that Dirichlet eigenfunctions can be equivalently computed by defining a Hamiltonian operator, whose potential energy and localization region can be controlled with ease. We vary with continuity the potential energy and study the relationship between the Dirichlet Laplacian and the Hamiltonian eigenbases with the functional map formalism. We develop a global analysis to capture the asymptotic behavior of the eigenpairs. We then focus on their local interactions, namely the veering patterns that arise between proximal eigenvalues. Armed with this knowledge, we are able to track the eigenfunctions in all possible configurations, shedding light on the nature of the functional maps. We exploit the Hamiltonian‐Dirichlet connection in a partial shape matching problem, obtaining state of the art results, and provide directions where our theoretical findings could be applied in future research. Emilian Postolache, Marco Fumero, Luca Cosmo, Emanuele Rodolà |
Comput. Graph. Forum | 4 |
| 2020 | 2-D Skeleton-Based Action Recognition via Two-Branch Stacked LSTM-RNNsabstractAction recognition in video sequences is an interesting field for many computer vision applications, including behavior analysis, event recognition, and video surveillance. In this article, a method based on 2D skeleton and two-branch stacked Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM) cells is proposed. Unlike 3D skeletons, usually generated by RGB-D cameras, the 2D skeletons adopted in this article are reconstructed starting from RGB video streams, therefore allowing the use of the proposed approach in both indoor and outdoor environments. Moreover, any case of missing skeletal data is managed by exploiting 3D-Convolutional Neural Networks (3D-CNNs). Comparative experiments with several key works on KTH and Weizmann datasets show that the method described in this paper outperforms the current state-of-the-art. Additional experiments on UCF Sports and IXMAS datasets demonstrate the effectiveness of our method in the presence of noisy data and perspective changes, respectively. Further investigations on UCF Sports, HMDB51, UCF101, and Kinetics400 highlight how the combination between the proposed two-branch stacked LSTM and the 3D-CNN-based network can manage missing skeleton information, greatly improving the overall accuracy. Moreover, additional tests on KTH and UCF Sports datasets also show the robustness of our approach in the presence of partial body occlusions. Finally, comparisons on UT-Kinect and NTU-RGB+D datasets show that the accuracy of the proposed method is fully comparable to that of works based on 3D skeletons. Danilo Avola, Marco Cascio, Luigi Cinque, Gian Luca Foresti, Cristiano Massaroni, Emanuele Rodolà |
IEEE Trans. Multim. | 6 |
| 2020 | Nonlinear spectral geometry processing via the TV transformabstractWe introduce a novel computational framework for digital geometry processing, based upon the derivation of a nonlinear operator associated to the total variation functional. Such an operator admits a generalized notion of spectral decomposition, yielding a convenient multiscale representation akin to Laplacian-based methods, while at the same time avoiding undesirable over-smoothing effects typical of such techniques. Our approach entails accurate, detail-preserving decomposition and manipulation of 3D shape geometry while taking an especially intuitive form: non-local semantic details are well separated into different bands, which can then be filtered and re-synthesized with a straightforward linear step. Our computational framework is flexible, can be applied to a variety of signals, and is easily adapted to different geometry representations, including triangle meshes and point clouds. We showcase our method through multiple applications in graphics, ranging from surface and signal denoising to enhancement, detail transfer, and cubic stylization. Marco Fumero, Michael Möller 0001, Emanuele Rodolà |
ACM Trans. Graph. | 3 |
| 2019 | High-Resolution Augmentation for Automatic Template-Based Matching of Human ModelsabstractWe propose a new approach for 3D shape matching of deformable human shapes. Our approach is based on the joint adoption of three different tools: an intrinsic spectral matching pipeline, a morphable model, and an extrinsic details refinement. By operating in conjunction, these tools allow us to greatly improve the quality of the matching while at the same time resolving the key issues exhibited by each tool individually. In this paper we present an innovative High-Resolution Augmentation (HRA) strategy that enables highly accurate correspondence even in the presence of significant mesh resolution mismatch between the input shapes. This augmentation provides an effective workaround for the resolution limitations imposed by the adopted morphable model. The HRA in its global and localized versions represents a novel refinement strategy for surface subdivision methods. We demonstrate the accuracy of the proposed pipeline on multiple challenging benchmarks, and showcase its effectiveness in surface registration and texture transfer. Riccardo Marin, Simone Melzi, Emanuele Rodolà, Umberto Castellani |
3DV | 3 |
| 2019 | Correspondence-Free Region Localization for Partial Shape Similarity via Hamiltonian Spectrum AlignmentabstractWe consider the problem of localizing relevant subsets of non-rigid geometric shapes given only a partial 3D query as the input. Such problems arise in several challenging tasks in 3D vision and graphics, including partial shape similarity, retrieval, and non-rigid correspondence. We phrase the problem as one of alignment between short sequences of eigenvalues of basic differential operators, which are constructed upon a scalar function defined on the 3D surfaces. Our method therefore seeks for a scalar function that entails this alignment. Differently from existing approaches, we do not require solving for a correspondence between the query and the target, therefore greatly simplifying the optimization process; our core technique is also descriptor-free, as it is driven by the geometry of the two objects as encoded in their operator spectra. We further show that our spectral alignment algorithm provides a remarkably simple alternative to the recent shape-from-spectrum reconstruction approaches. For both applications, we demonstrate improvement over the state-of-the-art either in terms of accuracy or computational cost. Arianna Rampini, Irene Tallini, Maks Ovsjanikov, Alexander M. Bronstein, Emanuele Rodolà |
3DV | 5 |
| 2019 | Isospectralization, or How to Hear Shape, Style, and CorrespondenceabstractThe question whether one can recover the shape of a geometric object from its Laplacian spectrum (`hear the shape of the drum') is a classical problem in spectral geometry with a broad range of implications and applications. While theoretically the answer to this question is negative (there exist examples of iso-spectral but non-isometric manifolds), little is known about the practical possibility of using the spectrum for shape reconstruction and optimization. In this paper, we introduce a numerical procedure called isospectralization, consisting of deforming one shape to make its Laplacian spectrum match that of another. We implement the isospectralization procedure using modern differentiable programming techniques and exemplify its applications in some of the classical and notoriously hard problems in geometry processing, computer vision, and graphics such as shape reconstruction, pose and style transfer, and dense deformable correspondence. Luca Cosmo, Mikhail Panine, Arianna Rampini, Maks Ovsjanikov, Michael M. Bronstein, Emanuele Rodolà |
CVPR | 6 |
| 2019 | Unsupervised Learning of Dense Shape CorrespondenceabstractWe introduce the first completely unsupervised correspondence learning approach for deformable 3D shapes. Key to our model is the understanding that natural deformations (such as changes in pose) approximately preserve the metric structure of the surface, yielding a natural criterion to drive the learning process toward distortion-minimizing predictions. On this basis, we overcome the need for annotated data and replace it by a purely geometric criterion. The resulting learning model is class-agnostic, and is able to leverage any type of deformable geometric data for the training phase. In contrast to existing supervised approaches which specialize on the class seen at training time, we demonstrate stronger generalization as well as applicability to a variety of challenging settings. We showcase our method on a wide selection of correspondence benchmarks, where we outperform other methods in terms of accuracy, generalization, and efficiency. Oshri Halimi, Or Litany, Emanuele Rodolà, Alexander M. Bronstein, Ron Kimmel |
CVPR | 3 |
| 2019 | GFrames: Gradient-Based Local Reference Frame for 3D Shape MatchingabstractWe introduce GFrames, a novel local reference frame (LRF) construction for 3D meshes and point clouds. GFrames are based on the computation of the intrinsic gradient of a scalar field defined on top of the input shape. The resulting tangent vector field defines a repeatable tangent direction of the local frame at each point; importantly, it directly inherits the properties and invariance classes of the underlying scalar function, making it remarkably robust under strong sampling artifacts, vertex noise, as well as non-rigid deformations. Existing local descriptors can directly benefit from our repeatable frames, as we showcase in a selection of 3D vision and shape analysis applications where we demonstrate state-of-the-art performance in a variety of challenging settings. Simone Melzi, Riccardo Spezialetti, Federico Tombari, Michael M. Bronstein, Luigi Di Stefano, Emanuele Rodolà |
CVPR | 6 |
| 2019 | Functional Maps Representation On Product ManifoldsabstractAbstract We consider the tasks of representing, analysing and manipulating maps between shapes. We model maps as densities over the product manifold of the input shapes; these densities can be treated as scalar functions and therefore are manipulable using the language of signal processing on manifolds. Being a manifold itself, the product space endows the set of maps with a geometry of its own, which we exploit to define map operations in the spectral domain; we also derive relationships with other existing representations (soft maps and functional maps). To apply these ideas in practice, we discretize product manifolds and their Laplace–Beltrami operators, and we introduce localized spectral analysis of the product manifold as a novel tool for map processing. Our framework applies to maps defined between and across 2D and 3D shapes without requiring special adjustment, and it can be implemented efficiently with simple operations on sparse matrices. Emanuele Rodolà, Zorah Lähner, Alexander M. Bronstein, Michael M. Bronstein, Justin Solomon 0001 |
Comput. Graph. Forum | 1 |
| 2019 | ZoomOut: spectral upsampling for efficient shape correspondenceabstractWe present a simple and efficient method for refining maps or correspondences by iterative upsampling in the spectral domain that can be implemented in a few lines of code. Our main observation is that high quality maps can be obtained even if the input correspondences are noisy or are encoded by a small number of coefficients in a spectral basis. We show how this approach can be used in conjunction with existing initialization techniques across a range of application scenarios, including symmetry detection, map refinement across complete shapes, non-rigid partial shape matching and function transfer. In each application we demonstrate an improvement with respect to both the quality of the results and the computational speed compared to the best competing methods, with up to two orders of magnitude speed-up in some applications. We also demonstrate that our method is both robust to noisy input and is scalable with respect to shape complexity. Finally, we present a theoretical justification for our approach, shedding light on structural properties of functional maps. Simone Melzi, Jing Ren 0004, Emanuele Rodolà, Abhishek Sharma 0013, Peter Wonka, Maks Ovsjanikov |
ACM Trans. Graph. | 3 |
| 2018 | Structured Prediction of Dense Maps between Geometric DomainsabstractWe introduce a new framework for learning dense correspondence between deformable geometric domains such as polygonal meshes and point clouds. Existing learning based approaches model correspondence as a labelling problem, where each point of a query domain receives a label identifying a point on some reference domain; the correspondence is then constructed a posteriori by composing the label predictions of two input geometries. We propose a paradigm shift and design a structured prediction model in the space of functional maps, linear operators that provide a compact representation of the correspondence. We model the learning process via a deep residual network which takes dense descriptor fields as input, and outputs a soft map between the two given objects. The resulting correspondence is shown to be accurate on several challenging shape correspondence benchmarks. Emanuele Rodolà |
ICASSP | 1 |
| 2018 | Localized Manifold Harmonics for Spectral Shape AnalysisabstractAbstract The use of Laplacian eigenfunctions is ubiquitous in a wide range of computer graphics and geometry processing applications. In particular, Laplacian eigenbases allow generalizing the classical Fourier analysis to manifolds. A key drawback of such bases is their inherently global nature, as the Laplacian eigenfunctions carry geometric and topological structure of the entire manifold. In this paper, we introduce a new framework for local spectral shape analysis. We show how to efficiently construct localized orthogonal bases by solving an optimization problem that in turn can be posed as the eigendecomposition of a new operator obtained by a modification of the standard Laplacian. We study the theoretical and computational aspects of the proposed framework and showcase our new construction on the classical problems of shape approximation and correspondence. We obtain significant improvement compared to classical Laplacian eigenbases as well as other alternatives for constructing localized bases. Simone Melzi, Emanuele Rodolà, Umberto Castellani, Michael M. Bronstein |
Comput. Graph. Forum | 2 |
| 2018 | Improved Functional Mappings via Product PreservationabstractAbstract In this paper, we consider the problem of information transfer across shapes and propose an extension to the widely used functional map representation. Our main observation is that in addition to the vector space structure of the functional spaces, which has been heavily exploited in the functional map framework, the functional algebra (i.e., the ability to take pointwise products of functions) can significantly extend the power of this framework. Equipped with this observation, we show how to improve one of the key applications of functional maps, namely transferring real‐valued functions without conversion to point‐to‐point correspondences. We demonstrate through extensive experiments that by decomposing a given function into a linear combination consisting not only of basis functions but also of their pointwise products, both the representation power and the quality of the function transfer can be improved significantly. Our modification, while computationally simple, allows us to achieve higher transfer accuracy while keeping the size of the basis and the functional map fixed. We also analyze the computational complexity of optimally representing functions through linear combinations of products in a given basis and prove NP‐completeness in some general cases. Finally, we argue that the use of function products can have a wide‐reaching effect in extending the power of functional maps in a variety of applications, in particular by enabling the transfer of high‐frequency functions without changing the representation size or complexity. Dorian Nogneng, Simone Melzi, Emanuele Rodolà, Umberto Castellani, Michael M. Bronstein, Maks Ovsjanikov |
Comput. Graph. Forum | 3 |
| 2017 | Spatial Maps: From Low Rank Spectral to Sparse Spatial Functional RepresentationsabstractFunctional representation is a well-established approach to represent dense correspondences between deformable shapes. The approach provides an efficient low rank representation of a continuous mapping between two shapes, however under that framework the correspondences are only intrinsically captured, which implies that the induced map is not guaranteed to map the whole surface, much less to form a continuous mapping. In this work, we define a novel approach to the computation of a continuous bijective map between two surfaces moving from the low rank spectral representation to a sparse spatial representation. Key to this is the observation that continuity and smoothness of the optimal map induces structure both on the spectral and the spatial domain, the former providing effective low rank approximations, while the latter exhibiting strong sparsity and locality that can be used in the solution of large-scale problems. We cast our approach in terms of the functional transfer through a fuzzy map between shapes satisfying infinitesimal mass transportation at each point. The result is that, not only the spatial map induces a sub-vertex correspondence between the surfaces, but also the transportation of the whole surface, and thus the bijectivity of the induced map is assured. The performance of the proposed method is assessed on several popular benchmarks. Andrea Gasparetto, Luca Cosmo, Emanuele Rodolà, Michael M. Bronstein, Andrea Torsello |
3DV | 3 |
| 2017 | Efficient Deformable Shape Correspondence via Kernel MatchingabstractWe present a method to match three dimensional shapes under non-isometric deformations, topology changes and partiality. We formulate the problem as matching between a set of pair-wise and point-wise descriptors, imposing a continuity prior on the mapping, and propose a projected descent optimization procedure inspired by difference of convex functions (DC) programming. Matthias Vestner, Zorah Lähner, Amit Boyarski, Or Litany, Ron Slossberg, Tal Remez, Emanuele Rodolà, Alexander M. Bronstein, Michael M. Bronstein, Ron Kimmel, Daniel Cremers |
3DV | 7 |
| 2017 | Geometric Deep Learning on Graphs and Manifolds Using Mixture Model CNNsabstractDeep learning has achieved a remarkable performance breakthrough in several fields, most notably in speech recognition, natural language processing, and computer vision. In particular, convolutional neural network (CNN) architectures currently produce state-of-the-art performance on a variety of image analysis tasks such as object detection and recognition. Most of deep learning research has so far focused on dealing with 1D, 2D, or 3D Euclidean-structured data such as acoustic signals, images, or videos. Recently, there has been an increasing interest in geometric deep learning, attempting to generalize deep learning methods to non-Euclidean structured data such as graphs and manifolds, with a variety of applications from the domains of network analysis, computational social science, or computer graphics. In this paper, we propose a unified framework allowing to generalize CNN architectures to non-Euclidean domains (graphs and manifolds) and learn local, stationary, and compositional task-specific features. We show that various non-Euclidean CNN methods previously proposed in the literature can be considered as particular instances of our framework. We test the proposed method on standard tasks from the realms of image-, graph-and 3D shape analysis and show that it consistently outperforms previous approaches. Federico Monti, Davide Boscaini, Jonathan Masci, Emanuele Rodolà, Jan Svoboda, Michael M. Bronstein |
CVPR | 4 |
| 2017 | Product Manifold Filter: Non-rigid Shape Correspondence via Kernel Density Estimation in the Product SpaceabstractMany algorithms for the computation of correspondences between deformable shapes rely on some variant of nearest neighbor matching in a descriptor space. Such are, for example, various point-wise correspondence recovery algorithms used as a post-processing stage in the functional correspondence framework. Such frequently used techniques implicitly make restrictive assumptions (e.g., nearisometry) on the considered shapes and in practice suffer from lack of accuracy and result in poor surjectivity. We propose an alternative recovery technique capable of guaranteeing a bijective correspondence and producing significantly higher accuracy and smoothness. Unlike other methods our approach does not depend on the assumption that the analyzed shapes are isometric. We derive the proposed method from the statistical framework of kernel density estimation and demonstrate its performance on several challenging deformable 3D shape matching datasets. Matthias Vestner, Roee Litman, Emanuele Rodolà, Alexander M. Bronstein, Daniel Cremers |
CVPR | 3 |
| 2017 | Deep Functional Maps: Structured Prediction for Dense Shape CorrespondenceabstractWe introduce a new framework for learning dense correspondence between deformable 3D shapes. Existing learning based approaches model shape correspondence as a labelling problem, where each point of a query shape receives a label identifying a point on some reference domain; the correspondence is then constructed a posteriori by composing the label predictions of two input shapes. We propose a paradigm shift and design a structured prediction model in the space of functional maps, linear operators that provide a compact representation of the correspondence. We model the learning process via a deep residual network which takes dense descriptor fields defined on two shapes as input, and outputs a soft map between the two given objects. The resulting correspondence is shown to be accurate on several challenging benchmarks comprising multiple categories, synthetic models, real scans with acquisition artifacts, topological noise, and partiality. Or Litany, Tal Remez, Emanuele Rodolà, Alexander M. Bronstein, Michael M. Bronstein |
ICCV | 3 |
| 2017 | Consistent Partial Matching of Shape Collections via Sparse ModelingabstractAbstract Recent efforts in the area of joint object matching approach the problem by taking as input a set of pairwise maps, which are then jointly optimized across the whole collection so that certain accuracy and consistency criteria are satisfied. One natural requirement is cycle‐consistency—namely the fact that map composition should give the same result regardless of the path taken in the shape collection. In this paper, we introduce a novel approach to obtain consistent matches without requiring initial pairwise solutions to be given as input. We do so by optimizing a joint measure of metric distortion directly over the space of cycle‐consistent maps; in order to allow for partially similar and extra‐class shapes, we formulate the problem as a series of quadratic programs with sparsity‐inducing constraints, making our technique a natural candidate for analysing collections with a large presence of outliers. The particular form of the problem allows us to leverage results and tools from the field of evolutionary game theory. This enables a highly efficient optimization procedure which assures accurate and provably consistent solutions in a matter of minutes in collections with hundreds of shapes. Luca Cosmo, Emanuele Rodolà, Andrea Albarelli, Facundo Mémoli, Daniel Cremers |
Comput. Graph. Forum | 2 |
| 2017 | Fully Spectral Partial Shape MatchingabstractWe propose an efficient procedure for calculating partial dense intrinsic correspondence between deformable shapes performed entirely in the spectral domain. Our technique relies on the recently introduced partial functional maps formalism and on the joint approximate diagonalization (JAD) of the Laplace-Beltrami operators previously introduced for matching non-isometric shapes. We show that a variant of the JAD problem with an appropriately modified coupling term (surprisingly) allows to construct quasi-harmonic bases localized on the latent corresponding parts. This circumvents the need to explicitly compute the unknown parts by means of the cumbersome alternating minimization used in the previous approaches, and allows performing all the calculations in the spectral domain with constant complexity independent of the number of shape vertices. We provide an extensive evaluation of the proposed technique on standard non-rigid correspondence benchmarks and show state-of-the-art performance in various settings, including partiality and the presence of topological noise. Or Litany, Emanuele Rodolà, Alexander M. Bronstein, Michael M. Bronstein |
Comput. Graph. Forum | 2 |
| 2017 | Partial Functional CorrespondenceabstractAbstract In this paper, we propose a method for computing partial functional correspondence between non‐rigid shapes. We use perturbation analysis to show how removal of shape parts changes the Laplace–Beltrami eigenfunctions, and exploit it as a prior on the spectral representation of the correspondence. Corresponding parts are optimization variables in our problem and are used to weight the functional correspondence; we are looking for the largest and most regular (in the Mumford–Shah sense) parts that minimize correspondence distortion. We show that our approach can cope with very challenging correspondence settings. Emanuele Rodolà, Luca Cosmo, Michael M. Bronstein, Andrea Torsello, Daniel Cremers |
Comput. Graph. Forum | 1 |
| 2017 | Regularized Pointwise Map Recovery from Functional CorrespondenceabstractAbstract The concept of using functional maps for representing dense correspondences between deformable shapes has proven to be extremely effective in many applications. However, despite the impact of this framework, the problem of recovering the point‐to‐point correspondence from a given functional map has received surprisingly little interest. In this paper, we analyse the aforementioned problem and propose a novel method for reconstructing pointwise correspondences from a given functional map. The proposed algorithm phrases the matching problem as a regularized alignment problem of the spectral embeddings of the two shapes. Opposed to established methods, our approach does not require the input shapes to be nearly‐isometric, and easily extends to recovering the point‐to‐point correspondence in part‐to‐whole shape matching problems. Our numerical experiments demonstrate that the proposed approach leads to a significant improvement in accuracy in several challenging cases. Emanuele Rodolà, Michael Möller 0001, Daniel Cremers |
Comput. Graph. Forum | 1 |
| 2016 | Matching Deformable Objects in ClutterabstractWe consider the problem of deformable object detection and dense correspondence in cluttered 3D scenes. Key ingredient to our method is the choice of representation: we formulate the problem in the spectral domain using the functional maps framework, where we seek for the most regular nearly-isometric parts in the model and the scene that minimize correspondence error. The problem is initialized by solving a sparse relaxation of a quadratic assignment problem on features obtained via data-driven metric learning. The resulting matching pipeline is solved efficiently, and yields accurate results in challenging settings that were previously left unexplored in the literature. Luca Cosmo, Emanuele Rodolà, Jonathan Masci, Andrea Torsello, Michael M. Bronstein |
3DV | 2 |
| 2016 | Coupled Functional MapsabstractClassical formulations of the shape matching problem involve the definition of a matching cost that directly depends on the action of the desired map when applied to some input data. Such formulations are typically one-sided - they seek for a mapping from one shape to the other, but not vice versa. In this paper we consider an unbiased formulation of this problem, in which we solve simultaneously for a low-distortion map relating the two given shapes and its inverse. We phrase the problem in the spectral domain using the language of functional maps, resulting in an especially compact and efficient optimization problem. The benefits of our proposed regularization are especially evident in the scarce data setting, where we demonstrate highly competitive results with respect to the state of the art. Davide Eynard, Emanuele Rodolà, Klaus Glashoff, Michael M. Bronstein |
3DV | 2 |
| 2016 | Shape Analysis with Anisotropic Windowed Fourier TransformabstractWe propose Anisotropic Windowed Fourier Transform (AWFT), a framework for localized space-frequency analysis of deformable 3D shapes. With AWFT, we are able to extract meaningful intrinsic localized orientation-sensitive structures on surfaces, and use them in applications such as shape segmentation, salient point detection, feature point description, and matching. Our method outperforms previous approaches in the considered applications. Simone Melzi, Emanuele Rodolà, Umberto Castellani, Michael M. Bronstein |
3DV | 2 |
| 2016 | Efficient Globally Optimal 2D-to-3D Deformable Shape MatchingabstractWe propose the first algorithm for non-rigid 2D-to-3D shape matching, where the input is a 2D query shape as well as a 3D target shape and the output is a continuous matching curve represented as a closed contour on the 3D shape. We cast the problem as finding the shortest circular path on the product 3-manifold of the two shapes. We prove that the optimal matching can be computed in polynomial time with a (worst-case) complexity of O(mn2 log(n)), wherem and n denote the number of vertices on the 2D and the 3D shape respectively. Quantitative evaluation confirms that the method provides excellent results for sketch-based deformable 3D shape retrieval. Zorah Lähner, Emanuele Rodolà, Frank R. Schmidt, Michael M. Bronstein, Daniel Cremers |
CVPR | 2 |
| 2016 | A game-theoretical approach for joint matching of multiple feature throughout unordered imagesabstractFeature matching is a key step in most Computer Vision tasks involving several views of the same subject. In fact, it plays a crucial role for a successful reconstruction of 3D information of the corresponding material points. Typical approaches to construct stable feature tracks throughout a sequence of images operate via a two-step process: First, feature matches are extracted among all pairs of points of view; these matches are then given in input to a regularizer that provides a final, globally consistent solution. In this paper, we formulate this matching problem as a simultaneous optimization over the entire image collection, without requiring previously computed pairwise matches to be given as input. As our formulation operates directly in the space of feature across multiple images, the final matches are consistent by construction. Our matching problem has a natural interpretation as a non-cooperative game, which allows us to leverage tools and results from Game Theory. We performed a specially crafted set of experiments demonstrating that our approach compares favorably with the state of the art, while retaining a high computational efficiency. Luca Cosmo, Andrea Albarelli, Filippo Bergamasco, Andrea Torsello, Emanuele Rodolà, Daniel Cremers |
ICPR | 5 |
| 2016 | Learning shape correspondence with anisotropic convolutional neural networksabstractConvolutional neural networks have achieved extraordinary results in many computer vision and pattern recognition applications; however, their adoption in the computer graphics and geometry processing communities is limited due to the non-Euclidean structure of their data. In this paper, we propose Anisotropic Convolutional Neural Network (ACNN), a generalization of classical CNNs to non-Euclidean domains, where classical convolutions are replaced by projections over a set of oriented anisotropic diffusion kernels. We use ACNNs to effectively learn intrinsic dense correspondences between deformable shapes, a fundamental problem in geometry processing, arising in a wide variety of applications. We tested ACNNs performance in very challenging settings, achieving state-of-the-art results on some of the most difficult recent correspondence benchmarks. Davide Boscaini, Jonathan Masci, Emanuele Rodolà, Michael M. Bronstein |
NIPS | 3 |
| 2016 | Anisotropic Diffusion DescriptorsabstractAbstract Spectral methods have recently gained popularity in many domains of computer graphics and geometry processing, especially shape processing, computation of shape descriptors, distances, and correspondence. Spectral geometric structures are intrinsic and thus invariant to isometric deformations, are efficiently computed, and can be constructed on shapes in different representations. A notable drawback of these constructions, however, is that they areisotropic, i.e., insensitive to direction. In this paper, we show how to construct direction‐sensitive spectral feature descriptors usinganisotropic diffusionon meshes and point clouds. The core of our construction are directed local kernels acting similarly to steerable filters, which are learned in a task‐specific manner. Remarkably, while being intrinsic, our descriptors allow to disambiguate reflection symmetries. We show the application of anisotropic descriptors for problems of shape correspondence on meshes and point clouds, achieving results significantly better than state‐of‐the‐art methods. Davide Boscaini, Jonathan Masci, Emanuele Rodolà, Michael M. Bronstein, Daniel Cremers |
Comput. Graph. Forum | 3 |
| 2016 | Non-Rigid PuzzlesabstractAbstract Shape correspondence is a fundamental problem in computer graphics and vision, with applications in various problems including animation, texture mapping, robotic vision, medical imaging, archaeology and many more. In settings where the shapes are allowed to undergo non‐rigid deformations and only partial views are available, the problem becomes very challenging. To this end, we present a non‐rigid multi‐part shape matching algorithm. We assume to be given a reference shape and its multiple parts undergoing a non‐rigid deformation. Each of these query parts can be additionally contaminated by clutter, may overlap with other parts, and there might be missing parts or redundant ones. Our method simultaneously solves for the segmentation of the reference model, and for a dense correspondence to (subsets of) the parts. Experimental results on synthetic as well as real scans demonstrate the effectiveness of our method in dealing with this challenging matching scenario. Or Litany, Emanuele Rodolà, Alexander M. Bronstein, Michael M. Bronstein, Daniel Cremers |
Comput. Graph. Forum | 2 |
| 2016 | An Accurate and Robust Artificial Marker Based on Cyclic CodesabstractArtificial markers are successfully adopted to solve several vision tasks, ranging from tracking to calibration. While most designs share the same working principles, many specialized approaches exist to address specific application domains. Some are specially crafted to boost pose recovery accuracy. Others are made robust to occlusion or easy to detect with minimal computational resources. The sheer amount of approaches available in recent literature is indeed a statement to the fact that no silver bullet exists. Furthermore, this is also a hint to the level of scholarly interest that still characterizes this research topic. With this paper we try to add a novel option to the offer, by introducing a general purpose fiducial marker which exhibits many useful properties while being easy to implement and fast to detect. The key ideas underlying our approach are three. The first one is to exploit the projective invariance of conics to jointly find the marker and set a reading frame for it. Moreover, the tag identity is assessed by a redundant cyclic coded sequence implemented using the same circular features used for detection. Finally, the specific design and feature organization of the marker are well suited for several practical tasks, ranging from camera calibration to information payload delivery. Filippo Bergamasco, Andrea Albarelli, Luca Cosmo, Emanuele Rodolà, Andrea Torsello |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Adopting an unconstrained ray model in light-field cameras for 3D shape reconstructionabstractGiven the raising interest in light-field technology and the increasing availability of professional devices, a feasible and accurate calibration method is paramount to unleash practical applications. In this paper we propose to embrace a fully non-parametric model for the imaging and we show that it can be properly calibrated with little effort using a dense active target. This process produces a dense set of independent rays that cannot be directly used to produce a conventional image. However, they are an ideal tool for 3D reconstruction tasks, since they are highly redundant, very accurate and they cover a wide range of different baselines. The feasibility and convenience of the process and the accuracy of the obtained calibration are comprehensively evaluated through several experiments. Filippo Bergamasco, Andrea Albarelli, Luca Cosmo, Andrea Torsello, Emanuele Rodolà, Daniel Cremers |
CVPR | 5 |
| 2015 | Realistic photometric stereo using partial differential irradiance equation ratios
Roberto Mecca, Emanuele Rodolà, Daniel Cremers |
Comput. Graph. | 2 |
| 2015 | Fast and accurate surface alignment through an isometry-enforcing game
Andrea Albarelli, Emanuele Rodolà, Andrea Torsello |
Pattern Recognit. | 2 |
| 2015 | A simple and effective relevance-based point sampling for 3D shapes
Emanuele Rodolà, Andrea Albarelli, Daniel Cremers, Andrea Torsello |
Pattern Recognit. Lett. | 1 |
| 2014 | Learning Similarities for Rigid and Non-rigid Object DetectionabstractIn this paper, we propose an optimization method for estimating the parameters that typically appear in graph-theoretical formulations of the matching problem for object detection. Although several methods have been proposed to optimize parameters for graph matching in a way to promote correct correspondences and to restrict wrong ones, our approach is novel in the sense that it aims at improving performance in the more general task of object detection. In our formulation, similarity functions are adjusted so as to increase the overall similarity among a reference model and the observed target, and at the same time reduce the similarity among reference and "non-target" objects. We evaluate the proposed method in two challenging scenarios, namely object detection using data captured with a Kinect sensor in a real environment, and intrinsic metric learning for deformable shapes, demonstrating substantial improvements in both settings. Asako Kanezaki, Emanuele Rodolà, Daniel Cremers, Tatsuya Harada |
3DV | 2 |
| 2014 | Optimal Intrinsic Descriptors for Non-Rigid Shape Analysis
Thomas Windheuser, Matthias Vestner, Emanuele Rodolà, Rudolph Triebel, Daniel Cremers |
BMVC | 3 |
| 2014 | Dense Non-rigid Shape Correspondence Using Random ForestsabstractWe propose a shape matching method that produces dense correspondences tuned to a specific class of shapes and deformations. In a scenario where this class is represented by a small set of example shapes, the proposed method learns a shape descriptor capturing the variability of the deformations in the given class. The approach enables the wave kernel signature to extend the class of recognized deformations from near isometries to the deformations appearing in the example set by means of a random forest classifier. With the help of the introduced spatial regularization, the proposed method achieves significant improvements over the baseline approach and obtains state-of-the-art results while keeping short computation times. Emanuele Rodolà, Samuel Rota Bulò, Thomas Windheuser, Matthias Vestner, Daniel Cremers |
CVPR | 1 |
| 2014 | Robust Region Detection via Consensus Segmentation of Deformable ShapesabstractAbstract We consider the problem of stable region detection and segmentation of deformable shapes. We pursue this goal by determining a consensus segmentation from a heterogeneous ensemble of putative segmentations, which are generated by a clustering process on an intrinsic embedding of the shape. The intuition is that the consensus segmentation, which relies on aggregate statistics gathered from the segmentations in the ensemble, can reveal components in the shape that are more stable to deformations than the single baseline segmentations. Compared to the existing approaches, our solution exhibits higher robustness and repeatability throughout a wide spectrum of non‐rigid transformations. It is computationally efficient, naturally extendible to point clouds, and remains semantically stable even across different object classes. A quantitative evaluation on standard datasets confirms the potentiality of our method as a valid tool for deformable shape analysis. Emanuele Rodolà, Samuel Rota Bulò, Daniel Cremers |
Comput. Graph. Forum | 1 |
| 2013 | Efficient Shape Matching using Vector ExtrapolationabstractWe propose the adoption of a vector extrapolation technique to accelerate convergence of correspondence problems under the quadratic assignment formulation for attributed graph matching (QAP). In order to capture a broad range of matching scenarios, we provide a class of relaxations of the QAP under elastic net constraints. This allows us to regulate the sparsity/complexity trade-off which is inherent to most instances of the matching problem, thus enabling us to study the application of the acceleration method over a family of problems of varying difficulty. The validity of the approach is assessed by considering three different matching scenarios; namely, rigid and non-rigid three-dimensional shape matching, and image matching for Structure from Motion. As demonstrated on both real and synthetic data, our approach leads to an increase in performance of up to one order of magnitude when compared to the standard methods. 1 Emanuele Rodolà, Tatsuya Harada, Yasuo Kuniyoshi, Daniel Cremers |
BMVC | 1 |
| 2013 | Can a Fully Unconstrained Imaging Model Be Applied Effectively to Central Cameras?abstractTraditional camera models are often the result of a compromise between the ability to account for non-linearities in the image formation model and the need for a feasible number of degrees of freedom in the estimation process. These considerations led to the definition of several ad hoc models that best adapt to different imaging devices, ranging from pinhole cameras with no radial distortion to the more complex catadioptric or polydioptric optics. In this paper we propose the use of an unconstrained model even in standard central camera settings dominated by the pinhole model, and introduce a novel calibration approach that can deal effectively with the huge number of free parameters associated with it, resulting in a higher precision calibration than what is possible with the standard pinhole model with correction for radial distortion. This effectively extends the use of general models to settings that traditionally have been ruled by parametric approaches out of practical considerations. The benefit of such an unconstrained model to quasi-pinhole central cameras is supported by an extensive experimental validation. Filippo Bergamasco, Andrea Albarelli, Emanuele Rodolà, Andrea Torsello |
CVPR | 3 |
| 2013 | Elastic Net Constraints for Shape MatchingabstractWe consider a parametrized relaxation of the widely adopted quadratic assignment problem (QAP) formulation for minimum distortion correspondence between deformable shapes. In order to control the accuracy/sparsity trade-off we introduce a weighting parameter on the combination of two existing relaxations, namely spectral and game-theoretic. This leads to the introduction of the elastic net penalty function into shape matching problems. In combination with an efficient algorithm to project onto the elastic net ball, we obtain an approach for deformable shape matching with controllable sparsity. Experiments on a standard benchmark confirm the effectiveness of the approach. Emanuele Rodolà, Andrea Torsello, Tatsuya Harada, Yasuo Kuniyoshi, Daniel Cremers |
ICCV | 1 |
| 2013 | A Scale Independent Selection Process for 3D Object Recognition in Cluttered Scenes
Emanuele Rodolà, Andrea Albarelli, Filippo Bergamasco, Andrea Torsello |
Int. J. Comput. Vis. | 1 |
| 2013 | Stable and fast techniques for unambiguous compound phase coding
Andrea Torsello, Andrea Albarelli, Emanuele Rodolà |
Image Vis. Comput. | 3 |
| 2012 | A game-theoretic approach to deformable shape matchingabstractWe consider the problem of minimum distortion intrinsic correspondence between deformable shapes, many useful formulations of which give rise to the NP-hard quadratic assignment problem (QAP). Previous attempts to use the spectral relaxation have had limited success due to the lack of sparsity of the obtained “fuzzy” solution. In this paper, we adopt the recently introduced alternative L1relaxation of the QAP based on the principles of game theory. We relate it to the Gromov and Lipschitz metrics between metric spaces and demonstrate on state-of-the-art benchmarks that the proposed approach is capable of finding very accurate sparse correspondences between deformable shapes. Emanuele Rodolà, Alexander M. Bronstein, Andrea Albarelli, Filippo Bergamasco, Andrea Torsello |
CVPR | 1 |
| 2012 | Imposing Semi-Local Geometric Constraints for Accurate Correspondences Selection in Structure from Motion: A Game-Theoretic Perspective
Andrea Albarelli, Emanuele Rodolà, Andrea Torsello |
Int. J. Comput. Vis. | 2 |
| 2011 | RUNE-Tag: A high accuracy fiducial marker with strong occlusion resilienceabstractOver the last decades fiducial markers have provided widely adopted tools to add reliable model-based features into an otherwise general scene. Given their central role in many computer vision tasks, countless different solutions have been proposed in the literature. Some designs are focused on the accuracy of the recovered camera pose with respect to the tag; some other concentrate on reaching high detection speed or on recognizing a large number of distinct markers in the scene. In such a crowded area both the researcher and the practitioner are licensed to wonder if there is any need to introduce yet another approach. Nevertheless, with this paper, we would like to present a general purpose fiducial marker system that can be deemed to add some valuable features to the pack. Specifically, by exploiting the projective properties of a circular set of sizeable dots, we propose a detection algorithm that is highly accurate. Further, applying a dot pattern scheme derived from error-correcting codes, allows for robustness with respect to very large occlusions. In addition, the design of the marker itself is flexible enough to accommodate different requirements in terms of pose accuracy and number of patterns. The overall performance of the marker system is evaluated in an extensive experimental section, where a comparison with a well-known baseline technique is presented. Filippo Bergamasco, Andrea Albarelli, Emanuele Rodolà, Andrea Torsello |
CVPR | 3 |
| 2011 | Multiview registration via graph diffusion of dual quaternionsabstractSurface registration is a fundamental step in the reconstruction of three-dimensional objects. While there are several fast and reliable methods to align two surfaces, the tools available to align multiple surfaces are relatively limited. In this paper we propose a novel multiview registration algorithm that projects several pairwise alignments onto a common reference frame. The projection is performed by representing the motions as dual quaternions, an algebraic structure that is related to the group of 3D rigid transformations, and by performing a diffusion along the graph of adjacent (i.e., pairwise alignable) views. The approach allows for a completely generic topology with which the pair-wise motions are diffused. An extensive set of experiments shows that the proposed approach is both orders of magnitude faster than the state of the art, and more robust to extreme positional noise and outliers. The dramatic speedup of the approach allows it to be alternated with pairwise alignment resulting in a smoother energy profile, reducing the risk of getting stuck at local minima. Andrea Torsello, Emanuele Rodolà, Andrea Albarelli |
CVPR | 2 |
| 2010 | Robust Camera Calibration using Inaccurate TargetsabstractAccurate intrinsic camera calibration is essential to any computer vision task that \ninvolves image based measurements. Given its crucial role with respect to precision, a \nlarge number of approaches have been proposed over the last decades. Despite this rich \nliterature, steady advancements in imaging hardware regularly push forward the need \nfor even more accurate techniques. Some authors suggest generalizations of the camera \nmodel itself, others propose novel designs for calibration targets or different optimization \nschemas. In this paper we take a completely different route by directly addressing one \nof the most overlooked problems in practical calibration scenarios. Specifically, we drop \nthe assumption that the target is known with enough precision and we adjust it in an \niterative way as part of the whole process. This is in fact the case with the typical target \nused in most of the calibration literature, which is usually printed on paper and stitched \non a flat surface. In the experimental section we show that even with such a cheaply \ncrafted target it is possible to obtain a very accurate camera calibration that outperforms \nthose obtained with well-known standard techniques. Andrea Albarelli, Emanuele Rodolà, Andrea Torsello |
BMVC | 2 |
| 2010 | A game-theoretic approach to fine surface registration without initial motion estimationabstractSurface registration is a fundamental step in the reconstruction of three-dimensional objects. This is typically a two step process where an initial coarse motion estimation is followed by a refinement. Most coarse registration algorithms exploit some local point descriptor that is intrinsic to the shape and does not depend on the relative position of the surfaces. By contrast, refinement techniques iteratively minimize a distance function measured between pairs of selected neighboring points and are thus strongly dependent on initial alignment. In this paper we propose a novel technique that allows to obtain a fine surface registration in a single step, without the need of an initial motion estimation. The main idea of our approach is to cast the selection of correspondences between points on the surfaces in a game-theoretic framework, where a natural selection process allows mating points that satisfy a mutual rigidity constraint to thrive, eliminating all the other correspondences. This process yields a very robust inlier selection scheme that does not depend on any particular technique for selecting the initial strategies as it relies only on the global geometric compatibility between correspondences. The practical effectiveness of the proposed approach is confirmed by an extensive set of experiments and comparisons with state-of-the-art techniques. Andrea Albarelli, Emanuele Rodolà, Andrea Torsello |
CVPR | 2 |
| 2010 | Loosely Distinctive Features for Robust Surface Alignment
Andrea Albarelli, Emanuele Rodolà, Andrea Torsello |
ECCV (5) | 2 |
| 2010 | Robust Figure Extraction on Textured Background: A Game-Theoretic ApproachabstractFeature-based image matching relies on the assumption that the features contained in the model are distinctive enough. When both model and data present a sizeable amount of clutter, the signal-to-noise ratio falls and the detection becomes more challenging. If such clutter exhibits a coherent structure, as it is the case for textured background, matching becomes even harder. In fact, the large amount of repeatable features extracted from the texture dims the strength of the relatively few interesting points of the object itself. In this paper we introduce a game-theoretic approach that allows to distinguish foreground features from background ones. In addition the same technique can be used to deal with the object matching itself. The whole procedure is validated by applying it to a practical scenario and by comparing it with a standard point-pattern matching technique. Andrea Albarelli, Emanuele Rodolà, Alberto Cavallarin, Andrea Torsello |
ICPR | 2 |
| 2010 | A Game-Theoretic Approach to Robust Selection of Multi-view Point CorrespondenceabstractIn this paper we introduce a robust matching technique that allows very accurate selection of corresponding feature points from multiple views. Robustness is achieved by enforcing global geometric consistency at an early stage of the matching process, without the need of subsequent verification through reprojection. The global consistency is reduced to a pairwise compatibility making use of the size and orientation information provided by common feature descriptors, thus projecting what is a high-order compatibility problem into a pairwise setting. Then a game-theoretic approach is used to select a maximally consistent set of candidate matches, where highly compatible matches are enforced while incompatible correspondences are driven to extinction. Emanuele Rodolà, Andrea Albarelli, Andrea Torsello |
ICPR | 1 |