VLDB 2026 Research / reviewers in the wild / expert
Roland Kwitt
dblp:60/4140
· DBLP profile ↗
55ranked-venue papers
18as first author
13since 2021 · last 2025
0000-0001-9947-4465ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 15 first-author · 5 since 2021Artificial intelligence and machine learning · 30 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 6 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CARL: A Framework for Equivariant Image RegistrationabstractImage registration estimates spatial correspondences between image pairs. These estimates are typically obtained via numerical optimization or regression by a deep network. A desirable property is that a correspondence estimate (e.g., the true oracle correspondence) for an image pair is maintained under deformations of the input images. Formally, the estimator should be equivariant to a desired class of image transformations. In this work, we present careful analyses of equivariance properties in the context of multi-step deep registration networks. Based on these analyses we 1) introduce the notions of [U, U] equivariance (network equivariance to the same deformations of the input images) and [W, U] equivariance (where input images can undergo different deformations); we 2) show that in a suitable multistep registration setup it is sufficient for overall [W, U] equivariance if the first step has [W, U] equivariance and all others have [U, U] equivariance; we 3) show that common displacement-predicting networks only exhibit [U, U] equivariance to translations instead of the more powerful [W, U ] equivariance; and we 4) show how to achieve multistep [W, U] equivariance via a coordinate-attention mechanism combined with displacement-predicting networks. Our approach obtains excellent practical performance for 3D abdomen, lung, and brain medical image registration. We match or outperform state-of-the-art (SOTA) registration approaches on all the datasets with a particularly strong performance for the challenging abdomen registration. Thomas Hastings Greer, Lin Tian 0001, François-Xavier Vialard, Roland Kwitt, Raúl San José Estépar, Marc Niethammer |
CVPR | 4 |
| 2025 | The Flood Complex: Large-Scale Persistent Homology on Millions of PointsabstractWe consider the problem of computing persistent homology (PH) for large-scale Euclidean point cloud data, aimed at downstream machine learning tasks, where the exponential growth of the most widely-used Vietoris-Rips complex imposes serious computational limitations. Although more scalable alternatives such as the Alpha complex or sparse Rips approximations exist, they often still result in a prohibitively large number of simplices. This poses challenges in the complex construction and in the subsequent PH computation, prohibiting their use on large-scale point clouds. To mitigate these issues, we introduce the Flood complex, inspired by the advantages of the Alpha and Witness complex constructions. Informally, at a given filtration value $r\geq 0$, the Flood complex contains all simplices from a Delaunay triangulation of a small subset of a point cloud $X$ that are fully covered by the union of balls of radius $r$ emanating from $X$, a process we call flooding. Our construction allows for efficient PH computation, possesses several desirable theoretical properties, and is amenable to GPU parallelization. Scaling experiments on 3D point cloud data show that we can compute PH of up to dimension 2 on several millions of points. Importantly, when evaluating object classification performance on real-world and synthetic data, we provide evidence that this scaling capability is needed, especially if objects are geometrically or topologically complex, yielding performance superior to other PH-based methods and neural networks for point cloud data.
Source code and datasets are available on GitHub: https://github.com/plus-rkwitt/flooder Florian Graf, Paolo Pellizzoni, Martin Uray, Stefan Huber 0001, Roland Kwitt |
NeurIPS | 5 |
| 2025 | Ph.D. Forum: Multi-Agent Reinforcement Learning in Wireless Network CommunicationabstractWireless network communication develops fast. Hence, data traffic is rising and more devices communicate like mobile users, autonomous cars and machines in factories [2]. These devices affect their communication by trying to access the same resources if they all use the full cellular spectrum. Thus, avoiding the overlap of used frequency bands and controlling the wireless network communication becomes more complicated in the future [1]. Sabrina Pochaba, Peter Dorfinger, Matthias Herlich, Roland Kwitt, Simon Hirlaender |
WoWMoM | 4 |
| 2024 | Position: Topological Deep Learning is the New Frontier for Relational LearningabstractTopological deep learning (TDL) is a rapidly evolving field that uses topological features to understand and design deep learning models. This paper posits that TDL is the new frontier for relational learning. TDL may complement graph representation learning and geometric deep learning by incorporating topological concepts, and can thus provide a natural choice for various machine learning settings. To this end, this paper discusses open problems in TDL, ranging from practical benefits to theoretical foundations. For each problem, it outlines potential solutions and future research opportunities. At the same time, this paper serves as an invitation to the scientific community to actively participate in TDL research to unlock the potential of this emerging field. Theodore Papamarkou, Tolga Birdal, Michael M. Bronstein, Gunnar E. Carlsson, Justin Curry, Yue Gao 0002, Mustafa Hajij, Roland Kwitt, Pietro Liò, Paolo Di Lorenzo, Vasileios Maroulas, Nina Miolane, Farzana Nasrin, Karthikeyan Natesan Ramamurthy, Bastian Rieck, Simone Scardapane, Michael T. Schaub, Petar Velickovic, Bei Wang 0001, Yusu Wang 0001, Guo-Wei Wei 0001, Ghada Zamzmi |
ICML | 8 |
| 2024 | uniGradICON: A Foundation Model for Medical Image Registration
Lin Tian 0001, Thomas Hastings Greer, Roland Kwitt, François-Xavier Vialard, Raúl San José Estépar, Sylvain Bouix, Richard J. Rushmore, Marc Niethammer |
MICCAI (2) | 3 |
| 2024 | Neural Persistence DynamicsabstractWe consider the problem of learning the dynamics in the topology of time-evolving point clouds, the prevalent spatiotemporal model for systems exhibiting collective behavior, such as swarms of insects and birds or particles in physics. In such systems, patterns emerge from (local) interactions among self-propelled entities. While several well-understood governing equations for motion and interaction exist, they are notoriously difficult to fit to data, as most prior work requires knowledge about individual motion trajectories, i.e., a requirement that is challenging to satisfy with an increasing number of entities. To evade such confounding factors, we investigate collective behavior from a _topological perspective_, but instead of summarizing entire observation sequences (as done previously), we propose learning a latent dynamical model from topological features _per time point_. The latter is then used to formulate a downstream regression task to predict the parametrization of some a priori specified governing equation. We implement this idea based on a latent ODE learned from vectorized (static) persistence diagrams and show that a combination of recent stability results for persistent homology justifies this modeling choice. Various (ablation) experiments not only demonstrate the relevance of each model component but provide compelling empirical evidence that our proposed model -- _Neural Persistence Dynamics_ -- substantially outperforms the state-of-the-art across a diverse set of parameter regression tasks. Sebastian Zeng, Florian Graf, Martin Uray, Stefan Huber 0001, Roland Kwitt |
NeurIPS | 5 |
| 2023 | GradICON: Approximate Diffeomorphisms via Gradient Inverse ConsistencyabstractWe present an approach to learning regular spatial transformations between image pairs in the context of medical image registration. Contrary to optimization-based registration techniques and many modern learning-based methods, we do not directly penalize transformation irregularities but instead promote transformation regularity via an inverse consistency penalty. We use a neural network to predict a map between a source and a target image as well as the map when swapping the source and target images. Different from existing approaches, we compose these two resulting maps and regularize deviations of the Jacobian of this composition from the identity matrix. This regularizer - GradICON - results in much better convergence when training registration models compared to promoting inverse consistency of the composition of maps directly while retaining the desirable implicit regularization effects of the latter. We achieve state-of-the-art registration performance on a variety of real-world medical image datasets using a single set of hyperparameters and a single non-dataset-specific training protocol. Code is available at https://github.com/uncbiag/ICON. Lin Tian 0001, Thomas Hastings Greer, François-Xavier Vialard, Roland Kwitt, Raúl San José Estépar, Richard J. Rushmore, Nikos Makris, Sylvain Bouix, Marc Niethammer |
CVPR | 4 |
| 2023 | Inverse Consistency by Construction for Multistep Deep Registration
Thomas Hastings Greer, Lin Tian 0001, François-Xavier Vialard, Roland Kwitt, Sylvain Bouix, Raúl San José Estépar, Richard J. Rushmore, Marc Niethammer |
MICCAI (10) | 4 |
| 2023 | Latent SDEs on Homogeneous SpacesabstractWe consider the problem of variational Bayesian inference in a latent variable model where a (possibly complex) observed stochastic process is governed by the unobserved solution of a latent stochastic differential equation (SDE). Motivated by the challenges that arise when trying to learn a latent SDE in $\mathbb{R}^n$ from large-scale data, such as efficient gradient computation, we take a step back and study a specific subclass instead. In our case, the SDE evolves inside a homogeneous latent space and is induced by stochastic dynamics of the corresponding (matrix) Lie group. In the context of learning problems, SDEs on the $n$-dimensional unit sphere are arguably the most relevant incarnation of this setup. For variational inference, the sphere not only facilitates using a uniform prior on the initial state of the SDE, but we also obtain a particularly simple and intuitive expression for the KL divergence between the approximate posterior and prior process in the evidence lower bound. We provide empirical evidence that a latent SDE of the proposed type can be learned efficiently by means of an existing one-step geometric Euler-Maruyama scheme. Despite restricting ourselves to a less diverse class of SDEs, we achieve competitive or even state-of-the-art performance on a collection of time series interpolation and classification benchmarks. Sebastian Zeng, Florian Graf, Roland Kwitt |
NeurIPS | 3 |
| 2022 | On Measuring Excess Capacity in Neural NetworksabstractWe study the excess capacity of deep networks in the context of supervised classification. That is, given a capacity measure of the underlying hypothesis class - in our case, empirical Rademacher complexity - to what extent can we (a priori) constrain this class while retaining an empirical error on a par with the unconstrained regime? To assess excess capacity in modern architectures (such as residual networks), we extend and unify prior Rademacher complexity bounds to accommodate function composition and addition, as well as the structure of convolutions. The capacity-driving terms in our bounds are the Lipschitz constants of the layers and a (2,1) group norm distance to the initializations of the convolution weights. Experiments on benchmark datasets of varying task difficulty indicate that (1) there is a substantial amount of excess capacity per task, and (2) capacity can be kept at a surprisingly similar level across tasks. Overall, this suggests a notion of compressibility with respect to weight norms, complementary to classic compression via weight pruning. Source code is available at https://github.com/rkwitt/excess_capacity. Florian Graf, Sebastian Zeng, Bastian Rieck, Marc Niethammer, Roland Kwitt |
NeurIPS | 5 |
| 2021 | ICON: Learning Regular Maps Through Inverse ConsistencyabstractLearning maps between data samples is fundamental. Applications range from representation learning, image translation and generative modeling, to the estimation of spatial deformations. Such maps relate feature vectors, or map between feature spaces. Well-behaved maps should be regular, which can be imposed explicitly or may emanate from the data itself. We explore what induces regularity for spatial transformations, e.g., when computing image registrations. Classical optimization-based models compute maps between pairs of samples and rely on an appropriate regularizer for well-posedness. Recent deep learning approaches have attempted to avoid using such regularizers altogether by relying on the sample population instead. We explore if it is possible to obtain spatial regularity using an inverse consistency loss only and elucidate what explains map regularity in such a context. We find that deep networks combined with an inverse consistency loss and randomized off-grid interpolation yield well behaved, approximately diffeomorphic, spatial transformations. Despite the simplicity of this approach, our experiments present compelling evidence, on both synthetic and real data, that regular maps can be obtained without carefully tuned explicit regularizers, while achieving competitive registration performance. Thomas Hastings Greer, Roland Kwitt, François-Xavier Vialard, Marc Niethammer |
ICCV | 2 |
| 2021 | Dissecting Supervised Constrastive Learning
Florian Graf, Christoph D. Hofer, Marc Niethammer, Roland Kwitt |
ICML | 4 |
| 2021 | Topological Attention for Time Series ForecastingabstractThe problem of (point) forecasting univariate time series is considered. Most approaches, ranging from traditional statistical methods to recent learning-based techniques with neural networks, directly operate on raw time series observations. As an extension, we study whether local topological properties, as captured via persistent homology, can serve as a reliable signal that provides complementary information for learning to forecast. To this end, we propose topological attention, which allows attending to local topological features within a time horizon of historical data. Our approach easily integrates into existing end-to-end trainable forecasting models, such as N-BEATS, and, in combination with the latter exhibits state-of-the-art performance on the large-scale M4 benchmark dataset of 100,000 diverse time series from different domains. Ablation experiments, as well as a comparison to recent techniques in a setting where only a single time series is available for training, corroborate the beneficial nature of including local topological information through an attention mechanism. Sebastian Zeng, Florian Graf, Christoph D. Hofer, Roland Kwitt |
NeurIPS | 4 |
| 2020 | Topologically Densified DistributionsabstractWe study regularization in the context of small sample-size learning with over-parametrized neural networks. Specifically, we shift focus from architectural properties, such as norms on the network weights, to properties of the internal representations before a linear classifier. Specifically, we impose a topological constraint on samples drawn from the probability measure induced in that space. This provably leads to mass concentration effects around the representations of training instances, i.e., a property beneficial for generalization. By leveraging previous work to impose topological constrains in a neural network setting, we provide empirical evidence (across various vision benchmarks) to support our claim for better generalization. Christoph D. Hofer, Florian Graf, Marc Niethammer, Roland Kwitt |
ICML | 4 |
| 2020 | Graph Filtration LearningabstractWe propose an approach to learning with graph-structured data in the problem domain of graph classification. In particular, we present a novel type of readout operation to aggregate node features into a graph-level representation. To this end, we leverage persistent homology computed via a real-valued, learnable, filter function. We establish the theoretical foundation for differentiating through the persistent homology computation. Empirically, we show that this type of readout operation compares favorably to previous techniques, especially when the graph connectivity structure is informative for the learning problem. Christoph D. Hofer, Florian Graf, Bastian Rieck, Marc Niethammer, Roland Kwitt |
ICML | 5 |
| 2020 | A shooting formulation of deep learningabstractA residual network may be regarded as a discretization of an ordinary differential equation (ODE) which, in the limit of time discretization, defines a continuous-depth network. Although important steps have been taken to realize the advantages of such continuous formulations, most current techniques assume identical layers. Indeed, existing works throw into relief the myriad difficulties of learning an infinite-dimensional parameter in a continuous-depth neural network. To this end, we introduce a shooting formulation which shifts the perspective from parameterizing a network layer-by-layer to parameterizing over optimal networks described only by a set of initial conditions. For scalability, we propose a novel particle-ensemble parameterization which fully specifies the optimal weight trajectory of the continuous-depth neural network. Our experiments show that our particle-ensemble shooting formulation can achieve competitive performance. Finally, though the current work is inspired by continuous-depth neural networks, the particle-ensemble shooting formulation also applies to discrete-time networks and may lead to a new fertile area of research in deep learning parameterization. François-Xavier Vialard, Roland Kwitt, Susan Wei, Marc Niethammer |
NeurIPS | 2 |
| 2019 | Metric Learning for Image RegistrationabstractImage registration is a key technique in medical image analysis to estimate deformations between image pairs. A good deformation model is important for high-quality estimates. However, most existing approaches use ad-hoc deformation models chosen for mathematical convenience rather than to capture observed data variation. Recent deep learning approaches learn deformation models directly from data. However, they provide limited control over the spatial regularity of transformations. Instead of learning the entire registration approach, we learn a spatially-adaptive regularizer within a registration model. This allows controlling the desired level of regularity and preserving structural properties of a registration model. For example, diffeomorphic transformations can be attained. Our approach is a radical departure from existing deep learning approaches to image registration by embedding a deep learning model in an optimization-based registration algorithm to parameterize and data-adapt the registration model itself. Source code is publicly-available at https://github.com/uncbiag/registration. Marc Niethammer, Roland Kwitt, François-Xavier Vialard |
CVPR | 2 |
| 2019 | Connectivity-Optimized Representation Learning via Persistent HomologyabstractWe study the problem of learning representations with controllable connectivity properties. This is beneficial in situations when the imposed structure can be leveraged upstream. In particular, we control the connectivity of an autoencoder’s latent space via a novel type of loss, operating on information from persistent homology. Under mild conditions, this loss is differentiable and we present a theoretical analysis of the properties induced by the loss. We choose one-class learning as our upstream task and demonstrate that the imposed structure enables informed parameter selection for modeling the in-class distribution via kernel density estimators. Evaluated on computer vision data, these one-class models exhibit competitive performance and, in a low sample size regime, outperform other methods by a large margin. Notably, our results indicate that a single autoencoder, trained on auxiliary (unlabeled) data, yields a mapping into latent space that can be reused across datasets for one-class learning. Christoph D. Hofer, Roland Kwitt, Marc Niethammer, Mandar Dixit |
ICML | 2 |
| 2019 | Learning Representations of Persistence BarcodesabstractWe consider the problem of supervised learning with summary representations of topological features in data. In particular, we focus on persistent homology, the prevalent tool used in topological data analysis. As the summary representations, referred to as barcodes or persistence diagrams, come in the unusual format of multi sets, equipped with computationally expensive metrics, they can not readily be processed with conventional learning techniques. While different approaches to address this problem have been proposed, either in the context of kernel-based learning, or via carefully designed vectorization techniques, it remains an open problem how to leverage advances in representation learning via deep neural networks. Appropriately handling topological summaries as input to neural networks would address the disadvantage of previous strategies which handle this type of data in a task-agnostic manner. In particular, we propose an approach that is designed to learn a task-specific representation of barcodes. In other words, we aim to learn a representation that adapts to the learning problem while, at the same time, preserving theoretical properties (such as stability). This is done by projecting barcodes into a finite dimensional vector space using a collection of parametrized functionals, so called structure elements, for which we provide a generic construction scheme. A theoretical analysis of this approach reveals sufficient conditions to preserve stability, and also shows that different choices of structure elements lead to great differences with respect to their suitability for numerical optimization. When implemented as a neural network input layer, our approach demonstrates compelling performance on various types of problems, including graph classification and eigenvalue prediction, the classification of 2D/3D object shapes and recognizing activities from EEG signals. Christoph D. Hofer, Roland Kwitt, Marc Niethammer |
J. Mach. Learn. Res. | 2 |
| 2019 | Fast predictive simple geodesic regression
Zhipeng Ding, Greg M. Fleishman, Xiao Yang 0002, Paul M. Thompson, Roland Kwitt, Marc Niethammer |
Medical Image Anal. | 5 |
| 2018 | Feature Space Transfer for Data AugmentationabstractThe problem of data augmentation in feature space is considered. A new architecture, denoted the FeATure TransfEr Network (FATTEN), is proposed for the modeling of feature trajectories induced by variations of object pose. This architecture exploits a parametrization of the pose manifold in terms of pose and appearance. This leads to a deep encoder/decoder network architecture, where the encoder factors into an appearance and a pose predictor. Unlike previous attempts at trajectory transfer, FATTEN can be efficiently trained end-to-end, with no need to train separate feature transfer functions. This is realized by supplying the decoder with information about a target pose and the use of a multi-task loss that penalizes category- and pose-mismatches. In result, FATTEN discourages discontinuous or non-smooth trajectories that fail to capture the structure of the pose manifold, and generalizes well on object recognition tasks involving large pose variation. Experimental results on the artificial ModelNet database show that it can successfully learn to map source features to target features of a desired pose, while preserving class identity. Most notably, by using feature space transfer for data augmentation (w.r.t. pose and depth) on SUN-RGBD objects, we demonstrate considerable performance improvements on one/few-shot object recognition in a transfer learning setup, compared to current state-of-the-art methods. Bo Liu 0043, Xudong Wang 0007, Mandar Dixit, Roland Kwitt, Nuno Vasconcelos |
CVPR | 4 |
| 2017 | Regression Uncertainty on the GrassmannianabstractTrends in longitudinal or cross-sectional studies over time are often captured through regression models. In their simplest manifestation, these regression models are formulated in $R^n$. However, in the context of imaging studies, the objects of interest which are to be regressed are frequently best modeled as elements of a Riemannian manifold. Regression on such spaces can be accomplished through geodesic regression. This paper develops an approach to compute confidence intervals for geodesic regression models. The approach is general, but illustrated and specifically developed for the Grassmann manifold, which allows us, e.g., to regress shapes or linear dynamical systems. Extensions to other manifolds can be obtained in a similar manner. We demonstrate our approach for regression with 2D/3D shapes using synthetic and real data. Yi Hong 0006, Xiao Yang 0002, Roland Kwitt, Martin Styner, Marc Niethammer |
AISTATS | 3 |
| 2017 | AGA: Attribute-Guided AugmentationabstractWe consider the problem of data augmentation, i.e., generating artificial samples to extend a given corpus of training data. Specifically, we propose attributed-guided augmentation (AGA) which learns a mapping that allows to synthesize data such that an attribute of a synthesized sample is at a desired value or strength. This is particularly interesting in situations where little data with no attribute annotation is available for learning, but we have access to a large external corpus of heavily annotated samples. While prior works primarily augment in the space of images, we propose to perform augmentation in feature space instead. We implement our approach as a deep encoder-decoder architecture that learns the synthesis function in an end-to-end manner. We demonstrate the utility of our approach on the problems of (1) one-shot object recognition in a transfer-learning setting where we have no prior knowledge of the new classes, as well as (2) object-based one-shot scene recognition. As external data, we leverage 3D depth and pose information from the SUN RGB-D dataset. Our experiments show that attribute-guided augmentation of high-level CNN features considerably improves one-shot recognition performance on both problems. Mandar Dixit, Roland Kwitt, Marc Niethammer, Nuno Vasconcelos |
CVPR | 2 |
| 2017 | Deep Learning with Topological SignaturesabstractInferring topological and geometrical information from data can offer an alternative perspective in machine learning problems. Methods from topological data analysis, e.g., persistent homology, enable us to obtain such information, typically in the form of summary representations of topological features. However, such topological signatures often come with an unusual structure (e.g., multisets of intervals) that is highly impractical for most machine learning techniques. While many strategies have been proposed to map these topological signatures into machine learning compatible representations, they suffer from being agnostic to the target learning task. In contrast, we propose a technique that enables us to input topological signatures to deep neural networks and learn a task-optimal representation during training. Our approach is realized as a novel input layer with favorable theoretical properties. Classification experiments on 2D object shapes and social network graphs demonstrate the versatility of the approach and, in case of the latter, we even outperform the state-of-the-art by a large margin. Christoph D. Hofer, Roland Kwitt, Marc Niethammer, Andreas Uhl |
NIPS | 2 |
| 2016 | One-Shot Learning of Scene Locations via Feature Trajectory TransferabstractThe appearance of (outdoor) scenes changes considerably with the strength of certain transient attributes, such as "rainy", "dark" or "sunny". Obviously, this also affects the representation of an image in feature space, e.g., as activations at a certain CNN layer, and consequently impacts scene recognition performance. In this work, we investigate the variability in these transient attributes as a rich source of information for studying how image representations change as a function of attribute strength. In particular, we leverage a recently introduced dataset with fine-grain annotations to estimate feature trajectories for a collection of transient attributes and then show how these trajectories can be transferred to new image representations. This enables us to synthesize new data along the transferred trajectories with respect to the dimensions of the space spanned by the transient attributes. Applicability of this concept is demonstrated on the problem of oneshot recognition of scene locations. We show that data synthesized via feature trajectory transfer considerably boosts recognition performance, (1) with respect to baselines and (2) in combination with state-of-the-art approaches in oneshot learning. Roland Kwitt, Sebastian Hegenbart, Marc Niethammer |
CVPR | 1 |
| 2016 | Parametric Regression on the GrassmannianabstractWe address the problem of fitting parametric curves on the Grassmann manifold for the purpose of intrinsic parametric regression. We start from the energy minimization formulation of linear least-squares in Euclidean space and generalize this concept to general nonflat Riemannian manifolds, following an optimal-control point of view. We then specialize this idea to the Grassmann manifold and demonstrate that it yields a simple, extensible and easy-to-implement solution to the parametric regression problem. In fact, it allows us to extend the basic geodesic model to (1) a "time-warped" variant and (2) cubic splines. We demonstrate the utility of the proposed solution on different vision problems, such as shape regression as a function of age, traffic-speed estimation and crowd-counting from surveillance video clips. Most notably, these problems can be conveniently solved within the same framework without any specifically-tailored steps along the processing pipeline. Yi Hong 0006, Roland Kwitt, Nikhil Singh 0002, Nuno Vasconcelos, Marc Niethammer |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | A stable multi-scale kernel for topological machine learningabstractTopological data analysis offers a rich source of valuable information to study vision problems. Yet, so far we lack a theoretically sound connection to popular kernel-based learning techniques, such as kernel SVMs or kernel PCA. In this work, we establish such a connection by designing a multi-scale kernel for persistence diagrams, a stable summary representation of topological features in data. We show that this kernel is positive definite and prove its stability with respect to the 1-Wasserstein distance. Experiments on two benchmark datasets for 3D shape classification/retrieval and texture recognition show considerable performance gains of the proposed method compared to an alternative approach that is based on the recently introduced persistence landscapes. Jan Reininghaus, Stefan Huber 0001, Ulrich Bauer, Roland Kwitt |
CVPR | 4 |
| 2015 | Model Criticism for Regression on the Grassmannian
Yi Hong 0006, Roland Kwitt, Marc Niethammer |
MICCAI (3) | 2 |
| 2015 | Statistical Topological Data Analysis - A Kernel PerspectiveabstractWe consider the problem of statistical computations with persistence diagrams, a summary representation of topological features in data. These diagrams encode persistent homology, a widely used invariant in topological data analysis. While several avenues towards a statistical treatment of the diagrams have been explored recently, we follow an alternative route that is motivated by the success of methods based on the embedding of probability measures into reproducing kernel Hilbert spaces. In fact, a positive definite kernel on persistence diagrams has recently been proposed, connecting persistent homology to popular kernel-based learning techniques such as support vector machines. However, important properties of that kernel which would enable a principled use in the context of probability measure embeddings remain to be explored. Our contribution is to close this gap by proving universality of a variant of the original kernel, and to demonstrate its effective use in two-sample hypothesis testing on synthetic as well as real-world data. Roland Kwitt, Stefan Huber 0001, Marc Niethammer, Weili Lin, Ulrich Bauer |
NIPS | 1 |
| 2014 | Geodesic Regression on the Grassmannian
Yi Hong 0006, Roland Kwitt, Nikhil Singh 0002, Bradley C. Davis, Nuno Vasconcelos, Marc Niethammer |
ECCV (2) | 2 |
| 2014 | Time-Warped Geodesic Regression
Yi Hong 0006, Nikhil Singh 0002, Roland Kwitt, Marc Niethammer |
MICCAI (2) | 3 |
| 2014 | Do We Need Annotation Experts? A Case Study in Celiac Disease Classification
Roland Kwitt, Sebastian Hegenbart, Nikhil Rasiwasia, Andreas Vécsei, Andreas Uhl |
MICCAI (2) | 1 |
| 2014 | Low-Rank to the Rescue - Atlas-Based Analyses in the Presence of Pathologies
Marc Niethammer, Roland Kwitt, Matt McCormick 0001, Stephen R. Aylward |
MICCAI (3) | 3 |
| 2014 | Statistical atlas construction via weighted functional boxplots
Yi Hong 0006, Bradley C. Davis, J. S. Marron, Roland Kwitt, Nikhil Singh 0002, Julia S. Kimbell, Elizabeth Pitkin, Richard Superfine, Stephanie Davis, Carlton J. Zdanski, Marc Niethammer |
Medical Image Anal. | 4 |
| 2013 | Weighted Functional Boxplot with Application to Statistical Atlas Construction
Yi Hong 0006, Bradley C. Davis, J. S. Marron, Roland Kwitt, Marc Niethammer |
MICCAI (3) | 4 |
| 2013 | Studying Cerebral Vasculature Using Structure Proximity and Graph Kernels
Roland Kwitt, Danielle F. Pace, Marc Niethammer, Stephen R. Aylward |
MICCAI (2) | 1 |
| 2013 | Localizing target structures in ultrasound video - A phantom study
Roland Kwitt, Nuno Vasconcelos, Sharif Razzaque, Stephen R. Aylward |
Medical Image Anal. | 1 |
| 2012 | Scene Recognition on the Semantic Manifold
Roland Kwitt, Nuno Vasconcelos, Nikhil Rasiwasia |
ECCV (4) | 1 |
| 2012 | Recognition in Ultrasound Videos: Where Am I?
Roland Kwitt, Nuno Vasconcelos, Sharif Razzaque, Stephen R. Aylward |
MICCAI (3) | 1 |
| 2012 | Endoscopic image analysis in semantic space
Roland Kwitt, Nuno Vasconcelos, Nikhil Rasiwasia, Andreas Uhl, Bradley C. Davis, Michael Häfner, Friedrich Wrba |
Medical Image Anal. | 1 |
| 2011 | Testing a multivariate model for wavelet coefficientsabstractIn this paper, we introduce a Goodness-of-Fit test for the Multivariate Exponential Power (MEP) distribution, a multivariate extension of the Generalized Gaussian, which has recently gained considerable interest as a model for wavelet coefficients in the context of color image retrieval and spread-spectrum watermarking. We present a size and power study of this test and show Goodness-of-Fit results for wavelet coefficients of natural and texture images from various popular databases. Roland Kwitt, Peter Meerwald-Stadler, Andreas Uhl, Geert Verdoolaege |
ICIP | 1 |
| 2011 | Infrared camera calibration for dense depth map constructionabstractIn this paper, we introduce a novel and cost effective approach to calibrate the geometric properties of a far-infrared (IR) sensor. We further demonstrate that fully automatic sensor-to-sensor calibration is feasible in a setup involving a laser range scanner, IR cameras as well as conventional cameras. The calibration result then serves as a basis for upsampling range measurements to the resolution of the IR or visible-light camera images. Since our approach allows to rely on IR information instead of visible-light information for upsampling, bad light conditions or even no visible light at all are no limitation. From a practical point of view, we only require one calibration board of relatively small size which facilitates application in outdoor environments and further allows seamless integration of the IR camera in an existing multi-sensor platform. Our experimental results demonstrate that IR images are particularly useful to obtain reasonable depth information for living objects, when visible-light cameras are either blind or require impractical exposure times. In fact, our approach provides a convenient solution to IR camera calibration and integration, an issue which is particularly important in scenarios where sensors are not permanently mounted on vehicles and consequently require on-site adjustment and calibration. Michael Gschwandtner, Roland Kwitt, Andreas Uhl, Wolfgang Pree |
Intelligent Vehicles Symposium | 2 |
| 2011 | Learning Pit Pattern Concepts for Gastroenterological Training
Roland Kwitt, Nikhil Rasiwasia, Nuno Vasconcelos, Andreas Uhl, Michael Häfner, Friedrich Wrba |
MICCAI (3) | 1 |
| 2011 | Lightweight Detection of Additive Watermarking in the DWT-DomainabstractThis article aims at lightweight, blind detection of additive spread-spectrum watermarks in the DWT domain. We focus on two host signal noise models and two types of hypothesis tests for watermark detection. As a crucial point of our work we take a closer look at the computational requirements of watermark detectors. This involves the computation of the detection response, parameter estimation and threshold selection. We show that by switching to approximate host signal parameter estimates or even fixed parameter settings we achieve a remarkable improvement in runtime performance without sacrificing detection performance. Our experimental results on a large number of images confirm the assumption that there is not necessarily a tradeoff between computation time and detection performance. Roland Kwitt, Peter Meerwald-Stadler, Andreas Uhl |
IEEE Trans. Image Process. | 1 |
| 2011 | Efficient Texture Image Retrieval Using Copulas in a Bayesian FrameworkabstractIn this paper, we investigate a novel joint statistical model for subband coefficient magnitudes of the dual-tree complex wavelet transform, which is then coupled to a Bayesian framework for content-based image retrieval. The joint model allows to capture the association among transform coefficients of the same decomposition scale and different color channels. It further facilitates to incorporate recent research work on modeling marginal coefficient distributions. We demonstrate the applicability of the novel model in the context of color texture retrieval on four texture image databases and compare retrieval performance to a collection of state-of-the-art approaches in the field. Our experiments further include a thorough computational analysis of the main building blocks, runtime measurements, and an analysis of storage requirements. Eventually, we identify a model configuration with low storage requirements, competitive retrieval accuracy, and a runtime behavior, which enables the deployment even on large image databases. Roland Kwitt, Peter Meerwald-Stadler, Andreas Uhl |
IEEE Trans. Image Process. | 1 |
| 2010 | Watermarking of 2D vector graphics with distortion constraintabstractWe study the watermarking of 2D vector data and introduce a framework which preserves topological properties of the input. Our framework is based on so-called maximum perturbation regions (MPR) of the input vertices, which is a concept similar to the just-noticeable-difference constraint. The MPRs are computed by means of the Voronoi diagram of the input and allow us to avoid (self-)intersections of input objects that might result from the embedding of the watermark. We demonstrate and analyze the applicability of this new framework by coupling it with a well-known approach to watermarking that is based on Fourier descriptors. However, our framework is general enough such that any robust scheme for the watermarking of vector data can be applied. Stefan Huber 0001, Roland Kwitt, Peter Meerwald-Stadler, Martin Held, Andreas Uhl |
ICME | 2 |
| 2010 | Lightweight Probabilistic Texture RetrievalabstractThis paper contemplates the framework of probabilistic image retrieval in the wavelet domain from a computational point of view. We not only focus on achieving high retrieval rates, but also discuss possible performance bottlenecks which might prevent practical application. We propose a novel retrieval approach which is motivated by previous research work on modeling the marginal distributions of wavelet transform coefficients. The building blocks of our work are the dual-tree complex wavelet transform and a number of statistical models for the coefficient magnitudes. Image similarity measurement is accomplished by using closed-form solutions for the Kullback-Leibler divergences between the statistical models. We provide an in-depth computational analysis regarding the number of arithmetic operations required for similarity measurement and model parameter estimation. The experimental retrieval results on a widely used texture image database show that we achieve competitive retrieval results at low computational cost. Roland Kwitt, Andreas Uhl |
IEEE Trans. Image Process. | 1 |
| 2009 | Color-image watermarking using multivariate power-exponential distributionabstractIn this paper we present a novel watermark detector for additive spread-spectrum watermarking in the wavelet transform domain of color images. We propose to model the highly correlated DWT subbands of the RGB color channels by multivariate power-exponential distributions. This statistical model is then exploited to derive a likelihood ratio test for watermark detection. Our results indicate that joint statistical modeling of color DWT detail subbands leads to increased detection performance compared to previous approaches, namely watermarking of the luminance channel only, decorrelating the color bands, or relying on a joint Gaussian host signal model. Roland Kwitt, Peter Meerwald-Stadler, Andreas Uhl |
ICIP | 1 |
| 2009 | A joint model of complex wavelet coefficients for texture retrievalabstractWe present a Copula-based statistical model of complex wavelet coefficient magnitudes for color texture image retrieval. Our model is based on two-parameter Weibull distributions and a multivariate Student t Copula. For similarity measurement we employ a Monte- Carlo approach to approximate the Kullback-Leibler divergence between two models. The experimental retrieval results show that the incorporation of the dependency structure between subbands significantly improves retrieval accuracy compared to previous approaches. Roland Kwitt, Andreas Uhl |
ICIP | 1 |
| 2009 | Improving Pit-Pattern Classification of Endoscopy Images by a Combination of Experts
Michael Häfner, Alfred Gangl, Roland Kwitt, Andreas Uhl, Andreas Vécsei, Friedrich Wrba |
MICCAI (1) | 3 |
| 2009 | Feature extraction from multi-directional multi-resolution image transformations for the classification of zoom-endoscopy images
Michael Häfner, Roland Kwitt, Andreas Uhl, Alfred Gangl, Friedrich Wrba, Andreas Vécsei |
Pattern Anal. Appl. | 2 |
| 2009 | Computer-assisted pit-pattern classification in different wavelet domains for supporting dignity assessment of colonic polyps
Michael Häfner, Roland Kwitt, Andreas Uhl, Friedrich Wrba, Alfred Gangl, Andreas Vécsei |
Pattern Recognit. | 2 |
| 2008 | Color eigen-subband features for endoscopy image classificationabstractThis paper presents a new image feature extraction approach in the wavelet domain. We incorporate color-channel information of the LAB color space into the feature extraction process by computing variances from decorrelated detail subbands of the stationary wavelet transform. We evaluate our approach on a medical image classification problem using a k-nearest neighbor classifier and sequential forward feature selection. our experimental results, which include a comparative study to the popular color wavelet energy correlation signatures show that we can produce highly discriminative feature sets in terms of leave-one-out classification accuracy. Roland Kwitt, Andreas Uhl |
ICASSP | 1 |
| 2008 | Image similarity measurement by Kullback-Leibler divergences between complex wavelet subband statistics for texture retrievalabstractIn this work, we present a texture-image retrieval approach, which is based on the idea of measuring the Kullback-Leibler divergence between the marginal distributions of complex wavelet coefficient magnitudes. We employ Kingsbury's dual-tree complex wavelet transform for image decomposition and propose to model the detail subband coefficient magnitudes by either two-parameter Weibull or Gamma distributions for which we provide closed-form solutions to the Kullback-Leibler divergence. The experimental results indicate that our approach can achieve higher retrieval rates than the classical approach of using the pyramidal discrete wavelet transform together with the generalized Gaussian model for detail subband coefficients. Roland Kwitt, Andreas Uhl |
ICIP | 1 |
| 2007 | Modeling the Marginal Distributions of Complex Wavelet Coefficient Magnitudes for the Classification of Zoom-Endoscopy ImagesabstractIn this paper, we propose a set of new image features for the classification of zoom-endoscopy images. The feature extraction step is based on fitting a two-parameter Weibull distribution to the wavelet coefficient magnitudes of sub-bands obtained from a complex wavelet transform variant. We show, that the shape and scale parameter possess more discriminative power than the classic mean and standard deviation based features for complex subband coefficient magnitudes. Furthermore, we discuss why the commonly used Rayleigh distribution model is suboptimal in our case. Roland Kwitt, Andreas Uhl |
ICCV | 1 |