Thomas Pock

dblp:48/6407 · DBLP profile ↗
← Back
79ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0001-6120-1058ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 63 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 52 · 6 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Bigger Isn't Always Better: Towards a General Prior for Medical Image Reconstruction
Lukas Glaszner, Martin Zach, Thomas Pock
Int. J. Comput. Vis.3
2026 Diffusion at Absolute Zero: Langevin Sampling Using Successive Moreau Envelopes
abstract
Abstract. We propose a method for sampling from Gibbs distributions of the form [Formula: see text] that leverages a family [Formula: see text] of approximations of the target density which is deliberately constructed such that [Formula: see text] exhibits favorable properties for sampling when [Formula: see text] is large and such that [Formula: see text] approaches [Formula: see text] as [Formula: see text] approaches 0. This sequence is obtained by replacing (parts of) the potential [Formula: see text] with its Moreau envelope. Through the sequential sampling from [Formula: see text] for decreasing values of [Formula: see text] by a Langevin algorithm with appropriate step size, the samples are guided from a simple starting density to the more complex target quickly. We prove that [Formula: see text] is Lipschitz continuous in the total variation distance and Hölder continuous in the Wasserstein-[Formula: see text] distance, that the sampling algorithm is ergodic, and that it converges to the target density without assuming convexity or differentiability of the potential [Formula: see text]. In addition to the theoretical analysis, we show experimental results that support the superiority of the method in terms of convergence speed and mode-coverage of multimodal densities over current algorithms. The experiments range from one-dimensional toy-problems to high-dimensional inverse imaging problems with learned potentials.
Andreas Habring, Alexander Falk, Martin Zach, Thomas Pock
SIAM J. Imaging Sci.4
2025 FlowSDF: Flow Matching for Medical Image Segmentation Using Distance Transforms
abstract
Abstract Medical image segmentation plays an important role in accurately identifying and isolating regions of interest within medical images. Generative approaches are particularly effective in modeling the statistical properties of segmentation masks that are closely related to the respective structures. In this work we introduce FlowSDF, an image-guided conditional flow matching framework, designed to represent the signed distance function (SDF), and, in turn, to represent an implicit distribution of segmentation masks. The advantage of leveraging the SDF is a more natural distortion when compared to that of binary masks. Through the learning of a vector field associated with the probability path of conditional SDF distributions, our framework enables accurate sampling of segmentation masks and the computation of relevant statistical measures. This probabilistic approach also facilitates the generation of uncertainty maps represented by the variance, thereby supporting enhanced robustness in prediction and further analysis. We qualitatively and quantitatively illustrate competitive performance of the proposed method on a public nuclei and gland segmentation data set, highlighting its utility in medical image segmentation applications.
Lea Bogensperger, Dominik Narnhofer, Alexander Falk, Konrad Schindler, Thomas Pock
Int. J. Comput. Vis.5
2024 Selective, Interpretable and Motion Consistent Privacy Attribute Obfuscation for Action Recognition
abstract
Concerns for the privacy of individuals captured in public imagery have led to privacy-preserving action recognition. Existing approaches often suffer from issues arising through obfuscation being applied globally and a lack of interpretability. Global obfuscation hides privacy sensitive regions, but also contextual regions important for action recognition. Lack of interpretability erodes trust in these new technologies. We highlight the limitations of current paradigms and propose a solution: Human selected privacy templates that yield interpretability by design, an ob-fuscation scheme that selectively hides attributes and also induces temporal consistency, which is important in action recognition. Our approach is architecture agnostic and directly modifies input imagery, while existing approaches generally require architecture training. Our approach offers more flexibility, as no training is required, and outperforms alternatives on three widely used datasets.
Filip Ilic, He Zhao 0004, Thomas Pock, Richard P. Wildes
CVPR3
2024 Diffusion-based generation of Histopathological Whole Slide Images at a Gigapixel scale
abstract
We present a novel diffusion-based approach to generate synthetic histopathological Whole Slide Images (WSIs) at an unprecedented gigapixel scale. Synthetic WSIs have many potential applications: They can augment training datasets to enhance the performance of many computational pathology applications. They allow the creation of synthesized copies of datasets that can be shared without violating privacy regulations. Or they can facilitate learning representations of WSIs without requiring data annotations. Despite this variety of applications, no existing deep-learning-based method generates WSIs at their typically high resolutions. Mainly due to the high computational complexity. Therefore, we propose a novel coarse-to-fine sampling scheme to tackle image generation of high-resolution WSIs. In this scheme, we increase the resolution of an initial low-resolution image to a high-resolution WSI. Particularly, a diffusion model sequentially adds fine details to images and increases their resolution. In our experiments, we train our method with WSIs from the TCGA-BRCA dataset. Additionally to quantitative evaluations, we also performed a user study with pathologists. The study results suggest that our generated WSIs resemble the structure of real WSIs.
Robert Harb, Thomas Pock, Heimo Müller
WACV2
2024 Total Generalized Variation on a Tree
abstract
Abstract. We consider a class of optimization problems defined over trees with unary cost terms and shifted pairwise cost terms. These problems arise when considering block coordinate descent (BCD) approaches for solving inverse problems with total generalized variation (TGV) regularizers or their nonconvex generalizations. We introduce a linear-time reduction that transforms the shifted problems into their nonshifted counterparts. However, combining existing continuous dynamic programming (DP) algorithms with the reduction does not lead to BCD iterations that compute TGV-like solutions. This problem can be overcome by considering a box-constrained modification of the subproblems or smoothing the cost terms of the TGV regularized problem. The former leads to shifted and box-constrained subproblems, for which we propose a linear-time reduction to their unconstrained counterpart. The latter naturally leads to problems with smooth unary and pairwise cost terms. With this in mind, we propose two novel continuous DP algorithms that can solve (convex and nonconvex) problems with piecewise quadratic unary and pairwise cost terms. We prove that the algorithm for the convex case has quadratic worst-case time and memory complexity, while the algorithm for the nonconvex case has exponential time and memory complexity, but works well in practice for smooth truncated total variation pairwise costs. Finally, we demonstrate the applicability of the proposed algorithms for solving inverse problems with first-order and higher-order regularizers.
Muhamed Kuric, Jan Ahmetspahic, Thomas Pock
SIAM J. Imaging Sci.3
2024 Posterior-Variance-Based Error Quantification for Inverse Problems in Imaging
abstract
Abstract. In this work, a method for obtaining pixelwise error bounds in Bayesian regularization of inverse imaging problems is introduced. The proposed method employs estimates of the posterior variance together with techniques from conformal prediction in order to obtain coverage guarantees for the error bounds, without making any assumption on the underlying data distribution. It is generally applicable to Bayesian regularization approaches, independent, e.g., of the concrete choice of the prior. Furthermore, the coverage guarantees can also be obtained in case only approximate sampling from the posterior is possible. With this in particular, the proposed framework is able to incorporate any learned prior in a black-box manner. Guaranteed coverage without assumptions on the underlying distributions is only achievable since the magnitude of the error bounds is, in general, unknown in advance. Nevertheless, experiments with multiple regularization approaches presented in the paper confirm that, in practice, the obtained error bounds are rather tight. For realizing the numerical experiments, a novel primal-dual Langevin algorithm for sampling from nonsmooth distributions is also introduced in this work, showing promising results in practice. While a proof of convergence for this primal-dual algorithm is still open, the theoretical guarantees of the proposed method do not require a guaranteed convergence of the sampling algorithm.
Dominik Narnhofer, Andreas Habring, Martin Holler, Thomas Pock
SIAM J. Imaging Sci.4
2023 Lightweight Video Denoising using Aggregated Shifted Window Attention
abstract
Video denoising is a fundamental problem in numerous computer vision applications. State-of-the-art attention-based denoising methods typically yield good results, but require vast amounts of GPU memory and usually suffer from very long computation times. Especially in the field of restoring digitized high-resolution historic films, these techniques are not applicable in practice. To overcome these issues, we introduce a lightweight video denoising network that combines efficient axial-coronal-sagittal (ACS) convolutions with a novel shifted window attention formulation (ASwin), which is based on the memory-efficient aggregation of self- and cross-attention across video frames. We numerically validate the performance and efficiency of our approach on synthetic Gaussian noise. Moreover, we train our network as a general-purpose blind denoising model for real-world videos, using a realistic noise synthesis pipeline to generate clean-noisy video pairs. A user study and non-reference quality assessment prove that our method outperforms the state-of-the-art on real-world historic videos in terms of denoising performance and temporal consistency.
Lydia Lindner, Alexander Effland, Filip Ilic, Thomas Pock, Erich Kobler
WACV4
2023 Stable Deep MRI Reconstruction Using Generative Priors
abstract
Data-driven approaches recently achieved remarkable success in magnetic resonance imaging (MRI) reconstruction, but integration into clinical routine remains challenging due to a lack of generalizability and interpretability. In this paper, we address these challenges in a unified framework based on generative image priors. We propose a novel deep neural network based regularizer which is trained in a generative setting on reference magnitude images only. After training, the regularizer encodes higher-level domain statistics which we demonstrate by synthesizing images without data. Embedding the trained model in a classical variational approach yields high-quality reconstructions irrespective of the sub-sampling pattern. In addition, the model shows stable behavior when confronted with out-of-distribution data in the form of contrast variation. Furthermore, a probabilistic interpretation provides a distribution of reconstructions and hence allows uncertainty quantification. To reconstruct parallel MRI, we propose a fast algorithm to jointly estimate the image and the sensitivity maps. The results demonstrate competitive performance, on par with state-of-the-art end-to-end deep learning methods, while preserving the flexibility with respect to sub-sampling patterns and allowing for uncertainty quantification.
Martin Zach, Florian Knoll, Thomas Pock
IEEE Trans. Medical Imaging3
2022 Learned Variational Video Color Propagation
Markus Hofinger, Erich Kobler, Alexander Effland, Thomas Pock
ECCV (23)4
2022 Is Appearance Free Action Recognition Possible?
Filip Ilic, Thomas Pock, Richard P. Wildes
ECCV (4)2
2022 Total Deep Variation: A Stable Regularization Method for Inverse Problems
abstract
Various problems in computer vision and medical imaging can be cast as inverse problems. A frequent method for solving inverse problems is the variational approach, which amounts to minimizing an energy composed of a data fidelity term and a regularizer. Classically, handcrafted regularizers are used, which are commonly outperformed by state-of-the-art deep learning approaches. In this work, we combine the variational formulation of inverse problems with deep learning by introducing the data-driven general-purpose total deep variation regularizer. In its core, a convolutional neural network extracts local features on multiple scales and in successive blocks. This combination allows for a rigorous mathematical analysis including an optimal control formulation of the training problem in a mean-field setting and a stability analysis with respect to the initial values and the parameters of the regularizer. In addition, we experimentally verify the robustness against adversarial attacks and numerically derive upper bounds for the generalization error. Finally, we achieve state-of-the-art results for several imaging tasks.
Erich Kobler, Alexander Effland, Karl Kunisch, Thomas Pock
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Bayesian Uncertainty Estimation of Learned Variational MRI Reconstruction
abstract
Recent deep learning approaches focus on improving quantitative scores of dedicated benchmarks, and therefore only reduce the observation-related (aleatoric) uncertainty. However, the model-immanent (epistemic) uncertainty is less frequently systematically analyzed. In this work, we introduce a Bayesian variational framework to quantify the epistemic uncertainty. To this end, we solve the linear inverse problem of undersampled MRI reconstruction in a variational setting. The associated energy functional is composed of a data fidelity term and the total deep variation (TDV) as a learned parametric regularizer. To estimate the epistemic uncertainty we draw the parameters of the TDV regularizer from a multivariate Gaussian distribution, whose mean and covariance matrix are learned in a stochastic optimal control problem. In several numerical experiments, we demonstrate that our approach yields competitive results for undersampled MRI reconstruction. Moreover, we can accurately quantify the pixelwise epistemic uncertainty, which can serve radiologists as an additional resource to visualize reconstruction reliability.
Dominik Narnhofer, Alexander Effland, Erich Kobler, Kerstin Hammernik, Florian Knoll, Thomas Pock
IEEE Trans. Medical Imaging6
2021 One-sided Frank-Wolfe algorithms for saddle problems
abstract
We study a class of convex-concave saddle-point problems of the form $\min_x\max_y ⟨Kx,y⟩+f_{\cal P}(x)-h^*(y)$ where $K$ is a linear operator, $f_{\cal P}$ is the sum of a convex function $f$ with a Lipschitz-continuous gradient and the indicator function of a bounded convex polytope ${\cal P}$, and $h^\ast$ is a convex (possibly nonsmooth) function. Such problem arises, for example, as a Lagrangian relaxation of various discrete optimization problems. Our main assumptions are the existence of an efficient {\em linear minimization oracle} ($lmo$) for $f_{\cal P}$ and an efficient {\em proximal map} ($prox$) for $h^*$ which motivate the solution via a blend of proximal primal-dual algorithms and Frank-Wolfe algorithms. In case $h^*$ is the indicator function of a linear constraint and function $f$ is quadratic, we show a $O(1/n^2)$ convergence rate on the dual objective, requiring $O(n \log n)$ calls of $lmo$. If the problem comes from the constrained optimization problem $\min_{x\in\mathbb R^d}\{f_{\cal P}(x)\:|\:Ax-b=0\}$ then we additionally get bound $O(1/n^2)$ both on the primal gap and on the infeasibility gap. In the most general case, we show a $O(1/n)$ convergence rate of the primal-dual gap again requiring $O(n\log n)$ calls of $lmo$. To the best of our knowledge, this improves on the known convergence rates for the considered class of saddle-point problems. We show applications to labeling problems frequently appearing in machine learning and computer vision.
Vladimir Kolmogorov, Thomas Pock
ICML2
2021 Learned Collaborative Stereo Refinement
abstract
Abstract In this work, we propose a learning-based method to denoise and refine disparity maps. The proposed variational network arises naturally from unrolling the iterates of a proximal gradient method applied to a variational energy defined in a joint disparity, color, and confidence image space. Our method allows to learn a robust collaborative regularizer leveraging the joint statistics of the color image, the confidence map and the disparity map. Due to the variational structure of our method, the individual steps can be easily visualized, thus enabling interpretability of the method. We can therefore provide interesting insights into how our method refines and denoises disparity maps. To this end, we can visualize and interpret the learned filters and activation functions and prove the increased reliability of the predicted pixel-wise confidence maps. Furthermore, the optimization based structure of our refinement module allows us to compute eigen disparity maps, which reveal structural properties of our refinement module. The efficiency of our method is demonstrated on the publicly available stereo benchmarks Middlebury 2014 and Kitti 2015.
Patrick Knöbelreiter, Thomas Pock
Int. J. Comput. Vis.2
2021 A Framework for the generation of digital twins of cardiac electrophysiology from clinical 12-leads ECGs
abstract
Cardiac digital twins (Cardiac Digital Twin (CDT)s) of human electrophysiology (Electrophysiology (EP)) are digital replicas of patient hearts derived from clinical data that match like-for-like all available clinical observations. Due to their inherent predictive potential, CDTs show high promise as a complementary modality aiding in clinical decision making and also in the cost-effective, safe and ethical testing of novel EP device therapies. However, current workflows for both the anatomical and functional twinning phases within CDT generation, referring to the inference of model anatomy and parameters from clinical data, are not sufficiently efficient, robust and accurate for advanced clinical and industrial applications. Our study addresses three primary limitations impeding the routine generation of high-fidelity CDTs by introducing; a comprehensive parameter vector encapsulating all factors relating to the ventricular EP; an abstract reference frame within the model allowing the unattended manipulation of model parameter fields; a novel fast-forward electrocardiogram (Electrocardiogram (ECG)) model for efficient and bio-physically-detailed simulation required for parameter inference. A novel workflow for the generation of CDTs is then introduced as an initial proof of concept. Anatomical twinning was performed within a reasonable time compatible with clinical workflows (<4h) for 12 subjects from clinically-attained magnetic resonance images. After assessment of the underlying fast forward ECG model against a gold standard bidomain ECG model, functional twinning of optimal parameters according to a clinically-attained 12 lead ECG was then performed using a forward Saltelli sampling approach for a single subject. The achieved results in terms of efficiency and fidelity demonstrate that our workflow is well-suited and viable for generating biophysically-detailed CDTs at scale.
Karli Gillette, Matthias A. F. Gsell, Anton J. Prassl, Elias Karabelas, Ursula Reiter, Gert Reiter, Thomas Grandits, Christian Payer, Darko Stern, Martin Urschler, Jason D. Bayer, Christoph M. Augustin, Aurel Neic, Thomas Pock, Edward J. Vigmond, Gernot Plank
Medical Image Anal.14
2021 Learning Consistent Discretizations of the Total Variation
abstract
In this work, we study a general framework of discrete approximations of the total variation for image reconstruction problems. The framework, for which we can show consistency in the sense of $\Gamma$-convergence, unifies and extends several existing discretization schemes. In addition, we propose algorithms for learning discretizations of the total variation in order to achieve the best possible reconstruction quality for particular image reconstruction tasks. Interestingly, the learned discretizations significantly differ between the tasks, illustrating that there is no universal best discretization of the total variation.
Antonin Chambolle, Thomas Pock
SIAM J. Imaging Sci.2
2021 Shared Prior Learning of Energy-Based Models for Image Reconstruction
abstract
We propose a novel learning-based framework for image reconstruction particularly designed for training without ground truth data, which has three major building blocks: energy-based learning, a patch-based Wasserstein loss functional, and shared prior learning. In energy-based learning, the parameters of an energy functional composed of a learned data fidelity term and a data-driven regularizer are computed in a mean-field optimal control problem. In the absence of ground truth data, we change the loss functional to a patch-based Wasserstein functional, in which local statistics of the output images are compared to uncorrupted reference patches. Finally, in shared prior learning, both aforementioned optimal control problems are optimized simultaneously with shared learned parameters of the regularizer to further enhance unsupervised image reconstruction. We derive several time discretization schemes of the gradient flow and verify their consistency in terms of Mosco convergence. In numerous numerical experiments, we demonstrate that the proposed method generates state-of-the-art results for various image reconstruction applications---even if no ground truth images are available for training.
Thomas Pinetz, Erich Kobler, Thomas Pock, Alexander Effland
SIAM J. Imaging Sci.3
2020 BP-MVSNet: Belief-Propagation-Layers for Multi-View-Stereo
abstract
In this work, we propose BP-MVSNet, a convolutional neural network (CNN)-based Multi-View-Stereo (MVS) method that uses a differentiable Conditional Random Field (CRF) layer for regularization. To this end, we propose to extend the BP layer [16] and add what is necessary to successfully use it in the MVS setting. We therefore show how we can calculate a normalization based on the expected 3D error, which we can then use to normalize the label jumps in the CRF. This is required to make the BP layer invariant to different scales in the MVS setting. In order to also enable fractional label jumps, we propose a differentiable interpolation step, which we embed into the computation of the pairwise term. These extensions allow us to integrate the BP layer into a multi-scale MVS network, where we continuously improve a rough initial estimate until we get high quality depth maps as a result. We evaluate the proposed BP-MVSNet in an ablation study and conduct extensive experiments on the DTU, Tanks and Temples and ETH3D data sets. The experiments show that we can significantly outperform the baseline and achieve state-of-the-art results.
Christian Sormann, Patrick Knöbelreiter, Andreas Kuhn 0005, Mattia Rossi, Thomas Pock, Friedrich Fraundorfer
3DV5
2020 Belief Propagation Reloaded: Learning BP-Layers for Labeling Problems
abstract
It has been proposed by many researchers that combining deep neural networks with graphical models can create more efficient and better regularized composite models. The main difficulties in implementing this in practice are associated with a discrepancy in suitable learning objectives as well as with the necessity of approximations for the inference. In this work we take one of the simplest inference methods, a truncated max-product Belief Propagation, and add what is necessary to make it a proper component of a deep learning model: connect it to learning formulations with losses on marginals and compute the backprop operation. This BP-Layer can be used as the final or an intermediate block in convolutional neural networks (CNNs), allowing us to design a hierarchical model composing BP inference and CNNs at different scale levels. The model is applicable to a range of dense prediction problems, is well-trainable and provides parameter-efficient and robust solutions in stereo, flow and semantic segmentation.
Patrick Knöbelreiter, Christian Sormann, Alexander Shekhovtsov 0001, Friedrich Fraundorfer, Thomas Pock
CVPR5
2020 Total Deep Variation for Linear Inverse Problems
abstract
Diverse inverse problems in imaging can be cast as variational problems composed of a task-specific data fidelity term and a regularization term. In this paper, we propose a novel learnable general-purpose regularizer exploiting recent architectural design patterns from deep learning. We cast the learning problem as a discrete sampled optimal control problem, for which we derive the adjoint state equations and an optimality condition. By exploiting the variational structure of our approach, we perform a sensitivity analysis with respect to the learned parameters obtained from different training datasets. Moreover, we carry out a nonlinear eigenfunction analysis, which reveals interesting properties of the learned regularizer. We show state-of-the-art performance for classical image restoration and medical image reconstruction problems.
Erich Kobler, Alexander Effland, Karl Kunisch, Thomas Pock
CVPR4
2020 Improving Optical Flow on a Pyramid Level
Markus Hofinger, Samuel Rota Bulò, Lorenzo Porzi, Arno Knapitsch, Thomas Pock, Peter Kontschieder
ECCV (28)5
2020 3D Fluid Flow Estimation with Integrated Particle Reconstruction
Katrin Lasinger, Christoph Vogel, Thomas Pock, Konrad Schindler
Int. J. Comput. Vis.3
2019 Fast Decomposable Submodular Function Minimization using Constrained Total Variation
abstract
We consider the problem of minimizing the sum of submodular set functions assuming minimization oracles of each summand function. Most existing approaches reformulate the problem as the convex minimization of the sum of the corresponding Lov\'asz extensions and the squared Euclidean norm, leading to algorithms requiring total variation oracles of the summand functions; without further assumptions, these more complex oracles require many calls to the simpler minimization oracles often available in practice. In this paper, we consider a modified convex problem requiring constrained version of the total variation oracles that can be solved with significantly fewer calls to the simple minimization oracles. We support our claims by showing results on graph cuts for 2D and 3D graphs.
Senanayak Sesh Kumar Karri, Francis R. Bach, Thomas Pock
NeurIPS3
2018 Learning Energy Based Inpainting for Optical Flow
Christoph Vogel, Patrick Knöbelreiter, Thomas Pock
ACCV (6)3
2018 Variational Deep Learning for Low-Dose Computed Tomography
abstract
In this work, we propose a learning-based variational network (VN) approach for reconstruction of low-dose 3D computed tomography data. We focus on two methods to decrease the radiation dose: (1) x-ray tube current reduction, which reduces the signal-to-noise ratio, and (2) x-ray beam interruption, which undersamples data and results in images with aliasing artifacts. While the learned VN denoises the current-reduced images in the first case, it reconstructs the undersampled data in the second case. Different VNs for denoising and reconstruction are trained on a single clinical 3D abdominal data set. The VNs are compared against state-of-the-art model-based denoising and sparse reconstruction techniques on a different clinical abdominal 3D data set with 4-fold dose reduction. Our results suggest that the proposed VNs enable higher radiation dose reductions and/or increase the image quality for a given dose.
Erich Kobler, Matthew J. Muckley, Florian Knoll, Kerstin Hammernik, Thomas Pock, Daniel K. Sodickson, Ricardo Otazo
ICASSP6
2018 Variational Fusion of Light Field and Photometric Stereo for Precise 3D Sensing within a Multi-Line Scan Framework
abstract
Recent work has shown the improved depth reconstruction by combining depth and surface normal information. In this paper, we build on the findings and introduce novel variational methods for a refined depth reconstruction for a multi-line scanner using light field and photometric stereo data. In this specific setup, the object is acquired while moving on a conveyor belt in a defined direction under the camera, which simultaneously captures light field and photometric stereo data as the object is transported. We perform our experiments on virtual and real-world data and achieve significantly improved results over state-of-the-art methods both in depth and surface normal accuracy.
Doris Antensteiner, Svorad Stolc, Thomas Pock
ICPR3
2018 Self-Supervised Learning for Stereo Reconstruction on Aerial Images
abstract
Recent developments established deep learning as an inevitable tool to boost the performance of dense matching and stereo estimation. On the downside, learning these networks requires a substantial amount of training data to be successful. Consequently, the application of these models outside of the laboratory is far from straight forward. In this work we propose a self-supervised training procedure that allows us to adapt our network to the specific (imaging) characteristics of the dataset at hand, without the requirement of external ground truth data. We instead generate interim training data by running our intermediate network on the whole dataset, followed by conservative outlier filtering. Bootstrapped from a pre-trained version of our hybrid CNN-CRF model, we alternate the generation of training data and network training. With this simple concept we are able to lift the completeness and accuracy of the pre-trained version significantly. We also show that our final model compares favorably to other popular stereo estimation algorithms on an aerial dataset.
Patrick Knöbelreiter, Christoph Vogel, Thomas Pock
IGARSS3
2018 Real-Time Intensity-Image Reconstruction for Event Cameras Using Manifold Regularisation
Gottfried Munda, Christian Reinbacher, Thomas Pock
Int. J. Comput. Vis.3
2017 Semantic 3D Reconstruction with Finite Element Bases
Audrey Richard, Christoph Vogel, Maros Blaha, Thomas Pock, Konrad Schindler
BMVC4
2017 End-to-End Training of Hybrid CNN-CRF Models for Stereo
abstract
We propose a novel and principled hybrid CNN+CRF model for stereo estimation. Our model allows to exploit the advantages of both, convolutional neural networks (CNNs) and conditional random fields (CRFs) in an unified approach. The CNNs compute expressive features for matching and distinctive color edges, which in turn are used to compute the unary and binary costs of the CRF. For inference, we apply a recently proposed highly parallel dual block descent algorithm which only needs a small fixed number of iterations to compute a high-quality approximate minimizer. As the main contribution of the paper, we propose a theoretically sound method based on the structured output support vector machine (SSVM) to train the hybrid CNN+CRF model on large-scale data end-to-end. Our trained models perform very well despite the fact that we are using shallow CNNs and do not apply any kind of post-processing to the final output of the CRF. We evaluate our combined models on challenging stereo benchmarks such as Middlebury 2014 and Kitti 2015 and also investigate the performance of each individual component.
Patrick Knöbelreiter, Christian Reinbacher, Alexander Shekhovtsov 0001, Thomas Pock
CVPR4
2017 Real-time panoramic tracking for event cameras
abstract
Event cameras are a paradigm shift in camera technology. Instead of full frames, the sensor captures a sparse set of events caused by intensity changes. Since only the changes are transferred, those cameras are able to capture quick movements of objects in the scene or of the camera itself. In this work we propose a novel method to perform camera tracking of event cameras in a panoramic setting with three degrees of freedom. We propose a direct camera tracking formulation, similar to state-of-the-art in visual odometry. We show that the minimal information needed for simultaneous tracking and mapping is the spatial position of events, without using the appearance of the imaged scene point. We verify the robustness to fast camera movements and dynamic objects in the scene on a recently proposed dataset [18] and self-recorded sequences.
Christian Reinbacher, Gottfried Munda, Thomas Pock
ICCP3
2017 Neural EPI-Volume Networks for Shape from Light Field
abstract
This paper presents a novel deep regression network to extract geometric information from Light Field (LF) data. Our network builds upon u-shaped network architectures. Those networks involve two symmetric parts, an encoding and a decoding part. In the first part the network encodes relevant information from the given input into a set of high-level feature maps. In the second part the generated feature maps are then decoded to the desired output. To predict reliable and robust depth information the proposed network examines 3D subsets of the 4D LF called Epipolar Plane Image (EPI) volumes. An important aspect of our network is the use of 3D convolutional layers, that allow to propagate information from two spatial dimensions and one directional dimension of the LF. Compared to previous work this allows for an additional spatial regularization, which reduces depth artifacts and simultaneously maintains clear depth discontinuities. Experimental results show that our approach allows to create high-quality reconstruction results, which outperform current state-of-the-art Shape from Light Field (SfLF) techniques. The main advantage of the proposed approach is the ability to provide those high-quality reconstructions at a low computation time.
Stefan Heber, Wei Yu 0014, Thomas Pock
ICCV3
2017 Trainable Nonlinear Reaction Diffusion: A Flexible Framework for Fast and Effective Image Restoration
abstract
Image restoration is a long-standing problem in low-level computer vision with many interesting applications. We describe a flexible learning framework based on the concept of nonlinear reaction diffusion models for various image restoration problems. By embodying recent improvements in nonlinear diffusion models, we propose a dynamic nonlinear reaction diffusion model with time-dependent parameters (i.e., linear filters and influence functions). In contrast to previous nonlinear diffusion models, all the parameters, including the filters and the influence functions, are simultaneously learned from training data through a loss based approach. We call this approach TNRD-Trainable Nonlinear Reaction Diffusion. The TNRD approach is applicable for a variety of image restoration tasks by incorporating appropriate reaction force. We demonstrate its capabilities with three representative applications, Gaussian image denoising, single image super resolution and JPEG deblocking. Experiments show that our trained nonlinear diffusion models largely benefit from the training of the parameters and finally lead to the best reported performance on common test datasets for the tested applications. Our trained models preserve the structural simplicity of diffusion models and take only a small number of diffusion steps, thus are highly efficient. Moreover, they are also well-suited for parallel computation on GPUs, which makes the inference procedure extremely fast.
Yunjin Chen, Thomas Pock
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 U-shaped Networks for Shape from Light Field
Stefan Heber, Wei Yu 0014, Thomas Pock
BMVC3
2016 Real-Time Intensity-Image Reconstruction for Event Cameras Using Manifold Regularisation
Christian Reinbacher, Gottfried Munda, Thomas Pock
BMVC3
2016 Large-Scale Semantic 3D Reconstruction: An Adaptive Multi-resolution Model for Multi-class Volumetric Labeling
abstract
We propose an adaptive multi-resolution formulation of semantic 3D reconstruction. Given a set of images of a scene, semantic 3D reconstruction aims to densely reconstruct both the 3D shape of the scene and a segmentation into semantic object classes. Jointly reasoning about shape and class allows one to take into account class-specific shape priors (e.g., building walls should be smooth and vertical, and vice versa smooth, vertical surfaces are likely to be building walls), leading to improved reconstruction results. So far, semantic 3D reconstruction methods have been limited to small scenes and low resolution, because of their large memory footprint and computational cost. To scale them up to large scenes, we propose a hierarchical scheme which refines the reconstruction only in regions that are likely to contain a surface, exploiting the fact that both high spatial resolution and high numerical precision are only required in those regions. Our scheme amounts to solving a sequence of convex optimizations while progressively removing constraints, in such a way that the energy, in each iteration, is the tightest possible approximation of the underlying energy at full resolution. In our experiments the method saves up to 98% memory and 95% computation time, without any loss of accuracy.
Maros Blaha, Christoph Vogel, Audrey Richard, Jan Dirk Wegner, Thomas Pock, Konrad Schindler
CVPR5
2016 Convolutional Networks for Shape from Light Field
abstract
Convolutional Neural Networks (CNNs) have recently been successfully applied to various Computer Vision (CV) applications. In this paper we utilize CNNs to predict depth information for given Light Field (LF) data. The proposed method learns an end-to-end mapping between the 4D light field and a representation of the corresponding 4D depth field in terms of 2D hyperplane orientations. The obtained prediction is then further refined in a post processing step by applying a higher-order regularization. Existing LF datasets are not sufficient for the purpose of the training scheme tackled in this paper. This is mainly due to the fact that the ground truth depth of existing datasets is inaccurate and/or the datasets are limited to a small number of LFs. This made it necessary to generate a new synthetic LF dataset, which is based on the raytracing software POV-Ray. This new dataset provides floating point accurate ground truth depth fields, and due to a random scene generator the dataset can be scaled as required.
Stefan Heber, Thomas Pock
CVPR2
2016 Learning joint demosaicing and denoising based on sequential energy minimization
abstract
Demosaicing is an important first step for color image acquisition. For practical reasons, demosaicing algorithms have to be both efficient and yield high quality results in the presence of noise. The demosaicing problem poses several challenges, e.g. zippering and false color artifacts as well as edge blur. In this work, we introduce a novel learning based method that can overcome these challenges. We formulate demosaicing as an image restoration problem and propose to learn efficient regularization inspired by a variational energy minimization framework that can be trained for different sensor layouts. Our algorithm performs joint demosaicing and denoising in close relation to the real physical mosaicing process on a camera sensor. This is achieved by learning a sequence of energy minimization problems composed of a set of RGB filters and corresponding activation functions. We evaluate our algorithm on the Microsoft Demosaicing data set in terms of peak signal to noise ratio (PSNR) and structured similarity index (SSIM). Our algorithm is highly efficient both in image quality and run time. We achieve an improvement of up to 2.6 dB over recent state-of-the-art algorithms.
Teresa Klatzer, Kerstin Hammernik, Patrick Knöbelreiter, Thomas Pock
ICCP4
2016 Total Variation on a Tree
abstract
We consider the problem of minimizing the continuous valued total variation subject to different unary terms on trees and propose fast direct algorithms based on dynamic programming to solve these problems. We treat both the convex and the nonconvex case and derive worst-case complexities that are equal to or better than existing methods. We show applications to total variation based two dimensional image processing and computer vision problems based on a Lagrangian decomposition approach. The resulting algorithms are very efficient, offer a high degree of parallelism, and come along with memory requirements which are only in the order of the number of image pixels.
Vladimir Kolmogorov, Thomas Pock, Michal Rolínek
SIAM J. Imaging Sci.2
2016 Inertial Proximal Alternating Linearized Minimization (iPALM) for Nonconvex and Nonsmooth Problems
abstract
In this paper we study nonconvex and nonsmooth optimization problems with semialgebraic data, where the variables vector is split into several blocks of variables. The problem consists of one smooth function of the entire variables vector and the sum of nonsmooth functions for each block separately. We analyze an inertial version of the proximal alternating linearized minimization algorithm and prove its global convergence to a critical point of the objective function at hand. We illustrate our theoretical findings by presenting numerical experiments on blind image deconvolution, on sparse nonnegative matrix factorization and on dictionary learning, which demonstrate the viability and effectiveness of the proposed method.
Thomas Pock, Shoham Sabach
SIAM J. Imaging Sci.1
2015 Depth Restoration via Joint Training of a Global Regression Model and CNNs
abstract
[1] Bredies, Kunisch and Pock. Total Generalized Variation, SIAM Journal on Imaging Sciences, 3(3):492-526, 2012 [2] Ferstl, Reinbacher, Ranftl, Ruther and Bischof. Image Guided Depth Upsampling using Anisotropic Total Generalized Variation, ICCV, 2013 [3} Kunisch and Pock. A Bilevel Optimization Approach for Parameter Learning in Variational Models. SIAM Journal on Imaging Sciences, 6(2):938-983, 2013 [4] Martull, Peris and Fukui. Realistic CG Stereo Image Dataset with Ground Truth Disparity Maps. ICPRW, 2012 [5] Ranftl and Pock. A Deep Variational Model for Image Segmentation, GCPR, 2014 References
Gernot Riegler, René Ranftl, Matthias Rüther, Thomas Pock, Horst Bischof
BMVC4
2015 On learning optimized reaction diffusion processes for effective image restoration
abstract
For several decades, image restoration remains an active research topic in low-level computer vision and hence new approaches are constantly emerging. However, many recently proposed algorithms achieve state-of-the-art performance only at the expense of very high computation time, which clearly limits their practical relevance. In this work, we propose a simple but effective approach with both high computational efficiency and high restoration quality. We extend conventional nonlinear reaction diffusion models by several parametrized linear filters as well as several parametrized influence functions. We propose to train the parameters of the filters and the influence functions through a loss based approach. Experiments show that our trained nonlinear reaction diffusion models largely benefit from the training of the parameters and finally lead to the best reported performance on common test datasets for image restoration. Due to their structural simplicity, our trained models are highly efficient and are also well-suited for parallel computation on GPUs.
Yunjin Chen, Wei Yu 0014, Thomas Pock
CVPR3
2015 Efficient minimal-surface regularization of perspective depth maps in variational stereo
abstract
We propose a method for dense three-dimensional surface reconstruction that leverages the strengths of shape-based approaches, by imposing regularization that respects the geometry of the surface, and the strength of depth-map-based stereo, by avoiding costly computation of surface topology. The result is a near real-time variational reconstruction algorithm free of the staircasing artifacts that affect depth-map and plane-sweeping approaches. This is made possible by exploiting the gauge ambiguity to design a novel representation of the regularizer that is linear in the parameters and hence amenable to be optimized with state-of-the-art primal-dual numerical schemes.
Gottfried Munda, Jonathan Balzer, Stefano Soatto, Thomas Pock
CVPR4
2015 On Iteratively Reweighted Algorithms for Nonsmooth Nonconvex Optimization in Computer Vision
abstract
Natural image statistics indicate that we should use nonconvex norms for most regularization tasks in image processing and computer vision. Still, they are rarely used in practice due to the challenge of optimization. Recently, iteratively reweighed $\ell_1$ minimization (IRL1) has been proposed as a way to tackle a class of nonconvex functions by solving a sequence of convex $\ell_2$-$\ell_1$ problems. We extend the problem class to the sum of a convex function and a (nonconvex) nondecreasing function applied to another convex function. The proposed algorithm sequentially optimizes suitably constructed convex majorizers. Convergence to a critical point is proved when the Kurdyka--Łojasiewicz property and additional mild restrictions hold for the objective function. The efficiency and practical importance of the algorithm are demonstrated in computer vision tasks such as image denoising and optical flow. Most applications seek smooth results with sharp discontinuities. These are achieved by combining nonconvexity with higher order regularization.
Peter Ochs, Alexey Dosovitskiy, Thomas Brox, Thomas Pock
SIAM J. Imaging Sci.4
2014 Shape from Light Field Meets Robust PCA
Stefan Heber, Thomas Pock
ECCV (6)2
2014 Non-local Total Generalized Variation for Optical Flow Estimation
René Ranftl, Kristian Bredies, Thomas Pock
ECCV (1)3
2014 iPiano: Inertial Proximal Algorithm for Nonconvex Optimization
abstract
In this paper we study an algorithm for solving a minimization problem composed of a differentiable (possibly nonconvex) and a convex (possibly nondifferentiable) function. The algorithm iPiano combines forward-backward splitting with an inertial force. It can be seen as a nonsmooth split version of the Heavy-ball method from Polyak. A rigorous analysis of the algorithm for the proposed class of problems yields global convergence of the function values and the arguments. This makes the algorithm robust for usage on nonconvex problems. The convergence result is obtained based on the Kurdyka--Łojasiewicz inequality. This is a very weak restriction, which was used to prove convergence for several other gradient methods. First, an abstract convergence theorem for a generic algorithm is proved, and then iPiano is shown to satisfy the requirements of this theorem. Furthermore, a convergence rate is established for the general problem class. We demonstrate iPiano on computer vision problems---image denoising with learned priors and diffusion based image compression.
Peter Ochs, Yunjin Chen, Thomas Brox, Thomas Pock
SIAM J. Imaging Sci.4
2014 A Higher-Order MRF Based Variational Model for Multiplicative Noise Reduction
abstract
The Fields of Experts (FoE) image prior model, a filter-based higher-order Markov Random Fields (MRF) model, has been shown to be effective for many image restoration problems. Motivated by the successes of FoE-based approaches, in this letter we propose a novel variational model for multiplicative noise reduction based on the FoE image prior model. The resulting model corresponds to a non-convex minimization problem, which can be efficiently solved by a recently published non-convex optimization algorithm. Experimental results based on synthetic speckle noise and real synthetic aperture radar (SAR) images suggest that the performance of our proposed method is on par with the best published despeckling algorithm. Besides, our proposed model comes along with an additional advantage, that the inference is extremely efficient. Our GPU based implementation takes less than 1s to produce state-of-the-art despeckling performance.
Yunjin Chen, WenSen Feng, René Ranftl, Hong Qiao, Thomas Pock
IEEE Signal Process. Lett.5
2014 Insights Into Analysis Operator Learning: From Patch-Based Sparse Models to Higher Order MRFs
abstract
This paper addresses a new learning algorithm for the recently introduced co-sparse analysis model. First, we give new insights into the co-sparse analysis model by establishing connections to filter-based MRF models, such as the field of experts model of Roth and Black. For training, we introduce a technique called bi-level optimization to learn the analysis operators. Compared with existing analysis operator learning approaches, our training procedure has the advantage that it is unconstrained with respect to the analysis operator. We investigate the effect of different aspects of the co-sparse analysis model and show that the sparsity promoting function (also called penalty function) is the most important factor in the model. In order to demonstrate the effectiveness of our training approach, we apply our trained models to various classical image restoration problems. Numerical experiments show that our trained models clearly outperform existing analysis operator learning approaches and are on par with state-of-the-art image denoising algorithms. Our approach develops a framework that is intuitive to understand and easy to implement.
Yunjin Chen, René Ranftl, Thomas Pock
IEEE Trans. Image Process.3
2013 An Iterated L1 Algorithm for Non-smooth Non-convex Optimization in Computer Vision
abstract
Natural image statistics indicate that we should use non-convex norms for most regularization tasks in image processing and computer vision. Still, they are rarely used in practice due to the challenge to optimize them. Recently, iteratively reweighed ℓ1minimization has been proposed as a way to tackle a class of non-convex functions by solving a sequence of convex ℓ2- ℓ1problems. Here we extend the problem class to linearly constrained optimization of a Lipschitz continuous function, which is the sum of a convex function and a function being concave and increasing on the non-negative orthant (possibly non-convex and non-concave on the whole space). This allows to apply the algorithm to many computer vision tasks. We show the effect of non-convex regularizers on image denoising, deconvolution, optical flow, and depth map fusion. Non-convexity is particularly interesting in combination with total generalized variation and learned image priors. Efficient optimization is made possible by some important properties that are shown to hold.
Peter Ochs, Alexey Dosovitskiy, Thomas Brox, Thomas Pock
CVPR4
2013 A Bilevel Optimization Approach for Parameter Learning in Variational Models
abstract
In this work we consider the problem of parameter learning for variational image denoising models. The learning problem is formulated as a bilevel optimization problem, where the lower-level problem is given by the variational model and the higher-level problem is expressed by means of a loss function that penalizes errors between the solution of the lower-level problem and the ground truth data. We consider a class of image denoising models incorporating $\ell_p$-norm--based analysis priors using a fixed set of linear operators. We devise semismooth Newton methods for solving the resulting nonsmooth bilevel optimization problems and show that the optimized image denoising models can achieve state-of-the-art performance.
Karl Kunisch, Thomas Pock
SIAM J. Imaging Sci.2
2012 Joint motion estimation and segmentation of complex scenes with label costs and occlusion modeling
abstract
We propose a unified variational formulation for joint motion estimation and segmentation with explicit occlusion handling. This is done by a multi-label representation of the flow field, where each label corresponds to a parametric representation of the motion. We use a convex formulation of the multi-label Potts model with label costs and show that the asymmetric map-uniqueness criterion can be integrated into our formulation by means of convex constraints. Explicit occlusion handling eliminates errors otherwise created by the regularization. As occlusions can occur only at object boundaries, a large number of objects may be required. By using a fast primal-dual algorithm we are able to handle several hundred motion segments. Results are shown on several classical motion segmentation and optical flow examples.
Markus Unger, Manuel Werlberger, Thomas Pock, Horst Bischof
CVPR3
2012 Dense reconstruction on-the-fly
abstract
We present a novel system that is capable of generating live dense volumetric reconstructions based on input from a micro aerial vehicle. The distributed reconstruction pipeline is based on state-of-the-art approaches to visual SLAM and variational depth map fusion, and is designed to exploit the individual capabilities of the system components. Results are visualized in real-time on a tablet interface, which gives the user the opportunity to interact. We demonstrate the performance of our approach by capturing several indoor and outdoor scenes on-the-fly and by evaluating our results with respect to a ground-truth model.
Andreas Wendel, Michael Maurer, Gottfried Munda, Thomas Pock, Horst Bischof
CVPR4
2012 Pushing the limits of stereo using variational stereo estimation
abstract
We examine high accuracy stereo estimation for binocular sequences that where obtained from a mobile platform. The ultimate goal is to improve the range of stereo systems without altering the setup. Based on a well-known variational optical flow model, we introduce a novel stereo model that features a second-order regularization, which both allows sub-pixel accurate solutions and piecewise planar disparity maps. The model incorporates a robust fidelity term to account for adverse illumination conditions that frequently arise in real-world scenes. Using several sequences that were taken from a mobile platform we show the robustness and accuracy of the proposed model.
René Ranftl, Stefan K. Gehrig, Thomas Pock, Horst Bischof
Intelligent Vehicles Symposium3
2012 A Convex Approach to Minimal Partitions
abstract
We describe a convex relaxation for a family of problems of minimal perimeter partitions. The minimization of the relaxed problem can be tackled numerically: we describe an algorithm and show some results. In most cases, our relaxed problem finds a correct numerical approximation of the optimal solution: we give some arguments to explain why it should be so and also discuss some situations where it fails.
Antonin Chambolle, Daniel Cremers, Thomas Pock
SIAM J. Imaging Sci.3
2011 Diagonal preconditioning for first order primal-dual algorithms in convex optimization
abstract
In this paper we study preconditioning techniques for the first-order primal-dual algorithm proposed in [5]. In particular, we propose simple and easy to compute diagonal preconditioners for which convergence of the algorithm is guaranteed without the need to compute any step size parameters. As a by-product, we show that for a certain instance of the preconditioning, the proposed algorithm is equivalent to the old and widely unknown alternating step method for monotropic programming [7]. We show numerical results on general linear programming problems and a few standard computer vision problems. In all examples, the preconditioned algorithm significantly outperforms the algorithm of [5].
Thomas Pock, Antonin Chambolle
ICCV1
2010 Interactive Multi-label Segmentation
Jakob Santner, Thomas Pock, Horst Bischof
ACCV (1)2
2010 Variational segmentation of elongated volumetric structures
abstract
We present an interactive approach for segmenting thin volumetric structures. The proposed segmentation model is based on an anisotropic weighted Total Variation energy with a global volumetric constraint and is minimized using an efficient numerical approach and a convex relaxation. The algorithm is globally optimal w.r.t. the relaxed problem for any volumetric constraint. The binary solution of the relaxed problem equals the globally optimal solution of the original problem. Implemented on today's user-programmable graphics cards, it allows real-time user interaction. The method is applied to and evaluated on the task of articular cartilage segmentation of human knee joints and segmentation of tubular structures like liver vessels and airway trees.
Christian Reinbacher, Thomas Pock, Christian Bauer 0001, Horst Bischof
CVPR2
2010 Online multi-class LPBoost
abstract
Online boosting is one of the most successful online learning algorithms in computer vision. While many challenging online learning problems are inherently multi-class, online boosting and its variants are only able to solve binary tasks. In this paper, we present Online Multi-Class LPBoost (OMCLP) which is directly applicable to multi-class problems. From a theoretical point of view, our algorithm tries to maximize the multi-class soft-margin of the samples. In order to solve the LP problem in online settings, we perform an efficient variant of online convex programming, which is based on primal-dual gradient descent-ascent update strategies. We conduct an extensive set of experiments over machine learning benchmark datasets, as well as, on Caltech 101 category recognition dataset. We show that our method is able to outperform other online multi-class methods. We also apply our method to tracking where, we present an intuitive way to convert the binary tracking by detection problem to a multi-class problem where background patterns which are similar to the target class, become virtual classes. Applying our novel model, we outperform or achieve the state-of-the-art results on benchmark tracking videos.
Amir Saffari, Martin Godec, Thomas Pock, Christian Leistner, Horst Bischof
CVPR3
2010 PROST: Parallel robust online simple tracking
abstract
Tracking-by-detection is increasingly popular in order to tackle the visual tracking problem. Existing adaptive methods suffer from the drifting problem, since they rely on self-updates of an on-line learning method. In contrast to previous work that tackled this problem by employing semi-supervised or multiple-instance learning, we show that augmenting an on-line learning method with complementary tracking approaches can lead to more stable results. In particular, we use a simple template model as a non-adaptive and thus stable component, a novel optical-flow-based mean-shift tracker as highly adaptive element and an on-line random forest as moderately adaptive appearance-based learner. We combine these three trackers in a cascade. All of our components run on GPUs or similar multi-core systems, which allows for real-time performance. We show the superiority of our system over current state-of-the-art tracking methods in several experiments on publicly available data.
Jakob Santner, Christian Leistner, Amir Saffari, Thomas Pock, Horst Bischof
CVPR4
2010 Motion estimation with non-local total variation regularization
abstract
State-of-the-art motion estimation algorithms suffer from three major problems: Poorly textured regions, occlusions and small scale image structures. Based on the Gestalt principles of grouping we propose to incorporate a low level image segmentation process in order to tackle these problems. Our new motion estimation algorithm is based on non-local total variation regularization which allows us to integrate the low level image segmentation process in a unified variational framework. Numerical results on the Middlebury optical flow benchmark data set demonstrate that we can cope with the aforementioned problems.
Manuel Werlberger, Thomas Pock, Horst Bischof
CVPR2
2010 Anisotropic Minimal Surfaces Integrating Photoconsistency and Normal Information for Multiview Stereo
Kalin Kolev, Thomas Pock, Daniel Cremers
ECCV (3)2
2010 Segmentation of interwoven 3d tubular tree structures utilizing shape priors and graph cuts
Christian Bauer 0001, Thomas Pock, Erich Sorantin, Horst Bischof, Reinhard Beichel
Medical Image Anal.2
2010 Total Generalized Variation
abstract
The novel concept of total generalized variation of a function u is introduced, and some of its essential properties are proved. Differently from the bounded variation seminorm, the new concept involves higher-order derivatives of u. Numerical examples illustrate the high quality of this functional as a regularization term for mathematical imaging problems. In particular this functional selectively regularizes on different regularity levels and, as a side effect, does not lead to a staircasing effect.
Kristian Bredies, Karl Kunisch, Thomas Pock
SIAM J. Imaging Sci.3
2010 Global Solutions of Variational Models with Convex Regularization
abstract
We propose an algorithmic framework for computing global solutions of variational models with convex regularity terms that permit quite arbitrary data terms. While the minimization of variational problems with convex data and regularity terms is straightforward (using, for example, gradient descent), this is no longer trivial for functionals with nonconvex data terms. Using the theoretical framework of calibrations, the original variational problem can be written as the maximum flux of a particular vector field going through the boundary of the subgraph of the unknown function. Upon relaxation this formulation turns the problem into a convex problem, although in a higher dimension. In order to solve this problem, we propose a fast primal-dual algorithm which significantly outperforms existing algorithms. In experimental results we show the application of our method to outlier filtering of range images and disparity estimation in stereo images using a variety of convex regularity terms.
Thomas Pock, Daniel Cremers, Horst Bischof, Antonin Chambolle
SIAM J. Imaging Sci.1
2009 Interactive Texture Segmentation using Random Forests and Total Variation
abstract
Common methods for interactive texture segmentation rely on probability maps based on low dimensional features such as e.g. intensity or color, that are usually modeled using basic learning algorithms such as histograms or Gaussian Mixture Models. The use of low level features allows for fast generation of these hypotheses but limits applicability to a small class of images. We address this problem by learning complex descriptors with Random Forests and exploiting their inherent parallelism in a GPU implementation. The segmentation itself is based on a convex energy functional that uses weighted Total Variation regularization and a point-wise data term allowing for continuous foreground/background membership hypotheses. Its globally optimal solution is obtained by a fast primal-dual algorithm providing a reasonable convergence criterion. As a result, we present a versatile interactive texture segmentation framework. We show experiments with natural, artificial and medical data and demonstrate superior results compared to two recent approaches.
Jakob Santner, Markus Unger, Thomas Pock, Christian Leistner, Amir Saffari, Horst Bischof
BMVC3
2009 Anisotropic Huber-L1 Optical Flow
abstract
TV regularization is an L1 penalization of the flow gradient magnitudes, and due to the tendency of the L1 norm to favor sparse solutions (i.e. lots of ‘zeros’), the fill-in effect caused by the regularizer leads to piecewise constant solutions in weakly textured areas. This effect, known as ‘staircasing’ in a 1D setting, can be reduced significantly by using a quadratic penalization for small gradient magnitudes while sticking to linear penalization for larger magnitudes to maintain the discontinuity preserving properties known from TV. A comparison of isotropic TV and isotropic Huber regularity is shown in Fig. 1 by means of rendering the disparities u1 of the Dimetrodon dataset. The color coded flow (cf. Fig. 1(a)) is superimposed as texture. Based on the two observations that motion discontinuities often occur along object boundaries and that in turn object boundaries often coincide
Manuel Werlberger, Werner Trobin, Thomas Pock, Andreas Wedel, Daniel Cremers, Horst Bischof
BMVC3
2009 A convex relaxation approach for computing minimal partitions
abstract
In this work we propose a convex relaxation approach for computing minimal partitions. Our approach is based on rewriting the minimal partition problem (also known as Potts model) in terms of a primal dual Total Variation functional. We show that the Potts prior can be incorporated by means of convex constraints on the dual variables. For minimization we propose an efficient primal dual projected gradient algorithm which also allows a fast implementation on parallel hardware. Although our approach does not guarantee to find global minimizers of the Potts model we can give a tight bound on the energy between the computed solution and the true minimizer. Furthermore we show that our relaxation approach dominates recently proposed relaxations. As a consequence, our approach allows to compute solutions closer to the true minimizer. For many practical problems we even find the global minimizer. We demonstrate the excellent performance of our approach on several multi-label image segmentation and stereo problems.
Thomas Pock, Antonin Chambolle, Daniel Cremers, Horst Bischof
CVPR1
2009 An algorithm for minimizing the Mumford-Shah functional
abstract
In this work we revisit the Mumford-Shah functional, one of the most studied variational approaches to image segmentation. The contribution of this paper is to propose an algorithm which allows to minimize a convex relaxation of the Mumford-Shah functional obtained by functional lifting. The algorithm is an efficient primal-dual projection algorithm for which we prove convergence. In contrast to existing algorithms for minimizing the full Mumford-Shah this is the first one which is based on a convex relaxation. As a consequence the computed solutions are independent of the initialization. Experimental results confirm that the proposed algorithm determines smooth approximations while preserving discontinuities of the underlying signal.
Thomas Pock, Daniel Cremers, Horst Bischof, Antonin Chambolle
ICCV1
2009 Large displacement optical flow computation withoutwarping
abstract
We propose an algorithm for large displacement optical flow estimation which does not require the commonly used coarse-to-fine warping strategy. It is based on a quadratic relaxation of the optical flow functional which decouples data term and regularizer in such a way that the non-linearized variational problem can be solved by an alternation of two globally optimal steps, one imposing optimal data consistency, the other imposing discontinuity-preserving regularity of the flow field. Experimental results confirm that the proposed algorithmic implementation outperforms the traditional warping strategy, in particular for the case of large displacements of small scale structures.
Frank Steinbrücker, Thomas Pock, Daniel Cremers
ICCV2
2009 Structure- and motion-adaptive regularization for high accuracy optic flow
abstract
The accurate estimation of motion in image sequences is of central importance to numerous computer vision applications. Most competitive algorithms compute flow fields by minimizing an energy made of a data and a regularity term. To date, the best performing methods rely on rather simple purely geometric regularizes favoring smooth motion. In this paper, we revisit regularization and show that appropriate adaptive regularization substantially improves the accuracy of estimated motion fields. In particular, we systematically evaluate regularizes which adoptively favor rigid body motion (if supported by the image data) and motion field discontinuities that coincide with discontinuities of the image structure. The proposed algorithm relies on sequential convex optimization, is real-time capable and outperforms all previously published algorithms by more than one average rank on the Middlebury optic flow benchmark.
Andreas Wedel, Daniel Cremers, Thomas Pock, Horst Bischof
ICCV3
2008 TVSeg - Interactive Total Variation Based Image Segmentation
abstract
Interactive object extraction is an important part in any image editing software. We present a two step segmentation algorithm that first obtains a binary segmentation and then applies matting on the border regions to obtain a smooth alpha channel. The proposed segmentation algorithm is based on the minimization of the Geodesic Active Contour energy. A fast Total Variation minimization algorithm is used to find the globally optimal solution. We show how user interaction can be incorporated and outline an efficient way to exploit color information. A novel matting approach, based on energy minimization, is presented. Experimental evaluations are discussed, and the algorithm is compared to state of the art object extraction algorithms. The GPU based binaries are available online.
Markus Unger, Thomas Pock, Werner Trobin, Daniel Cremers, Horst Bischof
BMVC2
2008 A Convex Formulation of Continuous Multi-label Problems
Thomas Pock, Thomas Schoenemann, Gottfried Munda, Horst Bischof, Daniel Cremers
ECCV (3)1
2008 Continuous Energy Minimization Via Repeated Binary Fusion
Werner Trobin, Thomas Pock, Daniel Cremers, Horst Bischof
ECCV (4)2
2007 Mumford-Shah Meets Stereo: Integration of Weak Depth Hypotheses
abstract
Recent results on stereo indicate that an accurate segmentation is crucial for obtaining faithful depth maps. Variational methods have successfully been applied to both image segmentation and computational stereo. In this paper we propose a combination in a unified framework. In particular, we use a Mumford-Shah-like functional to compute a piecewise smooth depth map of a stereo pair. Our approach has two novel features: First, the regularization term of the functional combines edge information obtained from the color segmentation with flow-driven depth discontinuities emerging during the optimization procedure. Second, we propose a robust data term which adoptively selects the best matches obtained from different weak stereo algorithms. We integrate these features in a theoretically consistent framework. The final depth map is the minimizer of the energy functional, which can be solved by the associated functional derivatives. The underlying numerical scheme allows an efficient implementation on modern graphics hardware. We illustrate the performance of our algorithm using the Middlebury database as well as on real imagery.
Thomas Pock, Christopher Zach, Horst Bischof
CVPR1
2007 A Globally Optimal Algorithm for Robust TV-L1 Range Image Integration
abstract
Robust integration of range images is an important task for building high-quality 3D models. Since range images, and in particular range maps from stereo vision, may have a substantial amount of outliers, any integration approach aiming at high-quality models needs an increased level of robustness. Additionally, a certain level of regularization is required to obtain smooth surfaces. Computational efficiency and global convergence are further preferable properties. The contribution of this paper is a unified framework to solve all these issues. Our method is based on minimizing an energy functional consisting of a total variation (TV) regularization force and an L1 data fidelity term. We present a novel and efficient numerical scheme, which combines the duality principle for the TV term with a point-wise optimization step. We demonstrate the superior performance of our algorithm on the well-known Middlebury multi-view database and additionally on real-world multi-view images.
Christopher Zach, Thomas Pock, Horst Bischof
ICCV2
2007 A Duality Based Algorithm for TV- L 1-Optical-Flow Image Registration
Thomas Pock, Martin Urschler, Christopher Zach, Reinhard Beichel, Horst Bischof
MICCAI (2)1
2007 Algorithmic Differentiation: Application to Variational Problems in Computer Vision
abstract
Many vision problems can be formulated as minimization of appropriate energy functionals. These energy functionals are usually minimized, based on the calculus of variations (Euler-Lagrange equation). Once the Euler-Lagrange equation has been determined, it needs to be discretized in order to implement it on a digital computer. This is not a trivial task and, is moreover, error-prone. In this paper, we propose a flexible alternative. We discretize the energy functional and, subsequently, apply the mathematical concept of algorithmic differentiation to directly derive algorithms that implement the energy functional's derivatives. This approach has several advantages: First, the computed derivatives are exact with respect to the implementation of the energy functional. Second, it is basically straightforward to compute second-order derivatives and, thus, the Hessian matrix of the energy functional. Third, algorithmic differentiation is a process which can be automated. We demonstrate this novel approach on three representative vision problems (namely, denoising, segmentation, and stereo) and show that state-of-the-art results are obtained with little effort.
Thomas Pock, Michael Pock, Horst Bischof
IEEE Trans. Pattern Anal. Mach. Intell.1