Christopher Zach

dblp:93/4824 · DBLP profile ↗
← Back
84ranked-venue papers
21as first author
19since 2021 · last 2026
0000-0003-2840-6187ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 66 · 18 first-author · 14 since 2021Artificial intelligence and machine learning · 65 · 20 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 5Human-computer interaction and ubiquitous computing · 3Systems, architecture and hardware · 2
YearPublicationVenuePosition
2026 Beyond Texture: Advanced Facial Privacy Protection via Hierarchical Diffusion Autoencoder
Ting-Yi Lu, Che-Tsung Lin, Christopher Zach, Shang-Hong Lai
ICPR (7)3
2025 Certifiably Optimal Anisotropic Rotation Averaging
Carl Olsson, Yaroslava Lochman, Johan Malmport, Christopher Zach
ICCV4
2025 Text in the dark: Extremely low-light text image enhancement
Che-Tsung Lin, Chun Chet Ng, Zhi Qin Tan, Wan Jun Nah, Xinyu Wang 0010, Jie-Long Kew, Po-Hao Hsu, Shang-Hong Lai, Chee Seng Chan, Christopher Zach
Signal Process. Image Commun.10
2024 Learned Trajectory Embedding for Subspace Clustering
abstract
Clustering multiple motions from observed point trajectories is a fundamental task in understanding dynamic scenes. Most motion models require multiple tracks to estimate their parameters, hence identifying clusters when multiple motions are observed is a very challenging task. This is even aggravated for high-dimensional motion models. The starting point of our work is that this high-dimensionality of motion model can actually be leveraged to our advantage as sufficiently long trajectories identify the underlying motion uniquely in practice. Consequently, we propose to learn a mapping from trajectories to embedding vectors that represent the generating motion. The obtained trajectory embeddings are useful for clustering multiple observed motions, but are also trained to contain sufficient information to recover the parameters of the underlying motion by utilizing a geometric loss. We therefore are able to use only weak supervision from given motion segmentation to train this mapping. The entire algorithm consisting of trajectory embedding, clustering and motion parameter estimation is highly efficient. We conduct experiments on the Hopkins155, Hopkins12, and KT3DMoSeg datasets and show state-of-the-art performance of our proposed method for trajectory-based motion segmentation on full sequences and its competitiveness on the occluded sequences. Project page: https://ylochman.github.io/trajectory-embedding.
Yaroslava Lochman, Carl Olsson, Christopher Zach
CVPR3
2024 TAG: Text Prompt Augmentation for Zero-Shot Out-of-Distribution Detection
Xixi Liu 0001, Christopher Zach
ECCV (73)2
2024 Two Tales of Single-Phase Contrastive Hebbian Learning
abstract
The search for "biologically plausible" learning algorithms has converged on the idea of representing gradients as activity differences. However, most approaches require a high degree of synchronization (distinct phases during learning) and introduce substantial computational overhead, which raises doubts regarding their biological plausibility as well as their potential utility for neuromorphic computing. Furthermore, they commonly rely on applying infinitesimal perturbations (nudges) to output units, which is impractical in noisy environments. Recently it has been shown that by modelling artificial neurons as dyads with two oppositely nudged compartments, it is possible for a fully local learning algorithm named ``dual propagation'' to bridge the performance gap to backpropagation, without requiring separate learning phases or infinitesimal nudging. However, the algorithm has the drawback that its numerical stability relies on symmetric nudging, which may be restrictive in biological and analog implementations. In this work we first provide a solid foundation for the objective underlying the dual propagation method, which also reveals a surpising connection with adversarial robustness. Second, we demonstrate how dual propagation is related to a particular adjoint state method, which is stable regardless of asymmetric nudging.
Rasmus Kjær Høier, Christopher Zach
ICML2
2024 When IC meets text: Towards a rich annotated integrated circuit text dataset
Chun Chet Ng, Che-Tsung Lin, Zhi Qin Tan, Xinyu Wang 0010, Jie-Long Kew, Chee Seng Chan, Christopher Zach
Pattern Recognit.7
2023 GEN: Pushing the Limits of Softmax-Based Out-of-Distribution Detection
abstract
Out-of-distribution (OOD) detection has been exten-sively studied in order to successfully deploy neural networks, in particular, for safety-critical applications. More-over, performing OOD detection on large-scale datasets is closer to reality, but is also more challenging. Sev-eral approaches need to either access the training data for score design or expose models to outliers during training. Some post-hoc methods are able to avoid the afore-mentioned constraints, but are less competitive. In this work, we propose Generalized ENtropy score (GEN), a simple but effective entropy-based score function, which can be applied to any pre-trained softmax-based classifier. Its performance is demonstrated on the large-scale ImageNet-lk OOD detection benchmark. It consistently improves the average AUROC across six commonly-used CNN-based and visual transformer classifiers over a num-ber of state-of-the-art post-hoc methods. The average AU- ROC improvement is at least 3.5%. Furthermore, we used GEN on top of feature-based enhancing methods as well as methods using training statistics to further improve the OOD detection performance. The code is available at: https://github.com/XixiLiu95/GEN.
Xixi Liu 0001, Yaroslava Lochman, Christopher Zach
CVPR3
2023 Rethinking Long-Tailed Visual Recognition with Dynamic Probability Smoothing and Frequency Weighted Focusing
abstract
Deep learning models trained on long-tailed (LT) datasets often exhibit bias towards head classes with high frequency. This paper highlights the limitations of existing solutions that combine class- and instance-level re-weighting loss in a naive manner. Specifically, we demonstrate that such solutions result in overfitting the training set, significantly impacting the rare classes. To address this issue, we propose a novel loss function that dynamically reduces the influence of outliers and assigns class-dependent focusing parameters. We also introduce a new long-tailed dataset, ICText-LT, featuring various image qualities and greater realism than artificially sampled datasets. Our method has proven effective, outperforming existing methods through superior quantitative results on CIFAR-LT, Tiny ImageNet-LT, and our new ICText-LT datasets. The source code and new dataset are available at https://github.com/nwjun/FFDS-Loss.
Wan Jun Nah, Chun Chet Ng, Che-Tsung Lin, Yeong Khang Lee, Jie-Long Kew, Zhi Qin Tan, Chee Seng Chan, Christopher Zach, Shang-Hong Lai
ICIP8
2023 Dual Propagation: Accelerating Contrastive Hebbian Learning with Dyadic Neurons
abstract
Activity difference based learning algorithms—such as contrastive Hebbian learning and equilibrium propagation—have been proposed as biologically plausible alternatives to error back-propagation. However, on traditional digital chips these algorithms suffer from having to solve a costly inference problem twice, making these approaches more than two orders of magnitude slower than back-propagation. In the analog realm equilibrium propagation may be promising for fast and energy efficient learning, but states still need to be inferred and stored twice. Inspired by lifted neural networks and compartmental neuron models we propose a simple energy based compartmental neuron model, termed dual propagation, in which each neuron is a dyad with two intrinsic states. At inference time these intrinsic states encode the error/activity duality through their difference and their mean respectively. The advantage of this method is that only a single inference phase is needed and that inference can be solved in layerwise closed-form. Experimentally we show on common computer vision datasets, including Imagenet32x32, that dual propagation performs equivalently to back-propagation both in terms of accuracy and runtime.
Rasmus Kjær Høier, D. Staudt, Christopher Zach
ICML3
2023 Decentralized Training of 3D Lane Detection with Automatic Labeling Using HD Maps
abstract
To have competent 3D lane detection for real-world driving, a massive amount of data from all over the world is needed, but data collection and manual annotation are costly and time-consuming. The diversity of data collected by developmental cars might still be limited compared to the data collected by a large fleet of customer cars. Federated learning enables training models on edge without transferring data out of devices. However, training supervised learning tasks at the edge is directly tied to having access to high-quality labels, which is limited at the edge.In this paper, we propose a fully automatic method to generate 3D lane labels at the edge using a pre-recorded HD map to enable the federated training of the 3D lane detection model. As a reference, a semi-automatic method is applied for creating a 3D-lane dataset used as ground truth. Our experimental results show that the model can achieve comparable performance when training on the same dataset in both a centralized and a decentralized manner. And the models trained on semi-automatic labeled datasets slightly outperform those trained on fully-automatically labeled datasets. This study shows that a well-performing 3D lane detection model can be trained in a supervised and fully decentralized manner, and most importantly, data privacy at the edge is guaranteed.
Yadong Mao, Zhuqi Xiao, Che-Tsung Lin, Pedro Porto Buarque de Gusmão, Nicholas D. Lane, Christopher Zach, Mina Alibeigi
VTC2023-Spring6
2023 Cycle-object consistency for image-to-image domain adaptation
Che-Tsung Lin, Jie-Long Kew, Chee Seng Chan, Shang-Hong Lai, Christopher Zach
Pattern Recognit.5
2022 AdaSTE: An Adaptive Straight-Through Estimator to Train Binary Neural Networks
abstract
We propose a new algorithm for training deep neural networks (DNNs) with binary weights. In particular, we first cast the problem of training binary neural networks (BiNNs) as a bilevel optimization instance and subsequently construct flexible relaxations of this bilevel program. The resulting training method shares its algorithmic simplicity with several existing approaches to train BiNNs, in particular with the straight-through gradient estimator successfully employed in BinaryConnect and subsequent methods. Infact, our proposed method can be interpreted as an adaptive variant of the original straight-through estimator that conditionally (but not always) acts like a linear mapping in the backward pass of error propagation. Experimental results demonstrate that our new algorithm offers favorable performance compared to existing approaches.11This work was partially supported by theWallenberg AI, Autonomous Systems and Software Program (WASP) funded by the Knut and Alice Wallenberg Foundation.
Huu Le, Rasmus Kjær Høier, Che-Tsung Lin, Christopher Zach
CVPR4
2022 CyEDA: Cycle-Object Edge Consistency Domain Adaptation
abstract
A difficulty of global-level translation is to preserve instance-level details in an image. Although some instance level translation methods can retain the details, most of them require either pre-trained object detection/segmentation network or annotation labels. In this work, we propose a novel method namely CyEDA to perform global level domain adaptation that can preserve image contents without any pre-trained networks integration or annotation labels. Specifically, we introduce blending masks and cycle-object edge consistency loss which exploit the preservation of image objects. We show that our approach can outperform other SOTAs in terms of image quality and FID score in both BDD100K and GTA datasets. The code and pre-trained models are publicly available at https://github.com/bjc1999/CyEDA.
Jing Chong Beh, Kam Woh Ng, Jie-Long Kew, Che-Tsung Lin, Chee Seng Chan, Shang-Hong Lai, Christopher Zach
ICIP7
2022 Extremely Low-Light Image Enhancement with Scene Text Restoration
abstract
Deep learning-based methods have made impressive progress in enhancing extremely low-light images - the image quality of the reconstructed images has generally improved. However, we found out that most of these methods could not sufficiently recover the image details, for instance, the texts in the scene. In this paper, a novel image enhancement framework is proposed to precisely restore the scene texts, as well as the overall quality of the image simultaneously under extremely low-light conditions. Mainly, we employed a self-regularised attention map, an edge map, and a novel text detection loss. In addition, leveraging the synthetic low-light images is beneficial for image enhancement on the genuine ones in terms of text detection. The quantitative and qualitative experimental results have shown that the proposed model outperforms state-of-the-art methods in image restoration, text detection, and text spotting on See In the Dark and ICDAR15 datasets.
Po-Hao Hsu, Che-Tsung Lin, Chun Chet Ng, Jie-Long Kew, Mei Yih Tan, Shang-Hong Lai, Chee Seng Chan, Christopher Zach
ICPR8
2022 Energy-based Models for Deep Probabilistic Regression
abstract
It is desirable that a deep neural network trained on a regression task not only achieves high prediction accuracy, but its prediction posteriors are also well-calibrated, especially in safety-critical settings. Recently, energy-based models specifically to enrich regression posteriors have been proposed and achieve state-of-art results in object detection tasks. However, applying these models at prediction time is not straightforward as the resulting inference methods require to minimize an underlying energy function. Furthermore, these methods empirically do not provide accurate prediction uncertainties. Inspired by recent joint energy-based models for classification, in this work, we propose to utilize a joint energy model for regression tasks and describe architectural differences needed in this setting. Within this framework, we apply our methods to three computer vision regression tasks. We demonstrate that joint energy-based models for deep probabilistic regression improve the calibration property, do not require expensive inference, and yield competitive accuracy in terms of the mean absolute error (MAE).
Xixi Liu 0001, Che-Tsung Lin, Christopher Zach
ICPR3
2022 Effortless Training of Joint Energy-Based Models with Sliced Score Matching
abstract
Standard discriminative classifiers can be upgraded to joint energy-based models (JEMs) by combining the classification loss with a log-evidence loss. Hence, such models intrinsically allow detection of out-of-distribution (OOD) samples, and empirically also provide better calibrated posteriors, i.e. prediction uncertainties. However, the training procedure suggested for JEMs (using stochastic gradient Langevin dynamics—or SGLD— to maximize the evidence) is reported to be brittle. In this work we propose to utilize score matching—in particular sliced score matching—to obtain a stable training method for JEMs. We observe empirically that the combination of score matching with the standard classification loss leads to improved OOD detection and better calibrated classifiers for otherwise identical DNN architectures. Additionally, we also analyze the impact of replacing the regular soft-max layer for classification with a gated soft-max one in order to improve the intrinsic transformation invariance and generalization ability.1
Xixi Liu 0001, D. Staudt, Che-Tsung Lin, Christopher Zach
ICPR4
2021 Robust Fitting with Truncated Least Squares: A Bilevel Optimization Approach
abstract
We tackle the problem of large-scale robust fitting using the truncated least squares (TLS) loss. Existing approaches commonly optimize this loss by employing a smooth surrogate, which allows the problem to be solved using well-known methods such as Iteratively Re-weighted Least Squares (IRLS). In this work, we present a new approach to optimize the TLS objective, where we propose to reformulate the original problem as a bi-level program. Then, by applying the Optimal Value Reformulation (OVR) technique to this new formulation, we derive a penalty approach to solve for the best fitting models, where the penalty parameters can be adaptively computed. Our final algorithm can be considered as a special instance of IRLS. As a result, we can incorporate our new algorithm into existing IRLS solvers, where we only need to modify the weight evaluation procedure. Our experimental results show promising results on several instances of large-scale bundle adjustment and non-linear refinement for essential matrix fitting.
Huu Le, Christopher Zach
3DV2
2021 BabelCalib: A Universal Approach to Calibrating Central Cameras
abstract
Existing calibration methods occasionally fail for large field-of-view cameras due to the non-linearity of the underlying problem and the lack of good initial values for all parameters of the used camera model. This might occur because a simpler projection model is assumed in an initial step, or a poor initial guess for the internal parameters is pre-defined. A lot of the difficulties of general camera calibration lie in the use of a forward projection model. We side-step these challenges by first proposing a solver to calibrate the parameters in terms of a back-projection model and then regress the parameters for a target forward model. These steps are incorporated in a robust estimation framework to cope with outlying detections. Extensive experiments demonstrate that our approach is very reliable and returns the most accurate calibration parameters as measured on the downstream task of absolute pose estimation on test sets. The code is released at https://github.com/ylochman/babelcalib.
Yaroslava Lochman, Kostiantyn Liepieshov, Michal Perdoch, Christopher Zach, James Pritts
ICCV5
2020 Progressive Batching for Efficient Non-linear Least Squares
Huu Le, Christopher Zach, Edward Rosten, Oliver J. Woodford
ACCV (3)2
2020 Lifted Regression/Reconstruction Networks
Rasmus Kjær Høier, Christopher Zach
BMVC2
2020 A Graduated Filter Method for Large Scale Robust Estimation
abstract
Due to the highly non-convex nature of large-scale robust parameter estimation, avoiding poor local minima is challenging in real-world applications where input data is contaminated by a large or unknown fraction of outliers. In this paper, we introduce a novel solver for robust estimation that possesses a strong ability to escape poor local minima. Our algorithm is built upon the class of traditional graduated optimization techniques, which are considered state-of-the-art local methods to solve problems having many poor minima. The novelty of our work lies in the introduction of an adaptive kernel (or residual) scaling scheme, which allows us to achieve faster convergence rates. Like other existing methods that aim to return good local minima for robust estimation tasks, our method relaxes the original robust problem, but adapts a filter framework from non-linear constrained optimization to automatically choose the level of relaxation. Experimental results on real large-scale datasets such as bundle adjustment instances demonstrate that our proposed method achieves competitive results.
Huu Le, Christopher Zach
CVPR2
2020 SG-VAE: Scene Grammar Variational Autoencoder to Generate New Indoor Scenes
Pulak Purkait, Christopher Zach, Ian D. Reid 0001
ECCV (24)2
2020 Truncated Inference for Latent Variable Optimization Problems: Application to Robust Estimation and Learning
Christopher Zach, Huu Le
ECCV (26)1
2019 Contrastive Learning for Lifted Networks
Christopher Zach, Virginia Estellers
BMVC1
2019 Pareto Meets Huber: Efficiently Avoiding Poor Minima in Robust Estimation
abstract
Robust cost optimization is the task of fitting parameters to data points containing outliers. In particular, we focus on large-scale computer vision problems, such as bundle adjustment, where Non-Linear Least Square (NLLS) solvers are the current workhorse. In this context, NLLS-based state of the art algorithms have been designed either to quickly improve the target objective and find a local minimum close to the initial value of the parameters, or to have a strong ability to escape poor local minima. In this paper, we propose a novel algorithm relying on multi-objective optimization which allows to match those two properties. We experimentally demonstrate that our algorithm has an ability to escape poor local minima that is on par with the best performing algorithms with a faster decrease of the target objective.
Christopher Zach, Guillaume Bourmaud
ICCV1
2019 Seeing Behind Things: Extending Semantic Segmentation to Occluded Regions
abstract
Semantic segmentation and instance level segmentation made substantial progress in recent years due to the emergence of deep neural networks (DNNs). A number of deep architectures with Convolution Neural Networks (CNNs) were proposed that surpass the traditional machine learning approaches for segmentation by a large margin. These architectures predict the directly observable semantic category of each pixel by usually optimizing a cross-entropy loss. In this work we push the limit of semantic segmentation towards predicting semantic labels of directly visible as well as occluded objects or objects parts, where the network's input is a single depth image. We group the semantic categories into one background and multiple foreground object groups, and we propose a modification of the standard cross-entropy loss to cope with the settings. In our experiments we demonstrate that a CNN trained by minimizing the proposed loss is able to predict semantic categories for visible and occluded object parts without requiring to increase the network size (compared to a standard segmentation task). The results are validated on a newly generated dataset (augmented from SUNCG) dataset.
Pulak Purkait, Christopher Zach, Ian D. Reid 0001
IROS2
2018 Weakly Supervised Learning of Indoor Geometry by Dual Warping
abstract
A major element of depth perception and 3D understanding is the ability to predict the 3D layout of a scene and its contained objects for a novel pose. Indoor environments are particularly suitable for novel view prediction, since the set of objects in such environments is relatively restricted. In this work we address the task of 3D prediction especially for indoor scenes by leveraging only weak supervision. In the literature 3D scene prediction is usually solved via a 3D voxel grid. However, such methods are limited to estimating rather coarse 3D voxel grids, since predicting entire voxel spaces has large computational costs. Hence, our method operates in image-space rather than in voxel space, and the task of 3D estimation essentially becomes a depth image completion problem. We propose a novel approach to easily generate training data containing depth maps with realistic occlusions, and subsequently train a network for completing those occluded regions. Using multiple publicly available datasets we benchmark our method against existing approaches and are able to obtain superior performance. We further demonstrate the flexibility of our method by presenting results for new view synthesis of RGB-D images.
Pulak Purkait, Ujwal Bonde, Christopher Zach
3DV3
2018 ContextNet: Exploring Context and Detail for Semantic Segmentation in Real-time
Rudra P. K. Poudel, Ujwal Bonde, Stephan Liwicki, Christopher Zach
BMVC4
2018 Synthetic View Generation for Absolute Pose Regression and Image Synthesis
Pulak Purkait, Cheng Zhao 0002, Christopher Zach
BMVC3
2018 Multiplicative vs. Additive Half-Quadratic Minimization for Robust Cost Optimization
Christopher Zach, Guillaume Bourmaud
BMVC1
2018 pOSE: Pseudo Object Space Error for Initialization-Free Bundle Adjustment
abstract
Bundle adjustment is a nonlinear refinement method for camera poses and 3D structure requiring sufficiently good initialization. In recent years, it was experimentally observed that useful minima can be reached even from arbitrary initialization for affine bundle adjustment problems (and fixed-rank matrix factorization instances in general). The key success factor lies in the use of the variable projection (VarPro) method, which is known to have a wide basin of convergence for such problems. In this paper, we propose the Pseudo Object Space Error (pOSE), which is an objective with cameras represented as a hybrid between the affine and projective models. This formulation allows us to obtain 3D reconstructions that are close to the true projective reconstructions while retaining a bilinear problem structure suitable for the VarPro method. Experimental results show that using pOSE has a high success rate to yield faithful 3D reconstructions from random initializations, taking one step towards initialization-free structure from motion.
Je Hyeong Hong, Christopher Zach
CVPR2
2018 Descending, Lifting or Smoothing: Secrets of Robust Cost Optimization
Christopher Zach, Guillaume Bourmaud
ECCV (12)1
2018 Minimal Solvers for Monocular Rolling Shutter Compensation Under Ackermann Motion
abstract
Modern automotive vehicles are often equipped with a budget commercial rolling shutter camera. These devices often produce distorted images due to the inter-row delay of the camera while capturing the image. Recent methods for monocular rolling shutter motion compensation utilize blur kernel and the straightness property of line segments. However, these methods are limited to handling rotational motion and also are not fast enough to operate in real time. In this paper, we propose a minimal solver for the rolling shutter motion compensation which assumes known vertical direction of the camera. Thanks to the Ackermann motion model of vehicles which consists of only two motion parameters, and two parameters for the simplified depth assumption that lead to a 4-line algorithm. The proposed minimal solver estimates the rolling shutter camera motion efficiently and accurately. The extensive experiments on real and simulated datasets demonstrate the benefits of our approach in terms of qualitative and quantitative results.
Pulak Purkait, Christopher Zach
WACV2
2018 Generalized fusion moves for continuous label optimization
Christopher Zach
Comput. Vis. Image Underst.1
2017 Scale Exploiting Minimal Solvers for Relative Pose with Calibrated Cameras
Stephan Liwicki, Christopher Zach
BMVC2
2017 Iterated Lifting for Robust Cost Optimization
Christopher Zach, Guillaume Bourmaud
BMVC1
2017 Revisiting the Variable Projection Method for Separable Nonlinear Least Squares Problems
abstract
Variable Projection (VarPro) is a framework to solve optimization problems efficiently by optimally eliminating a subset of the unknowns. It is in particular adapted for Separable Nonlinear Least Squares (SNLS) problems, a class of optimization problems including low-rank matrix factorization with missing data and affine bundle adjustment as instances. VarPro-based methods have received much attention over the last decade due to the experimentally observed large convergence basin for certain problem classes, where they have a clear advantage over standard methods based on Joint optimization over all unknowns. Yet no clear answers have been found in the literature as to why VarPro outperforms others and why Joint optimization, which has been successful in solving many computer vision tasks, fails on this type of problems. Also, the fact that VarPro has been mainly tested on small to medium-sized datasets has raised questions about its scalability. This paper intends to address these unsolved puzzles.
Je Hyeong Hong, Christopher Zach, Andrew W. Fitzgibbon
CVPR2
2017 Rolling Shutter Correction in Manhattan World
abstract
A vast majority of consumer cameras operate the rolling shutter mechanism, which often produces distorted images due to inter-row delay while capturing an image. Recent methods for monocular rolling shutter compensation utilize blur kernel, straightness of line segments, as well as angle and length preservation. However, they do not incorporate scene geometry explicitly for rolling shutter correction, therefore, information about the 3D scene geometry is often distorted by the correction process. In this paper we propose a novel method which leverages geometric properties of the scene-in particular vanishing directions-to estimate the camera motion during rolling shutter exposure from a single distorted image. The proposed method jointly estimates the orthogonal vanishing directions and the rolling shutter camera motion. We performed extensive experiments on synthetic and real datasets which demonstrate the benefits of our approach both in terms of qualitative and quantitative results (in terms of a geometric structure fitting) as well as with respect to computation time.
Pulak Purkait, Christopher Zach, Ales Leonardis
ICCV2
2017 Dense Semantic 3D Reconstruction
abstract
Both image segmentation and dense 3D modeling from images represent an intrinsically ill-posed problem. Strong regularizers are therefore required to constrain the solutions from being 'too noisy'. These priors generally yield overly smooth reconstructions and/or segmentations in certain regions while they fail to constrain the solution sufficiently in other areas. In this paper, we argue that image segmentation and dense 3D reconstruction contribute valuable information to each other's task. As a consequence, we propose a mathematical framework to formulate and solve a joint segmentation and dense reconstruction problem. On the one hand knowing about the semantic class of the geometry provides information about the likelihood of the surface direction. On the other hand the surface direction provides information about the likelihood of the semantic class. Experimental results on several data sets highlight the advantages of our joint formulation. We show how weakly observed surfaces are reconstructed more faithfully compared to a geometry only reconstruction. Thanks to the volumetric nature of our formulation we also infer surfaces which cannot be directly observed for example the surface between the ground and a building. Finally, our method returns a semantic segmentation which is consistent across the whole dataset.
Christian Häne, Christopher Zach, Andrea Cohen, Marc Pollefeys
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 Generalized Fusion Moves for Continuous Label Optimization
Christopher Zach
ACCV (5)1
2016 Projective Bundle Adjustment from Arbitrary Initialization Using the Variable Projection Method
Je Hyeong Hong, Christopher Zach, Andrew W. Fitzgibbon, Roberto Cipolla
ECCV (1)2
2016 Coarse-to-fine Planar Regularization for Dense Monocular Depth Estimation
Stephan Liwicki, Christopher Zach, Ondrej Miksik, Philip Torr 0001
ECCV (2)2
2015 A dynamic programming approach for fast and robust object pose recognition from range images
abstract
Joint object recognition and pose estimation solely from range images is an important task e.g. in robotics applications and in automated manufacturing environments. The lack of color information and limitations of current commodity depth sensors make this task a challenging computer vision problem, and a standard random sampling based approach is prohibitively time-consuming. We propose to address this difficult problem by generating promising inlier sets for pose estimation by early rejection of clear outliers with the help of local belief propagation (or dynamic programming). By exploiting data-parallelism our method is fast, and we also do not rely on a computationally expensive training phase. We demonstrate state-of-the art performance on a standard dataset and illustrate our approach on challenging real sequences.
Christopher Zach, Adrián Peñate Sánchez, Minh-Tri Pham
CVPR1
2015 The Likelihood-Ratio Test and Efficient Robust Estimation
abstract
Robust estimation of model parameters in the presence of outliers is a key problem in computer vision. RANSAC inspired techniques are widely used in this context, although their application might be limited due to the need of a priori knowledge on the inlier noise level. We propose a new approach for jointly optimizing over model parameters and the inlier noise level based on the likelihood ratio test. This allows control over the type I error incurred. We also propose an early bailout strategy for efficiency. Tests on both synthetic and real data show that our method outperforms the state-of-the-art in a fraction of the time.
Andrea Cohen, Christopher Zach
ICCV2
2014 Variational Regularization and Fusion of Surface Normal Maps
abstract
In this work we propose an optimization scheme for variational, vectorial denoising and fusion of surface normal maps. These are common outputs of shape from shading, photometric stereo or single image reconstruction methods, but tend to be noisy and request post-processing for further usage. Processing of normals maps, which do not provide knowledge about the underlying scene depth, is complicated due to their unit length constraint which renders the optimization non-linear and non-convex. The presented approach builds upon a linearization of the constraint to obtain a convex relaxation, while guaranteeing convergence. Experimental results demonstrate that our algorithm generates more consistent representations from estimated and potentially complementary normal maps.
Bernhard Zeisl, Christopher Zach, Marc Pollefeys
3DV2
2014 RetroDepth: 3D silhouette sensing for high-precision input on and above physical surfaces
abstract
We present RetroDepth, a new vision-based system for accurately sensing the 3D silhouettes of hands, styluses, and other objects, as they interact on and above physical surfaces. Our setup is simple, cheap, and easily reproducible, comprising of two infrared cameras, diffuse infrared LEDs, and any off-the-shelf retro-reflective material. The retro-reflector aids image segmentation, creating a strong contrast between the surface and any object in proximity. A new highly efficient stereo matching algorithm precisely estimates the 3D contours of interacting objects and the retro-reflective surfaces. A novel pipeline enables 3D finger, hand and object tracking, as well as gesture recognition, purely using these 3D contours. We demonstrate high-precision sensing, allowing robust disambiguation between a finger or stylus touching, pressing or interacting above the surface. This allows many interactive scenarios that seamlessly mix together freehand 3D interactions with touch, pressure and stylus input. As shown, these rich modalities of input are enabled on and above any retro-reflective surface, including custom "physical widgets" fabricated by users. We compare our system with Kinect and Leap Motion, and conclude with limitations and future work.
David Kim 0002, Shahram Izadi, Jakub Dostal, Christoph Rhemann, Cem Keskin, Christopher Zach, Jamie Shotton, Timothy A. Large, Steven Bathiche, Matthias Nießner, Alex Butler, Sean Ryan Fanello, Vivek Pradeep
CHI6
2014 A Principled Approach for Coarse-to-Fine MAP Inference
abstract
In this work we reconsider labeling problems with (virtually) continuous state spaces, which are of relevance in low level computer vision. In order to cope with such huge state spaces multi-scale methods have been proposed to approximately solve such labeling tasks. Although performing well in many cases, these methods do usually not come with any guarantees on the returned solution. A general and principled approach to solve labeling problems is based on the well-known linear programming relaxation, which appears to be prohibitive for large state spaces at the first glance. We demonstrate that a coarse-to-fine exploration strategy in the label space is able to optimize the LP relaxation for non-trivial problem instances with reasonable run-times and moderate memory requirements.
Christopher Zach
CVPR1
2014 Robust Bundle Adjustment Revisited
Christopher Zach
ECCV (5)1
2014 Multi-modal registration for correlative microscopy using image analogies
Tian Cao 0001, Christopher Zach, Shannon Modla, Debbie Powell, Kirk Czymmek, Marc Niethammer
Medical Image Anal.2
2014 Automatic atlas-based three-label cartilage segmentation from MR knee images
Liang Shan 0001, Christopher Zach, Cecil Charles, Marc Niethammer
Medical Image Anal.2
2014 What Is Optimized in Convex Relaxations for Multilabel Problems: Connecting Discrete and Continuously Inspired MAP Inference
abstract
In this work, we present a unified view on Markov random fields (MRFs) and recently proposed continuous tight convex relaxations for multilabel assignment in the image plane. These relaxations are far less biased toward the grid geometry than Markov random fields on grids. It turns out that the continuous methods are nonlinear extensions of the well-established local polytope MRF relaxation. In view of this result, a better understanding of these tight convex relaxations in the discrete setting is obtained. Further, a wider range of optimization methods is now applicable to find a minimizer of the tight formulation. We propose two methods to improve the efficiency of minimization. One uses a weaker, but more efficient continuously inspired approach as initialization and gradually refines the energy where it is necessary. The other one reformulates the dual energy enabling smooth approximations to be used for efficient optimization. We demonstrate the utility of our proposed minimization schemes in numerical experiments. Finally, we generalize the underlying energy formulation from isotropic metric smoothness costs to arbitrary nonmetric and orientation dependent smoothness terms.
Christopher Zach, Christian Häne, Marc Pollefeys
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 Real-time non-rigid reconstruction using an RGB-D camera
abstract
We present a combined hardware and software solution for markerless reconstruction of non-rigidly deforming physical objects with arbitrary shape in real-time . Our system uses a single self-contained stereo camera unit built from off-the-shelf components and consumer graphics hardware to generate spatio-temporally coherent 3D models at 30 Hz. A new stereo matching algorithm estimates real-time RGB-D data. We start by scanning a smooth template model of the subject as they move rigidly. This geometric surface prior avoids strong scene assumptions, such as a kinematic human skeleton or a parametric shape model. Next, a novel GPU pipeline performs non-rigid registration of live RGB-D data to the smooth template using an extended non-linear as-rigid-as-possible (ARAP) framework. High-frequency details are fused onto the final mesh using a linear deformation model. The system is an order of magnitude faster than state-of-the-art methods, while matching the quality and robustness of many offline algorithms. We show precise real-time reconstructions of diverse scenes, including: large deformations of users' heads, hands, and upper bodies; fine-scale wrinkles and folds of skin and clothing; and non-rigid interactions performed by users on flexible objects such as toys. We demonstrate how acquired models can be used for many interactive scenarios, including re-texturing, online performance capture and preview, and real-time shape and motion re-targeting.
Michael Zollhöfer, Matthias Nießner, Shahram Izadi, Christoph Rhemann, Christopher Zach, Matthew Fisher, Chenglei Wu, Andrew W. Fitzgibbon, Charles T. Loop, Christian Theobalt, Marc Stamminger
ACM Trans. Graph.5
2013 Dual Decomposition for Joint Discrete-Continuous Optimization
abstract
We analyse convex formulations for combined discrete-continuous MAP inference using the dual decomposition method. As a consquence we can provide a more intuitive derivation for the resulting convex relaxation than presented in the literature. Further, we show how to strengthen the relaxation by reparametrizing the potentials, hence convex relaxations for discrete-continuous inference does not share an important feature of LP relaxations for discrete labeling problems: incorporating unary potentials into higher order ones affects the quality of the relaxation. We argue that the convex model for discrete-continuous inference is very general and can be used as alternative for alternation-based methods often employed for such joint inference tasks.
Christopher Zach
AISTATS1
2013 Joint 3D Scene Reconstruction and Class Segmentation
abstract
Both image segmentation and dense 3D modeling from images represent an intrinsically ill-posed problem. Strong regularizers are therefore required to constrain the solutions from being 'too noisy'. Unfortunately, these priors generally yield overly smooth reconstructions and/or segmentations in certain regions whereas they fail in other areas to constrain the solution sufficiently. In this paper we argue that image segmentation and dense 3D reconstruction contribute valuable information to each other's task. As a consequence, we propose a rigorous mathematical framework to formulate and solve a joint segmentation and dense reconstruction problem. Image segmentations provide geometric cues about which surface orientations are more likely to appear at a certain location in space whereas a dense 3D reconstruction yields a suitable regularization for the segmentation problem by lifting the labeling from 2D images to 3D space. We show how appearance-based cues and 3D surface orientation priors can be learned from training data and subsequently used for class-specific regularization. Experimental results on several real data sets highlight the advantages of our joint formulation.
Christian Häne, Christopher Zach, Andrea Cohen, Roland Angst, Marc Pollefeys
CVPR2
2013 Scene Coordinate Regression Forests for Camera Relocalization in RGB-D Images
abstract
We address the problem of inferring the pose of an RGB-D camera relative to a known 3D scene, given only a single acquired image. Our approach employs a regression forest that is capable of inferring an estimate of each pixel's correspondence to 3D points in the scene's world coordinate frame. The forest uses only simple depth and RGB pixel comparison features, and does not require the computation of feature descriptors. The forest is trained to be capable of predicting correspondences at any pixel, so no interest point detectors are required. The camera pose is inferred using a robust optimization scheme. This starts with an initial set of hypothesized camera poses, constructed by applying the forest at a small fraction of image pixels. Preemptive RANSAC then iterates sampling more pixels at which to evaluate the forest, counting inliers, and refining the hypothesized poses. We evaluate on several varied scenes captured with an RGB-D camera and observe that the proposed technique achieves highly accurate relocalization and substantially out-performs two state of the art baselines.
Jamie Shotton, Ben Glocker, Christopher Zach, Shahram Izadi, Antonio Criminisi, Andrew W. Fitzgibbon
CVPR3
2013 MonoFusion: Real-time 3D reconstruction of small scenes with a single web camera
abstract
MonoFusion allows a user to build dense 3D reconstructions of their environment in real-time, utilizing only a single, off-the-shelf web camera as the input sensor. The camera could be one already available in a tablet, phone, or a standalone device. No additional input hardware is required. This removes the need for power intensive active sensors that do not work robustly in natural outdoor lighting. Using the input stream of the camera we first estimate the 6DoF camera pose using a sparse tracking method. These poses are then used for efficient dense stereo matching between the input frame and a key frame (extracted previously). The resulting dense depth maps are directly fused into a voxel-based implicit model (using a computationally inexpensive method) and surfaces are extracted per frame. The system is able to recover from tracking failures as well as filter out geometrically inconsistent noise from the 3D reconstruction. Our method is both simple to implement and efficient, making such systems even more accessible. This paper details the algorithmic components that make up our system and a GPU implementation of our approach. Qualitative results demonstrate high quality reconstructions even visually comparable to active depth sensor-based systems such as KinectFusion.
Vivek Pradeep, Christoph Rhemann, Shahram Izadi, Christopher Zach, Michael Bleyer, Steven Bathiche
ISMAR4
2013 Segmentation with area constraints
Marc Niethammer, Christopher Zach
Medical Image Anal.2
2012 Unknown Radial Distortion Centers in Multiple View Geometry Problems
José Brito 0001, Roland Angst, Kevin Köser, Christopher Zach, Pedro Branco 0001, Manuel João Ferreira, Marc Pollefeys
ACCV (4)4
2012 One-sided Radial-Fundamental Matrix Estimation
abstract
For modern consumer cameras often approximate calibration data is available, mak-ing applications such as 3D reconstruction or photo registration easier as compared to the pure uncalibrated setting. In this paper we address the setting with calibrated-uncalibrated image pairs: for one image intrinsic parameters are assumed to be known, whereas the second view has unknown distortion and calibration parameters. This sit-uation arises e.g. when one would like to register archive imagery to recently taken photos. A commonly adopted strategy for determining epipolar geometry is based on feature matching and minimal solvers inside a RANSAC framework. However, only very few existing solutions apply to the calibrated-uncalibrated setting. We propose a simple and numerically stable two-step scheme to first estimate radial distortion parameters and subsequently the focal length using novel solvers. We demonstrate the performance on synthetic and real datasets. 1
José Brito 0001, Christopher Zach, Kevin Köser, Manuel João Ferreira, Marc Pollefeys
BMVC2
2012 Discovering and exploiting 3D symmetries in structure from motion
abstract
Many architectural scenes contain symmetric or repeated structures, which can generate erroneous image correspondences during structure from motion (Sfm) computation. Prior work has shown that the detection and removal of these incorrect matches is crucial for accurate and robust recovery of scene structure. In this paper, we point out that these incorrect matches, in fact, provide strong cues to the existence of symmetries and structural regularities in the unknown 3D structure. We make two key contributions. First, we propose a method to recover various symmetry relations in the structure using geometric and appearance cues. A set of structural constraints derived from the symmetries are imposed within a new constrained bundle adjustment formulation, where symmetry priors are also incorporated. Second, we show that the recovered symmetries enable us to choose a natural coordinate system for the 3D structure where gauge freedom in rotation is held fixed. Furthermore, based on the symmetries, 3D structure completion is also performed. Our approach significantly reduces drift through ”structural” loop closures and improves the accuracy of reconstructions in urban scenes.
Andrea Cohen, Christopher Zach, Sudipta N. Sinha, Marc Pollefeys
CVPR2
2012 What is optimized in tight convex relaxations for multi-label problems?
abstract
In this work we present a unified view on Markov random fields and recently proposed continuous tight convex relaxations for multi-label assignment in the image plane. These relaxations are far less biased towards the grid geometry than Markov random fields. It turns out that the continuous methods are non-linear extensions of the local polytope MRF relaxation. In view of this result a better understanding of these tight convex relaxations in the discrete setting is obtained. Further, a wider range of optimization methods is now applicable to find a minimizer of the tight formulation. We propose two methods to improve the efficiency of minimization. One uses a weaker, but more efficient continuously inspired approach as initialization and gradually refines the energy where it is necessary. The other one reformulates the dual energy enabling smooth approximations to be used for efficient optimization. We demonstrate the utility of our proposed minimization schemes in numerical experiments.
Christopher Zach, Christian Häne, Marc Pollefeys
CVPR1
2012 A Convex Discrete-Continuous Approach for Markov Random Fields
Christopher Zach, Pushmeet Kohli
ECCV (6)1
2012 Large-scale, dense city reconstruction from user-contributed photos
Arnold Irschara, Christopher Zach, Manfred Klopschitz, Horst Bischof
Comput. Vis. Image Underst.2
2011 Adaptive random forest - How many "experts" to ask before making a decision?
abstract
How many people should you ask if you are not sure about your way? We provide an answer to this question for Random Forest classification. The presented method is based on the statistical formulation of confidence intervals and conjugate priors for binomial as well as multinomial distributions. We derive appealing decision rules to speed up the classification process by leveraging the fact that many samples can be clearly mapped to classes. Results on test data are provided, and we highlight the applicability of our method to a wide range of problems. The approach introduces only one non-heuristic parameter, that allows to trade-off accuracy and speed without any re-training of the classifier. The proposed method automatically adapts to the difficulty of the test data and makes classification significantly faster without deteriorating the accuracy.
Alexander G. Schwing, Christopher Zach, Yefeng Zheng 0001, Marc Pollefeys
CVPR2
2011 The generalized trace-norm and its application to structure-from-motion problems
abstract
In geometric computer vision, the structure from motion (SfM) problem can be formulated as a optimization problem with a rank constraint. It is well known that the trace norm of a matrix can act as a convex proxy for a low rank constraint. Hence, in recent work [7], the trace-norm relaxation has been applied to the SfM problem. However, SfM problems often exhibit a certain structure, for example a smooth camera path. Unfortunately, the trace norm relaxation can not make use of this additional structure. This observation motivates the main contribution of this paper. We present the so-called generalized trace norm which allows to encode prior knowledge about a specific problem into a convex regularization term which enforces a low rank solution while at the same time taking the problem structure into account. While deriving the generalized trace norm and stating its different formulations, we draw interesting connections to other fields, most importantly to the field of compressive sensing. Even though the generalized trace norm is a very general concept with a wide area of potential applications we are ultimately interested in applying it to SfM problems. Therefore, we also present an efficient algorithm to optimize the resulting generalized trace norm regularized optimization problems. Results show that the generalized trace norm indeed achieves its goals in providing a problem-dependent regularization.
Roland Angst, Christopher Zach, Marc Pollefeys
ICCV2
2011 Stereo depth map fusion for robot navigation
abstract
We present a method to reconstruct indoor environments from stereo image pairs, suitable for the navigation of robots. To enable a robot to navigate solely using visual cues it receives from a stereo camera, the depth information needs to be extracted from the image pairs and combined into a common representation. The initially determined raw depthmaps are fused into a two level heightmap representation which contains a floor and a ceiling height level. To reduce the noise in the height maps we employ a total variation regularized energy functional.With this 2.5D representation of the scene the computational complexity of the energy optimization is reduced by one dimension in contrast to other fusion techniques that work on the full 3D space such as volumetric fusion. While we show only results for indoor environments the approach can be extended to generate heightmaps for outdoor environments.
Christian Häne, Christopher Zach, Jongwoo Lim, Ananth Ranganathan, Marc Pollefeys
IROS2
2010 Disambiguating visual relations using loop constraints
abstract
Repetitive and ambiguous visual structures in general pose a severe problem in many computer vision applications. Identification of incorrect geometric relations between images solely based on low level features is not always possible, and a more global reasoning approach about the consistency of the estimated relations is required. We propose to utilize the typically observed redundancy in the hypothesized relations for such reasoning, and focus on the graph structure induced by those relations. Chaining the (reversible) transformations over cycles in this graph allows to build suitable statistics for identifying inconsistent loops in the graph. This data provides indirect evidence for conflicting visual relations. Inferring the set of likely false positive geometric relations from these non-local observations is formulated in a Bayesian framework. We demonstrate the utility of the proposed method in several applications, most prominently the computation of structure and motion from images.
Christopher Zach, Manfred Klopschitz, Marc Pollefeys
CVPR1
2010 Practical Methods for Convex Multi-view Reconstruction
Christopher Zach, Marc Pollefeys
ECCV (4)1
2010 Gated Softmax Classification
abstract
We describe a log-bilinear" model that computes class probabilities by combining an input vector multiplicatively with a vector of binary latent variables. Even though the latent variables can take on exponentially many possible combinations of values, we can efficiently compute the exact probability of each class by marginalizing over the latent variables. This makes it possible to get the exact gradient of the log likelihood. The bilinear score-functions are defined using a three-dimensional weight tensor, and we show that factorizing this tensor allows the model to encode invariances inherent in a task by learning a dictionary of invariant basis functions. Experiments on a set of benchmark problems show that this fully probabilistic model can achieve classification performance that is competitive with (kernel) SVMs, backpropagation, and deep belief nets."
Roland Memisevic, Christopher Zach, Geoffrey E. Hinton, Marc Pollefeys
NIPS2
2009 From structure-from-motion point clouds to fast location recognition
abstract
Efficient view registration with respect to a given 3D reconstruction has many applications like inside-out tracking in indoor and outdoor environments, and geo-locating images from large photo collections. We present a fast location recognition technique based on structure from motion point clouds. Vocabulary tree-based indexing of features directly returns relevant fragments of 3D models instead of documents from the images database. Additionally, we propose a compressed 3D scene representation which improves recognition rates while simultaneously reducing the computation time and the memory consumption. The design of our method is based on algorithms that efficiently utilize modern graphics processing units to deliver real-time performance for view registration. We demonstrate the approach by matching hand-held outdoor videos to known 3D urban models, and by registering images from online photo collections to the corresponding landmarks.
Arnold Irschara, Christopher Zach, Jan-Michael Frahm, Horst Bischof
CVPR2
2009 Continuous maximal flows and Wulff shapes: Application to MRFs
abstract
Convex and continuous energy formulations for low level vision problems enable efficient search procedures for the corresponding globally optimal solutions. In this work we extend the well-established continuous, isotropic capacity-based maximal flow framework to the anisotropic setting. By using powerful results from convex analysis, a very simple and efficient minimization procedure is derived. Further, we show that many important properties carry over to the new anisotropic framework, e.g. globally optimal binary results can be achieved simply by thresholding the continuous solution. In addition, we unify the anisotropic continuous maximal flow approach with a recently proposed convex and continuous formulation for Markov random fields, thereby allowing more general smoothness priors to be incorporated. Dense stereo results are included to illustrate the capabilities of the proposed approach.
Christopher Zach, Marc Niethammer, Jan-Michael Frahm
CVPR1
2009 A new minimal solution to the relative pose of a calibrated stereo camera with small field of view overlap
abstract
In this paper we present a new minimal solver for the relative pose of a calibrated stereo camera (i.e. a pair of rigidly mounted cameras). Our method is based on the fact that a feature visible in all four images (two image pairs acquired at two points in time) constrains the relative pose of the second stereo camera to lie on a sphere around this feature, which has a known, triangulated position in the first stereo camera coordinate frame. This constraint leaves three degrees of freedom; two for the location of the second camera on the sphere, and the third for the rotation in the respective tangent plane. We use three 2D correspondences, in particular two correspondences from the left (or right) camera and one correspondence from the other camera, to solve for these three remaining degrees of freedom. This approach is amenable to stereo cameras having a small overlap in their views. We present an efficient solution for this novel relative pose problem, describe the incorporation of our proposed solver into the RANSAC framework, evaluate its performance given noise and outliers, and demonstrate its use in a real-time structure from motion system.
Brian Clipp, Christopher Zach, Jan-Michael Frahm, Marc Pollefeys
ICCV2
2009 Towards Large-Scale Visual Mapping and Localization
Marc Pollefeys, Jan-Michael Frahm, Friedrich Fraundorfer, Christopher Zach, Changchang Wu, Brian Clipp, David Gallup
ISRR4
2009 Adaptive, real-time visual simultaneous localization and mapping
abstract
In this paper we present a real-time simultaneous localization and mapping system which uses a stereo camera as its only input. We combine the benefits of KLT feature tracking, which include high speed and robustness to repetitive features, with wide baseline features, which allow for feature matching after large camera motions. Updating the map of feature locations and camera poses is considerably more expensive than performing KLT tracking. For this reason we use the optical flow measured by the KLT tracker to adaptively select key frames for which we do a full map and camera pose update. In this way we limit the processing to only ¿interesting¿ parts of the video sequence. Additionally, we maintain a consistent scene scale at low cost by using a GPU implementation of multi-camera scene flow, a generalization of KLT to the motion of image features in three dimensions. The system uses multiple sub-maps; scalable, bag of features recognition and geometric verification to recover from motion estimation failure or ¿kidnapping¿. This architecture allows the robot to grow the existing map online and in real time while storing all of the data necessary for an off-line optimization to complete loops. We demonstrate the robustness of our system in a challenging indoor environment that includes semi-reflective glass walls and people moving in the scene.
Brian Clipp, Christopher Zach, Jongwoo Lim, Jan-Michael Frahm, Marc Pollefeys
WACV2
2008 What can missing correspondences tell us about 3D structure and motion?
abstract
Practically all existing approaches to structure and motion computation use only positive image correspondences to verify the camera pose hypotheses. Incorrect epipolar geometries are solely detected by identifying outliers among the found correspondences. Ambiguous patterns in the images are often incorrectly handled by these standard methods. In this work we propose two approaches to overcome such problems. First, we apply non-monotone reasoning on view triplets using a Bayesian formulation. In contrast to two-view epipolar geometry, image triplets allow the prediction of features in the third image. Absence of these features (i.e. missing correspondences) enables additional inference about the view triplet. Furthermore, we integrate these view triplet handling into an incremental procedure for structure and motion computation. Thus, our approach is able to refine the maintained 3D structure when additional image data is provided.
Christopher Zach, Arnold Irschara, Horst Bischof
CVPR1
2008 Modeling and Recognition of Landmark Image Collections Using Iconic Scene Graphs
Xiaowei Li 0007, Changchang Wu, Christopher Zach, Svetlana Lazebnik, Jan-Michael Frahm
ECCV (1)3
2007 Mumford-Shah Meets Stereo: Integration of Weak Depth Hypotheses
abstract
Recent results on stereo indicate that an accurate segmentation is crucial for obtaining faithful depth maps. Variational methods have successfully been applied to both image segmentation and computational stereo. In this paper we propose a combination in a unified framework. In particular, we use a Mumford-Shah-like functional to compute a piecewise smooth depth map of a stereo pair. Our approach has two novel features: First, the regularization term of the functional combines edge information obtained from the color segmentation with flow-driven depth discontinuities emerging during the optimization procedure. Second, we propose a robust data term which adoptively selects the best matches obtained from different weak stereo algorithms. We integrate these features in a theoretically consistent framework. The final depth map is the minimizer of the energy functional, which can be solved by the associated functional derivatives. The underlying numerical scheme allows an efficient implementation on modern graphics hardware. We illustrate the performance of our algorithm using the Middlebury database as well as on real imagery.
Thomas Pock, Christopher Zach, Horst Bischof
CVPR2
2007 Towards Wiki-based Dense City Modeling
abstract
This work reports on the advances and on the current status of a terrestrial city modeling approach, which uses images contributed by end-users as input. Hence, the Wiki principle well known from textual knowledge databases is transferred to the goal of incrementally building a virtual representation of the occupied habitat. In order to achieve this objective, many state-of-the-art computer vision methods must be applied and modified according to this task. We describe the utilized 3D vision methods and show initial results obtained from the current image database acquired by in-house participants.
Arnold Irschara, Christopher Zach, Horst Bischof
ICCV2
2007 A Globally Optimal Algorithm for Robust TV-L1 Range Image Integration
abstract
Robust integration of range images is an important task for building high-quality 3D models. Since range images, and in particular range maps from stereo vision, may have a substantial amount of outliers, any integration approach aiming at high-quality models needs an increased level of robustness. Additionally, a certain level of regularization is required to obtain smooth surfaces. Computational efficiency and global convergence are further preferable properties. The contribution of this paper is a unified framework to solve all these issues. Our method is based on minimizing an energy functional consisting of a total variation (TV) regularization force and an L1 data fidelity term. We present a novel and efficient numerical scheme, which combines the duality principle for the TV term with a point-wise optimization step. We demonstrate the superior performance of our algorithm on the well-known Middlebury multi-view database and additionally on real-world multi-view images.
Christopher Zach, Thomas Pock, Horst Bischof
ICCV1
2007 A Duality Based Algorithm for TV- L 1-Optical-Flow Image Registration
Thomas Pock, Martin Urschler, Christopher Zach, Reinhard Beichel, Horst Bischof
MICCAI (2)3
2007 Augmented Reality Scouting for Interactive 3D Reconstruction
abstract
This paper presents a first prototype of an interactive 3D reconstruction system for modeling urban scenes. An augmented reality scout is a person who is equipped with an ultra-mobile PC, an attached USB camera and a GPS receiver. The scout is exploring the urban environment and delivers a sequence of 2D images. These images are annotated with according GPS data and used iteratively as input for a 3D reconstruction engine which generates the 3D models on-the-fly. This turns modeling into an interactive and collaborative task
Bernhard Reitinger, Christopher Zach, Dieter Schmalstieg
VR2
2006 Automatic Point Landmark Matching for Regularizing Nonlinear Intensity Registration: Application to Thoracic CT Images
Martin Urschler, Christopher Zach, Hendrik Ditt, Horst Bischof
MICCAI (2)2
2002 Time-critical rendering of discrete and continuous levels of detail
abstract
We present a novel level of detail selection method for real-time rendering, that works on hierarchies of discrete and continuous representations. We integrate point rendered objects with polygonal geometry and demonstrate our approach in a terrain flyover application, where the digital elevation model is augmented with forests. The vegetation is rendered as continuous sequence of splats, which are organized in a hierarchy. Further we discuss enhancements to our basic method to improve its scalability.
Christopher Zach, Stephan Mantler, Konrad F. Karner
VRST1