EDBT 2026 Demo / reviewers in the wild / expert
Richard I. Hartley
dblp:h/RIHartley · also Richard Hartley 0001
· DBLP profile ↗
197ranked-venue papers
51as first author
25since 2021 · last 2026
0000-0002-5005-0191ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 159 · 43 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 127 · 28 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 10Systems, architecture and hardware · 8 · 4 first-author · 1 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning-Based Multi-View Stereo: A Surveyabstract3D reconstruction aims to recover the dense 3D structure of a scene. It plays an essential role in various applications such as Augmented/Virtual Reality (AR/VR), autonomous driving and robotics. Leveraging multiple views of a scene captured from different viewpoints, Multi-View Stereo (MVS) algorithms synthesize a comprehensive 3D representation, enabling precise reconstruction in complex environments. Due to its efficiency and effectiveness, MVS has become a pivotal method for image-based 3D reconstruction. Recently, with the success of deep learning, many learning-based MVS methods have been proposed, achieving impressive performance against traditional methods. We categorize these learning-based methods as: depth map-based, voxel-based, NeRF-based, 3D Gaussian Splatting-based, and large feed-forward methods. Among these, we focus significantly on depth map-based methods, which are the main family of MVS due to their conciseness, flexibility and scalability. In this survey, we provide a comprehensive review of the literature at the time of this writing. We investigate these learning-based methods, summarize their performances on popular benchmarks, and discuss promising future research directions in this area. Fangjinhua Wang, Qingtian Zhu, Di Chang, Quankai Gao, Junlin Han, Tong Zhang 0023, Richard I. Hartley, Marc Pollefeys |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Probability Density Geodesics in Image Diffusion Latent SpaceabstractDiffusion models indirectly estimate the probability density over a data space, which can be used to study its structure. In this work, we show that geodesics can be computed in diffusion latent space, where the norm induced by the spatially-varying inner product is inversely proportional to the probability density. In this formulation, a path that traverses a high density (that is, probable) region of image latent space is shorter than the equivalent path through a low density region. We present algorithms for solving the associated initial and boundary value problems and show how to compute the probability density along the path and the geodesic distance between two points. Using these techniques, we analyze how closely video clips approximate geodesics in a pre-trained image diffusion space. Finally, we demonstrate how these techniques can be applied to training-free image sequence interpolation and extrapolation, given a pre-trained image diffusion model. Qingtao Yu, Zhaoyuan Yang, Peter H. Tu, Jing Zhang 0052, Hongdong Li, Richard I. Hartley, Dylan Campbell |
CVPR | 7 |
| 2025 | FlashMo: Geometric Interpolants and Frequency-Aware Sparsity for Scalable Efficient Motion GenerationabstractDiffusion models have recently advanced 3D human motion generation by producing smoother and more realistic sequences from natural language. However, existing approaches face two major challenges: high computational cost during training and inference, and limited scalability due to reliance on U-Net inductive bias. To address these challenges, we propose **FlashMo**, a frequency-aware sparse motion diffusion model that prunes low-frequency tokens to enhance efficiency without custom kernel design. We further introduce *MotionSiT*, a scalable diffusion transformer based on a joint-temporal factorized interpolant with Lie group geodesics over $\mathrm{SO}(3)$ manifolds, enabling principled generation of joint rotations. Extensive experiments on the large-scale MotionHub V2 dataset and standard benchmarks including HumanML3D and KIT-ML demonstrate that our method significantly outperforms previous approaches in motion quality, efficiency, and scalability. Compared to the state-of-the-art 1-step distillation baseline, FlashMo reduces **12.9%** inference time and FID by **34.1%**. Project website: https://steve-zeyu-zhang.github.io/FlashMo. Zeyu Zhang 0006, Danning Li, Dong Gong, Ian D. Reid 0001, Richard I. Hartley |
NeurIPS | 6 |
| 2025 | Quantifying Bias in Text-to-Image Generative ModelsabstractBias in text-to-image (T2I) generation can propagate unfair social representations and may be exploited to push ulterior agendas. These biases raise concerns on the dependability and fairness of models that have become widely popular and readily available for public consumption. Existing works in T2I bias analysis typically focus on social biases. We look beyond that and instead propose an evaluation methodology to quantify general bias in T2I generative models without any preconceived notion. We introduce a suite of three metrics; namely, distribution bias, Jaccard hallucination and generative miss-rate, to extensively appraise general model bias. To validate the efficacy of these metrics, we also introduce a backdoor-inspired strategy, which provides a convenient handle over the extent of bias in a model for controlled analysis. We assess T2I models implementing six widely used pipelines in this domain. Our extensive analysis covers both general and task-oriented scenarios, employing over 105 K generated images. For prior art comparison, it also encompasses social bias analysis. Moreover, we also extend our technique to analyze bias in seven popular captioned image datasets. Our experiments establish that our approach is objective, domain-agnostic and it consistently measures different forms of T2I model biases. To further research efforts into T2I model biases, we have developed an open-source web application and practical implementation of this work, which is available onhttps://huggingface.co/spaces/JVice/try-before-you-biasHuggingFace. All relevant code is also publicly available onhttps://github.com/JJ-Vice/TryBeforeYouBiasGitHub. Jordan Vice, Naveed Akhtar, Richard I. Hartley, Ajmal Mian |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2024 | Adversarial Purification with the Manifold HypothesisabstractIn this work, we formulate a novel framework for adversarial robustness using the manifold hypothesis. This framework provides sufficient conditions for defending against adversarial examples. We develop an adversarial purification method with this framework. Our method combines manifold learning with variational inference to provide adversarial robustness without the need for expensive adversarial training. Experimentally, our approach can provide adversarial robustness even if attackers are aware of the existence of the defense. In addition, our method can also serve as a test-time defense mechanism for variational autoencoders. Zhaoyuan Yang, Jing Zhang 0052, Richard I. Hartley, Peter H. Tu |
AAAI | 4 |
| 2024 | LDP: Language-driven Dual-Pixel Image Defocus Deblurring NetworkabstractRecovering sharp images from dual-pixel (DP) pairs with disparity-dependent blur is a challenging task. Existing blur map-based deblurring methods have demonstrated promising results. In this paper, we propose, to the best of our knowledge, the first framework that introduces the contrastive language-image pre-training framework (CLIP) to accurately estimate the blur map from a DP pair unsu-pervisedly. To achieve this, we first carefully design text prompts to enable CLIP to understand blur-related geo-metric prior knowledge from the DP pair. Then, we pro-pose a format to input a stereo DP pair to CLIP without any fine-tuning, despite the fact that CLIP is pre-trained on monocular images. Given the estimated blur map, we intro-duce a blur-prior attention block, a blur-weighting loss, and a blur-aware loss to recover the all-in-focus image. Our method achieves state-of-the-art performance in extensive experiments (see Fig. 1). Hao Yang 0040, Liyuan Pan, Yan Yang 0011, Richard I. Hartley, Miaomiao Liu 0001 |
CVPR | 4 |
| 2024 | Motion Mamba: Efficient and Long Sequence Motion Generation
Zeyu Zhang 0006, Akide Liu, Ian D. Reid 0001, Richard I. Hartley, Bohan Zhuang, Hao Tang 0005 |
ECCV (1) | 4 |
| 2024 | Neural SDF Flow for 3D Reconstruction of Dynamic ScenesabstractIn this paper, we tackle the problem of 3D reconstruction of dynamic scenes from multi-view videos. Previous dynamic scene reconstruction works either attempt to model the motion of 3D points in space, which constrains them to handle a single articulated object or require depth maps as input. By contrast, we propose to directly estimate the change of Signed Distance Function (SDF), namely SDF
flow, of the dynamic scene. We show that the SDF flow captures the evolution of the scene surface. We further derive the mathematical relation between the SDF flow and the scene flow, which allows us to calculate the scene flow from the SDF flow analytically by solving linear equations. Our experiments on real-world multi-view video datasets show that our reconstructions are better than those of the state-of-the-art methods. Our code is available at https://github.com/wei-mao-2019/SDFFlow.git. Wei Mao 0001, Richard I. Hartley, Mathieu Salzmann, Miaomiao Liu 0001 |
ICLR | 2 |
| 2024 | IMPUS: Image Morphing with Perceptually-Uniform Sampling Using Diffusion ModelsabstractWe present a diffusion-based image morphing approach with perceptually-uniform sampling (IMPUS) that produces smooth, direct and realistic interpolations given an image pair. The embeddings of two images may lie on distinct conditioned distributions of a latent diffusion model, especially when they have significant semantic difference. To bridge this gap, we interpolate in the locally linear and continuous text embedding space and Gaussian latent space. We first optimize the endpoint text embeddings and then map the images to the latent space using a probability flow ODE. Unlike existing work that takes an indirect morphing path, we show that the model adaptation yields a direct path and suppresses ghosting artifacts in the interpolated images. To achieve this, we propose a heuristic bottleneck constraint based on a novel relative perceptual path diversity score that automatically controls the bottleneck size and balances the diversity along the path with its directness. We also propose a perceptually-uniform sampling technique that enables visually smooth changes between the interpolated images. Extensive experiments validate that our IMPUS can achieve smooth, direct, and realistic image morphing and is adaptable to several other generative tasks. Zhaoyuan Yang, Jing Zhang 0052, Dylan Campbell, Peter H. Tu, Richard I. Hartley |
ICLR | 8 |
| 2024 | NF-SLAM: Effective, Normalizing Flow-supported Neural Field representations for object-level visual SLAM in automotive applicationsabstractWe propose a novel, vision-only object-level SLAM framework for automotive applications representing 3D shapes by implicit signed distance functions. Our key innovation consists of augmenting the standard neural representation by a normalizing flow network. As a result, achieving strong representation power on the specific class of road vehicles is made possible by compact networks with only 16-dimensional latent codes. Furthermore, the newly proposed architecture exhibits a significant performance improvement in the presence of only sparse and noisy data, which is demonstrated through comparative experiments on synthetic data. The module is embedded into the back-end of a stereo-vision based framework for joint, incremental shape optimization. The loss function is given by a combination of a sparse 3D point-based SDF loss, a sparse rendering loss, and a semantic mask-based silhouette-consistency term. We furthermore leverage semantic information to determine keypoint extraction density in the front-end. Finally, experimental results on real-world data reveal accurate and reliable performance comparable to alternative frameworks that make use of direct depth readings. The proposed method performs well with only sparse 3D points obtained from bundle adjustment, and eventually continues to deliver stable results even under exclusive use of the mask-consistency term. Richard I. Hartley, Zirui Xie, Laurent Kneip, Zhenghua Yu |
IROS | 3 |
| 2024 | Weakly-Supervised Depth Estimation and Image Deblurring via Dual-Pixel SensorsabstractDual-pixel (DP) imaging sensors are getting more popularly adopted by modern cameras. A DP camera captures a pair of images in a single snapshot by splitting each pixel in half. Several previous studies show how to recover depth information by treating the DP pair as an approximate stereo pair. However, dual-pixel disparity occurs only in image regions with defocus blur which is unlike classic stereo disparity. Heavy defocus blur in DP pairs affects the performance of depth estimation approaches based on matching. Therefore, we treat the blur removal and the depth estimation as a joint problem. We investigate the formation of the DP pair, which links the blur and depth information, rather than blindly removing the blur effect. We propose a mathematical DP model that can improve depth estimation by the blur. This exploration motivated us to propose our previous work, an end-to-end DDDNet (DP-based Depth and Deblur Network), which jointly estimates depth and restores the image in a supervised fashion. However, collecting the ground-truth (GT) depth map for the DP pair is challenging and limits the depth estimation potential of the DP sensor. Therefore, we propose an extension of the DDDNet, called WDDNet (Weakly-supervised Depth and Deblur Network), which includes an efficient reblur solver that does not require GT depth maps for training. To achieve this, we convert all-in-focus images into supervisory signals for unsupervised depth estimation in our WDDNet. We jointly estimate an all-in-focus image and a disparity map, then use a Reblur and Fstack module to regularize the disparity estimation and image restoration. We conducted extensive experiments on synthetic and real data to demonstrate the competitive performance of our method when compared to state-of-the-art (SOTA) supervised approaches. Liyuan Pan, Richard I. Hartley, Liu Liu 0009, Shah Ariful Hoque Chowdhury, Yan Yang 0011, Hongdong Li, Miaomiao Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | BAGM: A Backdoor Attack for Manipulating Text-to-Image Generative ModelsabstractThe rise in popularity of text-to-image generative artificial intelligence (AI) has attracted widespread public interest. We demonstrate that this technology can be attacked to generate content that subtly manipulates its users. We propose a Backdoor Attack on text-to-image Generative Models (BAGM), which upon triggering, infuses the generated images with manipulative details that are naturally blended in the content. Our attack is the first to target three popular text-to-image generative models across three stages of the generative process by modifying the behaviour of the embedded tokenizer, the language model or the image generative model. Based on the penetration level, BAGM takes the form of a suite of attacks that are referred to assurface,shallowanddeepattacks in this article. Given the existing gap within this domain, we also contribute a comprehensive set of quantitative metrics designed specifically for assessing the effectiveness of backdoor attacks on text-to-image models. The efficacy of BAGM is established by attacking state-of-the-art generative models, using a marketing scenario as the target domain. To that end, we contribute a dataset of branded product images. Our embedded backdoors increase the bias towards the target outputs by more than five times the usual, without compromising the model robustness or the generated content utility. By exposing generative AI’s vulnerabilities, we encourage researchers to tackle these challenges and practitioners to exercise caution when using pre-trained models. Relevant code and input prompts can be found at https://github.com/JJ-Vice/BAGM, and the dataset is available at: https://ieee-dataport.org/documents/marketable-foods-mf-dataset. Jordan Vice, Naveed Akhtar, Richard I. Hartley, Ajmal Mian |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Manifold Learning Benefits GANsabstractIn this paper11Code: https://qithub.com/MaxwellYaoNi/LCSAGAN., we improve Generative Adversarial Net-works by incorporating a manifold learning step into the discriminator. We consider locality-constrained linear and subspace-based manifolds22The coding spaces considered in this paper are loosely termed man-ifolds. In most cases they are not manifolds in the strict mathematical sense, but rather topological spaces such as varieties, or simplicial com-plexes. The word will be used only in an informal sense., and locality-constrained non-linear manifolds. In our design, the manifold learning and coding steps are intertwined with layers of the discrimina-tor, with the goal of attracting intermediate feature repre-sentations onto manifolds. We adaptively balance the dis-crepancy between feature representations and their mani-fold view, which is a trade-off between denoising on the manifold and refining the manifold. We find that locality-constrained non-linear manifolds outperform linear mani-folds due to their non-uniform density and smoothness. We also substantially outperform state-of-the-art baselines. Yao Ni, Piotr Koniusz, Richard I. Hartley, Richard Nock |
CVPR | 3 |
| 2022 | ERA: Enhanced Rational Activations
Martin Trimmel, Mihai Zanfir, Richard I. Hartley, Cristian Sminchisescu |
ECCV (20) | 3 |
| 2022 | Contact-aware Human Motion ForecastingabstractIn this paper, we tackle the task of scene-aware 3D human motion forecasting, which consists of predicting future human poses given a 3D scene and a past human motion. A key challenge of this task is to ensure consistency between the human and the scene, accounting for human-scene interactions. Previous attempts to do so model such interactions only implicitly, and thus tend to produce artifacts such as ``ghost motion" because of the lack of explicit constraints between the local poses and the global motion. Here, by contrast, we propose to explicitly model the human-scene contacts. To this end, we introduce distance-based contact maps that capture the contact relationships between every joint and every 3D scene point at each time instant. We then develop a two-stage pipeline that first predicts the future contact maps from the past ones and the scene point cloud, and then forecasts the future human poses by conditioning them on the predicted contact maps. During training, we explicitly encourage consistency between the global motion and the local poses via a prior defined using the contact maps and future poses. Our approach outperforms the state-of-the-art human motion forecasting and human synthesis methods on both synthetic and real datasets. Our code is available at https://github.com/wei-mao-2019/ContAwareMotionPred. Wei Mao 0001, Miaomiao Liu 0001, Richard I. Hartley, Mathieu Salzmann |
NeurIPS | 3 |
| 2022 | Few-shot Weakly-Supervised Object Detection via Directional StatisticsabstractDetecting novel objects from few examples has become an emerging topic in computer vision recently. However, current methods need fully annotated training images to learn new object categories which limits their applicability in real world scenarios such as field robotics. In this work, we propose a probabilistic multiple-instance learning approach for few-shot Common Object Localization (COL) and few-shot Weakly Supervised Object Detection (WSOD). In these tasks, only image-level labels, which are much cheaper to acquire, are available. We find that operating on features extracted from the last layer of a pretrained Faster-RCNN is more effective compared to previous episodic learning based few-shot COL methods. Our model simultaneously learns the distribution of the novel objects and localizes them via expectation-maximization steps. As a probabilistic model, we employ von Mises-Fisher (vMF) distribution which captures the semantic information better than Gaussian distribution when applied to the pre-trained embedding space. When the novel objects are localized, we utilize them to learn a linear appearance model to detect novel classes in new images. Our extensive experiments show that the proposed method, despite being simple, outperforms strong baselines in few-shot COL and WSOD, as well as large-scale WSOD tasks. Amirreza Shaban, Amir Rahimi, Thalaiyasingam Ajanthan, Byron Boots, Richard I. Hartley |
WACV | 5 |
| 2022 | Deep Declarative NetworksabstractWe explore a class of end-to-end learnable models wherein data processing nodes (or network layers) are defined in terms of desired behavior rather than an explicit forward function. Specifically, the forward function is implicitly defined as the solution to a mathematical optimization problem. Consistent with nomenclature in the programming languages community, we name these models deep declarative networks. Importantly, it can be shown that the class of deep declarative networks subsumes current deep learning models. Moreover, invoking the implicit function theorem, we show how gradients can be back-propagated through many declaratively defined data processing nodes thereby enabling end-to-end learning. We discuss how these declarative processing nodes can be implemented in the popular PyTorch deep learning software library allowing declarative and imperative nodes to co-exist within the same network. We also provide numerous insights and illustrative examples of declarative nodes and demonstrate their application for image and point cloud classification tasks. Stephen Gould, Richard I. Hartley, Dylan Campbell |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | High Frame Rate Video Reconstruction Based on an Event CameraabstractEvent-based cameras measure intensity changes (called 'events') with microsecond accuracy under high-speed motion and challenging lighting conditions. With the 'active pixel sensor' (APS), the 'Dynamic and Active-pixel Vision Sensor' (DAVIS) allows the simultaneous output of intensity frames and events. However, the output images are captured at a relatively low frame rate and often suffer from motion blur. A blurred image can be regarded as the integral of a sequence of latent images, while events indicate changes between the latent images. Thus, we are able to model the blur-generation process by associating event data to a latent sharp image. Based on the abundant event data alongside a low frame rate, easily blurred images, we propose a simple yet effective approach to reconstruct high-quality and high frame rate sharp videos. Starting with a single blurred frame and its event data from DAVIS, we propose the Event-based Double Integral (EDI) model and solve it by adding regularization terms. Then, we extend it to multiple Event-based Double Integral (mEDI) model to get more smooth results based on multiple images and their events. Furthermore, we provide a new and more efficient solver to minimize the proposed energy model. By optimizing the energy function, we achieve significant improvements in removing blur and the reconstruction of a high temporal resolution video. The video generation is based on solving a simple non-convex optimization problem in a single scalar variable. Experimental results on both synthetic and real datasets demonstrate the superiority of our mEDI model and optimization method compared to the state-of-the-art. Liyuan Pan, Richard I. Hartley, Cedric Scheerlinck, Miaomiao Liu 0001, Xin Yu 0002, Yuchao Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Mirror Descent View for Neural Network QuantizationabstractQuantizing large Neural Networks (NN) while maintaining the performance is highly desirable for resource-limited devices due to reduced memory and time complexity. It is usually formulated as a constrained optimization problem and optimized via a modified version of gradient descent. In this work, by interpreting the continuous parameters (unconstrained) as the dual of the quantized ones, we introduce a Mirror Descent (MD) framework for NN quantization. Specifically, we provide conditions on the projections (i.e., mapping from continuous to quantized ones) which would enable us to derive valid mirror maps and in turn the respective MD updates. Furthermore, we present a numerically stable implementation of MD that requires storing an additional set of auxiliary variables (unconstrained), and show that it is strikingly analogous to the Straight Through Estimator (STE) based method which is typically viewed as a “trick” to avoid vanishing gradients issue. Our experiments on CIFAR-10/100, TinyImageNet, and ImageNet classification datasets with VGG-16, ResNet-18, and MobileNetV2 architectures show that our MD variants yield state-of-the-art performance. Thalaiyasingam Ajanthan, Kartik Gupta, Philip Torr 0001, Richard I. Hartley, Puneet K. Dokania |
AISTATS | 4 |
| 2021 | Learning Optical Flow From a Few MatchesabstractState-of-the-art neural network models for optical flow estimation require a dense correlation volume at high resolutions for representing per-pixel displacement. Although the dense correlation volume is informative for accurate estimation, its heavy computation and memory usage hinders the efficient training and deployment of the models. In this paper, we show that the dense correlation volume representation is redundant and accurate flow estimation can be achieved with only a fraction of elements in it. Based on this observation, we propose an alternative displacement representation, named Sparse Correlation Volume, which is constructed directly by computing the k closest matches in one feature map for each feature vector in the other feature map and stored in a sparse data structure. Experiments show that our method can reduce computational cost and memory use significantly, while maintaining high accuracy compared to previous approaches with dense correlation volumes. Shihao Jiang, Hongdong Li, Richard I. Hartley |
CVPR | 4 |
| 2021 | Dual Pixel Exploration: Simultaneous Depth Estimation and Image RestorationabstractThe dual-pixel (DP) hardware works by splitting each pixel in half and creating an image pair in a single snapshot. Several works estimate depth/inverse depth by treating the DP pair as a stereo pair. However, dual-pixel disparity only occurs in image regions with the defocus blur. The heavy defocus blur in DP pairs affects the performance of matching-based depth estimation approaches. Instead of removing the blur effect blindly, we study the formation of the DP pair which links the blur and the depth information. In this paper, we propose a mathematical DP model which can benefit depth estimation by the blur. These explorations motivate us to propose an end-to-end DDDNet (DP-based Depth and Deblur Network) to jointly estimate the depth and restore the image. Moreover, we define a re-blur loss, which reflects the relationship of the DP image formation process with depth information, to regularise our depth estimate in training. To meet the requirement of a large amount of data for learning, we propose the first DP image simulator which allows us to create datasets with DP pairs from any existing RGBD dataset. As a side contribution, we collect a real dataset for further research. Extensive experimental evaluation on both synthetic and real datasets shows that our approach achieves competitive performance compared to state-of-the-art approaches. Liyuan Pan, Shah Chowdhury, Richard I. Hartley, Miaomiao Liu 0001, Hongdong Li |
CVPR | 3 |
| 2021 | Learning to Estimate Hidden Motions with Global Motion AggregationabstractOcclusions pose a significant challenge to optical flow algorithms that rely on local evidences. We consider an occluded point to be one that is imaged in the reference frame but not in the next, a slight overloading of the standard definition since it also includes points that move out-of-frame. Estimating the motion of these points is extremely difficult, particularly in the two-frame setting. Previous work relies on CNNs to learn occlusions, without much success, or requires multiple frames to reason about occlusions using temporal smoothness. In this paper, we argue that the occlusion problem can be better solved in the two-frame case by modelling image self-similarities. We introduce a global motion aggregation module, a transformer-based approach to find long-range dependencies between pixels in the first image, and perform global aggregation on the corresponding motion features. We demonstrate that the optical flow estimates in the occluded regions can be significantly improved without damaging the performance in non-occluded regions. This approach obtains new state-of-the-art results on the challenging Sintel dataset, improving the average end-point error by 13.6% on Sintel Final and 13.7% on Sintel Clean. At the time of submission, our method ranks first on these benchmarks among all published and unpublished approaches. Code is available at https://github.com/zacjiang/GMA. Shihao Jiang, Dylan Campbell, Hongdong Li, Richard I. Hartley |
ICCV | 5 |
| 2021 | Calibration of Neural Networks using Splines
Kartik Gupta, Amir Rahimi, Thalaiyasingam Ajanthan, Thomas Mensink, Cristian Sminchisescu, Richard I. Hartley |
ICLR | 6 |
| 2021 | Fixed-Lens camera setup and calibrated image registration for multifocus multiview 3D reconstructionabstractImage-based 3D reconstruction or 3D photogrammetry of small-scale objects including insects and biological specimens is challenging due to the use of a high magnification lens with inherently limited depth of field, and the object’s fine structures. Therefore, the traditional 3D reconstruction techniques cannot be applied without additional image preprocessing. One such preprocessing technique is multifocus stacking/fusion that combines a set of partially focused images captured at different distances from the same viewing angle to create a single in-focus image. We found that the image formation is not properly considered by the traditional multifocus image capture and stacking techniques. The resulting in-focus images contain artifacts that violate the perspective projection. A 3D reconstruction using such images often fails to produce accurate 3D models of the captured objects. This paper shows how this problem can be solved effectively by a new multifocus multiview 3D reconstruction procedure which includes a new Fixed-Lens multifocus image capture and a calibrated image registration technique using analytic homography transformation. The experimental results using the real and synthetic images demonstrate the effectiveness of the proposed solutions by showing that both the fixed-lens image capture and multifocus stacking with calibrated image alignment significantly reduce the errors in the camera poses and produce more complete 3D reconstructed models as compared with those by the conventional moving lens image capture and multifocus stacking. Shah Ariful Hoque Chowdhury, Chuong V. Nguyen, Hengjia Li, Richard I. Hartley |
Neural Comput. Appl. | 4 |
| 2021 | Learning Saliency From Single Noisy Labelling: A Robust Model Fitting PerspectiveabstractThe advances made in predicting visual saliency using deep neural networks come at the expense of collecting large-scale annotated data. However, pixel-wise annotation is labor-intensive and overwhelming. In this paper, we propose to learn saliency prediction from a single noisy labelling, which is easy to obtain (e.g., from imperfect human annotation or from unsupervised saliency prediction methods). With this goal, we address a natural question: Can we learn saliency prediction while identifying clean labels in a unified framework? To answer this question, we call on the theory of robust model fitting and formulate deep saliency prediction from a single noisy labelling as robust network learning and exploit model consistency across iterations to identify inliers and outliers (i.e., noisy labels). Extensive experiments on different benchmark datasets demonstrate the superiority of our proposed framework, which can learn comparable saliency prediction with state-of-the-art fully supervised saliency methods. Furthermore, we show that simply by treating ground truth annotations as noisy labelling, our framework achieves tangible improvements over state-of-the-art methods. Jing Zhang 0052, Yuchao Dai, Tong Zhang 0023, Mehrtash Harandi, Nick Barnes, Richard I. Hartley |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | Joint Unsupervised Learning of Optical Flow and Egomotion with Bi-Level optimizationabstractWe address the problem of joint optical flow and camera motion estimation in rigid scenes by incorporating geometric constraints into an unsupervised deep learning framework. Unlike existing approaches which rely on brightness constancy and local smoothness for optical flow estimation, we exploit the global relationship between optical flow and camera motion using epipolar geometry. In particular, we formulate the prediction of optical flow and camera motion as a bi-level optimization problem, consisting of an upper-level problem to estimate the flow that conforms to the predicted camera motion, and a lower-level problem to estimate the camera motion given the predicted optical flow. We use implicit differentiation to enable backpropagation through the lower-level geometric optimization layer independent of its implementation, allowing end-toend training of the network. With globally-enforced geometric constraints, we are able to improve the quality of the estimated optical flow in challenging scenarios, and obtain better camera motion estimates compared to other unsupervised learning methods. Shihao Jiang, Dylan Campbell, Miaomiao Liu 0001, Stephen Gould, Richard I. Hartley |
3DV | 5 |
| 2020 | RANP: Resource Aware Neuron Pruning at Initialization for 3D CNNsabstractAlthough 3D Convolutional Neural Networks (CNNs) are essential for most learning based applications involving dense 3D data, their applicability is limited due to excessive memory and computational requirements. Compressing such networks by pruning therefore becomes highly desirable. However, pruning 3D CNNs is largely unexplored possibly because of the complex nature of typical pruning algorithms that embeds pruning into an iterative optimization paradigm. In this work, we introduce a Resource Aware Neuron Pruning (RANP) algorithm that prunes 3D CNNs at initialization to high sparsity levels. Specifically, the core idea is to obtain an importance score for each neuron based on their sensitivity to the loss function. This neuron importance is then reweighted according to the neuron resource consumption related to FLOPs or memory. We demonstrate the effectiveness of our pruning method on 3D semantic segmentation with widely used 3D-UNets on ShapeNet and BraTS'18 as well as on video classification with MobileNetV2 and I3D on UCF101 dataset. In these experiments, our RANP leads to roughly 50%-95% reduction in FLOPs and 35%-80% reduction in memory with negligible loss in accuracy compared to the unpruned networks. This significantly reduces the computational resources required to train 3D CNNs. The pruned network obtained by our algorithm can also be easily scaled up and transferred to another dataset for training. Thalaiyasingam Ajanthan, Vibhav Vineet, Richard I. Hartley |
3DV | 4 |
| 2020 | Fast and Differentiable Message Passing on Pairwise Markov Random Fields
Thalaiyasingam Ajanthan, Richard I. Hartley |
ACCV (3) | 3 |
| 2020 | Single Image Optical Flow Estimation With an Event CameraabstractEvent cameras are bio-inspired sensors that asynchronously report intensity changes in microsecond resolution. DAVIS can capture high dynamics of a scene and simultaneously output high temporal resolution events and low frame-rate intensity images. In this paper, we propose a single image (potentially blurred) and events based optical flow estimation approach. First, we demonstrate how events can be used to improve flow estimates. To this end, we encode the relation between flow and events effectively by presenting an event-based photometric consistency formulation. Then, we consider the special case of image blur caused by high dynamics in the visual environments and show that including the blur formation in our model further constrains flow estimation. This is in sharp contrast to existing works that ignore the blurred images while our formulation can naturally handle either blurred or sharp images to achieve accurate flow estimation. Finally, we reduce flow estimation, as well as image deblurring, to an alternative optimization problem of an objective function using the primal-dual algorithm. Experimental results on both synthetic and real data (with blurred and non-blurred images) show the superiority of our model in comparison to state-of-the-art approaches. Liyuan Pan, Miaomiao Liu 0001, Richard I. Hartley |
CVPR | 3 |
| 2020 | Pairwise Similarity Knowledge Transfer for Weakly Supervised Object Localization
Amir Rahimi, Amirreza Shaban, Thalaiyasingam Ajanthan, Richard I. Hartley, Byron Boots |
ECCV (24) | 4 |
| 2020 | Intra Order-preserving Functions for Calibration of Multi-Class Neural NetworksabstractPredicting calibrated confidence scores for multi-class deep networks is important for avoiding rare but costly mistakes. A common approach is to learn a post-hoc calibration function that transforms the output of the original network into calibrated confidence scores while maintaining the network's accuracy. However, previous post-hoc calibration techniques work only with simple calibration functions, potentially lacking sufficient representation to calibrate the complex function landscape of deep networks. In this work, we aim to learn general post-hoc calibration functions that can preserve the top-k predictions of any deep network. We call this family of functions intra order-preserving functions. We propose a new neural network architecture that represents a class of intra order-preserving functions by combining common neural network components. Additionally, we introduce order-invariant and diagonal sub-families, which can act as regularization for better generalization when the training data size is small. We show the effectiveness of the proposed method across a wide range of datasets and classifiers. Our method outperforms state-of-the-art post-hoc calibration methods, namely temperature scaling and Dirichlet calibration, in several evaluation metrics for the task. Amir Rahimi, Amirreza Shaban, Ching-An Cheng, Richard I. Hartley, Byron Boots |
NeurIPS | 4 |
| 2020 | Fast Postprocessing for Difficult Discrete Energy Minimization ProblemsabstractDespite the rapid progress in discrete energy minimization, certain problems involving high connectivity and a high number of labels are considered very hard but are still very relevant in computer vision. We propose a post-processing technique to improve the sub-optimal results of the existing methods on such problems. Our core contribution is a mapping between the binary min-cut problem and finding the shortest path in a directed acyclic graph. Using this mapping, we present an algorithm to find an approximate solution for the min-cut problem. We also extend the same idea for multi-label factor-graphs in the form of an iterative move-making algorithm. The proposed algorithm is extremely fast, yet outperforms the existing techniques in terms of accuracy as well as the computational time. We demonstrate competitive or better results on problems where already high-quality work is done. Ijaz Akhter, Loong Fah Cheong, Richard I. Hartley |
WACV | 3 |
| 2020 | EpO-Net: Exploiting Geometric Constraints on Dense Trajectories for Motion SaliencyabstractThe existing approaches for salient motion segmentation are unable to explicitly learn geometric cues and often give false detections on prominent static objects. We exploit multiview geometric constraints to avoid such shortcomings. To handle the nonrigid background like a sea, we also propose a robust fusion mechanism between motion and appearance-based features. We find dense trajectories, covering every pixel in the video, and propose trajectory-based epipolar distances to distinguish between background and foreground regions. Trajectory epipolar distances are dataindependent and can be readily computed given a few features' correspondences between the images. We show that by combining epipolar distances with optical flow, a powerful motion network can be learned. Enabling the network to leverage both of these features, we propose a simple mechanism, we call input-dropout. Comparing the motion-only networks, we outperform the previous state of the art on DAVIS-2016 dataset by 5.2% in the mean IoU score. By robustly fusing our motion network with an appearance network using the input-dropout mechanism, we also outperform the previous methods on DAVIS-2016, 2017 and Segtrackv2 dataset. Muhammad Faisal 0003, Ijaz Akhter, Mohsen Ali, Richard I. Hartley |
WACV | 4 |
| 2020 | Hallucinating Unaligned Face Images by Multiscale Transformative Discriminative Networks
Xin Yu 0002, Fatih Porikli, Basura Fernando, Richard I. Hartley |
Int. J. Comput. Vis. | 4 |
| 2020 | Semantic Face Hallucination: Super-Resolving Very Low-Resolution Face Images with Supplementary AttributesabstractGiven a tiny face image, existing face hallucination methods aim at super-resolving its high-resolution (HR) counterpart by learning a mapping from an exemplary dataset. Since a low-resolution (LR) input patch may correspond to many HR candidate patches, this ambiguity may lead to distorted HR facial details and wrong attributes such as gender reversal and rejuvenation. An LR input contains low-frequency facial components of its HR version while its residual face image, defined as the difference between the HR ground-truth and interpolated LR images, contains the missing high-frequency facial details. We demonstrate that supplementing residual images or feature maps with additional facial attribute information can significantly reduce the ambiguity in face super-resolution. To explore this idea, we develop an attribute-embedded upsampling network, which consists of an upsampling network and a discriminative network. The upsampling network is composed of an autoencoder with skip-connections, which incorporates facial attribute vectors into the residual features of LR inputs at the bottleneck of the autoencoder, and deconvolutional layers used for upsampling. The discriminative network is designed to examine whether super-resolved faces contain the desired attributes or not and then its loss is used for updating the upsampling network. In this manner, we can super-resolve tiny (16×16 pixels) unaligned face images with a large upscaling factor of 8× while reducing the uncertainty of one-to-many mappings remarkably. By conducting extensive evaluations on a large-scale dataset, we demonstrate that our method achieves superior face hallucination results and outperforms the state-of-the-art. Xin Yu 0002, Basura Fernando, Richard I. Hartley, Fatih Porikli |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Phase-Only Image Based Kernel Estimation for Single Image Blind DeblurringabstractThe image motion blurring process is generally modelled as the convolution of a blur kernel with a latent image. Therefore, the estimation of the blur kernel is essentially important for blind image deblurring. Unlike existing approaches which focus on approaching the problem by enforcing various priors on the blur kernel and the latent image, we are aiming at obtaining a high quality blur kernel directly by studying the problem in the frequency domain. We show that the auto-correlation of the absolute phase-only image 1 can provide faithful information about the motion (e.g., the motion direction and magnitude, we call it the motion pattern in this paper.) that caused the blur, leading to a new and efficient blur kernel estimation approach. The blur kernel is then refined and the sharp image is estimated by solving an optimization problem by enforcing a regularization on the blur kernel and the latent image. We further extend our approach to handle non-uniform blur, which involves spatially varying blur kernels. Our approach is evaluated extensively on synthetic and real data and shows good results compared to the state-of-the-art deblurring approaches. Liyuan Pan, Richard I. Hartley, Miaomiao Liu 0001, Yuchao Dai |
CVPR | 2 |
| 2019 | Bringing a Blurry Frame Alive at High Frame-Rate With an Event CameraabstractEvent-based cameras can measure intensity changes (called ‘events’) with microsecond accuracy under high-speed motion and challenging lighting conditions. With the active pixel sensor (APS), the event camera allows simultaneous output of the intensity frames. However, the output images are captured at a relatively low frame-rate and often suffer from motion blur. A blurry image can be regarded as the integral of a sequence of latent images, while the events indicate the changes between the latent images. Therefore, we are able to model the blur-generation process by associating event data to a latent image. In this paper, we propose a simple and effective approach, the Event-based Double Integral (EDI) model, to reconstruct a high frame-rate, sharp video from a single blurry frame and its event data. The video generation is based on solving a simple non-convex optimization problem in a single scalar variable. Experimental results on both synthetic and real images demonstrate the superiority of our EDI model and optimization method in comparison to the state-of-the-art. Liyuan Pan, Cedric Scheerlinck, Xin Yu 0002, Richard I. Hartley, Miaomiao Liu 0001, Yuchao Dai |
CVPR | 4 |
| 2019 | Proximal Mean-Field for Neural Network QuantizationabstractCompressing large Neural Networks (NN) by quantizing the parameters, while maintaining the performance is highly desirable due to reduced memory and time complexity. In this work, we cast NN quantization as a discrete labelling problem, and by examining relaxations, we design an efficient iterative optimization procedure that involves stochastic gradient descent followed by a projection. We prove that our simple projected gradient descent approach is, in fact, equivalent to a proximal version of the well-known mean-field method. These findings would allow the decades-old and theoretically grounded research on MRF optimization to be used to design better network quantization schemes. Our experiments on standard classification datasets (MNIST, CIFAR10/100, TinyImageNet) with convolutional and residual architectures show that our algorithm obtains fully-quantized networks with accuracies very close to the floating-point reference networks. Thalaiyasingam Ajanthan, Puneet K. Dokania, Richard I. Hartley, Philip Torr 0001 |
ICCV | 3 |
| 2019 | Siamese Networks: The Tale of Two ManifoldsabstractSiamese networks are non-linear deep models that have found their ways into a broad set of problems in learning theory, thanks to their embedding capabilities. In this paper, we study Siamese networks from a new perspective and question the validity of their training procedure. We show that in the majority of cases, the objective of a Siamese network is endowed with an invariance property. Neglecting the invariance property leads to a hindrance in training the Siamese networks. To alleviate this issue, we propose two Riemannian structures and generalize a well-established accelerated stochastic gradient descent method to take into account the proposed Riemannian structures. Our empirical evaluations suggest that by making use of the Riemannian geometry, we achieve state-of-the-art results against several algorithms for the challenging problem of fine-grained image classification. Soumava Kumar Roy, Mehrtash Harandi, Richard Nock, Richard I. Hartley |
ICCV | 4 |
| 2019 | Learning to Find Common Objects Across Few Image CollectionsabstractGiven a collection of bags where each bag is a set of images, our goal is to select one image from each bag such that the selected images are from the same object class. We model the selection as an energy minimization problem with unary and pairwise potential functions. Inspired by recent few-shot learning algorithms, we propose an approach to learn the potential functions directly from the data. Furthermore, we propose a fast greedy inference algorithm for energy minimization. We evaluate our approach on few-shot common object recognition as well as object co-localization tasks. Our experiments show that learning the pairwise and unary terms greatly improves the performance of the model over several well-known methods for these tasks. The proposed greedy optimization algorithm achieves performance comparable to state-of-the-art structured inference algorithms while being ~10 times faster. Amirreza Shaban, Amir Rahimi, Shray Bansal, Stephen Gould, Byron Boots, Richard I. Hartley |
ICCV | 6 |
| 2019 | Recovering Faces From Portraits with Auxiliary Facial AttributesabstractRecovering a photorealistic face from an artistic portrait is a challenging task since crucial facial details are often distorted or completely lost in artistic compositions. To handle this loss, we propose an Attribute-guided Face Recovery from Portraits (AFRP) that utilizes a Face Recovery Network (FRN) and a Discriminative Network (DN). FRN consists of an autoencoder with residual block-embedded skip-connections and incorporates facial attribute vectors into the feature maps of input portraits at the bottleneck of the autoencoder. DN has multiple convolutional and fully-connected layers, and its role is to enforce FRN to generate authentic face images with corresponding facial attributes dictated by the input attribute vectors. For the preservation of identities, we impose the recovered and ground-truth faces to share similar visual features. Specifically, DN determines whether the recovered image looks like a real face and checks if the facial attributes extracted from the recovered image are consistent with given attributes. Our method can recover photorealistic identity-preserving faces with desired attributes from unseen stylized portraits, artistic paintings, and hand-drawn sketches. On large-scale synthesized and sketch datasets, we demonstrate that our face recovery method achieves state-of-the-art results. Fatemeh Shiri, Xin Yu 0002, Fatih Porikli, Richard I. Hartley, Piotr Koniusz |
WACV | 4 |
| 2019 | Identity-Preserving Face Recovery from Stylized Portraits
Fatemeh Shiri, Xin Yu 0002, Fatih Porikli, Richard I. Hartley, Piotr Koniusz |
Int. J. Comput. Vis. | 4 |
| 2019 | Memory Efficient Max Flow for Multi-Label Submodular MRFsabstractMulti-label submodular Markov Random Fields (MRFs) have been shown to be solvable using max-flow based on an encoding of the labels proposed by Ishikawa, in which each variable$X_i$is represented by$\ell$nodes (where$\ell$is the number of labels) arranged in a column. However, this method in general requires$2\;\ell ^2$edges for each pair of neighbouring variables. This makes it inapplicable to realistic problems with many variables and labels, due to excessive memory requirement. In this paper, we introduce a variant of the max-flow algorithm that requires much less storage. Consequently, our algorithm makes it possible to optimally solve multi-label submodular problems involving large numbers of variables and labels on a standard computer. Thalaiyasingam Ajanthan, Richard I. Hartley, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Scalable Deep k-Subspace Clustering
Tong Zhang 0023, Pan Ji, Mehrtash Harandi, Richard I. Hartley, Ian D. Reid 0001 |
ACCV (5) | 4 |
| 2018 | Non-Linear Temporal Subspace Representations for Activity RecognitionabstractRepresentations that can compactly and effectively capture the temporal evolution of semantic content are important to computer vision and machine learning algorithms that operate on multi-variate time-series data. We investigate such representations motivated by the task of human action recognition. Here each data instance is encoded by a multivariate feature (such as via a deep CNN) where action dynamics are characterized by their variations in time. As these features are often non-linear, we propose a novel pooling method, kernelized rank pooling, that represents a given sequence compactly as the pre-image of the parameters of a hyperplane in a reproducing kernel Hilbert space, projections of data onto which captures their temporal order. We develop this idea further and show that such a pooling scheme can be cast as an order-constrained kernelized PCA objective. We then propose to use the parameters of a kernelized low-rank feature subspace as the representation of the sequences. We cast our formulation as an optimization problem on generalized Grassmann manifolds and then solve it efficiently using Riemannian optimization techniques. We present experiments on several action recognition datasets using diverse feature modalities and demonstrate state-of-the-art results. Anoop Cherian, Suvrit Sra, Stephen Gould, Richard I. Hartley |
CVPR | 4 |
| 2018 | Super-Resolving Very Low-Resolution Face Images With Supplementary AttributesabstractGiven a tiny face image, existing face hallucination methods aim at super-resolving its high-resolution (HR) counterpart by learning a mapping from an exemplar dataset. Since a low-resolution (LR) input patch may correspond to many HR candidate patches, this ambiguity may lead to distorted HR facial details and wrong attributes such as gender reversal. An LR input contains low-frequency facial components of its HR version while its residual face image, defined as the difference between the HR ground-truth and interpolated LR images, contains the missing high-frequency facial details. We demonstrate that supplementing residual images or feature maps with additional facial attribute information can significantly reduce the ambiguity in face super-resolution. To explore this idea, we develop an attribute-embedded upsampling network, which consists of an upsampling network and a discriminative network. The upsampling network is composed of an autoencoder with skip-connections, which incorporates facial attribute vectors into the residual features of LR inputs at the bottleneck of the autoencoder and deconvolutional layers used for upsampling. The discriminative network is designed to examine whether super-resolved faces contain the desired attributes or not and then its loss is used for updating the upsampling network. In this manner, we can super-resolve tiny (16×16 pixels) unaligned face images with a large upscaling factor of 8× while reducing the uncertainty of one-to-many mappings remarkably. By conducting extensive evaluations on a large-scale dataset, we demonstrate that our method achieves superior face hallucination results and outperforms the state-of-the-art. Xin Yu 0002, Basura Fernando, Richard I. Hartley, Fatih Porikli |
CVPR | 3 |
| 2018 | Deep Unsupervised Saliency Detection: A Multiple Noisy Labeling PerspectiveabstractThe success of current deep saliency detection methods heavily depends on the availability of large-scale supervision in the form of per-pixel labeling. Such supervision, while labor-intensive and not always possible, tends to hinder the generalization ability of the learned models. By contrast, traditional handcrafted features based unsupervised saliency detection methods, even though have been surpassed by the deep supervised methods, are generally dataset-independent and could be applied in the wild. This raises a natural question that "Is it possible to learn saliency maps without using labeled data while improving the generalization ability?". To this end, we present a novel perspective to unsupervised saliency detection through learning from multiple noisy labeling generated by "weak" and "noisy" unsupervised handcrafted saliency methods. Our end-to-end deep learning framework for unsupervised saliency detection consists of a latent saliency prediction module and a noise modeling module that work collaboratively and are optimized jointly. Explicit noise modeling enables us to deal with noisy saliency maps in a probabilistic way. Extensive experimental results on various benchmarking datasets show that our model not only outperforms all the unsupervised saliency methods with a large margin but also achieves comparable performance with the recent state-of-the-art supervised deep saliency methods. Jing Zhang 0052, Tong Zhang 0023, Yuchao Dai, Mehrtash Harandi, Richard I. Hartley |
CVPR | 5 |
| 2018 | Action Anticipation with RBF Kernelized Feature Mapping RNN
Yuge Shi, Basura Fernando, Richard I. Hartley |
ECCV (10) | 3 |
| 2018 | Face Super-Resolution Guided by Facial Component Heatmaps
Xin Yu 0002, Basura Fernando, Bernard Ghanem, Fatih Porikli, Richard I. Hartley |
ECCV (9) | 5 |
| 2018 | Identity-Preserving Face Recovery from PortraitsabstractRecovering the latent photorealistic faces from their artistic portraits aids human perception and facial analysis. However, a recovery process that can preserve identity is challenging because the fine details of real faces can be distorted or lost in stylized images. In this paper, we present a new Identity-preserving Face Recovery from Portraits (IFRP) to recover latent photorealistic faces from unaligned stylized portraits. Our IFRP method consists of two components: Style Removal Network (SRN) and Discriminative Network (DN). The SRN is designed to transfer feature maps of stylized images to the feature maps of the corresponding photorealistic faces. By embedding spatial transformer networks into the SRN, our method can compensate for misalignments of stylized faces automatically and output aligned realistic face images. The role of the DN is to enforce recovered faces to be similar to authentic faces. To ensure the identity preservation, we promote the recovered and ground-truth faces to share similar visual features via a distance measure which compares features of recovered and ground-truth faces extracted from a pre-trained VGG network. We evaluate our method on a large-scale synthesized dataset of real and stylized face pairs and attain state of the art results. In addition, our method can recover photorealistic faces from previously unseen stylized portraits, original paintings and human-drawn sketches. Fatemeh Shiri, Fatih Porikli, Richard I. Hartley, Piotr Koniusz |
WACV | 3 |
| 2018 | Dimensionality Reduction on SPD Manifolds: The Emergence of Geometry-Aware MethodsabstractRepresenting images and videos with Symmetric Positive Definite (SPD) matrices, and considering the Riemannian geometry of the resulting space, has been shown to yield high discriminative power in many visual recognition tasks. Unfortunately, computation on the Riemannian manifold of SPD matrices -especially of high-dimensional ones- comes at a high cost that limits the applicability of existing techniques. In this paper, we introduce algorithms able to handle high-dimensional SPD matrices by constructing a lower-dimensional SPD manifold. To this end, we propose to model the mapping from the high-dimensional SPD manifold to the low-dimensional one with an orthonormal projection. This lets us formulate dimensionality reduction as the problem of finding a projection that yields a low-dimensional manifold either with maximum discriminative power in the supervised scenario, or with maximum variance of the data in the unsupervised one. We show that learning can be expressed as an optimization problem on a Grassmann manifold and discuss fast solutions for special cases. Our evaluation on several classification tasks evidences that our approach leads to a significant accuracy gain over state-of-the-art methods. Mehrtash Harandi, Mathieu Salzmann, Richard I. Hartley |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Joint Dimensionality Reduction and Metric Learning: A Geometric TakeabstractTo be tractable and robust to data noise, existing metric learning algorithms commonly rely on PCA as a pre-processing step. How can we know, however, that PCA, or any other specific dimensionality reduction technique, is the method of choice for the problem at hand? The answer is simple: We cannot! To address this issue, in this paper, we develop a Riemannian framework to jointly learn a mapping performing dimensionality reduction and a metric in the induced space. Our experiments evidence that, while we directly work on high-dimensional features, our approach yields competitive runtimes with and higher accuracy than state-of-the-art metric learning algorithms. Mehrtash Harandi, Mathieu Salzmann, Richard I. Hartley |
ICML | 3 |
| 2017 | Automated detection and tracking of slalom paddlers from broadcast image sequences using cascade classifiers and discriminative correlation filters
Ami Drory, Gao Zhu, Hongdong Li, Richard I. Hartley |
Comput. Vis. Image Underst. | 4 |
| 2016 | Memory Efficient Max Flow for Multi-label Submodular MRFsabstractMulti-label submodular Markov Random Fields (MRFs) have been shown to be solvable using max-flow based on an encoding of the labels proposed by Ishikawa, in which each variable Xi is represented by l nodes (where l is the number of labels) arranged in a column. However, this method in general requires 2 l2 edges for each pair of neighbouring variables. This makes it inapplicable to realistic problems with many variables and labels, due to excessive memory requirement. In this paper, we introduce a variant of the max-flow algorithm that requires much less storage. Consequently, our algorithm makes it possible to optimally solve multi-label submodular problems involving large numbers of variables and labels on a standard computer. Thalaiyasingam Ajanthan, Richard I. Hartley, Mathieu Salzmann |
CVPR | 2 |
| 2016 | The generalized relative pose and scale problem: View-graph fusion via 2D-2D registrationabstractIt is well-known that the relative pose problem can be generalized to non-central cameras. We present a further generalization, denoted the generalized relative pose and scale problem. It has surprising importance for classical problems such as solving similarity transformations for view-graph concatenation in hierarchical structure from motion and loop-closure in visual SLAM, both posed as a 2D-2D registration problem. The relative pose problem and all its generalizations constitute a family of similar symmetric eigenvalue problems, which allow us to compress data and find a geometrically meaningful solution by an efficient search in the space of rotations. While the derivation of a completely general closed-form solver appears intractable, we make use of a simple heuristic global energy minimization scheme based on local minimum suppression, returning outstanding performance in practically relevant scenarios. Efficiency and reliability of our algorithm are demonstrated on both simulated and real data, supporting our claim of superior performance with respect to both generalized 2D-3D and 3D-3D registration approaches. By directly employing image information, we avoid the common noise in point clouds occurring especially along the depth direction. Laurent Kneip, Chris Sweeney, Richard I. Hartley |
WACV | 3 |
| 2016 | Sparse Coding on Symmetric Positive Definite Manifolds Using Bregman DivergencesabstractThis paper introduces sparse coding and dictionary learning for symmetric positive definite (SPD) matrices, which are often used in machine learning, computer vision, and related areas. Unlike traditional sparse coding schemes that work in vector spaces, in this paper, we discuss how SPD matrices can be described by sparse combination of dictionary atoms, where the atoms are also SPD matrices. We propose to seek sparse coding by embedding the space of SPD matrices into the Hilbert spaces through two types of the Bregman matrix divergences. This not only leads to an efficient way of performing sparse coding but also an online and iterative scheme for dictionary learning. We apply the proposed methods to several computer vision tasks where images are represented by region covariance matrices. Our proposed algorithms outperform state-of-the-art methods on a wide range of classification tasks, including face recognition, action recognition, material classification, and texture categorization. Mehrtash Harandi, Richard I. Hartley, Brian C. Lovell, Conrad Sanderson |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | Iteratively reweighted graph cut for multi-label MRFs with non-convex priorsabstractWhile widely acknowledged as highly effective in computer vision, multi-label MRFs with non-convex priors are difficult to optimize. To tackle this, we introduce an algorithm that iteratively approximates the original energy with an appropriately weighted surrogate energy that is easier to minimize. Our algorithm guarantees that the original energy decreases at each iteration. In particular, we consider the scenario where the global minimizer of the weighted surrogate energy can be obtained by a multi-label graph cut algorithm, and show that our algorithm then lets us handle of large variety of non-convex priors. We demonstrate the benefits of our method over state-of-the-art MRF energy minimization techniques on stereo and inpainting problems. Thalaiyasingam Ajanthan, Richard I. Hartley, Mathieu Salzmann, Hongdong Li |
CVPR | 2 |
| 2015 | A linear least-squares solution to elastic Shape-from-TemplateabstractWe cast SfT (Shape-from-Template) as the search of a vector field (X, δX), composed of the pose X and the displacement δX that produces the deformation. We propose the first fully linear least-squares SfT method modeling elastic deformations. It relies on a set of Solid Boundary Constraints (SBC) to position the template at X in the deformed frame. The displacement is mapped by the stiffness matrix to minimize the amount of force responsible for the deformation. This linear minimization is subjected to the Reprojection Boundary Constraints (RBC) of the deformed shape X + δX on the deformed image. Compared to state-of-the-art methods, this new formulation allows us to obtain accurate results at a low computation cost. Abed Malti, Adrien Bartoli, Richard I. Hartley |
CVPR | 3 |
| 2015 | LQ-bundle adjustmentabstractIn this paper we propose a method to solve for an Lqsolution of bundle adjustment, a non-linear parameter estimation problem. Given a set of images of a scene, bundle adjustment simultaneously estimates camera parameters and 3D structure of the scene. Generally, a least squares criterion is minimized by using the Levenberg-Marquardt (LM) method, a non-linear least squares optimization method. It is known that the least squares methods are not robust to outliers, even a single outlier can deviate the solution from its true value. Therefore, we propose a method to minimize an Lqcost function, for 1 ≤ qqcost function minimizes the sum of the q-th power of errors. The proposed method has an advantage of using the Levenberg-Marquardt (LM) method to find a robust solution of the problem. Our experimental results confirm that the proposed method is more robust to outliers than the standard least squares method. Khurrum Aftab, Richard I. Hartley |
ICIP | 2 |
| 2015 | Convergence of Iteratively Re-weighted Least Squares to Robust M-EstimatorsabstractThis paper presents a way of using the Iteratively Reweighted Least Squares (IRLS) method to minimize several robust cost functions such as the Huber function, the Cauchy function and others. It is known that IRLS (otherwise known as Weiszfeld) techniques are generally more robust to outliers than the corresponding least squares methods, but the full range of robust M-estimators that are amenable to IRLS has not been investigated. In this paper we address this question and show that IRLS methods can be used to minimize most common robust M-estimators. An exact condition is given and proved for decrease of the cost, from which convergence follows. In addition to the advantage of increased robustness, the proposed algorithm is far simpler than the standard L1 Weiszfeld algorithm. We show the applicability of the proposed algorithm to the rotation averaging, triangulation and point cloud alignment problems. Khurrum Aftab, Richard I. Hartley |
WACV | 2 |
| 2015 | Lq-Closest-Point to Affine Subspaces Using the Generalized Weiszfeld Algorithm
Khurrum Aftab, Richard I. Hartley, Jochen Trumpf |
Int. J. Comput. Vis. | 2 |
| 2015 | Extrinsic Methods for Coding and Dictionary Learning on Grassmann Manifolds
Mehrtash Harandi, Richard I. Hartley, Chunhua Shen, Brian C. Lovell, Conrad Sanderson |
Int. J. Comput. Vis. | 2 |
| 2015 | A Generalized Projective Reconstruction Theorem and Depth Constraints for Projective Factorization
Behrooz Nasihatkon, Richard I. Hartley, Jochen Trumpf |
Int. J. Comput. Vis. | 2 |
| 2015 | Generalized Weiszfeld Algorithms for Lq OptimizationabstractIn many computer vision applications, a desired model of some type is computed by minimizing a cost function based on several measurements. Typically, one may compute the model that minimizes the L2 cost, that is the sum of squares of measurement errors with respect to the model. However, the Lq solution which minimizes the sum of the qth power of errors usually gives more robust results in the presence of outliers for some values of q, for example, q = 1. The Weiszfeld algorithm is a classic algorithm for finding the geometric L1 mean of a set of points in Euclidean space. It is provably optimal and requires neither differentiation, nor line search. The Weiszfeld algorithm has also been generalized to find the L1 mean of a set of points on a Riemannian manifold of non-negative curvature. This paper shows that the Weiszfeld approach may be extended to a wide variety of problems to find an Lq mean for 1 ≤ q <; 2, while maintaining simplicity and provable convergence. We apply this problem to both single-rotation averaging (under which the algorithm provably finds the global Lq optimum) and multiple rotation averaging (for which no such proof exists). Experimental results of Lq optimization for rotations show the improved reliability and robustness compared to L2 optimization. Khurrum Aftab, Richard I. Hartley, Jochen Trumpf |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Kernel Methods on Riemannian Manifolds with Gaussian RBF KernelsabstractIn this paper, we develop an approach to exploiting kernel methods with manifold-valued data. In many computer vision problems, the data can be naturally represented as points on a Riemannian manifold. Due to the non-Euclidean geometry of Riemannian manifolds, usual Euclidean computer vision and machine learning algorithms yield inferior results on such data. In this paper, we define Gaussian radial basis function (RBF)-based positive definite kernels on manifolds that permit us to embed a given manifold with a corresponding metric in a high dimensional reproducing kernel Hilbert space. These kernels make it possible to utilize algorithms developed for linear spaces on nonlinear manifold-valued data. Since the Gaussian RBF defined with any given metric is not always positive definite, we present a unified framework for analyzing the positive definiteness of the Gaussian RBF on a generic metric space. We then use the proposed framework to identify positive definite kernels on two specific manifolds commonly encountered in computer vision: the Riemannian manifold of symmetric positive definite matrices and the Grassmann manifold, i.e., the Riemannian manifold of linear subspaces of a Euclidean space. We show that many popular algorithms designed for Euclidean spaces, such as support vector machines, discriminant analysis and principal component analysis can be generalized to Riemannian manifolds with the help of such positive definite Gaussian kernels. Sadeep Jayasumana, Richard I. Hartley, Mathieu Salzmann, Hongdong Li, Mehrtash Harandi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Mirror Surface Reconstruction from a Single ImageabstractThis paper tackles the problem of reconstructing the shape of a smooth mirror surface from a single image. In particular, we consider the case where the camera is observing the reflection of a static reference target in the unknown mirror. We first study the reconstruction problem given dense correspondences between 3D points on the reference target and image locations. In such conditions, our differential geometry analysis provides a theoretical proof that the shape of the mirror surface can be recovered if the pose of the reference target is known. We then relax our assumptions by considering the case where only sparse correspondences are available. In this scenario, we formulate reconstruction as an optimization problem, which can be solved using a nonlinear least-squares method. We demonstrate the effectiveness of our method on both synthetic and real images. We then provide a theoretical analysis of the potential degenerate cases with and without prior knowledge of the pose of the reference target. Finally we show that our theory can be similarly applied to the reconstruction of the surface of transparent object. Miaomiao Liu 0001, Richard I. Hartley, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Multi-Target Tracking With Time-Varying Clutter Rate and Detection Profile: Application to Time-Lapse Cell Microscopy SequencesabstractQuantitative analysis of the dynamics of tiny cellular and sub-cellular structures, known as particles, in time-lapse cell microscopy sequences requires the development of a reliable multi-target tracking method capable of tracking numerous similar targets in the presence of high levels of noise, high target density, complex motion patterns and intricate interactions. In this paper, we propose a framework for tracking these structures based on the random finite set Bayesian filtering framework. We focus on challenging biological applications where image characteristics such as noise and background intensity change during the acquisition process. Under these conditions, detection methods usually fail to detect all particles and are often followed by missed detections and many spurious measurements with unknown and time-varying rates. To deal with this, we propose a bootstrap filter composed of an estimator and a tracker. The estimator adaptively estimates the required meta parameters for the tracker such as clutter rate and the detection probability of the targets, while the tracker estimates the state of the targets. Our results show that the proposed approach can outperform state-of-the-art particle trackers on both synthetic and real data in this regime. Seyed Hamid Rezatofighi, Stephen Gould, Ba-Tuong Vo, Ba-Ngu Vo, Katarina Mele, Richard I. Hartley |
IEEE Trans. Medical Imaging | 6 |
| 2014 | Keynote lecture 2: "Riemannian manifolds, kernels and learning"abstractSummary form only given. I will talk about recent results from a number of people in my group on Riemannian manifolds in computer vision. In many Vision problems Riemannian manifolds come up as a natural model. Data related to a problem can be naturally represented as a point on a Riemannian manifold. This talk will give an intuitive introduction to Riemannian manifolds, and show how they can be applied in many situations. Manifolds of interest include the manifold of Positive Definite matrices and the Grassman Manifolds, which have a role in object recognition and classification, and the Kendall shape manifold, which represents the shape of 2D objects. Of particular interest is the question of when one can define positive-definite kernels on Riemannian manifolds. This would allow the application of kernel techniques of SVMs, Kernel FDA, dictionary learning etc directly on the manifold. Richard I. Hartley |
AVSS | 1 |
| 2014 | Optimizing over Radial Kernels on Compact ManifoldsabstractWe tackle the problem of optimizing over all possible positive definite radial kernels on Riemannian manifolds for classification. Kernel methods on Riemannian manifolds have recently become increasingly popular in computer vision. However, the number of known positive definite kernels on manifolds remain very limited. Furthermore, most kernels typically depend on at least one parameter that needs to be tuned for the problem at hand. A poor choice of kernel, or of parameter value, may yield significant performance drop-off. Here, we show that positive definite radial kernels on the unit n-sphere, the Grassmann manifold and Kendall's shape manifold can be expressed in a simple form whose parameters can be automatically optimized within a support vector machine framework. We demonstrate the benefits of our kernel learning algorithm on object, face, action and shape recognition. Sadeep Jayasumana, Richard I. Hartley, Mathieu Salzmann, Hongdong Li, Mehrtash Harandi |
CVPR | 2 |
| 2014 | On Projective Reconstruction in Arbitrary DimensionsabstractWe study the theory of projective reconstruction for multiple projections from an arbitrary dimensional projective space into lower-dimensional spaces. This problem is important due to its applications in the analysis of dynamical scenes. The current theory, due to Hartley and Schaffalitzky, is based on the Grassmann tensor, generalizing the ideas of fundamental matrix, trifocal tensor and quadrifocal tensor used in the well-studied case of 3D to 2D projections. We present a theory whose point of departure is the projective equations rather than the Grassmann tensor. This is a better fit for the analysis of approaches such as bundle adjustment and projective factorization which seek to directly solve the projective equations. In a first step, we prove that there is a unique Grassmann tensor corresponding to each set of image points, a question that remained open in the work of Hartley and Schaffalitzky. Then, we prove that projective equivalence follows from the set of projective equations given certain conditions on the estimated camera-point setup or the estimated projective depths. Finally, we demonstrate how wrong solutions to the projective factorization problem can happen, and classify such degenerate solutions based on the zero patterns in the estimated depth matrix. Behrooz Nasihatkon, Richard I. Hartley, Jochen Trumpf |
CVPR | 2 |
| 2014 | Globally Optimal Inlier Set Maximization with Unknown Rotation and Focal Length
Jean-Charles Bazin, Yongduek Seo, Richard I. Hartley, Marc Pollefeys |
ECCV (2) | 3 |
| 2014 | From Manifold to Manifold: Geometry-Aware Dimensionality Reduction for SPD Matrices
Mehrtash Harandi, Mathieu Salzmann, Richard I. Hartley |
ECCV (2) | 3 |
| 2014 | Expanding the Family of Grassmannian Kernels: An Embedding Perspective
Mehrtash Harandi, Mathieu Salzmann, Sadeep Jayasumana, Richard I. Hartley, Hongdong Li |
ECCV (7) | 4 |
| 2014 | Monocular template-based 3D surface reconstruction: Convex inextensible and nonconvex isometric methods
Florent Brunet, Adrien Bartoli, Richard I. Hartley |
Comput. Vis. Image Underst. | 3 |
| 2013 | Kernel Methods on the Riemannian Manifold of Symmetric Positive Definite MatricesabstractSymmetric Positive Definite (SPD) matrices have become popular to encode image information. Accounting for the geometry of the Riemannian manifold of SPD matrices has proven key to the success of many algorithms. However, most existing methods only approximate the true shape of the manifold locally by its tangent plane. In this paper, inspired by kernel methods, we propose to map SPD matrices to a high dimensional Hilbert space where Euclidean geometry applies. To encode the geometry of the manifold in the mapping, we introduce a family of provably positive definite kernels on the Riemannian manifold of SPD matrices. These kernels are derived from the Gaussian kernel, but exploit different metrics on the manifold. This lets us extend kernel-based algorithms developed for Euclidean spaces, such as SVM and kernel PCA, to the Riemannian manifold of SPD matrices. We demonstrate the benefits of our approach on the problems of pedestrian detection, object categorization, texture analysis, 2D motion segmentation and Diffusion Tensor Imaging (DTI) segmentation. Sadeep Jayasumana, Richard I. Hartley, Mathieu Salzmann, Hongdong Li, Mehrtash Harandi |
CVPR | 2 |
| 2013 | Mirror Surface Reconstruction from a Single ImageabstractThis paper tackles the problem of reconstructing the shape of a smooth mirror surface from a single image. In particular, we consider the case where the camera is ob-serving the reflection of a static reference target in the un-known mirror. We first study the reconstruction problem given dense correspondences between 3D points on the ref-erence target and image locations. In such conditions, our differential geometry analysis provides a theoretical proof that the shape of the mirror surface can be uniquely recov-ered if the pose of the reference target is known. We then relax our assumptions by considering the case where only sparse correspondences are available. In this scenario, we formulate reconstruction as an optimization problem, which can be solved using a nonlinear least-squares method. We demonstrate the effectiveness of our method on both syn-thetic and real images. 1. Miaomiao Liu 0001, Richard I. Hartley, Mathieu Salzmann |
CVPR | 2 |
| 2013 | Monocular Template-Based 3D Reconstruction of Extensible Surfaces with Local Linear ElasticityabstractWe propose a new approach for template-based extensible surface reconstruction from a single view. We extend the method of isometric surface reconstruction and more recent work on conformal surface reconstruction. Our approach relies on the minimization of a proposed stretching energy formalized with respect to the Poisson ratio parameter of the surface. We derive a patch-based formulation of this stretching energy by assuming local linear elasticity. This formulation unifies geometrical and mechanical constraints in a single energy term. We prevent local scale ambiguities by imposing a set of fixed boundary 3D points. We experimentally prove the sufficiency of this set of boundary points and demonstrate the effectiveness of our approach on different developable and non-developable surfaces with a wide range of extensibility. Abed Malti, Richard I. Hartley, Adrien Bartoli, Jae-Hak Kim |
CVPR | 2 |
| 2013 | Verifying Global Minima for L 2 Minimization Problems in Multiple View Geometry
Richard I. Hartley, Fredrik Kahl, Carl Olsson, Yongduek Seo |
Int. J. Comput. Vis. | 1 |
| 2013 | Rotation Averaging
Richard I. Hartley, Jochen Trumpf, Yuchao Dai, Hongdong Li |
Int. J. Comput. Vis. | 1 |
| 2013 | Image Set Based Face Recognition Using Self-Regularized Non-Negative Coding and Adaptive Distance Metric LearningabstractSimple nearest neighbor classification fails to exploit the additional information in image sets. We propose self-regularized nonnegative coding to define between set distance for robust face recognition. Set distance is measured between the nearest set points (samples) that can be approximated from their orthogonal basis vectors as well as from the set samples under the respective constraints of self-regularization and nonnegativity. Self-regularization constrains the orthogonal basis vectors to be similar to the approximated nearest point. The nonnegativity constraint ensures that each nearest point is approximated from a positive linear combination of the set samples. Both constraints are formulated as a single convex optimization problem and the accelerated proximal gradient method with linear-time Euclidean projection is adapted to efficiently find the optimal nearest points between two image sets. Using the nearest points between a query set and all the gallery sets as well as the active samples used to approximate them, we learn a more discriminative Mahalanobis distance for robust face recognition. The proposed algorithm works independently of the chosen features and has been tested on gray pixel values and local binary patterns. Experiments on three standard data sets show that the proposed method consistently outperforms existing state-of-the-art methods. Ajmal Mian, Yiqun Hu, Richard I. Hartley, Robyn A. Owens |
IEEE Trans. Image Process. | 3 |
| 2012 | Sparse Coding and Dictionary Learning for Symmetric Positive Definite Matrices: A Kernel Approach
Mehrtash Harandi, Conrad Sanderson, Richard I. Hartley, Brian C. Lovell |
ECCV (2) | 3 |
| 2012 | Application of the IMM-JPDA Filter to Multiple Target Tracking in Total Internal Reflection Fluorescence Microscopy Images
Seyed Hamid Rezatofighi, Stephen Gould, Richard I. Hartley, Katarina Mele, William E. Hughes |
MICCAI (1) | 3 |
| 2012 | An Efficient Hidden Variable Approach to Minimal-Case Camera Motion EstimationabstractIn this paper, we present an efficient new approach for solving two-view minimal-case problems in camera motion estimation, most notably the so-called five-point relative orientation problem and the six-point focal-length problem. Our approach is based on the hidden variable technique used in solving multivariate polynomial systems. The resulting algorithm is conceptually simple, which involves a relaxation which replaces monomials in all but one of the variables to reduce the problem to the solution of sets of linear equations, as well as solving a polynomial eigenvalue problem (polyeig). To efficiently find the polynomial eigenvalues, we make novel use of several numeric techniques, which include quotient-free Gaussian elimination, Levinson-Durbin iteration, and also a dedicated root-polishing procedure. We have tested the approach on different minimal cases and extensions, with satisfactory results obtained. Both the executables and source codes of the proposed algorithms are made freely downloadable. Richard I. Hartley, Hongdong Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | L1 rotation averaging using the Weiszfeld algorithmabstractWe consider the problem of rotation averaging under the L1norm. This problem is related to the classic Fermat-Weber problem for finding the geometric median of a set of points in IRn. We apply the classical Weiszfeld algorithm to this problem, adapting it iteratively in tangent spaces of SO(3) to obtain a provably convergent algorithm for finding the L1mean. This results in an extremely simple and rapid averaging algorithm, without the need for line search. The choice of L1mean (also called geometric median) is motivated by its greater robustness compared with rotation averaging under the L2norm (the usual averaging process). We apply this problem to both single-rotation averaging (under which the algorithm provably finds the global L1optimum) and multiple rotation averaging (for which no such proof exists). The algorithm is demonstrated to give markedly improved results, compared with L2averaging. We achieve a median rotation error of 0.82 degrees on the 595 images of the Notre Dame image set. Richard I. Hartley, Khurrum Aftab, Jochen Trumpf |
CVPR | 1 |
| 2011 | Graph connectivity in sparse subspace clusteringabstractSparse Subspace Clustering (SSC) is one of the recent approaches to subspace segmentation. In SSC a graph is constructed whose nodes are the data points and whose edges are inferred from the L1-sparse representation of each point by the others. It has been proved that if the points lie on a mixture of independent subspaces, the graphical structure of each subspace is disconnected from the others. However, the problem of connectivity within each subspace is still unanswered. This is important since the subspace segmentation in SSC is based on finding the connected components of the graph. Our analysis is built upon the connection between the sparse representation through L1-norm minimization and the geometry of convex poly-topes proposed by the compressed sensing community. After introduction of some assumptions to make the problem well-defined, it is proved that the connectivity within each subspace holds for 2- and 3-dimensional subspaces. The claim of connectivity for general d-dimensional case, even for generic configurations, is proved false by giving a counterexample in dimensions greater than 3. Behrooz Nasihatkon, Richard I. Hartley |
CVPR | 2 |
| 2011 | Superpixels via pseudo-Boolean optimizationabstractWe propose an algorithm for creating superpixels. The major step in our algorithm is simply minimizing two pseudo-Boolean functions. The processing time of our algorithm on images of moderate size is only half a second. Experiments on a benchmark dataset show that our method produces superpixels of comparable quality with existing algorithms. Last but not least, the speed of our algorithm is independent of the number of superpixels, which is usually the bottle-neck for the traditional algorithms of superpixel creation. Yuhang Zhang 0001, Richard I. Hartley, John Mashford, Stewart Burn |
ICCV | 2 |
| 2011 | Monocular Template-based Reconstruction of Inextensible Surfaces
Mathieu Perriollat, Richard I. Hartley, Adrien Bartoli |
Int. J. Comput. Vis. | 2 |
| 2010 | Monocular Template-Based Reconstruction of Smooth and Inextensible Surfaces
Florent Brunet, Richard I. Hartley, Adrien Bartoli, Nassir Navab, Rémy Malgouyres |
ACCV (3) | 2 |
| 2010 | Gradual Sampling and Mutual Information Maximisation for Markerless Motion Capture
Lei Wang 0001, Richard I. Hartley, Hongdong Li, Dan Xu 0001 |
ACCV (2) | 3 |
| 2010 | Compressive Evaluation in Human Motion Tracking
Lei Wang 0001, Richard I. Hartley, Hongdong Li, Dan Xu 0001 |
ACCV (4) | 3 |
| 2010 | Pyramid Center-Symmetric Local Binary/Trinary Patterns for Effective Pedestrian Detection
Yongbin Zheng, Chunhua Shen, Richard I. Hartley, Xinsheng Huang |
ACCV (4) | 3 |
| 2010 | Outlier removal using dualityabstractIn this paper we consider the problem of outlier removal for large scale multiview reconstruction problems. An efficient and very popular method for this task is RANSAC. However, as RANSAC only works on a subset of the images, mismatches in longer point tracks may go undetected. To deal with this problem we would like to have, as a post processing step to RANSAC, a method that works on the entire (or a larger) part of the sequence. In this paper we consider two algorithms for doing this. The first one is related to a method by Sim & Hartley where a quasiconvex problem is solved repeatedly and the error residuals with the largest error is removed. Instead of solving a quasiconvex problem in each step we show that it is enough to solve a single LP or SOCP which yields a significant speedup. Using duality we show that the same theoretical result holds for our method. The second algorithm is a faster version of the first, and it is related to the popular method of L1-optimization. While it is faster and works very well in practice, there is no theoretical guarantee of success. We show that these two methods are related through duality, and evaluate the methods on a number of data sets with promising results. Carl Olsson, Anders P. Eriksson, Richard I. Hartley |
CVPR | 3 |
| 2010 | Fast Multi-labelling for Stereo Matching
Yuhang Zhang 0001, Richard I. Hartley, Lei Wang 0001 |
ECCV (3) | 2 |
| 2010 | Motion Estimation for Nonoverlapping Multicamera Rigs: Linear Algebraic and {\rm L}_\infty Geometric SolutionsabstractWe investigate the problem of estimating the ego-motion of a multicamera rig from two positions of the rig. We describe and compare two new algorithms for finding the 6 degrees of freedom (3 for rotation and 3 for translation) of the motion. One algorithm gives a linear solution and the other is a geometric algorithm that minimizes the maximum measurement error-the optimal L{infinity} solution. They are described in the context of the General Camera Model (GCM), and we pay particular attention to multicamera systems in which the cameras have nonoverlapping or minimally overlapping field of view. Many nonlinear algorithms have been developed to solve the multicamera motion estimation problem. However, no linear solution or guaranteed optimal geometric solution has previously been proposed. We made two contributions: 1) a fast linear algebraic method using the GCM and 2) a guaranteed globally optimal algorithm based on the L{infinity} geometric error using the branch-and-bound technique. In deriving the linear method using the GCM, we give a detailed analysis of degeneracy of camera configurations. In finding the globally optimal solution, we apply a rotation space search technique recently proposed by Hartley and Kahl. Our experiments conducted on both synthetic and real data have shown excellent results. Jae-Hak Kim, Hongdong Li, Richard I. Hartley |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | Rotation Averaging with Application to Camera-Rig Calibration
Yuchao Dai, Jochen Trumpf, Hongdong Li, Nick Barnes, Richard I. Hartley |
ACCV (2) | 5 |
| 2009 | Projective least-squares: Global solutions with local optimizationabstractWork in multiple view geometry has focused on obtaining globally optimal solutions at the price of computational time efficiency. On the other hand, traditional bundle adjustment algorithms have been found to provide good solutions even though there may be multiple local minima. In this paper we justify this observation by giving a simple sufficient condition for global optimality that can be used to verify that a solution obtained from any local method is indeed global. The method is tested on numerous problem instances of both synthetic and real data sets. In the vast majority of cases we are able to verify that the solutions are optimal, in particular for small-scale problems. We also develop a branch and bound procedure that goes beyond verification. In cases where the sufficient condition does not hold, the algorithm returns either of the following two results: (i) a certificate of global optimality for the local solution or (ii) the global solution. Carl Olsson, Fredrik Kahl, Richard I. Hartley |
CVPR | 3 |
| 2009 | Minimizing energy functions on 4-connected lattices using eliminationabstractWe describe an energy minimization algorithm for functions defined on 4-connected lattices, of the type usually encountered in problems involving images. Such functions are often minimized using graph-cuts/max-flow, but this method is only applicable to submodular problems. In this paper, we describe an algorithm that will solve any binary problem, irrespective of whether it is submodular or not, and for multilabel problems we use alpha-expansion. The method is based on the elimination algorithm, which eliminates nodes from the graph until the remaining function is submodular. It can then be solved using max-flow. Values of eliminated variables are recovered using back-substitution. We compare the algorithm's performance against alternative methods for solving non-submodular problems, with favourable results. Peter Carr 0001, Richard I. Hartley |
ICCV | 2 |
| 2009 | Enforcing Monotonic Temporal Evolution in Dry Eye Images
Tamir Yedidya, Peter Carr 0001, Richard I. Hartley, Jean-Pierre Guillon |
MICCAI (1) | 3 |
| 2009 | Global Optimization through Rotation Space Search
Richard I. Hartley, Fredrik Kahl |
Int. J. Comput. Vis. | 1 |
| 2009 | Reconstruction from Projections Using Grassmann Tensors
Richard I. Hartley, Frederik Schaffalitzky |
Int. J. Comput. Vis. | 1 |
| 2009 | Identifying Anatomical Shape Difference by Regularized Discriminative DirectionabstractIdentifying the shape difference between two groups of anatomical objects is important for medical image analysis and computer-aided diagnosis. A method called "discriminative direction" in the literature has been proposed to solve this problem. In that method, the shape difference between groups is identified by deforming a shape along the discriminative direction. This paper conducts a thorough study about inferring this discriminative direction in an efficient and accurate way. First, finding the discriminative direction is reformulated as a preimage problem in kernel-based learning. This provides a complementary but conceptually simpler solution than the previous method. More importantly, we find that a shape deforming along the original discriminative direction cannot faithfully maintain its anatomical correctness. This unnecessarily introduces spurious shape differences and leads to inaccurate analysis. To overcome this problem, this paper further proposes a regularized discriminative direction by requiring a shape to conform to its underlying distribution when it deforms. Two different approaches are developed to impose the regularization, one from the perspective of probability distributions and the other from a geometric point of view, and their relationship is discussed. After verifying their superior performance through controlled experiments, we apply the proposed methods to detecting and localizing the hippocampal shape difference between sexes. We get results consistent with other independent research, providing a more compact representation of the shape difference compared with the established discriminative direction method. Luping Zhou, Richard I. Hartley, Lei Wang 0001, Paulette Lieby, Nick Barnes |
IEEE Trans. Medical Imaging | 2 |
| 2008 | Monocular Template-based Reconstruction of Inextensible SurfacesabstractWe present a monocular 3D reconstruction algorithm for inextensible deformable surfaces. It is based on point correspondences between the actual image and a template. Since the surface is inextensible, its deformations are isometric to the template, for which the surface shape is known. We exploit the underlying distance constraints to recover the 3D shape. Though these constraints have already been investigated in the literature, we propose a new way to handle them. As opposed to previous methods, ours does not require a known initial deformation. Spatial and temporal smoothness priors are easily incorporated. The reconstruction can be used for 3D augmented reality purposes thanks to a fast implementation. We report results on synthetic and real data. Some of them are faced to stereo-based 3D reconstructions to demonstrate the efficiency of our method. Mathieu Perriollat, Richard I. Hartley, Adrien Bartoli |
BMVC | 2 |
| 2008 | Verifying global minima for L2 minimization problemsabstractWe consider the least-squares (L2) triangulation problem and structure-and-motion with known rotatation, or known plane. Although optimal algorithms have been given for these algorithms under an L-infinity cost function, finding optimal least-squares (L2) solutions to these problems is difficult, since the cost functions are not convex, and in the worst case can have multiple minima. Iterative methods can usually be used to find a good solution, but this may be a local minimum. This paper provides a method for verifying whether a local-minimum solution is globally optimal, by providing a simple and rapid test involving the Hessian of the cost function. In tests of a data set involving 277,000 independent triangulation problems, it is shown that the test verifies the global optimality of an iterative solution in over 99.9% of the cases. Richard I. Hartley, Yongduek Seo |
CVPR | 1 |
| 2008 | Motion estimation for multi-camera systems using global optimizationabstractWe present a motion estimation algorithm for multi-camera systems consisting of more than one calibrated camera securely attached on a moving object. So, they move all together, but do not require to have overlapping views across the cameras. The geometrically optimal solution of the motion for the multi-camera systems under Linfinnorm is provided in this paper using a global optimization technique which has been introduced recently in the computer vision research field. Taking advantage of an optimal estimate of the essential matrix through searching rotation space, we provide the optimal solution for translation by using linear programming and branch & bound algorithm. Synthetic and real data experiments are conducted, and they show more robust and improved performance than the previous methods. Jae-Hak Kim, Hongdong Li, Richard I. Hartley |
CVPR | 3 |
| 2008 | A linear approach to motion estimation using generalized camera modelsabstractA well-known theoretical result for motion estimation using the generalized camera model is that 17 corresponding image rays can be used to solve linearly for the motion of a generalized camera. However, this paper shows that for many common configurations of the generalized camera models (e.g., multi-camera rig, catadioptric camera etc.), such a simple 17-point algorithm does not exist, due to some previously overlooked ambiguities. We further discover that, despite the above ambiguities, we are still able to solve the motion estimation problem effectively by a new algorithm proposed in this paper. Our algorithm is essentially linear, easy to implement, and the computational efficiency is very high. Experiments on both real and simulated data show that the new algorithm achieves reasonably high accuracy as well. Hongdong Li, Richard I. Hartley, Jae-Hak Kim |
CVPR | 2 |
| 2008 | Optimised KD-trees for fast image descriptor matchingabstractIn this paper, we look at improving the KD-tree for a specific usage: indexing a large number of SIFT and other types of image descriptors. We have extended priority search, to priority search among multiple trees. By creating multiple KD-trees from the same data set and simultaneously searching among these trees, we have improved the KD-treepsilas search performance significantly.We have also exploited the structure in SIFT descriptors (or structure in any data set) to reduce the time spent in backtracking. By using Principal Component Analysis to align the principal axes of the data with the coordinate axes, we have further increased the KD-treepsilas search performance. Chanop Silpa-Anan, Richard I. Hartley |
CVPR | 2 |
| 2008 | Perspective Nonrigid Shape and Motion Recovery
Richard I. Hartley, René Vidal |
ECCV (1) | 1 |
| 2008 | Real-Time Simulation of Medical Ultrasound from CT Images
Ramtin Shams, Richard I. Hartley, Nassir Navab |
MICCAI (2) | 2 |
| 2008 | Regularized Discriminative Direction for Shape Difference Analysis
Luping Zhou, Richard I. Hartley, Lei Wang 0001, Paulette Lieby, Nick Barnes |
MICCAI (1) | 2 |
| 2008 | Robust 6DOF Motion Estimation for Non-Overlapping, Multi-Camera SystemsabstractThis paper introduces a novel, robust approach for 6DOF motion estimation of a multi-camera system with non-overlapping views. The proposed approach is able to solve the pose estimation, including scale, for a two camera system with non-overlapping views. In contrast to previous approaches, it degrades gracefully if the motion is close to degenerate. For degenerate motions the technique estimates the remaining 5DOF. The proposed technique is evaluated on real and synthetic sequences. Brian Clipp, Jae-Hak Kim, Jan-Michael Frahm, Marc Pollefeys, Richard I. Hartley |
WACV | 5 |
| 2008 | Multiframe Motion Segmentation with Missing Data Using PowerFactorization and GPCA
René Vidal, Roberto Tron, Richard I. Hartley |
Int. J. Comput. Vis. | 3 |
| 2008 | Multiple-View Geometry Under the Linfinity-NormabstractThis paper presents a new framework for solving geometric structure and motion problems based on Linfinity-norm. Instead of using the common sum-of-squares cost-function, that is, the L2-norm, the model-fitting errors are measured using the L-norm. Unlike traditional methods based on L2, our framework allows for efficient computation of global estimates. We show that a variety of structure and motion problems, for example, triangulation, camera resectioning and homography estimation can be recast as quasi-convex optimization problems within this framework. These problems can be efficiently solved using Second-Order Cone Programming (SOCP) which is a standard technique in convex optimization. The methods have been implemented in Matlab and the resulting toolbox has been made publicly available. The algorithms have been validated on real data in different settings on problems with small and large dimensions and with excellent performance. Fredrik Kahl, Richard I. Hartley |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Three-View Multibody Structure from MotionabstractWe propose a geometric approach to 3-D motion segmentation from point correspondences in three perspective views. We demonstrate that after applying a polynomial embedding to the point correspondences they become related by the socalled multibody trilinear constraint and its associated multibody trifocal tensor, which are natural generalizations of the trilinear constraint and the trifocal tensor to multiple motions. We derive a rank constraint on the embedded correspondences, from which one can estimate the number of independent motions as well as linearly solve for the multibody trifocal tensor. We then show how to compute the epipolar lines associated with each image point from the common root of a set of univariate polynomials and the epipoles by solving a pair of plane clustering problems using Generalized PCA (GPCA). The individual trifocal tensors are then obtained from the second order derivatives of the multibody trilinear constraint. Given epipolar lines and epipoles, or trifocal tensors, one can immediately obtain an initial clustering of the correspondences. We use this clustering to initialize an iterative algorithm that alternates between the computation of the trifocal tensors and the segmentation of the correspondences. We test our algorithm on various synthetic and real scenes, and compare with other algebraic and iterative algorithms. René Vidal, Richard I. Hartley |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Optimal Algorithms in Multiview Geometry
Richard I. Hartley, Fredrik Kahl |
ACCV (1) | 1 |
| 2007 | Visual Odometry for Non-overlapping Views Using Second-Order Cone Programming
Jae-Hak Kim, Richard I. Hartley, Jan-Michael Frahm, Marc Pollefeys |
ACCV (2) | 2 |
| 2007 | A Fast Optimal Algorithm for L 2 Triangulation
Fangfang Lu, Richard I. Hartley |
ACCV (2) | 2 |
| 2007 | Sequential Linfinity Norm Minimization for Triangulation
Yongduek Seo, Richard I. Hartley |
ACCV (2) | 2 |
| 2007 | Where's the Weet-Bix?
Yuhang Zhang 0001, Lei Wang 0001, Richard I. Hartley, Hongdong Li |
ACCV (1) | 3 |
| 2007 | Minimal Solutions for Panoramic StitchingabstractThis paper presents minimal solutions for the geometric parameters of a camera rotating about its optical centre. In particular we present new 2 and 3 point solutions for the homography induced by a rotation with 1 and 2 unknown focal length parameters. Using tests on real data, we show that these algorithms outperform the standard 4 point linear homography solution in terms of accuracy of focal length estimation and image based projection errors. Matthew A. Brown, Richard I. Hartley, David Nistér |
CVPR | 2 |
| 2007 | Using Galois Theory to Prove Structure from Motion Algorithms are OptimalabstractThis paper presents a general method, based on Galois theory, for establishing that a problem can not be solved by a 'machine' that is capable of the standard arithmetic operations, extraction of radicals (that is, m-th roots for any m), as well as extraction of roots of polynomials of degree smaller than n, but no other numerical operations. The method is applied to two well known structure from motion problems: five point calibrated relative orientation, which can be realized by solving a tenth degree polynomial [6], and L2-optimal two-view triangulation, which can be realized by solving a sixth degree polynomial [3]. It is shown that both these solutions are optimal in the sense that an exact solution intrinsically requires the solution of a polynomial of the given degree (10 or 6 respectively), and cannot be solved by extracting roots of polynomials of any lesser degree. David Nistér, Richard I. Hartley, Henrik Stewénius |
CVPR | 2 |
| 2007 | An L Approach to Structure and Motion Problems in 1D-VisionabstractThe structure and motion problem of multiple one-dimensional projections of a two-dimensional environment is studied. One-dimensional cameras have proven useful in several different applications, most prominently for autonomous guided vehicles, but also in ordinary vision for analysing planar motion and the projection of lines. Previous results on one-dimensional vision are limited to classifying and solving minimal cases, bundle adjustment for finding local minima to the structure and motion problem and linear algorithms based on algebraic cost functions. In this paper, we present a method for finding the global minimum to the structure and motion problem using the max norm of reprojection errors. We show how the optimal solution can be computed efficiently using simple linear programming techniques. The algorithms have been tested on a variety of different scenarios, both real and synthetic, with good performance. In addition, we show how to solve the multiview triangulation problem, the camera pose problem and how to dualize the algorithm in the Carlsson duality sense, all within the same framework. Kalle Åström, Olof Enqvist, Carl Olsson, Fredrik Kahl, Richard I. Hartley |
ICCV | 5 |
| 2007 | Global Optimization through Searching Rotation Space and Optimal Estimation of the Essential MatrixabstractThis paper extends the set of problems for which a global solution can be found using modern optimization methods. In particular, the method is applied to estimation of the essential matrix, giving the first guaranteed optimal algorithm for estimating the relative pose under a geometric cost function, in this case, the L-infinity cost function. Convex optimization techniques has been shown to provide optimal solutions to many of the common problems in structure from motion. However, they do not apply to problems involving rotations. In this paper, we introduce a search method that allows such problems to be solved optimally. Apart from the essential matrix, the algorithm is applied to the camera pose problem, providing an optimal algorithm. Richard I. Hartley, Fredrik Kahl |
ICCV | 1 |
| 2007 | The 3D-3D Registration Problem RevisitedabstractWe describe a new framework for globally solving the 3D-3D registration problem with unknown point correspondences. This problem is significant as it is frequently encountered in many applications. Existing methods are not fully satisfactory, mainly due to the risk of local minima. Our framework is grounded on the Lipschitz global optimization theory. It achieves a guaranteed global optimality without any initialization. By exploiting the special structure of the problem itself and of the 3D rotation space SO(3), we propose a box-and-ball algorithm, which solves the problem efficiently. The main idea of the work can be applied to many other problems as well. Hongdong Li, Richard I. Hartley |
ICCV | 2 |
| 2007 | Convex Optimization for Deformable Surface 3-D Trackingabstract3-D shape recovery of non-rigid surfaces from 3-D to 2-D correspondences is an under-constrained problem that requires prior knowledge of the possible deformations. State-of-the-art solutions involve enforcing smoothness constraints that limit their applicability and prevent the recovery of sharply folding and creasing surfaces. Here, we propose a method that does not require such smoothness constraints. Instead, we represent surfaces as triangulated meshes and, assuming the pose in the first frame to be known, disallow large changes of edge orientation between consecutive frames, which is a generally applicable constraint when tracking surfaces in a 25 frames- per-second video sequence. We will show that tracking under these constraints can be formulated as a Second Order Cone Programming feasibility problem. This yields a convex optimization problem with stable solutions for a wide range of surfaces with very different physical properties. Mathieu Salzmann, Richard I. Hartley, Pascal Fua |
ICCV | 2 |
| 2007 | A Fast Method to Minimize L Error Norm for Geometric Vision ProblemsabstractMinimizing L∞error norm for some geometric vision problems provides global optimization using the well- developed algorithm called SOCP (second order cone programming). Because the error norm belongs to quasi- convex functions, bisection method is utilized to attain the global optimum. It tests the feasibility of the intersection of all the second order cones due to measurements, repeatedly adjusting the global error level. The computation time increases according to the size of measurement data since the number of second order cones for the feasibility test inflates correspondingly. We observe in this paper that not all the data need be included for the feasibility test because we minimize the maximum of the errors; we may use only a subset of the measurements to obtain the optimal estimate, and therefore we obtain a decreased computation time. In addition, by using L∞image error instead of L2Euclidean distance, we show that the problem is still a quasi-convex problem and can be solved by bisection method but with linear programming (LP). Our algorithm and experimental results are provided. Yongduek Seo, Richard I. Hartley |
ICCV | 2 |
| 2007 | Gradient Intensity-Based Registration of Multi-Modal Images of the BrainabstractWe present a fast and accurate framework for registration of multi-modal volumetric images based on decoupled estimation of registration parameters utilizing spatial information in the form of 'gradient intensity'. We introduce gradient intensity as a measure of spatial strength of an image in a given direction and show that it can be used to determine the rotational misalignment independent of translation between the images. The rotation parameters are obtained by maximizing the mutual information of 2D gradient intensity matrices obtained from 3D images, hence reducing the dimensionality of the problem and improving efficiency. The rotation parameters along with estimations of translation are then used to initialize an optimization step over a conventional pixel intensity-based method to achieve sub-voxel accuracy. Our optimization algorithm converges quickly and is less subject to the common problem of misregistration due to local extrema. Experiments show that our method significantly improves the robustness, performance and efficiency of registration compared to conventional pixel intensity-based methods. Ramtin Shams, Rodney A. Kennedy, Parastoo Sadeghi, Richard I. Hartley |
ICCV | 4 |
| 2007 | Automatic Dry Eye Detection
Tamir Yedidya, Richard I. Hartley, Jean-Pierre Guillon, Yogesan Kanagasingam |
MICCAI (1) | 2 |
| 2007 | A Study of Hippocampal Shape Difference Between Genders by Efficient Hypothesis Test and Discriminative Deformation
Luping Zhou, Richard I. Hartley, Paulette Lieby, Nick Barnes, Kaarin Anstey, Nicolas Cherbuin, Perminder S. Sachdev |
MICCAI (1) | 2 |
| 2007 | Critical Configurations for Projective Reconstruction from Multiple Views
Richard I. Hartley, Fredrik Kahl |
Int. J. Comput. Vis. | 1 |
| 2007 | Parameter-Free Radial Distortion Correction with Center of Distortion EstimationabstractWe propose a method of simultaneously calibrating the radial distortion function of a camera and the other internal calibration parameters. The method relies on the use of a planar (or, alternatively, nonplanar) calibration grid which is captured in several images. In this way, the determination of the radial distortion is an easy add-on to the popular calibration method proposed by Zhang [24]. The method is entirely noniterative and, hence, is extremely rapid and immune to the problem of local minima. Our method determines the radial distortion in a parameter-free way, not relying on any particular radial distortion model. This makes it applicable to a large range of cameras from narrow-angle to fish-eye lenses. The method also computes the center of radial distortion, which, we argue, is important in obtaining optimal results. Experiments show that this point may be significantly displaced from the center of the image or the principal point of the camera. Richard I. Hartley, Sing Bing Kang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Iterative Extensions of the Sturm/Triggs Algorithm: Convergence and NonconvergenceabstractWe give the first complete theoretical convergence analysis for the iterative extensions of the Sturm/Triggs algorithm. We show that the simplest extension, SIESTA, converges to nonsense results. Another proposed extension has similar problems, and experiments with "balanced" iterations show that they can fail to converge or become unstable. We present CIESTA, an algorithm which avoids these problems. It is identical to SIESTA except for one simple extra computation. Under weak assumptions, we prove that CIESTA iteratively decreases an error and approaches fixed points. With one more assumption, we prove it converges uniquely. Our results imply that CIESTA gives a reliable way of initializing other algorithms such as bundle adjustment. A descent method such as Gauss-Newton can be used to minimize the CIESTA error, combining quadratic convergence with the advantage of minimizing in the projective depths. Experiments show that CIESTA performs better than other iterations. John Oliensis, Richard I. Hartley |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Conformal spherical representation of 3D genus-zero meshes
Hongdong Li, Richard I. Hartley |
Pattern Recognit. | 2 |
| 2006 | Plane-Based Calibration and Auto-calibration of a Fish-Eye Camera
Hongdong Li, Richard I. Hartley |
ACCV (1) | 2 |
| 2006 | New 3D Fourier Descriptors for Genus-Zero Mesh Objects
Hongdong Li, Richard I. Hartley |
ACCV (1) | 2 |
| 2006 | Person Reidentification Using Spatiotemporal AppearanceabstractIn many surveillance applications it is desirable to determine if a given individual has been previously observed over a network of cameras. This is the person reidentification problem. This paper focuses on reidentification algorithms that use the overall appearance of an individual as opposed to passive biometrics such as face and gait. Person reidentification approaches have two aspects: (i) establish correspondence between parts, and (ii) generate signatures that are invariant to variations in illumination, pose, and the dynamic appearance of clothing. A novel spatiotemporal segmentation algorithm is employed to generate salient edgels that are robust to changes in appearance of clothing. The invariant signatures are generated by combining normalized color and salient edgel histograms. Two approaches are proposed to generate correspondences: (i) a model based approach that fits an articulated model to each individual to establish a correspondence map, and (ii) an interest point operator approach that nominates a large number of potential correspondences which are evaluated using a region growing scheme. Finally, the approaches are evaluated on a 44 person database across 3 disparate views. Niloofar Gheissari, Thomas B. Sebastian, Richard I. Hartley |
CVPR (2) | 3 |
| 2006 | Removing Outliers Using The L\infty NormabstractRecently, there has been interest in solving geometric vision problems such as triangulation and camera resectioning using L\infty minimization. One key advantage of using the L\infty norm rather than the L2 norm is that the L\infty cost function has a single minimum unlike the commonly used L2 cost function which typically has multiple local minima. However, one drawback of using L\infty minimization is that it is not robust to outliers. By minimizing the L\infty norm instead of the L2 norm, we are, in essence, fitting the outliers and not the good data. Therefore, before one can perform L\infty optimization on a problem, it is first necessary to remove outliers. A popular (but generally unsound) method of removing outliers is to minimize the cost function using standard optimization techniques; then if the residual error is too great, remove the offending measurements and continue. Although this method can fail for simple L2 optimization problems, we show in this paper that for a wide class of L\infty problems it is a valid technique. It is proved that the set of measurements with greatest residual must contain at least one outlier. Thus, if we keep throwing out the measurements with greatest residual, we will eventually remove all outliers in the data. We test this hypothesis on the multiview reconstruction problem and show that even simple strategies for throwing out these maximum residual measurements are effective in removing outliers. Kristy Sim, Richard I. Hartley |
CVPR (1) | 2 |
| 2006 | Recovering Camera Motion Using L\infty MinimizationabstractRecently, there has been interest in formulating various geometric problems in Computer Vision as L\infty optimization problems. The advantage of this approach is that under L\infty norm, such problems typically have a single minimum, and may be efficiently solved using Second-Order Cone Programming (SOCP). This paper shows that such techniques may be used effectively on the problem of determining the track of a camera given observations of features in the environment. The approach to this problem involves two steps: determination of the orientation of the camera by estimation of relative orientation between pairs of views, followed by determination of the translation of the camera. This paper focusses on the second step, that of determining the motion of the camera. It is shown that it may be solved effectively by using SOCP to reconcile translation estimates obtained for pairs or triples of views. In addition, it is observed that the individual translation estimates are not known with equal certainty in all directions. To account for this anisotropy in uncertainty, we introduce the use of covariances into the L\infty optimization framework. Kristy Sim, Richard I. Hartley |
CVPR (1) | 2 |
| 2006 | Iterative Extensions of the Sturm/Triggs Algorithm: Convergence and Nonconvergence
John Oliensis, Richard I. Hartley |
ECCV (4) | 2 |
| 2006 | Inverse tensor transfer with applications to novel view synthesis and multi-baseline stereo
Hongdong Li, Richard I. Hartley |
Signal Process. Image Commun. | 2 |
| 2005 | Shape from Non-homogeneous, Non-stationary, Anisotropic, Perspective TextureabstractWe present a method for Shape-from-Texture in one of its most general forms. Previous Shape-from-Texture papers assume that the texture is constrained by one or more of the following properties: homogeneity, isotropy, stationarity, or viewed orthographically. We make none of these assumptions. We do not presume that the frontal texture is known a priori, or from a known set, or even present in the image. Instead, surface smoothness is assumed, and the surface is recovered via a consistency constraint. The key idea is that the frontal texture is estimated, and a correct estimation leads to the most consistent surface. In addition to surface shape, a frontal view of the texture is also recovered. Results are given for synthetic and real examples. 1 Angeline M. Loh, Richard I. Hartley |
BMVC | 2 |
| 2005 | Parameter-Free Radial Distortion Correction with Centre of Distortion EstimationabstractWe propose a method of simultaneously calibrating the radial distortion function of a camera along with the other internal calibration parameters. The method relies on the use of a planar (or alternatively nonplanar) calibration grid, which is captured in several images. In this way, the determination of the radial distortion is an easy add-on to the popular calibration method proposed by Zhang [1999]. The method is entirely noniterative, and hence is extremely rapid and immune from the problem of local minima. Our method determines the radial distortion in a parameter-free way, not relying on any particular radial distortion model. This makes it applicable to a large range of cameras from narrow-angle to fish-eye lenses. The method also computes the centre of radial distortion, which we argue is important in obtaining optimal results. Experiments show that this point may be significantly displaced from the centre of the image, or the principal point of the camera. Richard I. Hartley, Sing Bing Kang |
ICCV | 1 |
| 2005 | Inverse tensor transfer for novel view synthesisabstractThis paper provides a new transfer based novel view synthesis method. This method does not need a pre-computed dense depth map, therefore overcomes most common problems associated with conventional dense correspondence algorithms, yet still produce very photo-realistic novel images. The power of the method comes from the introducing and using of a novel inverse tensor transfer technique, which offers a simple mechanism to exploit both photometric constraints and geometric constraints across multiple input images. Our method works equally well for both calibrated images and un-calibrated images. Experiments on real sequences show promising results. Hongdong Li, Richard I. Hartley |
ICIP (2) | 2 |
| 2004 | L-8Minimization in Geometric Reconstruction Problems
Richard I. Hartley, Frederik Schaffalitzky |
CVPR (1) | 1 |
| 2004 | The Multibody Trifocal Tensor: Motion Segmentation from 3 Perspective Views
Richard I. Hartley, René Vidal |
CVPR (1) | 1 |
| 2004 | Motion Segmentation with Missing Data Using PowerFactorization and GPCA
René Vidal, Richard I. Hartley |
CVPR (2) | 2 |
| 2004 | Reconstruction from Projections Using Grassmann Tensors
Richard I. Hartley, Frederik Schaffalitzky |
ECCV (1) | 1 |
| 2004 | A new and compact algorithm for simultaneously matching and estimationabstractFeature matching and transformation estimation are two fundamental problems in computer vision research. These two problems are often related and even interlocked; solving one is solving the other's precondition. This makes them hard to solve. In order to overcome this difficulty, the paper presents a new compact algorithm requiring less than 10 lines of Matlab code. We show that the solutions of correspondence and transformation are merely two factors of two grammian matrices, and can be worked out by a factorization method. A Newton-Schulz numerical iteration algorithm is used for the factorization. The two interlocked problems are solved in an alternate (flip-flop) way. The effectiveness and efficiency are illustrated by experiments on both synthetic and real images. Global and fast convergence is attained, even starting from randomly chosen initial guesses. Hongdong Li, Richard I. Hartley |
ICASSP (3) | 2 |
| 2003 | Motion From 3D Line Correspondences: Linear and Non-Linear SolutionsabstractWe address the problem of aligning two reconstructions of lines and cameras in projective, affine, metric or Euclidean space. We propose several 3D (three-dimensional) and image-related linear algorithms. The result can be used to initialize the nonlinear minimization of several proposed error functions, as well as the maximum likelihood estimator that we derive. We evaluate and compare our algorithms to existing ones using simulated and real data. Adrien Bartoli, Richard I. Hartley, Fredrik Kahl |
CVPR (1) | 2 |
| 2003 | A Critical Configuration for Reconstruction from Rectilinear MotionabstractThis paper investigates critical configurations for projective reconstruction from multiple images taken by a camera moving in a straight line. Projective reconstruction refers to a determination of the 3D (three-dimensional) geometrical configuration of a set of 3D points and cameras, given only correspondences between points in the images. A configuration of points and cameras is critical if it cannot be determined uniquely (up to a projective transform) from the image coordinates of the points. It is shown that a configuration consisting of any number of cameras lying on a straight line, and any number of points lying on a twisted cubic constitutes a critical configuration. An alternative configuration consisting of a set of points and cameras all lying on a rational quartic curve exists. Richard I. Hartley, Fredrik Kahl |
CVPR (1) | 1 |
| 2002 | Sensitivity of Calibration to Principal Point Position
Richard I. Hartley, Robert Kaucic |
ECCV (2) | 1 |
| 2002 | Critical Curves and Surfaces for Euclidean Reconstruction
Fredrik Kahl, Richard I. Hartley |
ECCV (2) | 2 |
| 2001 | Critical Configurations for N-view Projective ReconstructionabstractIn this paper we give a characterization of critical configurations for projective reconstruction with any number of points and views. A set of cameras and points is said to be critical if the projected image points are insufficient to determine the placement of the points and the cameras uniquely, up to a projective transformation. For two views, the critical configurations are well-known. In this paper it is shown that a configuration of n 3 cameras and in points all lying on the intersection of two distinct ruled quadrics is critical. In distinction to the two-view case, which in general allows two alternative solutions, there is a family of ambiguous reconstructions for the n-view case. As a partial converse, it Is shown that for any critical configuration, all the points lie on the intersection of two ruled quadrics. Fredrik Kahl, Richard I. Hartley, Kalle Åström |
CVPR (2) | 2 |
| 2001 | Plane-based Projective Reconstruction
Robert Kaucic, Richard I. Hartley, Nicolas Y. Dano |
ICCV | 2 |
| 2001 | Visual navigation in a plane using the conformal point
Richard I. Hartley, Chanop Silpa-Anan |
ISRR | 1 |
| 2000 | Reconstruction from Six-Point SequencesabstractAn algorithm is given for computing projective structure from a set of six points seen in a sequence of many images. The method is based on the notion of duality between cameras and points first pointed out by Carlsson and Weinshall. The current implementation avoids the weakness inherent in previous implementations of this method in which numerical accuracy is compromised by the distortion of image point error distributions under projective transformation. It is shown in this paper that one may compute the dual fundamental matrix by minimizing a cost function giving a first-order approximation to geometric distance error in the original untransformed image measurements. This is done by a modification of a standard near-optimal method for computing the fundamental matrix. Subsequently, the error measurements are adjusted optimally to conform with exact imaging geometry by application of the triangulation method of Hartley-Sturm. Richard I. Hartley, Nicolas Y. Dano |
CVPR | 1 |
| 2000 | Ambiguous Configurations for 3-View Projective Reconstruction
Richard I. Hartley |
ECCV (1) | 1 |
| 2000 | A Six Point Solution for Structure and Motion
Frederik Schaffalitzky, Andrew Zisserman, Richard I. Hartley, Philip Torr 0001 |
ECCV (1) | 3 |
| 2000 | Statistical Significance as an Aid to System Performance Evaluation
Peter H. Tu, Richard I. Hartley |
ECCV (2) | 2 |
| 1999 | Linear Self-Calibration of a Rotating and Zooming CameraabstractA linear self-calibration method is given for computing the calibration of a stationary but rotating camera. The internal parameters of the camera are allowed to vary from image to image, allowing for zooming (change of focal length) and possible variation of the principal point of the camera. In order for calibration to be possible some constraints must be placed on the calibration of each image. The method works under the minimal assumption of zero-skew (rectangular pixels), or the more restrictive but reasonable conditions of square pixels, known pixel aspect ratio, and known principal point. Being linear the algorithm is extremely rapid, and avoids the convergence problems characteristic of iterative algorithms. Lourdes Agapito, Eric Hayman, Richard I. Hartley |
CVPR | 3 |
| 1999 | Camera Calibration and the Search for InfinityabstractThis paper considers the problem of self-calibration of a camera from an image sequence in the case where the camera's internal parameters (most notably focal length) may change. The problem of camera self-calibration from a sequence of images has proven to be a difficult one in practice, due to the need ultimately to resort to non-linear methods, which have often proven to be unreliable. In a stratified approach to self-calibration, a projective reconstruction is obtained first and this is successively refined first to an affine and then to a Euclidean (or metric) reconstruction. It has been observed that the difficult step is to obtain the affine reconstruction, or equivalently to locate the plane at infinity in the projective coordinate frame. The problem is inherently non-linear and requires iterative methods that risk not finding the optimal solution. The present paper overcomes this difficulty by imposing chirality constraints to limit the search for the plane at infinity to a 3-dimensional cubic region of parameter space. It is then possible to carry out a dense search over this cube in reasonable time. For each hypothesised placement of the plane at infinity, the calibration problem is reduced to one of calibration of a nontranslating camera, for which fast non-iterative algorithms exist. A cost function based on the result of the trial calibration is used to determine the best placement of the plane at infinity. Because of the simplicity of each trial, speeds of over 10,000 trials per second are achieved on a 256 MHz processor. It is shown that this dense search allows one to avoid areas of local minima effectively and find global minima of the cost function. Richard I. Hartley, Lourdes Agapito, Ian D. Reid 0001, Eric Hayman |
ICCV | 1 |
| 1999 | Theory and Practice of Projective Rectification
Richard I. Hartley |
Int. J. Comput. Vis. | 1 |
| 1998 | Computation of the Quadrifocal Tensor
Richard I. Hartley |
ECCV (1) | 1 |
| 1998 | Minimizing Algebraic Error in Geometric Estimation ProblemsabstractThis paper gives a widely applicable technique for solving many of the parameter estimation problems encountered in geometric computer vision. A commonly used approach is to minimize an algebraic error function instead of a possibly preferable geometric error function. It is claimed in this paper that minimizing algebraic error will usually give excellent results, and in fact the main problem with most algorithms minimizing algebraic distance is that they do not take account of mathematical constraints that should be imposed on the quantity being estimated. This paper gives an efficient method of minimizing algebraic distance while taking account of the constraints. This provides new algorithms for the problems of resectioning a pinhole camera, computing the fundamental matrix, and computing the tri-focal tensor. Evaluation results are given for the resectioning and tri-focal tensor estimation algorithms. Richard I. Hartley |
ICCV | 1 |
| 1998 | Chirality
Richard I. Hartley |
Int. J. Comput. Vis. | 1 |
| 1998 | High precision X-ray stereo for automated 3D CAD-based inspectionabstractAn important challenge in industrial metrology is to provide rapid measurement of critical 3D internal object geometry for either inspecting high volume parts or controlling a machining process. Existing metrological techniques are typically too slow to meet this need or can not measure small features with high precision. In this paper, we present a new method that achieves fast, accurate, internal 3D geometry measurement based on 3D reconstruction from a few X-ray views of a part. Our approach utilizes an accurate camera model for the X-ray sensor, calibration using in situ ground truth and geometry-guided X-ray feature extraction to achieve this goal and has been fully implemented in a prototype 3D measurement system. We describe a novel application of the system to CAD-based verification of drilled hole positioning. Experimental results are given to illustrate the precision of the system and 3D measurement on real industrial parts. J. Alison Noble, Rajiv Gupta 0002, Joseph L. Mundy, Andrea Schmitz, Richard I. Hartley |
IEEE Trans. Robotics Autom. | 5 |
| 1997 | How Useful is Projective Geometry?
Patrick Gros, Richard I. Hartley, Roger Mohr, Long Quan |
Comput. Vis. Image Underst. | 2 |
| 1997 | Reply to Pizlo, Rosenfeld, and Weiss
Richard I. Hartley, Roger Mohr |
Comput. Vis. Image Underst. | 1 |
| 1997 | Triangulation
Richard I. Hartley, Peter F. Sturm |
Comput. Vis. Image Underst. | 1 |
| 1997 | Self-Calibration of Stationary Cameras
Richard I. Hartley |
Int. J. Comput. Vis. | 1 |
| 1997 | Lines and Points in Three Views and the Trifocal Tensor
Richard I. Hartley |
Int. J. Comput. Vis. | 1 |
| 1997 | Kruppa's Equations Derived from the Fundamental MatrixabstractThe purpose of this paper is to give a specific form for Kruppa's equations in terms of the fundamental matrix. Kruppa's equations can be written explicitly in terms of the singular value decomposition (SVD) of the fundamental matrix. Richard I. Hartley |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1997 | In Defense of the Eight-Point AlgorithmabstractThe fundamental matrix is a basic tool in the analysis of scenes taken with two uncalibrated cameras, and the eight-point algorithm is a frequently cited method for computing the fundamental matrix from a set of eight or more point matches. It has the advantage of simplicity of implementation. The prevailing view is, however, that it is extremely susceptible to noise and hence virtually useless for most purposes. This paper challenges that view, by showing that by preceding the algorithm with a very simple normalization (translation and scaling) of the coordinates of the matched points, results are obtained comparable with the best iterative algorithms. This improved performance is justified by theory and verified by extensive experiments on real images. Richard I. Hartley |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1996 | Self-Calibration from Image Triplets
Martin Armstrong, Andrew Zisserman, Richard I. Hartley |
ECCV (1) | 3 |
| 1995 | Triangulation
Richard I. Hartley, Peter F. Sturm |
CAIP | 1 |
| 1995 | A Linear Method for Reconstruction from Lines and PointsabstractDiscusses the basic role of the trifocal tensor in scene reconstruction. This 3/spl times/3/spl times/3 tensor plays a role in the analysis of scenes from three views analogous to the role played by the fundamental matrix in the two-view case. In particular, the trifocal tensor maybe computed by a linear algorithm from a set of 13 line correspondences in three views. It is further shown in this paper to be essentially identical to a set of coefficients introduced by Shashua (1994) to effect point transfer in the three-view case. This observation means that the 13-line algorithm may be extended to allow for the computation of the trifocal tensor given any mixture of sufficiently many line and point correspondences. From the trifocal tensor, the camera image matrices may be computed, and the scene may be reconstructed. For unrelated uncalibrated cameras, this reconstruction is unique up to projectivity. Thus, projective reconstruction of a set of lines and points may be reconstructed linearly from three views.> Richard I. Hartley |
ICCV | 1 |
| 1995 | In Defence of the 8-Point AlgorithmabstractThe fundamental matrix is a basic tool in the analysis of scenes taken with two uncalibrated cameras, and the 8 point algorithm is a frequently cited method for computing the fundamental matrix from a set of 8 or more point matches. It has the advantage of simplicity of implementation. The prevailing view is, however, that it is extremely susceptible to noise and hence virtually useless for most purposes. The paper challenges that view, by showing that by preceding the algorithm with a very simple normalization (translation and scaling) of the coordinates of the matched points, results are obtained comparable with the best iterative algorithms. This improved performance is justified by theory and verified by extensive experiments on real images.> Richard I. Hartley |
ICCV | 1 |
| 1995 | Camera calibration for 2.5-D X-ray metrologyabstractThis paper presents a new methodology for camera calibration of stereo X-ray projections. The image acquisition systems typically used in nondestructive evaluation result in views that are orthographic along one image axis and perspective along the other. A camera model for this sensing geometry, called the linear pushbroom model, is described. Four methods of calibration, which make different assumptions about what is known about the camera parameters, are presented and compared. Rajiv Gupta 0002, J. Alison Noble, Richard I. Hartley, Joseph L. Mundy, Andrea Schmitz |
ICIP (3) | 3 |
| 1995 | CAD-Based Inspection Using X-Ray StereoabstractAn important challenge in industrial metrology is to provide rapid measurement of critical 3D internal object geometry for either inspecting high volume parts or controlling a machining process. Existing metrological techniques are typically too slow to meet this need or can not measure small features with high precision. In this paper, we present an X-ray stereo system which aims to achieve fast 3D geometry measurement from a few X-ray views of a part. We describe the key algorithms in our system and a novel application of it to CAD-based verification of drilled hole positioning. Experimental results are given to illustrate the accuracy of the current system and inspection on a real part. J. Alison Noble, Rajiv Gupta 0002, Joseph L. Mundy, Andrea Schmitz, Richard I. Hartley, W. Hoffman |
ICRA | 5 |
| 1994 | Projective reconstruction from line correspondencesabstractThe paper gives a practical rapid algorithm for doing projective reconstruction of a scene consisting of a set of lines seen in three or more images with uncalibrated cameras. The algorithm is evaluated on real and ideal data to determine its performance in the presence of varying degrees of noise. By carefully consideration of sources of error, it is possible to get accurate reconstruction with realistic levels of noise. The algorithm can be applied to images from different cameras or the same camera. For images with the same camera with unknown calibration, it is possible to do a complete Euclidean reconstruction of the image. This extends to the case of uncalibrated cameras previous results on scene reconstruction from lines,.> Richard I. Hartley |
CVPR | 1 |
| 1994 | An algorithm for self calibration from several viewsabstractThis paper gives a practical algorithm for the self-calibration of a camera from several views. The method involves non-iterative methods for finding an initial calibration for the camera, followed by least-squares iteration to an optimum solution. At the same time, a scaled Euclidean reconstruction of the scene appearing in the images is computed.> Richard I. Hartley |
CVPR | 1 |
| 1994 | Self-Calibration from Multiple Views with a Rotating Camera
Richard I. Hartley |
ECCV (1) | 1 |
| 1994 | Linear Pushbroom Cameras
Richard I. Hartley, Rajiv Gupta 0002 |
ECCV (1) | 1 |
| 1994 | Quantitative Measurement of Manufactured Diamond Shape
Richard I. Hartley, J. Alison Noble, James Grande, Jane Liu |
ECCV (1) | 1 |
| 1994 | X-Ray Metrology for Quality AssuranceabstractThere is considerable current interest in deriving accurate dimensional measurements of the internal geometry of complex manufactured parts, particularly castings. This paper describes an approach to the reconstruction of 3D part geometry from multiple digital X-ray images. A novel method for radiographic stereo is described which takes into account the special imaging geometry of the digital X-ray sensor modeled by a linear moving array, or pushbroom, camera. The 3D reconstruction algorithm employs a nominal geometric model which is perturbed by X-ray image constraints. Manufacturing applications are discussed and illustrated by experimental results on synthetic phantoms and actual casting images.> J. Alison Noble, Richard I. Hartley, Joseph L. Mundy, J. Farley |
ICRA | 2 |
| 1994 | Projective Reconstruction and Invariants from Multiple ImagesabstractThis correspondence investigates projective reconstruction of geometric configurations seen in two or more perspective views, and the computation of projective invariants of these configurations from their images. A basic tool in this investigation is the fundamental matrix that describes the epipolar correspondence between image pairs. It is proven that once the epipolar geometry is known, the configurations of many geometric structures (for instance sets of points or lines) are determined up to a collineation of projective 3-space /spl Pscrsup 3/ by their projection in two independent images. This theorem is the key to a method for the computation of invariants of the geometry. Invariants of six points in /spl Pscrsup 3/ and of four lines in /spl Pscrsup 3/ are defined and discussed. An example with real images shows that they are effective in distinguishing different geometrical configurations. Since the fundamental matrix is a basic tool in the computation of these invariants, new methods of computing the fundamental matrix from seven-point correspondences in two images or six-point correspondences in three images are given.> Richard I. Hartley |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1994 | Optimizing pipelined networks of associative and commutative operatorsabstractA method of tree-height minimization of networks of commutative and associative operators is described. The algorithm aims at minimizing latency and shimming delays in a synchronous data flow architecture such as that used in bit/digit-serial computation. The algorithm rearranges adder/subtractor trees to meet the joint goals, often allowing otherwise impossible scheduling constraints to be met. The algorithmic methods are found to apply also to trees of adders and shifters, such as those found in shift/add multipliers.> Richard I. Hartley, Albert E. Casavant |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1993 | Computing matched-epipolar projectionsabstractA new method is given for image rectification, the process of resampling pairs of stereo images taken from widely differing viewpoints in order to produce a pair of matched epipolar projections. These are projections in which the epipolar lines run parallel with the x-axis and disparities between the images are in the x-direction only. The method is based on an examination of the essential matrix of Longuet-Higgins (1981), which describes the epipolar geometry of the image pair. The approach taken is consistent with that advocated by O. Faugeras (1992) of avoiding camera calibration. A matrix called the epipolar transformation matrix is defined. It is used to determine a pair of 2-D projective transforms to be applied to the two images in order to match the epipolar lines. The advantages include the simplicity of the 2-D projective transformation, which allows very fast resampling, as well as subsequent simplification in identifying matched points and in scene reconstruction.> Richard I. Hartley, Rajiv Gupta 0002 |
CVPR | 1 |
| 1992 | Stereo from uncalibrated camerasabstractThe problem of computing placement of points in 3-D space, given two uncalibrated perspective views, is considered. The main theorem shows that the placement of the points is determined only up to an arbitrary projective transformation of 3-space. Given additional ground control points, however, the location of the points and the camera parameters may be determined. The method is linear and noniterative, whereas previously known methods for solving the camera calibration and placement problem to take proper account of both ground-control points and image correspondences are unsatisfactory in requiring either iterative methods or model restrictions. As a result of the main theorem, it is possible to determine projective invariants of 3-D geometric configurations from two perspective views.> Richard I. Hartley, Rajiv Gupta 0002, Tom Chang |
CVPR | 1 |
| 1992 | Estimation of Relative Camera Positions for Uncalibrated Cameras
Richard I. Hartley |
ECCV | 1 |
| 1990 | A New Simultaneous Circuit Partitioning and Chip Placement Approach Based on Simulated AnnealingabstractThe problems of circuit partitioning and chip placement have been studied in the past. Given a circuit partitioned into chips, one can optimize the placement of the chips on a printed circuit board with regard to a given cost function. Conversely, given a placement of the chips on the board, one can optimize the partitioning of the circuit into the chips with regard to the same cost function. However, given neither the circuit partitioning nor the chip placement, we are faced with a difficult optimization problem. Our target technology is one in which the chips are unpackaged chips placed on a substrate, analogous to the printed circuit board and interconnected together with high density interconnect to realize a complex system. We propose a new approach in which the circuit is both partitioned and placed simultaneously by a simulated annealing based algorithm. Our approach is seen to yield excellent results in reasonable run times. Abhijit Chatterjee, Richard I. Hartley |
DAC | 2 |
| 1990 | A synthesis, test and debug environment for rapid prototyping of DSP designsabstractRecently, considerable progress has been made in the design of digital signal processing (DSP) integrated circuits and systems. In order to address the need for rapid and economical production and testing of hardware prototypes, a hardware and software system called DIODES is being developed for the rapid prototyping, testing and debugging of DSP designs. DIODES will allow the user to design a DSP system, have it partitioned into predefined function blocks, have it assembled using advanced packing technology and then thoroughly test and debug the design both stand-alone and in a larger electronic system environment. This paper gives an overview of the whole system focussing particularly on the debugging hardware and software support environment. The algorithmic description is translated by the DIODES synthesis software into a structural specification suitable for High-Density Interconnect (HDI) fabrication. The synthesis process includes the insertion of test capabilities into the hardware to allow for debugging the design. The rapid turnaround of HDI fabrication means that the user can have a prototype DIODES module in hand ready for testing within at most a day or two, and at moderate cost.> Richard I. Hartley, Kenneth Welles II, Michael J. Hartman |
RSP | 1 |
| 1990 | Rapid prototyping of electronic systemsabstractDescribes a system for the rapid prototyping, testing and debugging of DSP designs. The DSP algorithm is first coded in a high level algorithmic language. This description is translated by the synthesis software into a structural specification suitable for HDI fabrication. The synthesis process includes the insertion of test capabilities into the hardware to allow for debugging the design. The rapid turnaround of HDI fabrication means that the user can have a prototype HDI module (DSP accelerator) in his hands ready for testing within at most a day or two, and at moderate cost. The DSP accelerator is then placed in a socket on a specially designed board (called the mother-board), connected to a Sun or other workstation, where it may be thoroughly tested and debugged. The debugging capabilities include structural verification, debugging and in-system test. Structural verification is testing the DSP accelerator to see that it is connected together correctly and that all the parts are working. Debugging involves sending test vectors through the chip, stepping, observing internal node values and limited reconfigurability with the purpose of debugging the algorithm. In-system test is done by connecting the DSP accelerator via a cable into the target system. This allows the system to be run at full speed (up to 20 MHz) with the DSP accelerator in place, but still retaining full observability of internal nodes.> Michael J. Hartman, Richard I. Hartley, Kenneth Welles II, Paul Delano, Arani Chatterjee |
RSP | 2 |
| 1989 | Tree-height minimization in pipelined architecturesabstractA method of tree-height minimization for networks of commutative and associative operators is proposed. An algorithm is described for minimizing latency and shimming delays in a synchronous data-flow architecture such as that used in pipelined or bit- or digit-serial computation. The algorithm rearranges operator trees to meet the joint goals, often allowing otherwise impossible scheduling constraints to be met. It may also be applied to word-parallel pipelined architectures to optimize operator trees within pipelined stages. The method is evaluated by testing it on several filter examples for which it finds optimal network topologies and schedules.> Richard I. Hartley, Albert E. Casavant |
ICCAD | 1 |
| 1989 | Drawing Polygons Given Angle Sequences
Richard I. Hartley |
Inf. Process. Lett. | 1 |
| 1988 | A Digit-Serial Silicon Compiler
Richard I. Hartley, Peter F. Corbett |
DAC | 1 |
| 1988 | Behavioral to structural translation in a bit-serial silicon compilerabstractAfter a brief discussion of previous work in the field of behavioral-to-structural translation, the bit-serial architecture which is the target of the bit-serial compiler is discussed. The bit-serial language (BSL) is then defined, and the details of the behavioral-to-structural translation are given.> Richard I. Hartley, Jeffrey R. Jasica |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1987 | A silicon compiler for digital signal processing: Methodology, implementation, and applicationsabstractThis paper describes a fully integrated silicon compilation tool geared towards digital signal processing applications. The silicon compiler presented here uses a bit-serial architecture with a 1.25- µm CMOS cell library. It accepts as its input a high-level description language tailored for digital signal processing algorithms. The language supports the basic signal processing constructs such as multiplication, addition, subtraction, sample delays, logical operators, relational operators, as well as a conditional assignment. The compiler is equipped with behavioral, logic, and fault simulators, and performs placement and routing. The paper also details the use of the silicon compiler for the implementation of classical DSP algorithms: digital filters, FFT, programmable filters, as well as other more specialized applications such as adaptive algorithms and waveform synthesis. Moreover, some techniques are presented to implement more complex mathematical functions commonly used in DSP. The results of chip designs using the compiler and its impact on future designs are highlighted. Fathy F. Yassa, Jeffrey R. Jasica, Richard I. Hartley, Sharbel E. Noujaim |
Proc. IEEE | 3 |