David Suter

dblp:37/5634 · DBLP profile ↗
← Back
153ranked-venue papers
5as first author
24since 2021 · last 2025
0000-0001-6306-3023ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 114 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 92 · 1 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 2Computer networks · 1
YearPublicationVenuePosition
2025 From Pixels to Prognosis: A Multi-modal Attention-Based Framework for Visceral Adipose Tissue Estimation
Arooba Maqsood, Afsah Saleem, Marc Sim, David Suter, Simone Radavelli-Bagatini, Jonathan M. Hodgson, Richard Prince, Kun Zhu 0028, William D. Leslie, John T. Schousboe, Joshua R. Lewis, Syed Zulqarnain Gilani
MICCAI (15)4
2025 Adaptive Graph Attention-Guided Parallel Sampling and Embedded Selection for Multi-Model Fitting
abstract
Multi-model fitting is a fundamental challenge in computer vision, where real-world data often contains severe gross outliers and pseudo-outliers. Existing methods rely on inefficient sequential hypothesize-and-verify frameworks that require a predefined number of models and inlier thresholds-parameters that are difficult to determine in practical scenes. To overcome these limitations, we propose a novel Adaptive Graph Attention-guided parallel multi-model fitting method (AGASAC) that jointly learns local and global features, performs parallel hypothesis sampling, and executes confidence-embedded model selection. Specifically, we design a dual-confidence graph attention module that models data relationships using an adaptive graph attention network. This module computes minimal-set confidence and quality confidence to guide the multi-model fitting process, eliminating manual parameter tuning. Additionally, we propose a parallel discriminative sampling module that leverages minimal-set confidence to concurrently sample hypotheses. By enforcing a quantized consensus constraint, this module maximizes inter-model variance while minimizing intra-model discrepancy. It enables computationally efficient hypothesis generation and pseudo-outlier suppression. To obtain high-quality models, we present a quality-embedded selection module that integrates quality confidence into the joint optimization of model selection and data clustering. Extensive experiments show that the proposed method achieves a lower transfer error of 0.39 pixels and a 36.92% runtime reduction, surpassing state-of-the-art methods. The code is available at https://github.com/YWY-Vivian/AGASAC.
Wenyu Yin, Shuyuan Lin, David Suter, Hanzi Wang
ACM Multimedia3
2024 Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
abstract
Embodied AI aims to develop robots that can understand and execute human language instructions, as well as communicate in natural languages. On this front, we study the task of generating highly detailed navigational instructions for the embodied robots to follow. Although recent studies have demonstrated significant leaps in the generation of step-by-step instructions from sequences of images, the generated instructions lack variety in terms of their referral to objects and landmarks. Existing speaker models learn strategies to evade the evaluation metrics and obtain higher scores even for low-quality sentences. In this work, we propose SAS (Spatially-Aware Speaker), an instruction generator or Speaker model that utilises both structural and semantic knowledge of the environment to produce richer instructions. For training, we employ a reward learning method in an adversarial setting to avoid systematic bias introduced by language evaluation metrics. Empirically, our method outperforms existing instruction generation models, evaluated using standard metrics. Our code is available at https://github.com/gmuraleekrishna/SAS.
Muraleekrishna Gopinathan, Martin Masek, Jumana M. Abu-Khalaf, David Suter
ACL (1)4
2024 An Exploration of Diabetic Foot Osteomyelitis X-ray Data for Deep Learning Applications
Brandon Abela, Martin Masek, Jumana M. Abu-Khalaf, David Suter, Ashu Gupta
AIME (2)4
2024 Segment Any Object Model (SAOM): Real-To-Simulation Fine-Tuning Strategy For Multi-Class Multi-Instance Segmentation
abstract
Multi-class multi-instance segmentation is the task of identifying masks for multiple object classes and multiple instances of the same class within an image. The foundational Segment Anything Model (SAM) is designed for promptable multi-class multi-instance segmentation but tends to output part or sub-part masks in the “everything” mode for various real-world applications. Whole object segmentation masks play a crucial role for indoor scene understanding, especially in robotics applications. We propose a new domain invariant Real-to-Simulation (Real-Sim) fine-tuning strategy for SAM. We use object images and ground truth data collected from Ai2Thor simulator during fine-tuning (real-to-sim). To allow our Segment Any Object Model (SAOM) to work in the “everything” mode, we propose the novel nearest neighbour assignment method, updating point embeddings for each ground-truth mask. SAOM is evaluated on our own dataset collected from Ai2Thor simulator. SAOM significantly improves on SAM, with a $28 \%$ increase in mIoU and a $25 \%$ increase in mAcc for 54 frequently-seen indoor object classes. Moreover, our Real-to-Simulation fine-tuning strategy demonstrates promising generalization performance in real environments without being trained on the real-world data (sim-to-real). The dataset and the code are available here.
Mariia Khan, Yue Qiu 0001, Yuren Cong, Bodo Rosenhahn, Jumana M. Abu-Khalaf, David Suter
ICIP6
2024 StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
abstract
Embodied navigation requires robots to understand and interact with the environment based on given tasks. Vision-Language Navigation (VLN) is an embodied navigation task, where a robot navigates within a previously seen and unseen environment, based on linguistic instruction and visual inputs. VLN agents need access to both local and global action spaces; former for immediate decision making and the latter for recovering from navigational mistakes. Prior VLN agents rely only on instruction-viewpoint alignment for local and global decision making and back-track to a previously visited viewpoint, if the instruction and its current viewpoint mismatches. These methods are prone to mistakes, due to the complexity of the instruction and partial observability of the environment. We posit that, back-tracking is sub-optimal and agent that is aware of its mistakes can recover efficiently. For optimal recovery, exploration should be extended to unexplored viewpoints (or frontiers). The optimal frontier is a recently observed but unexplored viewpoint that aligns with the instruction and is novel. We introduce a memory-based and mistake-aware path planning strategy for VLN agents, called StratXplore, that presents global and local action planning to select the optimal frontier for path correction. The proposed method collects all past actions and viewpoint features during navigation and then selects the optimal frontier suitable for recovery. Experimental results show this simple yet effective strategy improves the success rate on two VLN datasets with different task complexities.
Muraleekrishna Gopinathan, Jumana M. Abu-Khalaf, David Suter, Martin Masek
IROS3
2024 Indoor Scene Change Understanding (SCU): Segment, Describe, and Revert Any Change
abstract
Understanding of scene changes is crucial for embodied AI applications, such as visual room rearrangement, where the agent must revert changes by restoring the objects to their original locations or states. Visual changes between two scenes, pre- and post-rearrangement, encompass two tasks: scene change detection (locating changes) and image difference captioning (describing changes). While previous methods, focused on sequential 2D images, have addressed these tasks separately, it is essential to emphasize the significance of their combination. Therefore, we propose a new Scene Change Understanding (SCU) task for simultaneous change detection and description. Moreover, we go beyond change language description generation and aim to generate rearrangement instructions for the robotic agent to revert changes. To solve this task, we propose a novel method - EmbSCU, which allows to compare instance-level change object masks (for 53 frequently-seen indoor object classes) before and after changes and generate rearrangement language instructions for the agent. EmbSCU is built on our Segment Any Object Model (SAOMv2) - a fine-tuned version of Segment Anything Model (SAM), adapted to obtain instance-level object masks for both foreground and background objects in indoor embodied environments. EmbSCU is evaluated on our own dataset of sequential 2D image pairs before and after changes, collected from the Ai2Thor simulator. The proposed framework achieves promising results in both change detection and change description. Moreover, EmbSCU demonstrates positive generalization results on real-world scenes without using any real-life data during training. The dataset and the code are available here.
Mariia Khan, Yue Qiu 0001, Yuren Cong, Bodo Rosenhahn, David Suter, Jumana M. Abu-Khalaf
IROS5
2024 A Hybrid CNN-Transformer Feature Pyramid Network for Granular Abdominal Aortic Calcification Detection from DXA Images
Zaid Ilyas, Afsah Saleem, David Suter, John T. Schousboe, William D. Leslie, Joshua R. Lewis, Syed Zulqarnain Gilani
MICCAI (11)3
2024 Single Domain Generalization via Normalised Cross-correlation Based Convolutions
abstract
Deep learning techniques often perform poorly in the presence of domain shift, where the test data follows a different distribution than the training data. The most practically desirable approach to address this issue is Single Domain Generalization (S-DG), which aims to train robust models using data from a single source. Prior work on S-DG has primarily focused on using data augmentation techniques to generate diverse training data. In this paper, we explore an alternative approach by investigating the robustness of linear operators, such as convolution and dense layers commonly used in deep learning. We propose a novel operator called "XCNorm" that computes the normalized cross-correlation between weights and an input feature patch. This approach is invariant to both affine shifts and changes in energy within a local feature patch and eliminates the need for commonly used non-linear activation functions. We show that deep neural networks composed of this operator are robust to common semantic distribution shifts. Furthermore, our empirical results on single-domain generalization benchmarks demonstrate that our proposed technique performs comparably to the stateof-the-art methods.1
Weiqin Chuah, Ruwan B. Tennakoon, Reza Hoseinnezhad, David Suter, Alireza Bab-Hadiashar
WACV4
2024 Estimating Blood Alcohol Level Through Facial Features for Driver Impairment Assessment
abstract
Drunk driving-related road accidents contribute significantly to the global burden of road injuries. Addressing alcohol-related harm, particularly during safety-critical activities like driving, requires real-time monitoring of an individual’s blood alcohol concentration (BAC). We devise an in-vehicle machine learning system that harnesses standard commercial RGB cameras to predict critical levels of BAC. Our system can detect instances of alcohol intoxication impairment as subtle as 0.05 g/dL (WHO recommended legal limit for driving), with an accuracy of 75%, by leveraging the physiological manifestations of alcohol intoxication on a driver’s face. This system holds great promise for improving road safety. In tandem, we have compiled a data set of 60 subjects engaged in simulated driving scenarios, spanning three levels of alcohol intoxication. These scenarios were captured and divided into video segments labeled "sober","low’, and "severe" Alcohol Intoxication Impairment (AII), constituting the basis for evaluating our system’s performance. To the best of our knowledge, this study is the first to create a large-scale real-life dataset of alcohol intoxication and assess intoxication levels using an off-the-shelf RGB camera to detect drunk driving.
Ensiyeh Keshtkaran, Brodie von Berg, Grant Regan, David Suter, Syed Zulqarnain Gilani
WACV4
2023 SCOL: Supervised Contrastive Ordinal Loss for Abdominal Aortic Calcification Scoring on Vertebral Fracture Assessment Scans
Afsah Saleem, Zaid Ilyas, David Suter, Ghulam M. Hassan, Siobhan Reid, John T. Schousboe, Richard Prince, William D. Leslie, Joshua R. Lewis, Syed Zulqarnain Gilani
MICCAI (6)3
2023 Generalized framework for image and video object segmentation using affinity learning and message passing GNNS
abstract
Despite significant amount of work reported in the computer vision literature, segmenting images or videos based on multiple cues such as objectness, texture and motion, is still a challenge. This is particularly true when the number of objects to be segmented is not known or there are objects that are not classified in the training data (unknown objects). A possible remedy to this problem is to utiize graph-based clustering techniques such as Correlation Clustering. It is known that using long range affinities (Lifted multicut), makes correlation clustering more accurate than using only adjacent affinities (Multicut). However, the former is computationally expensive and hard to use. In this paper, we introduce a new framework to perform image/motion segmentation using an affinity learning module and a Message Passing Graph Neural Network (MPGNN). The affinity learning module uses a permutation invariant affinity representation to overcome the multi-object problem. The paper shows, both theoretically and empirically, that the proposed MPGNN aggregates higher order information and thereby converts the Lifted Multicut Problem (LMP) to a Multicut Problem (MP), which is easier and faster to solve. Importantly, the proposed method can be generalized to deal with different clustering problems with the same MPGNN architecture. For instance, our method produces competitive results for single image segmentation (on BSDS dataset) as well as unsupervised video object segmentation (on DAVIS17 dataset), by only changing the feature extraction part. In addition, using an ablation study on the proposed MPGNN architecture, we show that the way we update the parameterized affinities directly contributes to the accuracy of the results.
Sundaram Muthu, Ruwan B. Tennakoon, Tharindu Rathnayake, Reza Hoseinnezhad, David Suter, Alireza Bab-Hadiashar
Comput. Vis. Image Underst.5
2023 An Information-Theoretic Method to Automatic Shortcut Avoidance and Domain Generalization for Dense Prediction Tasks
abstract
Deep convolutional neural networks for dense prediction tasks are commonly optimized using synthetic data, as generating pixel-wise annotations for real-world data is laborious. However, the synthetically trained models do not generalize well to real-world environments. This poor "synthetic to real" (S2R) generalization we address through the lens of shortcut learning. We demonstrate that the learning of feature representations in deep convolutional networks is heavily influenced by synthetic data artifacts (shortcut attributes). To mitigate this issue, we propose an Information-Theoretic Shortcut Avoidance (ITSA) approach to automatically restrict shortcut-related information from being encoded into the feature representations. Specifically, our proposed method minimizes the sensitivity of latent features to input variations: to regularize the learning of robust and shortcut-invariant features in synthetically trained models. To avoid the prohibitive computational cost of direct input sensitivity optimization, we propose a practical yet feasible algorithm to achieve robustness. Our results show that the proposed method can effectively improve S2R generalization in multiple distinct dense prediction tasks, such as stereo matching, optical flow, and semantic segmentation. Importantly, the proposed method enhances the robustness of the synthetically trained networks and outperforms their fine-tuned counterparts (on real data) for challenging out-of-domain applications.
Weiqin Chuah, Ruwan B. Tennakoon, Reza Hoseinnezhad, David Suter, Alireza Bab-Hadiashar
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Unsupervised Learning for Maximum Consensus Robust Fitting: A Reinforcement Learning Approach
abstract
Robust model fitting is a core algorithm in several computer vision applications. Despite being studied for decades, solving this problem efficiently for datasets that are heavily contaminated by outliers is still challenging: due to the underlying computational complexity. A recent focus has been on learning-based algorithms. However, most of these approaches are supervised (which require a large amount of labelled training data). In this paper, we introduce a novel unsupervised learning framework: that learns to directly (without labelled data) solve robust model fitting. Moreover, unlike other learning-based methods, our work is agnostic to the underlying input features, and can be easily generalized to a wide variety of LP-type problems with quasi-convex residuals. We empirically show that our method outperforms existing (un)supervised learning approaches, and also achieves competitive results compared to traditional (non-learning-based) methods. Our approach is designed to try to maximise consensus (MaxCon), similar to the popular RANSAC. The basis of our approach, is to adopt a Reinforcement Learning framework. This requires designing appropriate reward functions, and state encodings. We provide a family of reward functions, tunable by choice of a parameter. We also investigate the application of different basic and enhanced Q-learning components.
Giang Truong, Huu Le, Erchuan Zhang, David Suter, Syed Zulqarnain Gilani
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 ITSA: An Information-Theoretic Approach to Automatic Shortcut Avoidance and Domain Generalization in Stereo Matching Networks
abstract
State-of-the-art stereo matching networks trained only on synthetic data often fail to generalize to more challenging real data domains. In this paper, we attempt to unfold an important factor that hinders the networks from generalizing across domains: through the lens of shortcut learning. We demonstrate that the learning of feature representations in stereo matching networks is heavily influenced by synthetic data artefacts (shortcut attributes). To mitigate this issue, we propose an Information-Theoretic Shortcut Avoidance (ITSA) approach to automatically restrict shortcut-related information from being encoded into the feature representations. As a result, our proposed method learns robust and shortcut-invariant features by minimizing the sensitivity of latent features to input variations. To avoid the prohibitive computational cost of direct input sensitivity optimization, we propose an effective yet feasible algorithm to achieve robustness. We show that using this method, state-of-the-art stereo matching networks that are trained purely on synthetic data can effectively generalize to challenging and previously unseen real data scenarios. Importantly, the proposed method enhances the robustness of the synthetic trained networks to the point that they outperform their fine-tuned counterparts (on real data) for challenging out-of-domain stereo datasets.
Weiqin Chuah, Ruwan B. Tennakoon, Reza Hoseinnezhad, Alireza Bab-Hadiashar, David Suter
CVPR5
2022 A Hybrid Quantum-Classical Algorithm for Robust Fitting
abstract
Fitting geometric models onto outlier contaminated data is provably intractable. Many computer vision systems rely on random sampling heuristics to solve robust fitting, which do not provide optimality guarantees and error bounds. It is therefore critical to develop novel approaches that can bridge the gap between exact solutions that are costly, and fast heuristics that offer no quality assurances. In this paper, we propose a hybrid quantum-classical algorithm for robust fitting. Our core contribution is a novel robust fitting formulation that solves a sequence of integer programs and terminates with a global solution or an error bound. The combinatorial subproblems are amenable to a quantum annealer, which helps to tighten the bound efficiently. While our usage of quantum computing does not surmount the fundamental intractability of robust fitting, by providing error bounds our algorithm is a practical improvement over randomised heuristics. Moreover, our work represents a concrete application of quantum computing in computer vision. We present results obtained using an actual quantum computer (D-Wave Advantage) and via simulation11Source code: https://github.com/dadung/HQC-robust-fitting.
Anh-Dzung Doan, Michele Sasdelli, David Suter, Tat-Jun Chin
CVPR3
2022 Maximum Consensus by Weighted Influences of Monotone Boolean Functions
abstract
Maximisation of Consensus (MaxCon) is one of the most widely used robust criteria in computer vision. Tennakoon et al. (CVPR2021), made a connection between MaxCon and estimation of influences of a Monotone Boolean function. In such, there are two distributions involved: the distribution defining the influence measure; and the distribution used for sampling to estimate the influence measure. This paper studies the concept of weighted influences for solving MaxCon. In particular, we study the Bernoulli measures. Theoretically, we prove the weighted influences, under this measure, of points belonging to larger structures are smaller than those of points belonging to smaller structures in general. We also consider another “natural” family of weighting strategies: sampling with uniform measure concentrated on a particular (Hamming) level of the cube. One can choose to have matching distributions: the same for defining the measure as for implementing the sampling. This has the advantage that the sampler is an unbiased estimator of the measure. Based on weighted sampling, we modify the algorithm of Tennakoon et al., and test on both synthetic and real datasets. We show some modest gains of Bernoulli sampling, and we illuminate some of the interactions between structure in data and weighted measures and weighted sampling.
Erchuan Zhang, David Suter, Ruwan B. Tennakoon, Tat-Jun Chin, Alireza Bab-Hadiashar, Giang Truong, Syed Zulqarnain Gilani
CVPR2
2022 Show, Attend and Detect: Towards Fine-Grained Assessment of Abdominal Aortic Calcification on Vertebral Fracture Assessment Scans
Syed Zulqarnain Gilani, Naeha Sharif, David Suter, John T. Schousboe, Siobhan Reid, William D. Leslie, Joshua R. Lewis
MICCAI (3)3
2022 Sparse Hypergraph Community Detection Thresholds in Stochastic Block Model
abstract
Community detection in random graphs or hypergraphs is an interesting fundamental problem in statistics, machine learning and computer vision. When the hypergraphs are generated by a {\em stochastic block model}, the existence of a sharp threshold on the model parameters for community detection was conjectured by Angelini et al. 2015. In this paper, we confirm the positive part of the conjecture, the possibility of non-trivial reconstruction above the threshold, for the case of two blocks. We do so by comparing the hypergraph stochastic block model with its Erd{\"o}s-R{\'e}nyi counterpart. We also obtain estimates for the parameters of the hypergraph stochastic block model. The methods developed in this paper are generalised from the study of sparse random graphs by Mossel et al. 2015 and are motivated by the work of Yuan et al. 2022. Furthermore, we present some discussion on the negative part of the conjecture, i.e., non-reconstruction of community structures.
Erchuan Zhang, David Suter, Giang Truong, Syed Zulqarnain Gilani
NeurIPS2
2022 Semantic Guided Long Range Stereo Depth Estimation for Safer Autonomous Vehicle Applications
abstract
Autonomous vehicles in intelligent transportation systems must be able to perform reliable and safe navigation. This necessitates accurate object detection, which is commonly achieved by high-precision depth perception. Existing stereo vision-based depth estimation systems generally involve computation of pixel correspondences and estimation of disparities between rectified image pairs. The estimated disparity values will be converted into depth values in downstream applications. As most applications often work in the depth domain, the accuracy of depth estimation is often more compelling than disparity estimation. However, at large distances (> 50m), the accuracy of disparity estimation does not directly translate to the accuracy of depth estimation. In the context of learning-based stereo systems, this is mainly due to biases imposed by the choices of the disparity-based loss function and the training data. Consequently, the learning algorithms often produce unreliable depth estimates of under-represented foreground objects, particularly at large distances. To resolve this issue, we first analyze the effect of those biases and then propose a pair of depth-based loss functions for foreground objects and background separately. These loss functions can be tuned and can balance the inherent bias of the stereo learning algorithms. The efficacy of our solution is demonstrated by an extensive set of experiments, which are benchmarked against state of the art. We show on the KITTI 2015 benchmark that our proposed solution yields substantial improvements in disparity and depth estimation, particularly for objects located at distances beyond 50 meters, outperforming the previous state of the art by 10%.
Weiqin Chuah, Ruwan B. Tennakoon, Reza Hoseinnezhad, David Suter, Alireza Bab-Hadiashar
IEEE Trans. Intell. Transp. Syst.4
2021 Consensus Maximisation Using Influences of Monotone Boolean Functions
abstract
Consensus maximisation (MaxCon), which is widely used for robust fitting in computer vision, aims to find the largest subset of data that fits the model within some tolerance level. In this paper, we outline the connection between MaxCon problem and the abstract problem of finding the maximum upper zero of a Monotone Boolean Function (MBF) defined over the Boolean Cube. Then, we link the concept of influences (in a MBF) to the concept of outlier (in MaxCon) and show that influences of points belonging to the largest structure in data would generally be smaller under certain conditions. Based on this observation, we present an iterative algorithm to perform consensus maximisation. Results for both synthetic and real visual data experiments show that the MBF based algorithm is capable of generating a near optimal solution relatively quickly. This is particularly important where there are large number of outliers (gross or pseudo) in the observed data.
Ruwan B. Tennakoon, David Suter, Erchuan Zhang, Tat-Jun Chin, Alireza Bab-Hadiashar
CVPR2
2021 Unsupervised Learning for Robust Fitting: A Reinforcement Learning Approach
abstract
Robust model fitting is a core algorithm in a large number of computer vision applications. Solving this problem efficiently for datasets highly contaminated with outliers is, however, still challenging due to the underlying computational complexity. Recent literature has focused on learning-based algorithms. However, most approaches are supervised (which require a large amount of labelled training data). In this paper, we introduce a novel unsupervised learning framework that learns to directly solve robust model fitting. Unlike other methods, our work is agnostic to the underlying input features, and can be easily generalized to a wide variety of LP-type problems with quasi-convex residuals. We empirically show that our method out-performs existing unsupervised learning approaches, and achieves competitive results compared to traditional methods on several important computer vision problems1.
Giang Truong, Huu Le, David Suter, Erchuan Zhang, Syed Zulqarnain Gilani
CVPR3
2021 Segmentation by Continuous Latent Semantic Analysis for Multi-structure Model Fitting
Guobao Xiao, Hanzi Wang, Jiayi Ma 0001, David Suter
Int. J. Comput. Vis.4
2021 Deterministic Approximate Methods for Maximum Consensus Robust Fitting
abstract
Maximum consensus estimation plays a critically important role in several robust fitting problems in computer vision. Currently, the most prevalent algorithms for consensus maximization draw from the class of randomized hypothesize-and-verify algorithms, which are cheap but can usually deliver only rough approximate solutions. On the other extreme, there are exact algorithms which are exhaustive search in nature and can be costly for practical-sized inputs. This paper fills the gap between the two extremes by proposing deterministic algorithms to approximately optimize the maximum consensus criterion. Our work begins by reformulating consensus maximization with linear complementarity constraints. Then, we develop two novel algorithms: one based on non-smooth penalty method with a Frank-Wolfe style optimization scheme, the other based on the Alternating Direction Method of Multipliers (ADMM). Both algorithms solve convex subproblems to efficiently perform the optimization. We demonstrate the capability of our algorithms to greatly improve a rough initial estimate, such as those obtained using least squares or a randomized algorithm. Compared to the exact algorithms, our approach is much more practical on realistic input sizes. Further, our approach is naturally applicable to estimation problems with geometric residuals. Matlab code and demo program for our methods can be downloaded from https://goo.gl/FQcxpi.
Huu Le, Tat-Jun Chin, Anders P. Eriksson, Thanh-Toan Do, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.5
2020 End-to-End Learning of Object Motion Estimation from Retinal Events for Event-Based Object Tracking
abstract
Event cameras, which are asynchronous bio-inspired vision sensors, have shown great potential in computer vision and artificial intelligence. However, the application of event cameras to object-level motion estimation or tracking is still in its infancy. The main idea behind this work is to propose a novel deep neural network to learn and regress a parametric object-level motion/transform model for event-based object tracking. To achieve this goal, we propose a synchronous Time-Surface with Linear Time Decay (TSLTD) representation, which effectively encodes the spatio-temporal information of asynchronous retinal events into TSLTD frames with clear motion patterns. We feed the sequence of TSLTD frames to a novel Retinal Motion Regression Network (RMRNet) to perform an end-to-end 5-DoF object motion regression. Our method is compared with state-of-the-art object tracking methods, that are based on conventional cameras or event cameras. The experimental results show the superiority of our method in handling various challenging environments such as fast motion and low illumination conditions.
Haosheng Chen 0001, David Suter, Qiangqiang Wu, Hanzi Wang
AAAI2
2020 Quantum Robust Fitting
Tat-Jun Chin, David Suter, Shin-Fang Ch'ng, James Quach
ACCV (1)2
2020 Motion Segmentation of RGB-D Sequences: Combining Semantic and Motion Information Using Statistical Inference
abstract
This paper presents an innovative method for motion segmentation in RGB-D dynamic videos with multiple moving objects. The focus is on finding static, small or slow moving objects (often overlooked by other methods) that their inclusion can improve the motion segmentation results. In our approach, semantic object based segmentation and motion cues are combined to estimate the number of moving objects, their motion parameters and perform segmentation. Selective object-based sampling and correspondence matching are used to estimate object specific motion parameters. The main issue with such an approach is the over segmentation of moving parts due to the fact that different objects can have the same motion (e.g. background objects). To resolve this issue, we propose to identify objects with similar motions by characterizing each motion by a distribution of a simple metric and using a statistical inference theory to assess their similarities. To demonstrate the significance of the proposed statistical inference, we present an ablation study, with and without static objects inclusion, on SLAM accuracy using the TUM-RGBD dataset. To test the effectiveness of the proposed method for finding small or slow moving objects, we applied the method to RGB-D MultiBody and SBM-RGBD motion segmentation datasets. The results showed that we can improve the accuracy of motion segmentation for small objects while remaining competitive on overall measures.
Sundaram Muthu, Ruwan B. Tennakoon, Tharindu Rathnayake, Reza Hoseinnezhad, David Suter, Alireza Bab-Hadiashar
IEEE Trans. Image Process.5
2019 Hypergraph Optimization for Multi-Structural Geometric Model Fitting
abstract
Recently, some hypergraph-based methods have been proposed to deal with the problem of model fitting in computer vision, mainly due to the superior capability of hypergraph to represent the complex relationship between data points. However, a hypergraph becomes extremely complicated when the input data include a large number of data points (usually contaminated with noises and outliers), which will significantly increase the computational burden. In order to overcome the above problem, we propose a novel hypergraph optimization based model fitting (HOMF) method to construct a simple but effective hypergraph. Specifically, HOMF includes two main parts: an adaptive inlier estimation algorithm for vertex optimization and an iterative hyperedge optimization algorithm for hyperedge optimization. The proposed method is highly efficient, and it can obtain accurate model fitting results within a few iterations. Moreover, HOMF can then directly apply spectral clustering, to achieve good fitting performance. Extensive experimental results show that HOMF outperforms several state-of-the-art model fitting methods on both synthetic data and real images, especially in sampling efficiency and in handling data with severe outliers.
Shuyuan Lin, Guobao Xiao, Yan Yan 0001, David Suter, Hanzi Wang
AAAI4
2019 Superpixel-Guided Two-View Deterministic Geometric Model Fitting
Guobao Xiao, Hanzi Wang, Yan Yan 0001, David Suter
Int. J. Comput. Vis.4
2019 Searching for Representative Modes on Hypergraphs for Robust Geometric Model Fitting
abstract
In this paper, we propose a simple and effective geometric model fitting method to fit and segment multi-structure data even in the presence of severe outliers. We cast the task of geometric model fitting as a representative mode-seeking problem on hypergraphs. Specifically, a hypergraph is first constructed, where the vertices represent model hypotheses and the hyperedges denote data points. The hypergraph involves higher-order similarities (instead of pairwise similarities used on a simple graph), and it can characterize complex relationships between model hypotheses and data points. In addition, we develop a hypergraph reduction technique to remove "insignificant" vertices while retaining as many "significant" vertices as possible in the hypergraph. Based on the simplified hypergraph, we then propose a novel mode-seeking algorithm to search for representative modes within reasonable time. Finally, the proposed mode-seeking algorithm detects modes according to two key elements, i.e., the weighting scores of vertices and the similarity analysis between vertices. Overall, the proposed fitting method is able to efficiently and effectively estimate the number and the parameters of model instances in the data simultaneously. Experimental results demonstrate that the proposed method achieves significant superiority over several state-of-the-art model fitting methods on both synthetic data and real images.
Hanzi Wang, Guobao Xiao, Yan Yan 0001, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Non-smooth M-estimator for Maximum Consensus Estimation
Huu Le, Anders P. Eriksson, Thanh-Toan Do, Tat-Jun Chin, David Suter
BMVC5
2018 Deterministic Consensus Maximization with Biconvex Programming
Zhipeng Cai 0003, Tat-Jun Chin, Huu Le, David Suter
ECCV (12)4
2017 An Exact Penalty Method for Locally Convergent Maximum Consensus
abstract
Maximum consensus estimation plays a critically important role in computer vision. Currently, the most prevalent approach draws from the class of non-deterministic hypothesize-and-verify algorithms, which are cheap but do not guarantee solution quality. On the other extreme, there are global algorithms which are exhaustive search in nature and can be costly for practical-sized inputs. This paper aims to fill the gap between the two extremes by proposing a locally convergent maximum consensus algorithm. Our method is based on a formulating the problem with linear complementarity constraints, then defining a penalized version which is provably equivalent to the original problem. Based on the penalty problem, we develop a Frank-Wolfe algorithm that can deterministically solve the maximum consensus problem. Compared to the randomized techniques, our method is deterministic and locally convergent, relative to the global algorithms, our method is much more practical on realistic input sizes. Further, our approach is naturally applicable to problems with geometric residuals.
Huu Le, Tat-Jun Chin, David Suter
CVPR3
2017 Quasiconvex Plane Sweep for Triangulation with Outliers
abstract
Triangulation is a fundamental task in 3D computer vision. Unsurprisingly, it is a well-investigated problem with many mature algorithms. However, algorithms for robust triangulation, which are necessary to produce correct results in the presence of egregiously incorrect measurements (i.e., outliers), have received much less attention. The default approach to deal with outliers in triangulation is by random sampling. The randomized heuristic is not only suboptimal, it could, in fact, be computationally inefficient on large-scale datasets. In this paper, we propose a novel locally optimal algorithm for robust triangulation. A key feature of our method is to efficiently derive the local update step by plane sweeping a set of quasiconvex functions. Underpinning our method is a new theory behind quasiconvex plane sweep, which has not been examined previously in computational geometry. Relative to the random sampling heuristic, our algorithm not only guarantees deterministic convergence to a local minimum, it typically achieves higher quality solutions in similar runtimes.
Qianggong Zhang, Tat-Jun Chin, David Suter
ICCV3
2017 Efficient guided hypothesis generation for multi-structure epipolar geometry estimation
Taotao Lai, Hanzi Wang, Yan Yan 0001, Guobao Xiao, David Suter
Comput. Vis. Image Underst.5
2017 Efficient Globally Optimal Consensus Maximisation with Tree Search
abstract
Maximum consensus is one of the most popular criteria for robust estimation in computer vision. Despite its widespread use, optimising the criterion is still customarily done by randomised sample-and-test techniques, which do not guarantee optimality of the result. Several globally optimal algorithms exist, but they are too slow to challenge the dominance of randomised methods. Our work aims to change this state of affairs by proposing an efficient algorithm for global maximisation of consensus. Under the framework of LP-type methods, we show how consensus maximisation for a wide variety of vision tasks can be posed as a tree search problem. This insight leads to a novel algorithm based on A* search. We propose efficient heuristic and support set updating routines that enable A* search to efficiently find globally optimal results. On common estimation problems, our algorithm is much faster than previous exact methods. Our work identifies a promising direction for globally optimal consensus maximisation.
Tat-Jun Chin, Pulak Purkait, Anders P. Eriksson, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.4
2017 Clustering with Hypergraphs: The Case for Large Hyperedges
abstract
The extension of conventional clustering to hypergraph clustering, which involves higher order similarities instead of pairwise similarities, is increasingly gaining attention in computer vision. This is due to the fact that many clustering problems require an affinity measure that must involve a subset of data of size more than two. In the context of hypergraph clustering, the calculation of such higher order similarities on data subsets gives rise to hyperedges. Almost all previous work on hypergraph clustering in computer vision, however, has considered the smallest possible hyperedge size, due to a lack of study into the potential benefits of large hyperedges and effective algorithms to generate them. In this paper, we show that large hyperedges are better from both a theoretical and an empirical standpoint. We then propose a novel guided sampling strategy for large hyperedges, based on the concept of random cluster models. Our method can generate large pure hyperedges that significantly improve grouping accuracy without exponential increases in sampling costs. We demonstrate the efficacy of our technique on various higher-order grouping problems. In particular, we show that our approach improves the accuracy and efficiency of motion segmentation from dense, long-term, trajectories.
Pulak Purkait, Tat-Jun Chin, Alireza Sadri, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Conformal Surface Alignment with Optimal Möbius Search
abstract
Deformations of surfaces with the same intrinsic shape can often be described accurately by a conformal model. A major focus of computational conformal geometry is the estimation of the conformal mapping that aligns a given pair of object surfaces. The uniformization theorem enables this task to be acccomplished in a canonical 2D domain, wherein the surfaces can be aligned using a Möbius transformation. Current algorithms for estimating Möbius transformations, however, often cannot provide satisfactory alignment or are computationally too costly. This paper introduces a novel globally optimal algorithm for estimating Möbius transformations to align surfaces that are topological discs. Unlike previous methods, the proposed algorithm deterministically calculates the best transformation, without requiring good initializations. Further, our algorithm is also much faster than previous techniques in practice. We demonstrate the efficacy of our algorithm on data commonly used in computational conformal geometry.
Huu Le, Tat-Jun Chin, David Suter
CVPR3
2016 Superpixel-Based Two-View Deterministic Fitting for Multiple-Structure Data
Guobao Xiao, Hanzi Wang, Yan Yan 0001, David Suter
ECCV (6)4
2016 Fast Rotation Search with Stereographic Projections for 3D Registration
abstract
Registering two 3D point clouds involves estimating the rigid transform that brings the two point clouds into alignment. Recently there has been a surge of interest in using branch-and-bound (BnB) optimisation for point cloud registration. While BnB guarantees globally optimal solutions, it is usually too slow to be practical. A fundamental source of difficulty lies in the search for the rotational parameters. In this work, first by assuming that the translation is known, we focus on constructing a fast rotation search algorithm. With respect to an inherently robust geometric matching criterion, we propose a novel bounding function for BnB that is provably tighter than previously proposed bounds. Further, we also propose a fast algorithm to evaluate our bounding function. Our idea is based on using stereographic projections to precompute and index all possible point matches in spatial R-trees for rapid evaluations. The result is a fast and globally optimal rotation search algorithm. To conduct full 3D registration, we co-optimise the translation by embedding our rotation search kernel in a nested BnB algorithm. Since the inner rotation search is very efficient, the overall 6DOF optimisation is speeded up significantly without losing global optimality. On various challenging point clouds, including those taken out of lab settings, our approach demonstrates superior efficiency.
Álvaro Parra Bustos, Tat-Jun Chin, Anders P. Eriksson, Hongdong Li, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.5
2016 Robust Model Fitting Using Higher Than Minimal Subset Sampling
abstract
Identifying the underlying model in a set of data contaminated by noise and outliers is a fundamental task in computer vision. The cost function associated with such tasks is often highly complex, hence in most cases only an approximate solution is obtained by evaluating the cost function on discrete locations in the parameter (hypothesis) space. To be successful at least one hypothesis has to be in the vicinity of the solution. Due to noise hypotheses generated by minimal subsets can be far from the underlying model, even when the samples are from the said structure. In this paper we investigate the feasibility of using higher than minimal subset sampling for hypothesis generation. Our empirical studies showed that increasing the sample size beyond minimal size ( p ), in particular up to p+2, will significantly increase the probability of generating a hypothesis closer to the true model when subsets are selected from inliers. On the other hand, the probability of selecting an all inlier sample rapidly decreases with the sample size, making direct extension of existing methods unfeasible. Hence, we propose a new computationally tractable method for robust model fitting that uses higher than minimal subsets. Here, one starts from an arbitrary hypothesis (which does not need to be in the vicinity of the solution) and moves until either a structure in data is found or the process is re-initialized. The method also has the ability to identify when the algorithm has reached a hypothesis with adequate accuracy and stops appropriately, thereby saving computational time. The experimental analysis carried out using synthetic and real data shows that the proposed method is both accurate and efficient compared to the state-of-the-art robust model fitting techniques.
Ruwan B. Tennakoon, Alireza Bab-Hadiashar, Zhenwei Cao, Reza Hoseinnezhad, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.5
2016 Hypergraph modelling for geometric model fitting
Guobao Xiao, Hanzi Wang, Taotao Lai, David Suter
Pattern Recognit.4
2015 Efficient globally optimal consensus maximisation with tree search
abstract
Maximum consensus is one of the most popular criteria for robust estimation in computer vision. Despite its widespread use, optimising the criterion is still customarily done by randomised sample-and-test techniques, which do not guarantee optimality of the result. Several globally optimal algorithms exist, but they are too slow to challenge the dominance of randomised methods. We aim to change this state of affairs by proposing a very efficient algorithm for global maximisation of consensus. Under the framework of LP-type methods, we show how consensus maximisation for a wide variety of vision tasks can be posed as a tree search problem. This insight leads to a novel algorithm based on A* search. We propose efficient heuristic and support set updating routines that enable A* search to rapidly find globally optimal results. On common estimation problems, our algorithm is several orders of magnitude faster than previous exact methods. Our work identifies a promising solution for globally optimal consensus maximisation.
Tat-Jun Chin, Pulak Purkait, Anders P. Eriksson, David Suter
CVPR4
2015 Mode-Seeking on Hypergraphs for Robust Geometric Model Fitting
abstract
In this paper, we propose a novel geometric model fitting method, called Mode-Seeking on Hypergraphs (MSH), to deal with multi-structure data even in the presence of severe outliers. The proposed method formulates geometric model fitting as a mode seeking problem on a hypergraph in which vertices represent model hypotheses and hyperedges denote data points. MSH intuitively detects model instances by a simple and effective mode seeking algorithm. In addition to the mode seeking algorithm, MSH includes a similarity measure between vertices on the hypergraph and a "weight-aware sampling" technique. The proposed method not only alleviates sensitivity to the data distribution, but also is scalable to large scale problems. Experimental results further demonstrate that the proposed method has significant superiority over the state-of-the-art fitting methods on both synthetic data and real images.
Hanzi Wang, Guobao Xiao, Yan Yan 0001, David Suter
ICCV4
2014 Fast Rotation Search with Stereographic Projections for 3D Registration
abstract
Recently there has been a surge of interest to use branch-and-bound (bnb) optimisation for 3D point cloud registration. While bnb guarantees globally optimal solutions, it is usually too slow to be practical. A fundamental source of difficulty is the search for the rotation parameters in the 3D rigid transform. In this work, assuming that the translation parameters are known, we focus on constructing a fast rotation search algorithm. With respect to an inherently robust geometric matching criterion, we propose a novel bounding function for bnb that allows rapid evaluation. Underpinning our bounding function is the usage of stereographic projections to precompute and spatially index all possible point matches. This yields a robust and global algorithm that is significantly faster than previous methods. To conduct full 3D registration, the translation can be supplied by 3D feature matching, or by another optimisation framework that provides the translation. On various challenging point clouds, including those taken out of lab settings, our approach demonstrates superior efficiency.
Álvaro Parra Bustos, Tat-Jun Chin, David Suter
CVPR3
2014 Fast Supervised Hashing with Decision Trees for High-Dimensional Data
abstract
Supervised hashing aims to map the original features to compact binary codes that are able to preserve label based similarity in the Hamming space. Non-linear hash functions have demonstrated their advantage over linear ones due to their powerful generalization capability. In the literature, kernel functions are typically used to achieve non-linearity in hashing, which achieve encouraging retrieval perfor- mance at the price of slow evaluation and training time. Here we propose to use boosted decision trees for achieving non-linearity in hashing, which are fast to train and evaluate, hence more suitable for hashing with high dimensional data. In our approach, we first propose sub-modular formulations for the hashing binary code inference problem and an efficient GraphCut based block search method for solving large-scale inference. Then we learn hash func- tions by training boosted decision trees to fit the binary codes. Experiments demonstrate that our proposed method significantly outperforms most state-of-the-art methods in retrieval precision and training time. Especially for high- dimensional data, our method is orders of magnitude faster than many methods in terms of training time.
Guosheng Lin, Chunhua Shen, Qinfeng Shi, Anton van den Hengel, David Suter
CVPR5
2014 Clustering with Hypergraphs: The Case for Large Hyperedges
Pulak Purkait, Tat-Jun Chin, Hanno Ackermann, David Suter
ECCV (4)4
2014 Fast rotation search for real-time interactive point cloud registration
abstract
Our goal is the registration of multiple 3D point clouds obtained from LIDAR scans of underground mines. Such a capability is crucial to the surveying and planning operations in mining. Often, the point clouds only partially overlap and initial alignment is unavailable. Here, we propose an interactive user-assisted point cloud registration system. Guided by the system, the user's role is simply to identify and search for overlapping regions across the point clouds. Specifically, given two point sets, the user clicks on a point in one set, then simply hovers the mouse on the other set to find a matching point. Each mouse position gives rise to a translation, and our system instantly optimises the rotation that aligns the point clouds.
Tat-Jun Chin, Álvaro Parra Bustos, Michael S. Brown, David Suter
I3D4
2014 Sampling Minimal Subsets with Large Spans for Robust Estimation
Quoc-Huy Tran, Tat-Jun Chin, Wojciech Chojnacki, David Suter
Int. J. Comput. Vis.4
2014 The Random Cluster Model for Robust Geometric Fitting
abstract
Random hypothesis generation is central to robust geometric model fitting in computer vision. The predominant technique is to randomly sample minimal subsets of the data, and hypothesize the geometric models from the selected subsets. While taking minimal subsets increases the chance of successively "hitting" inliers in a sample, hypotheses fitted on minimal subsets may be severely biased due to the influence of measurement noise, even if the minimal subsets contain purely inliers. In this paper we propose Random Cluster Models, a technique used to simulate coupled spin systems, to conduct hypothesis generation using subsets larger than minimal. We show how large clusters of data from genuine instances of the model can be efficiently harvested to produce accurate hypotheses that are less affected by the vagaries of fitting on minimal subsets. A second aspect of the problem is the optimization of the set of structures that best fit the data. We show how our novel hypothesis sampler can be integrated seamlessly with graph cuts under a simple annealing framework to optimize the fitting efficiently. Unlike previous methods that conduct hypothesis sampling and fitting optimization in two disjoint stages, our algorithm performs the two subtasks alternatingly and in a mutually reinforcing manner. Experimental results show clear improvements in overall efficiency.
Trung-Thanh Pham, Tat-Jun Chin, Jin Yu 0001, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.4
2014 As-Projective-As-Possible Image Stitching with Moving DLT
abstract
The success of commercial image stitching tools often leads to the impression that image stitching is a "solved problem". The reality, however, is that many tools give unconvincing results when the input photos violate fairly restrictive imaging assumptions; the main two being that the photos correspond to views that differ purely by rotation, or that the imaged scene is effectively planar. Such assumptions underpin the usage of 2D projective transforms or homographies to align photos. In the hands of the casual user, such conditions are often violated, yielding misalignment artifacts or "ghosting" in the results. Accordingly, many existing image stitching tools depend critically on post-processing routines to conceal ghosting. In this paper, we propose a novel estimation technique called Moving Direct Linear Transformation (Moving DLT) that is able to tweak or fine-tune the projective warp to accommodate the deviations of the input data from the idealized conditions. This produces as-projective-as-possible image alignment that significantly reduces ghosting without compromising the geometric realism of perspective image stitching. Our technique thus lessens the dependency on potentially expensive postprocessing algorithms. In addition, we describe how multiple as-projective-as-possible warps can be simultaneously refined via bundle adjustment to accurately align multiple images for large panorama creation.
Julio Zaragoza, Tat-Jun Chin, Quoc-Huy Tran, Michael S. Brown, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.5
2014 Multi-subregion based correlation filter bank for robust face recognition
Yan Yan 0001, Hanzi Wang, David Suter
Pattern Recognit.3
2014 Interacting Geometric Priors For Robust Multimodel Fitting
abstract
Recent works on multimodel fitting are often formulated as an energy minimization task, where the energy function includes fitting error and regularization terms, such as low-level spatial smoothness and model complexity. In this paper, we introduce a novel energy with high-level geometric priors that consider interactions between geometric models, such that certain preferred model configurations may be induced.We argue that in many applications, such prior geometric properties are available and should be fruitfully exploited. For example, in surface fitting to point clouds, the building walls are usually either orthogonal or parallel to each other. Our proposed energy function is useful in dealing with unknown distributions of data errors and outliers, which are often the factors leading to biased estimation. Furthermore, the energy can be efficiently minimized using the expansion move method. We evaluate the performance on several vision applications using real data sets. Experimental results show that our method outperforms the state-of-the-art methods without significant increase in computation.
Trung-Thanh Pham, Tat-Jun Chin, Konrad Schindler, David Suter
IEEE Trans. Image Process.4
2013 As-Projective-As-Possible Image Stitching with Moving DLT
abstract
We investigate projective estimation under model inadequacies, i.e., when the underpinning assumptions of the projective model are not fully satisfied by the data. We focus on the task of image stitching which is customarily solved by estimating a projective warp - a model that is justified when the scene is planar or when the views differ purely by rotation. Such conditions are easily violated in practice, and this yields stitching results with ghosting artefacts that necessitate the usage of deghosting algorithms. To this end we propose as-projective-as-possible warps, i.e., warps that aim to be globally projective, yet allow local non-projective deviations to account for violations to the assumed imaging conditions. Based on a novel estimation technique called Moving Direct Linear Transformation (Moving DLT), our method seamlessly bridges image regions that are inconsistent with the projective model. The result is highly accurate image stitching, with significantly reduced ghosting effects, thus lowering the dependency on post hoc deghosting.
Julio Zaragoza, Tat-Jun Chin, Michael S. Brown, David Suter
CVPR4
2013 Improved wireless tracking using radio frequency and video sensors
Thuraiappah Sathyan, Tat-Jun Chin, David Suter, Mark Hedley
FUSION3
2013 A General Two-Step Approach to Learning-Based Hashing
abstract
Most existing approaches to hashing apply a single form of hash function, and an optimization process which is typically deeply coupled to this specific form. This tight coupling restricts the flexibility of the method to respond to the data, and can result in complex optimization problems that are difficult to solve. Here we propose a flexible yet simple framework that is able to accommodate different types of loss functions and hash functions. This framework allows a number of existing approaches to hashing to be placed in context, and simplifies the development of new problem-specific hashing methods. Our framework decomposes the hashing learning problem into two steps: hash bit learning and hash function learning based on the learned bits. The first step can typically be formulated as binary quadratic problems, and the second step can be accomplished by training standard binary classifiers. Both problems have been extensively studied in the literature. Our extensive experiments demonstrate that the proposed framework is effective, flexible and outperforms the state-of-the-art.
Guosheng Lin, Chunhua Shen, David Suter, Anton van den Hengel
ICCV3
2013 A simultaneous sample-and-filter strategy for robust multi-structure model fitting
Hoi Sim Wong, Tat-Jun Chin, Jin Yu 0001, David Suter
Comput. Vis. Image Underst.4
2013 Mode seeking over permutations for rapid geometric model fitting
Hoi Sim Wong, Tat-Jun Chin, Jin Yu 0001, David Suter
Pattern Recognit.4
2012 Fast Training of Effective Multi-class Boosting Using Coordinate Descent Optimization
Guosheng Lin, Chunhua Shen, Anton van den Hengel, David Suter
ACCV (2)4
2012 The Random Cluster Model for robust geometric fitting
abstract
Random hypothesis generation is central to robust geometric model fitting in computer vision. The predominant technique is to randomly sample minimal or elemental subsets of the data, and hypothesize the geometric model from the selected subsets. While taking minimal subsets increases the chance of simultaneously “hitting” inliers in a sample, it amplifies the noise of the underlying model, and hypotheses fitted on minimal subsets may be severely biased even if they contain purely inliers. In this paper we propose to use Random Cluster Models, a technique used to simulate coupled spin systems, to conduct hypothesis generation using subsets larger than minimal. We show how large clusters of data from genuine instances of the geometric model can be efficiently harvested to produce more accurate hypotheses. To take advantage of our hypothesis generator, we construct a simple annealing method based on graph cuts to fit multiple instances of the geometric model in the data. Experimental results show clear improvements in efficiency over other methods based on minimal subset samplers.
Trung-Thanh Pham, Tat-Jun Chin, Jin Yu 0001, David Suter
CVPR4
2012 In Defence of RANSAC for Outlier Rejection in Deformable Registration
Quoc-Huy Tran, Tat-Jun Chin, Gustavo Carneiro 0001, Michael S. Brown, David Suter
ECCV (4)5
2012 Adaptive human silhouette reconstruction based on the exploration of temporal information
abstract
Human silhouette reconstruction has a wide range of applications in motion analysis, object segmentation and tracking, etc. In this paper, we propose a human silhouette reconstruction method based on the exploration of temporal information. Given a test silhouette, the proposed method aims to find its reliable templates for reconstruction by using the intrinsic temporal relationship among different frames. To effectively obtain such templates, we propose an adaptive criterion based on the non-negative least square optimization. Experimental results on two challenging datasets demonstrate the effectiveness of our method.
Xi Li 0001, Tat-Jun Chin, David Suter
ICASSP4
2012 Superpixel-driven level set tracking
abstract
In this paper, we propose a superpixel-driven method for level set tracking. In particular, by taking a superpixel-based speed function, the level set evolution is accelerated greatly. We define a mutual information based speed function using a superpixel-unit as the underlying representation, which captures the correlation of a superpixel with object/background. In order to enhance the robustness of our method, a shape prior is incorporated to constrain the contour evolution. Experimental results on a number of challenging sequences demonstrate the effectiveness and robustness of our method.
Xi Li 0001, Tat-Jun Chin, David Suter
ICIP4
2012 Accelerated Hypothesis Generation for Multistructure Data via Preference Analysis
abstract
Random hypothesis generation is integral to many robust geometric model fitting techniques. Unfortunately, it is also computationally expensive, especially for higher order geometric models and heavily contaminated data. We propose a fundamentally new approach to accelerate hypothesis sampling by guiding it with information derived from residual sorting. We show that residual sorting innately encodes the probability of two points having arisen from the same model, and is obtained without recourse to domain knowledge (e.g., keypoint matching scores) typically used in previous sampling enhancement methods. More crucially, our approach encourages sampling within coherent structures and thus can very rapidly generate all-inlier minimal subsets that maximize the robust criterion. Sampling within coherent structures also affords a natural ability to handle multistructure data, a condition that is usually detrimental to other methods. The result is a sampling scheme that offers substantial speed-ups on common computer vision tasks such as homography and fundamental matrix estimation. We show on many computer vision data, especially those with multiple structures, that ours is the only method capable of retrieving satisfactory results within realistic time budgets.
Tat-Jun Chin, Jin Yu 0001, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 Simultaneously Fitting and Segmenting Multiple-Structure Data with Outliers
abstract
We propose a robust fitting framework, called Adaptive Kernel-Scale Weighted Hypotheses (AKSWH), to segment multiple-structure data even in the presence of a large number of outliers. Our framework contains a novel scale estimator called Iterative Kth Ordered Scale Estimator (IKOSE). IKOSE can accurately estimate the scale of inliers for heavily corrupted multiple-structure data and is of interest by itself since it can be used in other robust estimators. In addition to IKOSE, our framework includes several original elements based on the weighting, clustering, and fusing of hypotheses. AKSWH can provide accurate estimates of the number of model instances and the parameters and the scale of each model instance simultaneously. We demonstrate good performance in practical applications such as line fitting, circle fitting, range image segmentation, homography estimation, and two--view-based motion segmentation, using both synthetic data and real images.
Hanzi Wang, Tat-Jun Chin, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 Visual tracking of numerous targets via multi-Bernoulli filtering of image data
Reza Hoseinnezhad, Ba-Ngu Vo, Ba-Tuong Vo, David Suter
Pattern Recognit.4
2011 A global optimization approach to robust multi-model fitting
abstract
We present a novel Quadratic Program (QP) formulation for robust multi-model fitting of geometric structures in vision data. Our objective function enforces both the fidelity of a model to the data and the similarity between its associated inliers. Departing from most previous optimization-based approaches, the outcome of our method is a ranking of a given set of putative models, instead of a pre-specified number of “good” candidates (or an attempt to decide the right number of models). This is particularly useful when the number of structures in the data is a priori unascertainable due to unknown intent and purposes. Another key advantage of our approach is that it operates in a unified optimization framework, and the standard QP form of our problem formulation permits globally convergent optimization techniques. We tested our method on several geometric multi-model fitting problems on both synthetic and real data. Experiments show that our method consistently achieves state-of-the-art results.
Jin Yu 0001, Tat-Jun Chin, David Suter
CVPR3
2011 Bayesian integration of audio and visual information for multi-target tracking using a CB-member filter
abstract
A new method is presented for integration of audio and visual information in multiple target tracking applications. The proposed approach uses a Bayesian filtering formulation and exploits multi-Bernoulli random finite set approximations. The work presented in this paper is the first principled Bayesian estimation approach to solve the sensor fusion problems that involve intermittent sensory data (e.g. audio data for a person who occasionally speaks.) We have examined our method with case studies from the SPEVI database. The results show nearly perfect tracking of people not only when they are silent but also when they are not visible to the camera (but speaking).
Reza Hoseinnezhad, Ba-Ngu Vo, Ba-Tuong Vo, David Suter
ICASSP4
2011 Dynamic and hierarchical multi-structure geometric model fitting
abstract
The ability to generate good model hypotheses is instrumental to accurate and robust geometric model fitting. We present a novel dynamic hypothesis generation algorithm for robust fitting of multiple structures. Underpinning our method is a fast guided sampling scheme enabled by analysing correlation of preferences induced by data and hypothesis residuals. Our method progressively accumulates evidence in the search space, and uses the information to dynamically (1) identify outliers, (2) filter unpromising hypotheses, and (3) bias the sampling for active discovery of multiple structures in the data-All achieved without sacrificing the speed associated with sampling-based methods. Our algorithm yields a disproportionately higher number of good hypotheses among the sampling outcomes, i.e., most hypotheses correspond to the genuine structures in the data. This directly supports a novel hierarchical model fitting algorithm that elicits the underlying stratified manner in which the structures are organized, allowing more meaningful results than traditional “flat” multi-structure fitting.
Hoi Sim Wong, Tat-Jun Chin, Jin Yu 0001, David Suter
ICCV4
2011 An adversarial optimization approach to efficient outlier removal
abstract
This paper proposes a novel adversarial optimization approach to efficient outlier removal in computer vision. We characterize the outlier removal problem as a game that involves two players of conflicting interests, namely, optimizer and outlier. Such an adversarial view not only brings new insights into various existing methods, but also gives rise to a general optimization framework that provably unifies them. Under the proposed framework, we develop a new outlier removal approach that is able to offer a much needed control over the trade-off between reliability and speed, which is otherwise not available in previous methods. The proposed approach is driven by a mixed-integer minmax (convex-concave) optimization process. Although a minmax problem is generally not amenable to efficient optimization, we show that for some commonly used vision objective functions, an equivalent Linear Program reformulation exists. We demonstrate our method on two representative multiview geometry problems. Experiments on real image data illustrate superior practical performance of our method over recent techniques.
Jin Yu 0001, Anders P. Eriksson, Tat-Jun Chin, David Suter
ICCV4
2011 Simultaneous Sampling and Multi-Structure Fitting with Adaptive Reversible Jump MCMC
abstract
Multi-structure model fitting has traditionally taken a two-stage approach: First, sample a (large) number of model hypotheses, then select the subset of hypotheses that optimise a joint fitting and model selection criterion. This disjoint two-stage approach is arguably suboptimal and inefficient - if the random sampling did not retrieve a good set of hypotheses, the optimised outcome will not represent a good fit. To overcome this weakness we propose a new multi-structure fitting approach based on Reversible Jump MCMC. Instrumental in raising the effectiveness of our method is an adaptive hypothesis generator, whose proposal distribution is learned incrementally and online. We prove that this adaptive proposal satisfies the diminishing adaptation property crucial for ensuring ergodicity in MCMC. Our method effectively conducts hypothesis sampling and optimisation simultaneously, and gives superior computational efficiency over other methods.
Trung-Thanh Pham, Tat-Jun Chin, Jin Yu 0001, David Suter
NIPS4
2011 Boosting histograms of descriptor distances for scalable multiclass specific scene recognition
Tat-Jun Chin, David Suter, Hanzi Wang
Image Vis. Comput.2
2011 Recognition of adult images, videos, and web page bags
abstract
In this article, we develop an integrated adult-content recognition system which can detect adult images, adult videos, and adult Web page bags, where a Web page bag consists of a Web page and a predefined number of Web pages linked to it through hyperlinks. In our adult image-recognition algorithm, we model skin patches rather than skin pixels, resulting in better results than state-of-the-art algorithms which model skin pixels. In our adult video-recognition algorithm, information from the accompanying audio section around an image in an adult video is used to obtain a prior classification of the image. The algorithm achieves a better performance than the ones which use image information alone or audio information alone. The adult Web page bag recognition is carried out using multi-instance learning based on the combination of classifying texts, images and videos in Web pages. Both the speed and the accuracy for recognizing the Web adult content are increased, in contrast to recognizing Web pages one-by-one.
Weiming Hu 0004, Haiqiang Zuo, Ou Wu 0001, Yunfei Chen 0002, Zhongfei Zhang, David Suter
ACM Trans. Multim. Comput. Commun. Appl.6
2010 Efficient Multi-structure Robust Fitting with Incremental Top-k Lists Comparison
Hoi Sim Wong, Tat-Jun Chin, Jin Yu 0001, David Suter
ACCV (4)4
2010 Multi-structure model selection via kernel optimisation
abstract
Our goal is to fit the multiple instances (or structures) of a generic model existing in data. Here we propose a novel model selection scheme to estimate the number of genuine structures present. In contrast to conventional model selection approaches, our method is driven by kernel-based learning. The input data is first clustered based on their potential to have emerged from the same structure. However the number of clusters is deliberately overestimated to obtain a set of initial model fits onto the data. We then resolve the oversegmentation via a series of kernel optimisation conducted through multiple kernel learning, and the concept of kernel-target alignment is used as a model selection criterion. Experiments on synthetic and real data show that our method outperforms previous model selection schemes. We also focus on the application of multi-body motion segmentation. In particular we demonstrate success on estimating the number of motions on sequences with more than 3 unique motions.
Tat-Jun Chin, David Suter, Hanzi Wang
CVPR2
2010 Accelerated Hypothesis Generation for Multi-structure Robust Fitting
Tat-Jun Chin, Jin Yu 0001, David Suter
ECCV (5)3
2010 Multi-object filtering from image sequence without detection
abstract
Almost every single-view visual multi-target tracking method presented in the literature includes a detection routine that maps the image data to point measurements relevant to the target states. These measurements are commonly further processed by a filter to estimate the number of targets and their states. This paper presents a novel visual tracking technique based on a multi-object filtering algorithm that operates directly on the image observations without the need for any detection. Experimental results on tracking sport players show that our proposed method can automatically track numerous interacting targets and quickly finds players entering or leaving the scene.
Reza Hoseinnezhad, Ba-Ngu Vo, David Suter, Ba-Tuong Vo
ICASSP3
2010 Visual localization and segmentation based on foreground/background modeling
abstract
In this paper, we propose a novel method to localize (or track) a foreground object and segment the foreground object from the surrounding background with occlusions for a moving camera. We measure the likelihood of a target position by using a combination of a generative model and a discriminative model, considering not only the foreground similarity to the target model but also the dissimilarity between the foreground and the background appearances. Object segmentation is treated as a binary labeling problem. A Markov Random Field (MRF) is employed to add a spatial smooth prior on the foreground/background patterns. We demonstrate the advantages of the proposed method on several challenging videos and compare our results with the results of several other popular methods. The proposed method has achieved good results.
Hanzi Wang, Tat-Jun Chin, David Suter
ICASSP3
2010 BoostML: An Adaptive Metric Learning for Nearest Neighbor Classification
Nayyar Abbas Zaidi, David McG. Squire, David Suter
PAKDD (1)3
2009 Keypoint induced distance profiles for visual recognition
abstract
We show that histograms of keypoint descriptor distances can make useful features for visual recognition. Descriptor distances are often exhaustively computed between sets of keypoints, but besides finding the k-smallest distances the structure of the distribution of these distances has been largely overlooked. We highlight the potential of such information in the task of particular scene recognition. Discriminative scene signatures in the form of histograms of keypoint descriptor distances are constructed in a supervised manner. The distances are computed between properly selected reference keypoints and the keypoints detected in the input image. The signature is low dimensional, computationally cheap to obtain, and can distinguish a large number of scenes. We introduce a scheme based on multiclass AdaBoost to select the appropriate reference keypoints. The resulting system is capable of handling a large number of scene classes at a fraction of the time required for exhaustively matching sets of keypoints. This supports supports a coarse-to-fine search strategy for approaches reliant on keypoint matching. We test the idea on 3 datasets for particular scene recognition and report the obtained results.
Tat-Jun Chin, David Suter
CVPR2
2009 Bayesian multi-object estimation from image observations
Ba-Ngu Vo, Ba-Tuong Vo, Nam-Trung Pham, David Suter
FUSION4
2009 Robust fitting of multiple structures: The statistical learning approach
abstract
We propose an unconventional but highly effective approach to robust fitting of multiple structures by using statistical learning concepts. We design a novel Mercer kernel for the robust estimation problem which elicits the potential of two points to have emerged from the same underlying structure. The Mercer kernel permits the application of well-grounded statistical learning methods, among which nonlinear dimensionality reduction, principal component analysis and spectral clustering are applied for robust fitting. Our method can remove gross outliers and in parallel discover the multiple structures present. It functions well under severe outliers (more than 90% of the data) and considerable inlier noise without requiring elaborate manual tuning or unrealistic prior information. Experiments on synthetic and real problems illustrate the superiority of the proposed idea over previous methods.
Tat-Jun Chin, Hanzi Wang, David Suter
ICCV3
2009 The Ordered Residual Kernel for Robust Motion Subspace Clustering
abstract
We present a novel and highly effective approach for multi-body motion segmentation. Drawing inspiration from robust statistical model fitting, we estimate putative subspace hypotheses from the data. However, instead of ranking them we encapsulate the hypotheses in a novel Mercer kernel which elicits the potential of two point trajectories to have emerged from the same subspace. The kernel permits the application of well-established statistical learning methods for effective outlier rejection, automatic recovery of the number of motions and accurate segmentation of the point trajectories. The method operates well under severe outliers arising from spurious trajectories or mistracks. Detailed experiments on a recent benchmark dataset (Hopkins 155) show that our method is superior to other state-of-the-art approaches in terms of recovering the number of motions, segmentation accuracy, robustness against gross outliers and computational efficiency.
Tat-Jun Chin, Hanzi Wang, David Suter
NIPS3
2009 3D terrestrial LIDAR classifications with super-voxels and multi-scale Conditional Random Fields
Ee Hui Lim, David Suter
Comput. Aided Des.2
2009 Rank Constraints for Homographies over Two Views: Revisiting the Rank Four Constraint
Pei Chen 0001, David Suter
Int. J. Comput. Vis.2
2009 Human action recognition by feature-reduced Gaussian process classification
Hang Zhou 0005, Liang Wang 0001, David Suter
Pattern Recognit. Lett.3
2009 Simultaneously Estimating the Fundamental Matrix and Homographies
abstract
The estimation of the fundamental matrix (FM) and/or one or more homographies between two views is of great interest for a number of computer vision and robotics tasks. We consider the joint estimation of the FM and one or more homographies. Given point matches between two views (and assuming rigid geometry of the camera-scene displacement), it is well known that all of the matched points satisfy the epipolar constraint that is usually characterized by the FM. Subsets of these point matches may also obey a constraint characterized by a homography (all matches in the subset coming from three-dimensional (3-D) points lying on a 3-D plane). The estimations of homographies and the FM are well-studied problems, and therefore, the (separate) estimation of the FM, or the homography matrices, can be considered as effectively solved problems with mature algorithms. However, the homographies and FM are not independent of each other: therefore, separate estimation of each is likely to be suboptimal. In this paper, we propose to simultaneously estimate the FM and homographies by employing the compatibility constraint between them. This is done by first concentrating on a set of parameters that (jointly) parameterize the entire set of homographies and FM (simultaneously) and that also implicitly enforce the compatibility between the estimates of each set. We then derive a reduced form with the purpose of improving the speed. We propose a solution method in which the Sampson error for the FM and homographies is minimized by the Levenberg-Marquardt (LM) algorithm. Experiments show that the gains can be compared with separate estimates (the FM and/or the homographies).
Pei Chen 0001, David Suter
IEEE Trans. Robotics2
2008 Improved building detection by Gaussian processes classification via feature space rescale and spectral kernel selection
abstract
We use spectral analysis to facilitate Gaussian processes (GP) classification. Our solution provides two improvements: scaling of the data to achieve a more isotropic nature, as well as a method to choose the kernel to match certain data characteristics. Given the dataset, from the Fourier transform of the training data we compare the frequency domain features of each dimension to estimate a rescaling (towards making the data isotropic). Also, the spectrum of the training data is compared with several candidate kernel spectrums. From this comparison the best matching kernel is chosen. In these ways, the training data matches better the GP classification kernel function (and hence the underlying assumed correlation characteristics), resulting in a better GP classification result. Test results on both non image and image data show the efficiency and effectiveness of our approach.
Hang Zhou 0005, David Suter
CVPR2
2008 Manifold optimisation for motion factorisation
abstract
This paper presents a novel formulation for the popular factorisation based solution for Structure from Motion. Since our measurement matrices are populated with incomplete and inaccurate data, SVD based total least squares solution are less than appropriate. Instead, we approach the problem as a non-linear unconstrained minimisation problem on the product manifold of the Special Euclidean Group (SE3). The restriction of the domain of optimisation to the SE3product manifold not only implies that each intermediate solution is a plausible object motion, but also ensures better intrinsic stability for the minimisation algorithm. We compare our method with existing state of art, and show that our algorithm exhibits superior performance.
Appu Shaji, Sharat Chandran, David Suter
ICPR3
2008 Confidence rated boosting algorithm for generic object detection
abstract
In this paper we propose a confidence rated boosting algorithm based on Ada-boost for generic object detection. Confidence rated Ada-boost algorithm has not been applied to generic object detection problem; in that sense our work is novel. We represent images as bag of words, where the words are SIFT descriptors extracted over some interest points. We compare our boosting algorithm to another version of boosting algorithm called Gentle-boost. Our approach generalizes well and performs equal or better than Gentle-boost. We show our results on four categories from the Caltech data sets, in terms of ROC curves.
Nayyar Abbas Zaidi, David Suter
ICPR2
2008 Improving Gaussian processes classification by spectral data reorganizing
abstract
We improve Gaussian processes (GP) classification by reorganizing the (non-stationary and anisotropic) data to better fit to the isotropic GP kernel. First, the data is partitioned into two parts: along the feature with the highest frequency bandwidth. Secondly, for each part of the data, only the spectrally homogeneous features are chosen and used (the rest discarded) for GP classification. In this way, anisotropy of the data is lessened from the frequency point of view. Tests on synthetic data as well as real datasets show that our approach is effective and outperforms Automatic Relevance Determination (ARD).
Hang Zhou 0005, David Suter
ICPR2
2008 Human motion recognition using Gaussian Processes classification
abstract
This paper investigates the applicability of Gaussian Processes (GP) classification for recognition of articulated and deformable human motions from image sequences. Using Tensor Subspace Analysis (TSA), space-time human silhouettes (extracted from motion videos) are transformed to low-dimensional multivariate time series, based on which structure-based statistical features are calculated to summarize the motion properties. GP classification is then used to learn and predict motion categories. Experimental results on two real-world state-of-the-art datasets show that the proposed approach is effective, and outperforms Support Vector Machine (SVM).
Hang Zhou 0005, Liang Wang 0001, David Suter
ICPR3
2008 Visual learning and recognition of sequential data manifolds with applications to human movement analysis
Liang Wang 0001, David Suter
Comput. Vis. Image Underst.2
2008 A Model-Selection Framework for Multibody Structure-and-Motion of Image Sequences
Konrad Schindler, David Suter, Hanzi Wang
Int. J. Comput. Vis.2
2008 Out-of-Sample Extrapolation of Learned Manifolds
abstract
We investigate the problem of extrapolating the embedding of a manifold learned from finite samples to novel out-of-sample data. We concentrate on the manifold learning method called Maximum Variance Unfolding (MVU) for which the extrapolation problem is still largely unsolved. Taking the perspective of MVU learning being equivalent to Kernel PCA, our problem reduces to extending a kernel matrix generated from an unknown kernel function to novel points. Leveraging on previous developments, we propose a novel solution which involves approximating the kernel eigenfunction using Gaussian basis functions. We also show how the width of the Gaussian can be tuned to achieve extrapolation. Experimental results which demonstrate the effectiveness of the proposed approach are also included.
Tat-Jun Chin, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Object detection by global contour shape
Konrad Schindler, David Suter
Pattern Recognit.2
2007 Human Pose Extraction from Monocular Videos using Constrained Non-Rigid Factorization
abstract
We focus on the problem of automatically extracting the 3D configuration of human poses from 2D image features tracked over a finite interval of time . This problem is highly non-linear in nature and confounds standard regression techniques. Our approach effectively marries a non-rigid factorization algorithm with prior learned statistical models from archival motion capture database. We show that a stand alone non-rigid factorization algorithm is highly unsuitable for this problem. However, when coupled with the learned statistical model in the form of a constrained non- linear programming method, it yields a substantially better solution.
Appu Shaji, Behjat Siddiquie, Sharat Chandran, David Suter
BMVC4
2007 Recognizing Human Activities from Silhouettes: Motion Subspace and Factorial Discriminative Graphical Model
abstract
We describe a probabilistic framework for recognizing human activities in monocular video based on simple silhouette observations in this paper. The methodology combines kernel principal component analysis (KPCA) based feature extraction and factorial conditional random field (FCRF) based motion modeling. Silhouette data is represented more compactly by nonlinear dimensionality reduction that explores the underlying structure of the articulated action space and preserves explicit temporal orders in projection trajectories of motions. FCRF models temporal sequences in multiple interacting ways, thus increasing joint accuracy by information sharing, with the ideal advantages of discriminative models over generative ones (e.g., relaxing independence assumption between observations and the ability to effectively incorporate both overlapping features and long-range dependencies). The experimental results on two recent datasets have shown that the proposed framework can not only accurately recognize human activities with temporal, intra-and inter-person variations, but also is considerably robust to noise and other factors such as partial occlusion and irregularities in motion styles.
Liang Wang 0001, David Suter
CVPR2
2007 Fast Sparse Gaussian Processes Learning for Man-Made Structure Classification
abstract
Informative Vector Machine (IVM) is an efficient fast sparse Gaussian process's (GP) method previously suggested for active learning. It greatly reduces the computational cost of GP classification and makes the GP learning close to real time. We apply IVM for man-made structure classification (a two class problem). Our work includes the investigation of the performance of IVM with varied active data points as well as the effects of different choices of GP kernels. Satisfactory results have been obtained, showing that the approach keeps full GP classification performance and yet is significantly faster (by virtue if using a subset of the whole training data points).
Hang Zhou 0005, David Suter
CVPR2
2007 Conditional Random Field for 3D Point Clouds with Adaptive Data Reduction
abstract
We proposed using Conditional Random Fields with adaptive data reduction for the classification of 3D point clouds acquired from a Riegl Terrestrial laser scanner. The training and inference of the acquired large outdoor urban data can be time consuming. We approach the problem by computing an adaptive support region for each data point using 3D scale theory. For training and inference of the discriminative Conditional Random Fields, smaller set of data samples that contains relevant information within the support region is selected instead of using all point cloud data. We tested the algorithm on synthetically generated data and urban point clouds data acquired from the laser scanner. The computed support region is also used in feature extraction for urban point clouds data. The results showed improvement in the training and inference rate while maintaining comparable classification accuracy.
Ee Hui Lim, David Suter
CW2
2007 Extrapolating Learned Manifolds for Human Activity Recognition
abstract
The problem of human activity recognition via visual stimuli can be approached using manifold learning, since the silhouette (binary) images of a person undergoing a smooth motion can be represented as a manifold in the image space. While manifold learning methods allow the characterization of the activity manifolds, performing activity recognition requires distinguishing between manifolds. This invariably involves the extrapolation of learned activity manifolds to new silhouettes -a task that is not fully addressed in the literature. This paper investigates and compares methods for the extrapolation of learned manifolds within the context of activity recognition. Also, the problem of obtaining dense samples for learning human silhouette manifolds is addressed.
Tat-Jun Chin, Liang Wang 0001, Konrad Schindler, David Suter
ICIP (1)4
2007 Man-Made Structure Segmentation using Gaussian Processes and Wavelet Features
abstract
We apply Gaussian process classification (GPC) to man-made structure segmentation, treated as a two class problem. GPC is a discriminative approach, and thus focuses on modelling the posterior directly. It relaxes the strong assumption of conditional independence of the observed data (generally used in a generative model). In addition, wavelet transform features, which are effective in describing directional textures, are incorporated in the feature vector. Satisfactory results have been obtained which show the effectiveness of our approach.
Hang Zhou 0005, David Suter
ICIP (4)2
2007 Adaptive Object Tracking Based on an Effective Appearance Filter
abstract
We propose a similarity measure based on a Spatial-color Mixture of Gaussians (SMOG) appearance model for particle filters. This improves on the popular similarity measure based on color histograms because it considers not only the colors in a region but also the spatial layout of the colors. Hence, the SMOG-based similarity measure is more discriminative. To efficiently compute the parameters for SMOG, we propose a new technique, with which the computational time is greatly reduced. We also extend our method by integrating multiple cues to increase the reliability and robustness. Experiments show that our method can successfully track objects in many difficult situations.
Hanzi Wang, David Suter, Konrad Schindler, Chunhua Shen
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 A consensus-based method for tracking: Modelling background scenario and foreground appearance
Hanzi Wang, David Suter
Pattern Recognit.2
2007 Incremental Kernel Principal Component Analysis
abstract
The kernel principal component analysis (KPCA) has been applied in numerous image-related machine learning applications and it has exhibited superior performance over previous approaches, such as PCA. However, the standard implementation of KPCA scales badly with the problem size, making computations for large problems infeasible. Also, the "batch" nature of the standard KPCA computation method does not allow for applications that require online processing. This has somewhat restricted the domains in which KPCA can potentially be applied. This paper introduces an incremental computation algorithm for KPCA to address these two problems. The basis of the proposed solution lies in computing incremental linear PCA in the kernel induced feature space, and constructing reduced-set expansions to maintain constant update speed and memory usage. We also provide experimental results which demonstrate the effectiveness of the approach.
Tat-Jun Chin, David Suter
IEEE Trans. Image Process.2
2007 Learning and Matching of Dynamic Shape Manifolds for Human Action Recognition
abstract
In this paper, we learn explicit representations for dynamic shape manifolds of moving humans for the task of action recognition. We exploit locality preserving projections (LPP) for dimensionality reduction, leading to a low-dimensional embedding of human movements. Given a sequence of moving silhouettes associated to an action video, by LPP, we project them into a low-dimensional space to characterize the spatiotemporal property of the action, as well as to preserve much of the geometric structure. To match the embedded action trajectories, the median Hausdorff distance or normalized spatiotemporal correlation is used for similarity measures. Action classification is then achieved in a nearest-neighbor framework. To evaluate the proposed method, extensive experiments have been carried out on a recent dataset including ten actions performed by nine different subjects. The experimental results show that the proposed method is able to not only recognize human actions effectively, but also considerably tolerate some challenging conditions, e.g., partial occlusion, low-quality videos, changes in viewpoints, scales, and clothes; within-class variations caused by different subjects with different physical build; styles of motion; etc.
Liang Wang 0001, David Suter
IEEE Trans. Image Process.2
2006 A New Distance Criterion for Face Recognition Using Image Sets
Tat-Jun Chin, David Suter
ACCV (1)2
2006 Feature Detection with an Improved Anisotropic Filter
Mohamed Gobara, David Suter
ACCV (2)2
2006 A Novel Robust Statistical Method for Background Initialization and Visual Surveillance
Hanzi Wang, David Suter
ACCV (1)2
2006 Improving the Speed of Kernel PCA on Large Scale Datasets
abstract
This paper concerns making large scale Kernel Principal Component Analysis (KPCA) feasible on regular hardware. The KPCA has been proven a useful non-linear feature extractor in several computer vision applications. The standard computation method for KPCA, however, scales badly with the problem size, thus limiting the potential of the technique for large scale data. We propose a novel method to alleviate this problem. The essence of our solution lies in partitioning the data and greedily filtering each partition for a sparse representation. Incremental KPCA is then utilized to merge each partition to arrive at the overall KPCA. We also provide experimental results which demonstrate the effectiveness of the approach.
Tat-Jun Chin, David Suter
AVSS2
2006 Analyzing Human Movements from Silhouettes Using Manifold Learning
abstract
A novel method for learning and recognizing sequential image data is proposed, and promising applications to vision-based human movement analysis are demonstrated. To find more compact representations of high-dimensional silhouette data, we exploit locality preserving projections (LPP) to achieve low-dimensional manifold embedding. Further, we present two kinds of methods to analyze and recognize learned motion manifolds. One is correlation matching based on the Hausdorrf distance, and the other is a probabilistic method using continuous hidden Markov models (HMM). Encouraging results are obtained in two representative experiments in the areas of human activity recognition and gait-based human identification.
Liang Wang 0001, David Suter
AVSS2
2006 A Compact Architecture for Wireless Video Surveillance over CDMA Network
abstract
Wireless video surveillance over CDMA network enables transfer of high quality live video via public wireless broadband, which is accessible anywhere with CDMA mobile phone network coverage. Real time video can be received on an internet connected PC, laptop or even mobile phone. This makes possible surveillance from moving vehicles, or quick deployment in almost unrestricted locations. While many general purpose solutions exist for sending video over CDMA, they have not taken into consideration certain specific features of the network nor the requirements of the related applications. This paper presents a new architecture for video transmission over CDMA with simple hardware complexity, high reliability and rich functions. The compact and flexible system has been under trial in China, Australia, Singapore, Italy, Egypt and Brazil, amongst others.
Hang Zhou 0005, David Suter
AVSS2
2006 Incremental Kernel PCA for Efficient Non-linear Feature Extraction
abstract
The Kernel Principal Component Analysis (KPCA) has been effectively applied as an unsupervised non-linear feature extractor in many machine learning applications. However, with a time complexity of O(n3), the practicality of KPCA on large datasets is minimal. In this paper, we propose an approximate incremental KPCA algorithm which allows efficient processing of large datasets. We extend a linear PCA updating algorithm to the non-linear case by utilizing the kernel trick, and apply a reduced set construction method to compress expressions for the derived KPCA basis at each update. In addition, we show how multiple feature space vectors can be compressed efficiently, and how approximated KPCA bases can be re-orthogonalized using the kernel trick. The proposed method is justified through experimental validations.
Tat-Jun Chin, David Suter
BMVC2
2006 3D Object Pose Inference via Kernel Principal Component Analysis with Image Euclidian Distance (IMED)
abstract
Kernel Principal Component Analysis (KPCA) is a powerful non-linear unsupervised learning technique for high dimensional pattern analysis. KPCA on images, however, usually considers each image pixel as an independent dimension and does not take into account the spatial relationship of nearby pixels. In this paper, we show how the Image Euclidian Distance (IMED), which takes into account local pixel intensities, can efficiently be embedded into KPCA via the Kronecker product and Eigenvector projections, whilst still retaining desirable properties of Euclidian distance (such as kernel positive definitiveness and effective image de-noising). We demonstrate that KPCA with embedded IMED is a more intuitive and accurate technique than standard KPCA through a 3D object pose estimation application. 1
Therdsak Tangkuampien, David Suter
BMVC2
2006 Real-Time Human Pose Inference using Kernel Principal Component Pre-image Approximations
abstract
We present a real-time markerless human motion capture technique based on un-calibrated synchronized cameras. Training sets of real motions captured from marker based systems are used to learn an optimal pose manifold of human motion via Kernel Principal Component Analysis (KPCA). Similarly, a synthetic silhouette manifold is also learnt, and markerless motion capture can then be viewed as the problem of mapping from the silhouette manifold to the pose manifold. After training, novel silhouettes of previously unseen actors are projected through the two manifolds using Locally Linear Embedding (LLE) reconstruction. The output pose is generated by approximating the pre-image (inverse mapping) of the LLE reconstructed vector from the pose manifold. 1
Therdsak Tangkuampien, David Suter
BMVC2
2006 Effective Appearance Model and Similarity Measure for Particle Filtering and Visual Tracking
Hanzi Wang, David Suter, Konrad Schindler
ECCV (3)2
2006 Parametric model-based motion segmentation using surface selection criterion
Niloofar Gheissari, Alireza Bab-Hadiashar, David Suter
Comput. Vis. Image Underst.3
2006 An Analysis of Linear Subspace Approaches for Computer Vision and Pattern Recognition
Pei Chen 0001, David Suter
Int. J. Comput. Vis.2
2006 Two-View Multibody Structure-and-Motion with Outliers through Model Selection
abstract
Multibody structure-and-motion (MSaM) is the problem to establish the multiple-view geometry of several views of a 3D scene taken at different times, where the scene consists of multiple rigid objects moving relative to each other. We examine the case of two views. The setting is the following: Given are a set of corresponding image points in two images, which originate from an unknown number of moving scene objects, each giving rise to a motion model. Furthermore, the measurement noise is unknown, and there are a number of gross errors, which are outliers to all models. The task is to find an optimal set of motion models for the measurements. It is solved through Monte-Carlo sampling, careful statistical analysis of the sampled set of motion models, and simultaneous selection of multiple motion models to best explain the measurements. The framework is not restricted to any particular model selection mechanism because it is developed from a Bayesian viewpoint: Different model selection criteria are seen as different priors for the set of moving objects, which allow one to bias the selection procedure for different purposes.
Konrad Schindler, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Two-View Multibody Structure-and-Motion with Outliers
abstract
Multi-body structure-and-motion (MSaM) is the problem to establish the multiple-view geometry of several views of a 3D scene taken at different times, where the scene consists of multiple rigid objects moving relative to each other. We examine the case of two views. The setting is the following: given are a set of corresponding image points in two images, which originate from an unknown number of moving scene objects, each giving rise to a motion model. Furthermore, the measurement noise is unknown, and there are a number of gross errors, which are outliers to all models. The task to find an optimal set of motion models for the measurements is solved through Monte-Carlo sampling, careful statistical analysis of the data and simultaneous selection of multiple motion models.
Konrad Schindler, David Suter
CVPR (2)2
2005 A re-evaluation of mixture of Gaussian background modeling [video signal processing applications]
abstract
The mixture of Gaussians (MOG) has been widely used for robustly modeling complicated backgrounds, especially those with small repetitive movements (such as leaves, bushes, rotating fan, ocean waves, rain). The performance of MOG can be greatly improved by tackling several practical issues. In this paper, we quantitatively evaluate (using the Wallflower benchmarks) the performance of the MOG with and without our modifications. The experimental results show that the MOG, with our modifications, can achieve much better results - even outperforming other state-of-the-art methods.
Hanzi Wang, David Suter
ICASSP (2)2
2005 Tracking and segmenting people with occlusions by a sample consensus based method
abstract
One of the most difficult issues in visual tracking is to track people in groups, especially under occlusions. In this paper, we present a novel sample consensus based method, which utilizes both color and spatial information of human bodies, to model the appearance of people. We use this appearance model to segment and track people through occlusions. We show experimental results in several video sequences to validate the effectiveness of the proposed method.
Hanzi Wang, David Suter
ICIP (2)2
2005 Subspace-based face recognition: outlier detection and a new distance criterion
abstract
Illumination effects, including shadows and varying lighting, make the problem of face recognition challenging. Experimental and theoretical results show that the face images under different illumination conditions approximately lie in a low-dimensional subspace, hence principal component analysis (PCA) or low-dimensional subspace techniques have been used. Following this spirit, we propose new techniques for the face recognition problem, including an outlier detection strategy (mainly for those points not following the Lambertian reflectance model), and a new error criterion for the recognition algorithm. Experiments using the Yale-B face database show the effectiveness of the new strategies.
Pei Chen 0001, David Suter
Int. J. Pattern Recognit. Artif. Intell.2
2005 Object tracking in image sequences using point features
Prithiraj Tissainayagam, David Suter
Pattern Recognit.2
2004 Robust Fitting by Adaptive-Scale Residual Consensus
Hanzi Wang, David Suter
ECCV (3)2
2004 Shift-invariant wavelet denoising using interscale dependency
abstract
Using statistical modeling in the wavelet domain, we address the problem of image denoising. Despite being effective, the denoised images can suffer from the Gibbs-like artifacts, like ringing around the edges and speckles in the smooth regions. We employ shift-invariant (SI) wavelet denoising in order to reduce these unpleasant artifacts. Not only is the visual quality greatly improved but also a PSNR gain of about 0.7/spl sim/0.9 dB is obtained. The proposed approach, siPAB, outperforms siHMT, which is a competitive SI wavelet denoising approach, by 0.1/spl sim/0.5 dB.
Pei Chen 0001, David Suter
ICIP2
2004 MDPE: A Very Robust Estimator for Model Fitting and Range Image Segmentation
Hanzi Wang, David Suter
Int. J. Comput. Vis.2
2004 Special issue on statistical methods in video processing
David Suter, Dorin Comaniciu, Kenichi Kanatani
Image Vis. Comput.1
2004 Assessing the performance of corner detectors for point feature tracking applications
Prithiraj Tissainayagam, David Suter
Image Vis. Comput.2
2004 Recovering the Missing Components in a Large Noisy Low-Rank Matrix: Application to SFM
abstract
In computer vision, it is common to require operations on matrices with "missing data," for example, because of occlusion or tracking failures in the Structure from Motion (SFM) problem. Such a problem can be tackled, allowing the recovery of the missing values, if the matrix should be of low rank (when noise free). The filling in of missing values is known as imputation. Imputation can also be applied in the various subspace techniques for face and shape classification, online "recommender" systems, and a wide variety of other applications. However, iterative imputation can lead to the "recovery" of data that is seriously in error. In this paper, we provide a method to recover the most reliable imputation, in terms of deciding when the inclusion of extra rows or columns, containing significant numbers of missing entries, is likely to lead to poor recovery of the missing parts. Although the proposed approach can be equally applied to a wide range of imputation methods, this paper addresses only the SFM problem. The performance of the proposed method is compared with Jacobs' and Shum's methods for SFM.
Pei Chen 0001, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Robust Adaptive-Scale Parametric Model Estimation for Computer Vision
abstract
Robust model fitting essentially requires the application of two estimators. The first is an estimator for the values of the model parameters. The second is an estimator for the scale of the noise in the (inlier) data. Indeed, we propose two novel robust techniques: the Two-Step Scale estimator (TSSE) and the Adaptive Scale Sample Consensus (ASSC) estimator. TSSE applies nonparametric density estimation and density gradient estimation techniques, to robustly estimate the scale of the inliers. The ASSC estimator combines Random Sample Consensus (RANSAC) and TSSE: using a modified objective function that depends upon both the number of inliers and the corresponding scale. ASSC is very robust to discontinuous signals and data with multiple structures, being able to tolerate more than 80 percent outliers. The main advantage of ASSC over RANSAC is that prior knowledge about the scale of inliers is not needed. ASSC can simultaneously estimate the parameters of a model and the scale of the inliers belonging to that model. Experiments on synthetic data show that ASSC has better robustness to heavily corrupted data than Least Median Squares (LMedS), Residual Consensus (RESC), and Adaptive Least Kth order Squares (ALKS). We also apply ASSC to two fundamental computer vision tasks: range image segmentation and robust fundamental matrix estimation. Experiments show very promising results.
Hanzi Wang, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Variable Bandwidth QMDPE and Its Application in Robust Optical Flow Estimation
abstract
Robust estimators, such as least median of squared (LMedS) residuals, M-estimators, the least trimmed squares (LTS) etc., have been employed to estimate optical flow from image sequences in recent years. However, these robust estimators have a breakdown point of no more than 50%. We propose a novel robust estimator, called variable bandwidth quick maximum density power estimator (vbQMDPE), which can tolerate more than 50% outliers. We apply the novel proposed estimator to robust optical flow estimation. Our method yields better results than most other recently proposed methods, and it has the potential to better handle multiple motion effects.
Hanzi Wang, David Suter
ICCV2
2003 Contour tracking with automatic motion model switching
Prithiraj Tissainayagam, David Suter
Pattern Recognit.2
2003 Using symmetry in robust model fitting
Hanzi Wang, David Suter
Pattern Recognit. Lett.2
2002 A novel robust method for large numbers of gross errors
abstract
In computer vision tasks, it frequently happens that gross noise occupies the absolute majority of the data. Most robust estimators can tolerate no more than 50% gross errors. In this article, we propose a highly robust estimator, called MDPE (maximum density power estimator), employing density estimation and density gradient estimation techniques in the residual space. This estimator can tolerate more than 85% outliers. Experiments illustrate that the MDPE has a higher breakdown point and less errors than other recently proposed similar estimators: least median of squares (LMedS), residual consensus (RESC), and adaptive least kth order squares (ALKS).
Hanzi Wang, David Suter
ICARCV2
2002 LTSD: a highly efficient symmetry-based robust estimator
abstract
Although the least median of squares (LMedS) method and the least trimmed squares (LTS) method are said to hive a high breakdown point (50%), they can break down at unexpectedly lower percentages of outliers when those outliers are clustered. In this paper, we investigate the breakdown of LMedS and the LTS when a large percentage of clustered outliers exist in the data. We introduce the concept of symmetry distance (SD) and propose an improved method, called the least trimmed symmetry distance (LTSD). The experimental results show the LTSD gives better results than the LMedS method and the LTS method particularly when there is a large percentage of clustered outliers and/or a large standard variance in the inlier population.
Hanzi Wang, David Suter
ICARCV2
2001 Performance Prediction Analysis of a Point Feature Tracker Based on Different Motion Models
Prithiraj Tissainayagam, David Suter
Comput. Vis. Image Underst.2
2001 Visual tracking with automatic motion model switching
Prithiraj Tissainayagam, David Suter
Pattern Recognit.2
2000 Visual Tracking of Multiple Objects with Automatic Motion Model Switching
abstract
In this paper we present an efficient contour tracking algorithm which can track 2D silhouettes of multiple objects in extended image sequences captured by a static camera. We represent contours using cubic B-splines, and our tracking algorithm is based on tracking a lower dimensional shape-space. The tracker is coupled with a multiple model filtering algorithm which caters for objects that move with variable motion. The model based technique we provide is capable of tracking rigid and nonrigid object contours with good accuracy.
Prithiraj Tissainayagam, David Suter
ICPR2
2000 Left Ventricular Motion Reconstruction Based on Elastic Vector Splines
abstract
In medical imaging it is common to reconstruct dense motion estimates, from sparse measurements of that motion, using some form of elastic spline (thin-plate spline, snakes and other deformable models, etc.). Usually the elastic spline uses only bending energy (second-order smoothness constraint) or stretching energy (first-order smoothness constraint), or a combination of the two. These elastic splines belong to a family of elastic vector splines called the Laplacian splines. This spline family is derived from an energy minimization functional, which is composed of multiple-order smoothness constraints. These splines can be explicitly tuned to vary the smoothness of the solution according to the deformation in the modeled material/tissue. In this context, it is natural to question which members of the family will reconstruct the motion more accurately. We compare different members of this spline family to assess how well these splines reconstruct human cardiac motion. We find that the commonly used splines (containing first-order and/or second-order smoothness terms only) are not the most accurate for modeling human cardiac motion.
David Suter
IEEE Trans. Medical Imaging1
1998 Robust Total least Squares Based Optic Flow Computation
Alireza Bab-Hadiashar, David Suter
ACCV (1)2
1998 Robust Motion Segmentation Using Rank Ordering Estimatiors
Alireza Bab-Hadiashar, David Suter
ACCV (2)2
1998 Multiscale Image Representation and Edge Detection
David Suter
ACCV (2)2
1998 Robust range segmentation
abstract
This paper proposes a robust estimator which is capable of segmenting multi-structural data. The proposed estimation technique is used to segment range data into linear (planar) and quadratic surfaces. The performance of the proposed range segmentation algorithm is tested on a number of real data sets.
Alireza Bab-Hadiashar, David Suter
ICPR2
1998 Image coordinate transformation based on DIV-CURL vector splines
abstract
We present a vector spline technique for vector field reconstruction. These vector splines are based on an energy minimization functional, which involves the divergence and the rotational fields of the approximated vector. This technique can be used to determine the underlying coordinate transformation in image mapping.
David Suter
ICPR2
1998 Visual tracking and motion determination using the IMM algorithm
abstract
We present a feature tracking system with automatic motion determination of features in an image sequence. The positions of features (corners) extracted in the first frame of a sequence are estimated and predicted in the subsequent frames by using an extension of Bayesian multiple hypothesis technique (MHT) based on different motion models. The tracking of features is based on the interacting multiple model (IMM). The paper shows how the IMM algorithm combined with a MHT framework can be used in a visual tracking scenario. We considered different order (types) velocity and acceleration models for the IMM algorithm and applied them to two image sequences, the PUMA sequence and toy car sequence. The study shows that the method proposed can distinguish between different motions depicted in an image sequence with very good tracking results.
Prithiraj Tissainayagam, David Suter
ICPR2
1998 Robust Optic Flow Computation
Alireza Bab-Hadiashar, David Suter
Int. J. Comput. Vis.2
1997 Optic flow calculation using robust statistics
abstract
A method for calculating optic flow, using robust statistics, is developed. The method generally out-performs all competing methods in terms of accuracy. One of the key features in the success of this method, is that we use least median of squares, which is known to be robust to outliers. The computational cost is kept very low by using an approximate solution to the least median of squares only in a first stage that detects outliers. The essential ingredients of our method should be applicable in a wide range of other computer vision problems.
Alireza Bab-Hadiashar, David Suter
CVPR2
1996 Robust optic flow estimation using least median of squares
abstract
A new approach to optic flow calculation, based on a highly robust statistical technique, is presented. In this algorithm, the optic flow problem is first formulated as a standard least squares problem. Then, its associated closest point problem is introduced and the transformation which takes this problem to a standard regression problem is provided. The least median of squares technique is used to solve the resulting regression problem. Some experimental results for both synthetic and real image sequences are also presented.
Alireza Bab-Hadiashar, David Suter
ICIP (1)2
1995 Restoration of historic film for digital compression: a case study
abstract
With the advent of compressed video standards such as MPEG, film archives around the world are looking at using the medium as a new means of distributing historically significant film material. However, we find that a compression standard such as MPEG, which is optimised for the statistics of modern motion picture film and video, is poorly suited to coding motion pictures recorded with early technologies and suffering from severe age related degradation. Indeed restoration of some historically significant film is necessary, not just so that it may be returned to its former visual quality, but to prevent the film from looking significantly worse after compression. We describe the degradation artifacts encountered-in a historically significant film made in 1906, the strategies employed to improve its compressed quality and the results of these efforts.
P. Richardson, David Suter
ICIP2
1994 Motion estimation and vector splines
abstract
Many formulations of visual reconstruction problems (e.g. optic flow, shape from shading, biomedical motion estimation from CT data) involve the recovery of a vector field. Often the solution is characterized via a generalized spline or regularization formulation using a smoothness constraint. This paper introduces a decomposition of the smoothness constraint into two parts: one related to the divergence of the vector field and one related to the curl or vorticity. This allows one to "tune" the smoothness to the properties of the data. One can, for example, use a high weighting on the smoothness imposed upon the curl in order to preserve the divergent parts of the field. For a particular spline within the family introduced by this decomposition process, we derive an exact solution and demonstrate the approach on examples.>
David Suter
CVPR1
1992 Mixed Finite Element Based Neural Networks in Visual Reconstruction
abstract
This paper shows how visual reconstruction problems can be solved using analog networks resulting from the reformulation of the reconstruction problem into a series of coupled sub-problems. A novel analog implementation based upon Platt's constraint networks is suggested for neural network implementation and the effectiveness is demonstrated with examples. The complete approach bears a philosophical similarity to the Harris Coupled Depth-Slope model of visual reconstruction. Significant differences appear, though, in the Lagrangian formulation, the mixed finite element discretization, and the analog implementation. It is also shown that the resulting networks can be structurally different from those suggested by Harris.
David Suter
Int. J. Pattern Recognit. Artif. Intell.1
1991 Constraint Networks in Vision
abstract
Applications in machine vision of constraint networks based on an augmented Lagrangian formulation are discussed. Only those applications that have a fundamental significance are addressed. The first of these provides a generalization of the Harris coupled depth-slope analog model of visual reconstruction. Because of the generality of the approach, one can derive many more alternative structures, and the mathematical setting places this approach within the bounds of mixed finite element theory. This offers many advantages in terms of the associated mathematical theory and implementation on digital machines. The second use is in data fusion, which is a crucial task for systems using multiple sensors or methods of analysis of data.>
David Suter
IEEE Trans. Computers1