VLDB 2026 Research / reviewers in the wild / expert
David Suter
dblp:37/5634
· DBLP profile ↗
153ranked-venue papers
5as first author
24since 2021 · last 2025
0000-0001-6306-3023ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 114 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 92 · 1 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 2Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From Pixels to Prognosis: A Multi-modal Attention-Based Framework for Visceral Adipose Tissue Estimation
Arooba Maqsood, Afsah Saleem, Marc Sim, David Suter, Simone Radavelli-Bagatini, Jonathan M. Hodgson, Richard Prince, Kun Zhu 0028, William D. Leslie, John T. Schousboe, Joshua R. Lewis, Syed Zulqarnain Gilani |
MICCAI (15) | 4 |
| 2025 | Adaptive Graph Attention-Guided Parallel Sampling and Embedded Selection for Multi-Model FittingabstractMulti-model fitting is a fundamental challenge in computer vision, where real-world data often contains severe gross outliers and pseudo-outliers. Existing methods rely on inefficient sequential hypothesize-and-verify frameworks that require a predefined number of models and inlier thresholds-parameters that are difficult to determine in practical scenes. To overcome these limitations, we propose a novel Adaptive Graph Attention-guided parallel multi-model fitting method (AGASAC) that jointly learns local and global features, performs parallel hypothesis sampling, and executes confidence-embedded model selection. Specifically, we design a dual-confidence graph attention module that models data relationships using an adaptive graph attention network. This module computes minimal-set confidence and quality confidence to guide the multi-model fitting process, eliminating manual parameter tuning. Additionally, we propose a parallel discriminative sampling module that leverages minimal-set confidence to concurrently sample hypotheses. By enforcing a quantized consensus constraint, this module maximizes inter-model variance while minimizing intra-model discrepancy. It enables computationally efficient hypothesis generation and pseudo-outlier suppression. To obtain high-quality models, we present a quality-embedded selection module that integrates quality confidence into the joint optimization of model selection and data clustering. Extensive experiments show that the proposed method achieves a lower transfer error of 0.39 pixels and a 36.92% runtime reduction, surpassing state-of-the-art methods. The code is available at https://github.com/YWY-Vivian/AGASAC. Wenyu Yin, Shuyuan Lin, David Suter, Hanzi Wang |
ACM Multimedia | 3 |
| 2024 | Spatially-Aware Speaker for Vision-and-Language Navigation Instruction GenerationabstractEmbodied AI aims to develop robots that can understand and execute human language instructions, as well as communicate in natural languages. On this front, we study the task of generating highly detailed navigational instructions for the embodied robots to follow. Although recent studies have demonstrated significant leaps in the generation of step-by-step instructions from sequences of images, the generated instructions lack variety in terms of their referral to objects and landmarks. Existing speaker models learn strategies to evade the evaluation metrics and obtain higher scores even for low-quality sentences. In this work, we propose SAS (Spatially-Aware Speaker), an instruction generator or Speaker model that utilises both structural and semantic knowledge of the environment to produce richer instructions. For training, we employ a reward learning method in an adversarial setting to avoid systematic bias introduced by language evaluation metrics. Empirically, our method outperforms existing instruction generation models, evaluated using standard metrics. Our code is available at https://github.com/gmuraleekrishna/SAS. Muraleekrishna Gopinathan, Martin Masek, Jumana M. Abu-Khalaf, David Suter |
ACL (1) | 4 |
| 2024 | An Exploration of Diabetic Foot Osteomyelitis X-ray Data for Deep Learning Applications
Brandon Abela, Martin Masek, Jumana M. Abu-Khalaf, David Suter, Ashu Gupta |
AIME (2) | 4 |
| 2024 | Segment Any Object Model (SAOM): Real-To-Simulation Fine-Tuning Strategy For Multi-Class Multi-Instance SegmentationabstractMulti-class multi-instance segmentation is the task of identifying masks for multiple object classes and multiple instances of the same class within an image. The foundational Segment Anything Model (SAM) is designed for promptable multi-class multi-instance segmentation but tends to output part or sub-part masks in the “everything” mode for various real-world applications. Whole object segmentation masks play a crucial role for indoor scene understanding, especially in robotics applications. We propose a new domain invariant Real-to-Simulation (Real-Sim) fine-tuning strategy for SAM. We use object images and ground truth data collected from Ai2Thor simulator during fine-tuning (real-to-sim). To allow our Segment Any Object Model (SAOM) to work in the “everything” mode, we propose the novel nearest neighbour assignment method, updating point embeddings for each ground-truth mask. SAOM is evaluated on our own dataset collected from Ai2Thor simulator. SAOM significantly improves on SAM, with a $28 \%$ increase in mIoU and a $25 \%$ increase in mAcc for 54 frequently-seen indoor object classes. Moreover, our Real-to-Simulation fine-tuning strategy demonstrates promising generalization performance in real environments without being trained on the real-world data (sim-to-real). The dataset and the code are available here. Mariia Khan, Yue Qiu 0001, Yuren Cong, Bodo Rosenhahn, Jumana M. Abu-Khalaf, David Suter |
ICIP | 6 |
| 2024 | StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language NavigationabstractEmbodied navigation requires robots to understand and interact with the environment based on given tasks. Vision-Language Navigation (VLN) is an embodied navigation task, where a robot navigates within a previously seen and unseen environment, based on linguistic instruction and visual inputs. VLN agents need access to both local and global action spaces; former for immediate decision making and the latter for recovering from navigational mistakes. Prior VLN agents rely only on instruction-viewpoint alignment for local and global decision making and back-track to a previously visited viewpoint, if the instruction and its current viewpoint mismatches. These methods are prone to mistakes, due to the complexity of the instruction and partial observability of the environment. We posit that, back-tracking is sub-optimal and agent that is aware of its mistakes can recover efficiently. For optimal recovery, exploration should be extended to unexplored viewpoints (or frontiers). The optimal frontier is a recently observed but unexplored viewpoint that aligns with the instruction and is novel. We introduce a memory-based and mistake-aware path planning strategy for VLN agents, called StratXplore, that presents global and local action planning to select the optimal frontier for path correction. The proposed method collects all past actions and viewpoint features during navigation and then selects the optimal frontier suitable for recovery. Experimental results show this simple yet effective strategy improves the success rate on two VLN datasets with different task complexities. Muraleekrishna Gopinathan, Jumana M. Abu-Khalaf, David Suter, Martin Masek |
IROS | 3 |
| 2024 | Indoor Scene Change Understanding (SCU): Segment, Describe, and Revert Any ChangeabstractUnderstanding of scene changes is crucial for embodied AI applications, such as visual room rearrangement, where the agent must revert changes by restoring the objects to their original locations or states. Visual changes between two scenes, pre- and post-rearrangement, encompass two tasks: scene change detection (locating changes) and image difference captioning (describing changes). While previous methods, focused on sequential 2D images, have addressed these tasks separately, it is essential to emphasize the significance of their combination. Therefore, we propose a new Scene Change Understanding (SCU) task for simultaneous change detection and description. Moreover, we go beyond change language description generation and aim to generate rearrangement instructions for the robotic agent to revert changes. To solve this task, we propose a novel method - EmbSCU, which allows to compare instance-level change object masks (for 53 frequently-seen indoor object classes) before and after changes and generate rearrangement language instructions for the agent. EmbSCU is built on our Segment Any Object Model (SAOMv2) - a fine-tuned version of Segment Anything Model (SAM), adapted to obtain instance-level object masks for both foreground and background objects in indoor embodied environments. EmbSCU is evaluated on our own dataset of sequential 2D image pairs before and after changes, collected from the Ai2Thor simulator. The proposed framework achieves promising results in both change detection and change description. Moreover, EmbSCU demonstrates positive generalization results on real-world scenes without using any real-life data during training. The dataset and the code are available here. Mariia Khan, Yue Qiu 0001, Yuren Cong, Bodo Rosenhahn, David Suter, Jumana M. Abu-Khalaf |
IROS | 5 |
| 2024 | A Hybrid CNN-Transformer Feature Pyramid Network for Granular Abdominal Aortic Calcification Detection from DXA Images
Zaid Ilyas, Afsah Saleem, David Suter, John T. Schousboe, William D. Leslie, Joshua R. Lewis, Syed Zulqarnain Gilani |
MICCAI (11) | 3 |
| 2024 | Single Domain Generalization via Normalised Cross-correlation Based ConvolutionsabstractDeep learning techniques often perform poorly in the presence of domain shift, where the test data follows a different distribution than the training data. The most practically desirable approach to address this issue is Single Domain Generalization (S-DG), which aims to train robust models using data from a single source. Prior work on S-DG has primarily focused on using data augmentation techniques to generate diverse training data. In this paper, we explore an alternative approach by investigating the robustness of linear operators, such as convolution and dense layers commonly used in deep learning. We propose a novel operator called "XCNorm" that computes the normalized cross-correlation between weights and an input feature patch. This approach is invariant to both affine shifts and changes in energy within a local feature patch and eliminates the need for commonly used non-linear activation functions. We show that deep neural networks composed of this operator are robust to common semantic distribution shifts. Furthermore, our empirical results on single-domain generalization benchmarks demonstrate that our proposed technique performs comparably to the stateof-the-art methods.1 Weiqin Chuah, Ruwan B. Tennakoon, Reza Hoseinnezhad, David Suter, Alireza Bab-Hadiashar |
WACV | 4 |
| 2024 | Estimating Blood Alcohol Level Through Facial Features for Driver Impairment AssessmentabstractDrunk driving-related road accidents contribute significantly to the global burden of road injuries. Addressing alcohol-related harm, particularly during safety-critical activities like driving, requires real-time monitoring of an individual’s blood alcohol concentration (BAC). We devise an in-vehicle machine learning system that harnesses standard commercial RGB cameras to predict critical levels of BAC. Our system can detect instances of alcohol intoxication impairment as subtle as 0.05 g/dL (WHO recommended legal limit for driving), with an accuracy of 75%, by leveraging the physiological manifestations of alcohol intoxication on a driver’s face. This system holds great promise for improving road safety. In tandem, we have compiled a data set of 60 subjects engaged in simulated driving scenarios, spanning three levels of alcohol intoxication. These scenarios were captured and divided into video segments labeled "sober","low’, and "severe" Alcohol Intoxication Impairment (AII), constituting the basis for evaluating our system’s performance. To the best of our knowledge, this study is the first to create a large-scale real-life dataset of alcohol intoxication and assess intoxication levels using an off-the-shelf RGB camera to detect drunk driving. Ensiyeh Keshtkaran, Brodie von Berg, Grant Regan, David Suter, Syed Zulqarnain Gilani |
WACV | 4 |
| 2023 | SCOL: Supervised Contrastive Ordinal Loss for Abdominal Aortic Calcification Scoring on Vertebral Fracture Assessment Scans
Afsah Saleem, Zaid Ilyas, David Suter, Ghulam M. Hassan, Siobhan Reid, John T. Schousboe, Richard Prince, William D. Leslie, Joshua R. Lewis, Syed Zulqarnain Gilani |
MICCAI (6) | 3 |
| 2023 | Generalized framework for image and video object segmentation using affinity learning and message passing GNNSabstractDespite significant amount of work reported in the computer vision literature, segmenting images or videos based on multiple cues such as objectness, texture and motion, is still a challenge. This is particularly true when the number of objects to be segmented is not known or there are objects that are not classified in the training data (unknown objects). A possible remedy to this problem is to utiize graph-based clustering techniques such as Correlation Clustering. It is known that using long range affinities (Lifted multicut), makes correlation clustering more accurate than using only adjacent affinities (Multicut). However, the former is computationally expensive and hard to use. In this paper, we introduce a new framework to perform image/motion segmentation using an affinity learning module and a Message Passing Graph Neural Network (MPGNN). The affinity learning module uses a permutation invariant affinity representation to overcome the multi-object problem. The paper shows, both theoretically and empirically, that the proposed MPGNN aggregates higher order information and thereby converts the Lifted Multicut Problem (LMP) to a Multicut Problem (MP), which is easier and faster to solve. Importantly, the proposed method can be generalized to deal with different clustering problems with the same MPGNN architecture. For instance, our method produces competitive results for single image segmentation (on BSDS dataset) as well as unsupervised video object segmentation (on DAVIS17 dataset), by only changing the feature extraction part. In addition, using an ablation study on the proposed MPGNN architecture, we show that the way we update the parameterized affinities directly contributes to the accuracy of the results. Sundaram Muthu, Ruwan B. Tennakoon, Tharindu Rathnayake, Reza Hoseinnezhad, David Suter, Alireza Bab-Hadiashar |
Comput. Vis. Image Underst. | 5 |
| 2023 | An Information-Theoretic Method to Automatic Shortcut Avoidance and Domain Generalization for Dense Prediction TasksabstractDeep convolutional neural networks for dense prediction tasks are commonly optimized using synthetic data, as generating pixel-wise annotations for real-world data is laborious. However, the synthetically trained models do not generalize well to real-world environments. This poor "synthetic to real" (S2R) generalization we address through the lens of shortcut learning. We demonstrate that the learning of feature representations in deep convolutional networks is heavily influenced by synthetic data artifacts (shortcut attributes). To mitigate this issue, we propose an Information-Theoretic Shortcut Avoidance (ITSA) approach to automatically restrict shortcut-related information from being encoded into the feature representations. Specifically, our proposed method minimizes the sensitivity of latent features to input variations: to regularize the learning of robust and shortcut-invariant features in synthetically trained models. To avoid the prohibitive computational cost of direct input sensitivity optimization, we propose a practical yet feasible algorithm to achieve robustness. Our results show that the proposed method can effectively improve S2R generalization in multiple distinct dense prediction tasks, such as stereo matching, optical flow, and semantic segmentation. Importantly, the proposed method enhances the robustness of the synthetically trained networks and outperforms their fine-tuned counterparts (on real data) for challenging out-of-domain applications. Weiqin Chuah, Ruwan B. Tennakoon, Reza Hoseinnezhad, David Suter, Alireza Bab-Hadiashar |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Unsupervised Learning for Maximum Consensus Robust Fitting: A Reinforcement Learning ApproachabstractRobust model fitting is a core algorithm in several computer vision applications. Despite being studied for decades, solving this problem efficiently for datasets that are heavily contaminated by outliers is still challenging: due to the underlying computational complexity. A recent focus has been on learning-based algorithms. However, most of these approaches are supervised (which require a large amount of labelled training data). In this paper, we introduce a novel unsupervised learning framework: that learns to directly (without labelled data) solve robust model fitting. Moreover, unlike other learning-based methods, our work is agnostic to the underlying input features, and can be easily generalized to a wide variety of LP-type problems with quasi-convex residuals. We empirically show that our method outperforms existing (un)supervised learning approaches, and also achieves competitive results compared to traditional (non-learning-based) methods. Our approach is designed to try to maximise consensus (MaxCon), similar to the popular RANSAC. The basis of our approach, is to adopt a Reinforcement Learning framework. This requires designing appropriate reward functions, and state encodings. We provide a family of reward functions, tunable by choice of a parameter. We also investigate the application of different basic and enhanced Q-learning components. Giang Truong, Huu Le, Erchuan Zhang, David Suter, Syed Zulqarnain Gilani |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | ITSA: An Information-Theoretic Approach to Automatic Shortcut Avoidance and Domain Generalization in Stereo Matching NetworksabstractState-of-the-art stereo matching networks trained only on synthetic data often fail to generalize to more challenging real data domains. In this paper, we attempt to unfold an important factor that hinders the networks from generalizing across domains: through the lens of shortcut learning. We demonstrate that the learning of feature representations in stereo matching networks is heavily influenced by synthetic data artefacts (shortcut attributes). To mitigate this issue, we propose an Information-Theoretic Shortcut Avoidance (ITSA) approach to automatically restrict shortcut-related information from being encoded into the feature representations. As a result, our proposed method learns robust and shortcut-invariant features by minimizing the sensitivity of latent features to input variations. To avoid the prohibitive computational cost of direct input sensitivity optimization, we propose an effective yet feasible algorithm to achieve robustness. We show that using this method, state-of-the-art stereo matching networks that are trained purely on synthetic data can effectively generalize to challenging and previously unseen real data scenarios. Importantly, the proposed method enhances the robustness of the synthetic trained networks to the point that they outperform their fine-tuned counterparts (on real data) for challenging out-of-domain stereo datasets. Weiqin Chuah, Ruwan B. Tennakoon, Reza Hoseinnezhad, Alireza Bab-Hadiashar, David Suter |
CVPR | 5 |
| 2022 | A Hybrid Quantum-Classical Algorithm for Robust FittingabstractFitting geometric models onto outlier contaminated data is provably intractable. Many computer vision systems rely on random sampling heuristics to solve robust fitting, which do not provide optimality guarantees and error bounds. It is therefore critical to develop novel approaches that can bridge the gap between exact solutions that are costly, and fast heuristics that offer no quality assurances. In this paper, we propose a hybrid quantum-classical algorithm for robust fitting. Our core contribution is a novel robust fitting formulation that solves a sequence of integer programs and terminates with a global solution or an error bound. The combinatorial subproblems are amenable to a quantum annealer, which helps to tighten the bound efficiently. While our usage of quantum computing does not surmount the fundamental intractability of robust fitting, by providing error bounds our algorithm is a practical improvement over randomised heuristics. Moreover, our work represents a concrete application of quantum computing in computer vision. We present results obtained using an actual quantum computer (D-Wave Advantage) and via simulation11Source code: https://github.com/dadung/HQC-robust-fitting. Anh-Dzung Doan, Michele Sasdelli, David Suter, Tat-Jun Chin |
CVPR | 3 |
| 2022 | Maximum Consensus by Weighted Influences of Monotone Boolean FunctionsabstractMaximisation of Consensus (MaxCon) is one of the most widely used robust criteria in computer vision. Tennakoon et al. (CVPR2021), made a connection between MaxCon and estimation of influences of a Monotone Boolean function. In such, there are two distributions involved: the distribution defining the influence measure; and the distribution used for sampling to estimate the influence measure. This paper studies the concept of weighted influences for solving MaxCon. In particular, we study the Bernoulli measures. Theoretically, we prove the weighted influences, under this measure, of points belonging to larger structures are smaller than those of points belonging to smaller structures in general. We also consider another “natural” family of weighting strategies: sampling with uniform measure concentrated on a particular (Hamming) level of the cube. One can choose to have matching distributions: the same for defining the measure as for implementing the sampling. This has the advantage that the sampler is an unbiased estimator of the measure. Based on weighted sampling, we modify the algorithm of Tennakoon et al., and test on both synthetic and real datasets. We show some modest gains of Bernoulli sampling, and we illuminate some of the interactions between structure in data and weighted measures and weighted sampling. Erchuan Zhang, David Suter, Ruwan B. Tennakoon, Tat-Jun Chin, Alireza Bab-Hadiashar, Giang Truong, Syed Zulqarnain Gilani |
CVPR | 2 |
| 2022 | Show, Attend and Detect: Towards Fine-Grained Assessment of Abdominal Aortic Calcification on Vertebral Fracture Assessment Scans
Syed Zulqarnain Gilani, Naeha Sharif, David Suter, John T. Schousboe, Siobhan Reid, William D. Leslie, Joshua R. Lewis |
MICCAI (3) | 3 |
| 2022 | Sparse Hypergraph Community Detection Thresholds in Stochastic Block ModelabstractCommunity detection in random graphs or hypergraphs is an interesting fundamental problem in statistics, machine learning and computer vision. When the hypergraphs are generated by a {\em stochastic block model}, the existence of a sharp threshold on the model parameters for community detection was conjectured by Angelini et al. 2015. In this paper, we confirm the positive part of the conjecture, the possibility of non-trivial reconstruction above the threshold, for the case of two blocks. We do so by comparing the hypergraph stochastic block model with its Erd{\"o}s-R{\'e}nyi counterpart. We also obtain estimates for the parameters of the hypergraph stochastic block model. The methods developed in this paper are generalised from the study of sparse random graphs by Mossel et al. 2015 and are motivated by the work of Yuan et al. 2022. Furthermore, we present some discussion on the negative part of the conjecture, i.e., non-reconstruction of community structures. Erchuan Zhang, David Suter, Giang Truong, Syed Zulqarnain Gilani |
NeurIPS | 2 |
| 2022 | Semantic Guided Long Range Stereo Depth Estimation for Safer Autonomous Vehicle ApplicationsabstractAutonomous vehicles in intelligent transportation systems must be able to perform reliable and safe navigation. This necessitates accurate object detection, which is commonly achieved by high-precision depth perception. Existing stereo vision-based depth estimation systems generally involve computation of pixel correspondences and estimation of disparities between rectified image pairs. The estimated disparity values will be converted into depth values in downstream applications. As most applications often work in the depth domain, the accuracy of depth estimation is often more compelling than disparity estimation. However, at large distances (> 50m), the accuracy of disparity estimation does not directly translate to the accuracy of depth estimation. In the context of learning-based stereo systems, this is mainly due to biases imposed by the choices of the disparity-based loss function and the training data. Consequently, the learning algorithms often produce unreliable depth estimates of under-represented foreground objects, particularly at large distances. To resolve this issue, we first analyze the effect of those biases and then propose a pair of depth-based loss functions for foreground objects and background separately. These loss functions can be tuned and can balance the inherent bias of the stereo learning algorithms. The efficacy of our solution is demonstrated by an extensive set of experiments, which are benchmarked against state of the art. We show on the KITTI 2015 benchmark that our proposed solution yields substantial improvements in disparity and depth estimation, particularly for objects located at distances beyond 50 meters, outperforming the previous state of the art by 10%. Weiqin Chuah, Ruwan B. Tennakoon, Reza Hoseinnezhad, David Suter, Alireza Bab-Hadiashar |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Consensus Maximisation Using Influences of Monotone Boolean FunctionsabstractConsensus maximisation (MaxCon), which is widely used for robust fitting in computer vision, aims to find the largest subset of data that fits the model within some tolerance level. In this paper, we outline the connection between MaxCon problem and the abstract problem of finding the maximum upper zero of a Monotone Boolean Function (MBF) defined over the Boolean Cube. Then, we link the concept of influences (in a MBF) to the concept of outlier (in MaxCon) and show that influences of points belonging to the largest structure in data would generally be smaller under certain conditions. Based on this observation, we present an iterative algorithm to perform consensus maximisation. Results for both synthetic and real visual data experiments show that the MBF based algorithm is capable of generating a near optimal solution relatively quickly. This is particularly important where there are large number of outliers (gross or pseudo) in the observed data. Ruwan B. Tennakoon, David Suter, Erchuan Zhang, Tat-Jun Chin, Alireza Bab-Hadiashar |
CVPR | 2 |
| 2021 | Unsupervised Learning for Robust Fitting: A Reinforcement Learning ApproachabstractRobust model fitting is a core algorithm in a large number of computer vision applications. Solving this problem efficiently for datasets highly contaminated with outliers is, however, still challenging due to the underlying computational complexity. Recent literature has focused on learning-based algorithms. However, most approaches are supervised (which require a large amount of labelled training data). In this paper, we introduce a novel unsupervised learning framework that learns to directly solve robust model fitting. Unlike other methods, our work is agnostic to the underlying input features, and can be easily generalized to a wide variety of LP-type problems with quasi-convex residuals. We empirically show that our method out-performs existing unsupervised learning approaches, and achieves competitive results compared to traditional methods on several important computer vision problems1. Giang Truong, Huu Le, David Suter, Erchuan Zhang, Syed Zulqarnain Gilani |
CVPR | 3 |
| 2021 | Segmentation by Continuous Latent Semantic Analysis for Multi-structure Model Fitting
Guobao Xiao, Hanzi Wang, Jiayi Ma 0001, David Suter |
Int. J. Comput. Vis. | 4 |
| 2021 | Deterministic Approximate Methods for Maximum Consensus Robust FittingabstractMaximum consensus estimation plays a critically important role in several robust fitting problems in computer vision. Currently, the most prevalent algorithms for consensus maximization draw from the class of randomized hypothesize-and-verify algorithms, which are cheap but can usually deliver only rough approximate solutions. On the other extreme, there are exact algorithms which are exhaustive search in nature and can be costly for practical-sized inputs. This paper fills the gap between the two extremes by proposing deterministic algorithms to approximately optimize the maximum consensus criterion. Our work begins by reformulating consensus maximization with linear complementarity constraints. Then, we develop two novel algorithms: one based on non-smooth penalty method with a Frank-Wolfe style optimization scheme, the other based on the Alternating Direction Method of Multipliers (ADMM). Both algorithms solve convex subproblems to efficiently perform the optimization. We demonstrate the capability of our algorithms to greatly improve a rough initial estimate, such as those obtained using least squares or a randomized algorithm. Compared to the exact algorithms, our approach is much more practical on realistic input sizes. Further, our approach is naturally applicable to estimation problems with geometric residuals. Matlab code and demo program for our methods can be downloaded from https://goo.gl/FQcxpi. Huu Le, Tat-Jun Chin, Anders P. Eriksson, Thanh-Toan Do, David Suter |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | End-to-End Learning of Object Motion Estimation from Retinal Events for Event-Based Object TrackingabstractEvent cameras, which are asynchronous bio-inspired vision sensors, have shown great potential in computer vision and artificial intelligence. However, the application of event cameras to object-level motion estimation or tracking is still in its infancy. The main idea behind this work is to propose a novel deep neural network to learn and regress a parametric object-level motion/transform model for event-based object tracking. To achieve this goal, we propose a synchronous Time-Surface with Linear Time Decay (TSLTD) representation, which effectively encodes the spatio-temporal information of asynchronous retinal events into TSLTD frames with clear motion patterns. We feed the sequence of TSLTD frames to a novel Retinal Motion Regression Network (RMRNet) to perform an end-to-end 5-DoF object motion regression. Our method is compared with state-of-the-art object tracking methods, that are based on conventional cameras or event cameras. The experimental results show the superiority of our method in handling various challenging environments such as fast motion and low illumination conditions. Haosheng Chen 0001, David Suter, Qiangqiang Wu, Hanzi Wang |
AAAI | 2 |
| 2020 | Quantum Robust Fitting
Tat-Jun Chin, David Suter, Shin-Fang Ch'ng, James Quach |
ACCV (1) | 2 |
| 2020 | Motion Segmentation of RGB-D Sequences: Combining Semantic and Motion Information Using Statistical InferenceabstractThis paper presents an innovative method for motion segmentation in RGB-D dynamic videos with multiple moving objects. The focus is on finding static, small or slow moving objects (often overlooked by other methods) that their inclusion can improve the motion segmentation results. In our approach, semantic object based segmentation and motion cues are combined to estimate the number of moving objects, their motion parameters and perform segmentation. Selective object-based sampling and correspondence matching are used to estimate object specific motion parameters. The main issue with such an approach is the over segmentation of moving parts due to the fact that different objects can have the same motion (e.g. background objects). To resolve this issue, we propose to identify objects with similar motions by characterizing each motion by a distribution of a simple metric and using a statistical inference theory to assess their similarities. To demonstrate the significance of the proposed statistical inference, we present an ablation study, with and without static objects inclusion, on SLAM accuracy using the TUM-RGBD dataset. To test the effectiveness of the proposed method for finding small or slow moving objects, we applied the method to RGB-D MultiBody and SBM-RGBD motion segmentation datasets. The results showed that we can improve the accuracy of motion segmentation for small objects while remaining competitive on overall measures. Sundaram Muthu, Ruwan B. Tennakoon, Tharindu Rathnayake, Reza Hoseinnezhad, David Suter, Alireza Bab-Hadiashar |
IEEE Trans. Image Process. | 5 |
| 2019 | Hypergraph Optimization for Multi-Structural Geometric Model FittingabstractRecently, some hypergraph-based methods have been proposed to deal with the problem of model fitting in computer vision, mainly due to the superior capability of hypergraph to represent the complex relationship between data points. However, a hypergraph becomes extremely complicated when the input data include a large number of data points (usually contaminated with noises and outliers), which will significantly increase the computational burden. In order to overcome the above problem, we propose a novel hypergraph optimization based model fitting (HOMF) method to construct a simple but effective hypergraph. Specifically, HOMF includes two main parts: an adaptive inlier estimation algorithm for vertex optimization and an iterative hyperedge optimization algorithm for hyperedge optimization. The proposed method is highly efficient, and it can obtain accurate model fitting results within a few iterations. Moreover, HOMF can then directly apply spectral clustering, to achieve good fitting performance. Extensive experimental results show that HOMF outperforms several state-of-the-art model fitting methods on both synthetic data and real images, especially in sampling efficiency and in handling data with severe outliers. Shuyuan Lin, Guobao Xiao, Yan Yan 0001, David Suter, Hanzi Wang |
AAAI | 4 |
| 2019 | Superpixel-Guided Two-View Deterministic Geometric Model Fitting
Guobao Xiao, Hanzi Wang, Yan Yan 0001, David Suter |
Int. J. Comput. Vis. | 4 |
| 2019 | Searching for Representative Modes on Hypergraphs for Robust Geometric Model FittingabstractIn this paper, we propose a simple and effective geometric model fitting method to fit and segment multi-structure data even in the presence of severe outliers. We cast the task of geometric model fitting as a representative mode-seeking problem on hypergraphs. Specifically, a hypergraph is first constructed, where the vertices represent model hypotheses and the hyperedges denote data points. The hypergraph involves higher-order similarities (instead of pairwise similarities used on a simple graph), and it can characterize complex relationships between model hypotheses and data points. In addition, we develop a hypergraph reduction technique to remove "insignificant" vertices while retaining as many "significant" vertices as possible in the hypergraph. Based on the simplified hypergraph, we then propose a novel mode-seeking algorithm to search for representative modes within reasonable time. Finally, the proposed mode-seeking algorithm detects modes according to two key elements, i.e., the weighting scores of vertices and the similarity analysis between vertices. Overall, the proposed fitting method is able to efficiently and effectively estimate the number and the parameters of model instances in the data simultaneously. Experimental results demonstrate that the proposed method achieves significant superiority over several state-of-the-art model fitting methods on both synthetic data and real images. Hanzi Wang, Guobao Xiao, Yan Yan 0001, David Suter |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Non-smooth M-estimator for Maximum Consensus Estimation
Huu Le, Anders P. Eriksson, Thanh-Toan Do, Tat-Jun Chin, David Suter |
BMVC | 5 |
| 2018 | Deterministic Consensus Maximization with Biconvex Programming
Zhipeng Cai 0003, Tat-Jun Chin, Huu Le, David Suter |
ECCV (12) | 4 |
| 2017 | An Exact Penalty Method for Locally Convergent Maximum ConsensusabstractMaximum consensus estimation plays a critically important role in computer vision. Currently, the most prevalent approach draws from the class of non-deterministic hypothesize-and-verify algorithms, which are cheap but do not guarantee solution quality. On the other extreme, there are global algorithms which are exhaustive search in nature and can be costly for practical-sized inputs. This paper aims to fill the gap between the two extremes by proposing a locally convergent maximum consensus algorithm. Our method is based on a formulating the problem with linear complementarity constraints, then defining a penalized version which is provably equivalent to the original problem. Based on the penalty problem, we develop a Frank-Wolfe algorithm that can deterministically solve the maximum consensus problem. Compared to the randomized techniques, our method is deterministic and locally convergent, relative to the global algorithms, our method is much more practical on realistic input sizes. Further, our approach is naturally applicable to problems with geometric residuals. Huu Le, Tat-Jun Chin, David Suter |
CVPR | 3 |
| 2017 | Quasiconvex Plane Sweep for Triangulation with OutliersabstractTriangulation is a fundamental task in 3D computer vision. Unsurprisingly, it is a well-investigated problem with many mature algorithms. However, algorithms for robust triangulation, which are necessary to produce correct results in the presence of egregiously incorrect measurements (i.e., outliers), have received much less attention. The default approach to deal with outliers in triangulation is by random sampling. The randomized heuristic is not only suboptimal, it could, in fact, be computationally inefficient on large-scale datasets. In this paper, we propose a novel locally optimal algorithm for robust triangulation. A key feature of our method is to efficiently derive the local update step by plane sweeping a set of quasiconvex functions. Underpinning our method is a new theory behind quasiconvex plane sweep, which has not been examined previously in computational geometry. Relative to the random sampling heuristic, our algorithm not only guarantees deterministic convergence to a local minimum, it typically achieves higher quality solutions in similar runtimes. Qianggong Zhang, Tat-Jun Chin, David Suter |
ICCV | 3 |
| 2017 | Efficient guided hypothesis generation for multi-structure epipolar geometry estimation
Taotao Lai, Hanzi Wang, Yan Yan 0001, Guobao Xiao, David Suter |
Comput. Vis. Image Underst. | 5 |
| 2017 | Efficient Globally Optimal Consensus Maximisation with Tree SearchabstractMaximum consensus is one of the most popular criteria for robust estimation in computer vision. Despite its widespread use, optimising the criterion is still customarily done by randomised sample-and-test techniques, which do not guarantee optimality of the result. Several globally optimal algorithms exist, but they are too slow to challenge the dominance of randomised methods. Our work aims to change this state of affairs by proposing an efficient algorithm for global maximisation of consensus. Under the framework of LP-type methods, we show how consensus maximisation for a wide variety of vision tasks can be posed as a tree search problem. This insight leads to a novel algorithm based on A* search. We propose efficient heuristic and support set updating routines that enable A* search to efficiently find globally optimal results. On common estimation problems, our algorithm is much faster than previous exact methods. Our work identifies a promising direction for globally optimal consensus maximisation. Tat-Jun Chin, Pulak Purkait, Anders P. Eriksson, David Suter |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Clustering with Hypergraphs: The Case for Large HyperedgesabstractThe extension of conventional clustering to hypergraph clustering, which involves higher order similarities instead of pairwise similarities, is increasingly gaining attention in computer vision. This is due to the fact that many clustering problems require an affinity measure that must involve a subset of data of size more than two. In the context of hypergraph clustering, the calculation of such higher order similarities on data subsets gives rise to hyperedges. Almost all previous work on hypergraph clustering in computer vision, however, has considered the smallest possible hyperedge size, due to a lack of study into the potential benefits of large hyperedges and effective algorithms to generate them. In this paper, we show that large hyperedges are better from both a theoretical and an empirical standpoint. We then propose a novel guided sampling strategy for large hyperedges, based on the concept of random cluster models. Our method can generate large pure hyperedges that significantly improve grouping accuracy without exponential increases in sampling costs. We demonstrate the efficacy of our technique on various higher-order grouping problems. In particular, we show that our approach improves the accuracy and efficiency of motion segmentation from dense, long-term, trajectories. Pulak Purkait, Tat-Jun Chin, Alireza Sadri, David Suter |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | Conformal Surface Alignment with Optimal Möbius SearchabstractDeformations of surfaces with the same intrinsic shape can often be described accurately by a conformal model. A major focus of computational conformal geometry is the estimation of the conformal mapping that aligns a given pair of object surfaces. The uniformization theorem enables this task to be acccomplished in a canonical 2D domain, wherein the surfaces can be aligned using a Möbius transformation. Current algorithms for estimating Möbius transformations, however, often cannot provide satisfactory alignment or are computationally too costly. This paper introduces a novel globally optimal algorithm for estimating Möbius transformations to align surfaces that are topological discs. Unlike previous methods, the proposed algorithm deterministically calculates the best transformation, without requiring good initializations. Further, our algorithm is also much faster than previous techniques in practice. We demonstrate the efficacy of our algorithm on data commonly used in computational conformal geometry. Huu Le, Tat-Jun Chin, David Suter |
CVPR | 3 |
| 2016 | Superpixel-Based Two-View Deterministic Fitting for Multiple-Structure Data
Guobao Xiao, Hanzi Wang, Yan Yan 0001, David Suter |
ECCV (6) | 4 |
| 2016 | Fast Rotation Search with Stereographic Projections for 3D RegistrationabstractRegistering two 3D point clouds involves estimating the rigid transform that brings the two point clouds into alignment. Recently there has been a surge of interest in using branch-and-bound (BnB) optimisation for point cloud registration. While BnB guarantees globally optimal solutions, it is usually too slow to be practical. A fundamental source of difficulty lies in the search for the rotational parameters. In this work, first by assuming that the translation is known, we focus on constructing a fast rotation search algorithm. With respect to an inherently robust geometric matching criterion, we propose a novel bounding function for BnB that is provably tighter than previously proposed bounds. Further, we also propose a fast algorithm to evaluate our bounding function. Our idea is based on using stereographic projections to precompute and index all possible point matches in spatial R-trees for rapid evaluations. The result is a fast and globally optimal rotation search algorithm. To conduct full 3D registration, we co-optimise the translation by embedding our rotation search kernel in a nested BnB algorithm. Since the inner rotation search is very efficient, the overall 6DOF optimisation is speeded up significantly without losing global optimality. On various challenging point clouds, including those taken out of lab settings, our approach demonstrates superior efficiency. Álvaro Parra Bustos, Tat-Jun Chin, Anders P. Eriksson, Hongdong Li, David Suter |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2016 | Robust Model Fitting Using Higher Than Minimal Subset SamplingabstractIdentifying the underlying model in a set of data contaminated by noise and outliers is a fundamental task in computer vision. The cost function associated with such tasks is often highly complex, hence in most cases only an approximate solution is obtained by evaluating the cost function on discrete locations in the parameter (hypothesis) space. To be successful at least one hypothesis has to be in the vicinity of the solution. Due to noise hypotheses generated by minimal subsets can be far from the underlying model, even when the samples are from the said structure. In this paper we investigate the feasibility of using higher than minimal subset sampling for hypothesis generation. Our empirical studies showed that increasing the sample size beyond minimal size ( p ), in particular up to p+2, will significantly increase the probability of generating a hypothesis closer to the true model when subsets are selected from inliers. On the other hand, the probability of selecting an all inlier sample rapidly decreases with the sample size, making direct extension of existing methods unfeasible. Hence, we propose a new computationally tractable method for robust model fitting that uses higher than minimal subsets. Here, one starts from an arbitrary hypothesis (which does not need to be in the vicinity of the solution) and moves until either a structure in data is found or the process is re-initialized. The method also has the ability to identify when the algorithm has reached a hypothesis with adequate accuracy and stops appropriately, thereby saving computational time. The experimental analysis carried out using synthetic and real data shows that the proposed method is both accurate and efficient compared to the state-of-the-art robust model fitting techniques. Ruwan B. Tennakoon, Alireza Bab-Hadiashar, Zhenwei Cao, Reza Hoseinnezhad, David Suter |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2016 | Hypergraph modelling for geometric model fitting
Guobao Xiao, Hanzi Wang, Taotao Lai, David Suter |
Pattern Recognit. | 4 |
| 2015 | Efficient globally optimal consensus maximisation with tree searchabstractMaximum consensus is one of the most popular criteria for robust estimation in computer vision. Despite its widespread use, optimising the criterion is still customarily done by randomised sample-and-test techniques, which do not guarantee optimality of the result. Several globally optimal algorithms exist, but they are too slow to challenge the dominance of randomised methods. We aim to change this state of affairs by proposing a very efficient algorithm for global maximisation of consensus. Under the framework of LP-type methods, we show how consensus maximisation for a wide variety of vision tasks can be posed as a tree search problem. This insight leads to a novel algorithm based on A* search. We propose efficient heuristic and support set updating routines that enable A* search to rapidly find globally optimal results. On common estimation problems, our algorithm is several orders of magnitude faster than previous exact methods. Our work identifies a promising solution for globally optimal consensus maximisation. Tat-Jun Chin, Pulak Purkait, Anders P. Eriksson, David Suter |
CVPR | 4 |
| 2015 | Mode-Seeking on Hypergraphs for Robust Geometric Model FittingabstractIn this paper, we propose a novel geometric model fitting method, called Mode-Seeking on Hypergraphs (MSH), to deal with multi-structure data even in the presence of severe outliers. The proposed method formulates geometric model fitting as a mode seeking problem on a hypergraph in which vertices represent model hypotheses and hyperedges denote data points. MSH intuitively detects model instances by a simple and effective mode seeking algorithm. In addition to the mode seeking algorithm, MSH includes a similarity measure between vertices on the hypergraph and a "weight-aware sampling" technique. The proposed method not only alleviates sensitivity to the data distribution, but also is scalable to large scale problems. Experimental results further demonstrate that the proposed method has significant superiority over the state-of-the-art fitting methods on both synthetic data and real images. Hanzi Wang, Guobao Xiao, Yan Yan 0001, David Suter |
ICCV | 4 |
| 2014 | Fast Rotation Search with Stereographic Projections for 3D RegistrationabstractRecently there has been a surge of interest to use branch-and-bound (bnb) optimisation for 3D point cloud registration. While bnb guarantees globally optimal solutions, it is usually too slow to be practical. A fundamental source of difficulty is the search for the rotation parameters in the 3D rigid transform. In this work, assuming that the translation parameters are known, we focus on constructing a fast rotation search algorithm. With respect to an inherently robust geometric matching criterion, we propose a novel bounding function for bnb that allows rapid evaluation. Underpinning our bounding function is the usage of stereographic projections to precompute and spatially index all possible point matches. This yields a robust and global algorithm that is significantly faster than previous methods. To conduct full 3D registration, the translation can be supplied by 3D feature matching, or by another optimisation framework that provides the translation. On various challenging point clouds, including those taken out of lab settings, our approach demonstrates superior efficiency. Álvaro Parra Bustos, Tat-Jun Chin, David Suter |
CVPR | 3 |
| 2014 | Fast Supervised Hashing with Decision Trees for High-Dimensional DataabstractSupervised hashing aims to map the original features to compact binary codes that are able to preserve label based similarity in the Hamming space. Non-linear hash functions have demonstrated their advantage over linear ones due to their powerful generalization capability. In the literature, kernel functions are typically used to achieve non-linearity in hashing, which achieve encouraging retrieval perfor- mance at the price of slow evaluation and training time. Here we propose to use boosted decision trees for achieving non-linearity in hashing, which are fast to train and evaluate, hence more suitable for hashing with high dimensional data. In our approach, we first propose sub-modular formulations for the hashing binary code inference problem and an efficient GraphCut based block search method for solving large-scale inference. Then we learn hash func- tions by training boosted decision trees to fit the binary codes. Experiments demonstrate that our proposed method significantly outperforms most state-of-the-art methods in retrieval precision and training time. Especially for high- dimensional data, our method is orders of magnitude faster than many methods in terms of training time. Guosheng Lin, Chunhua Shen, Qinfeng Shi, Anton van den Hengel, David Suter |
CVPR | 5 |
| 2014 | Clustering with Hypergraphs: The Case for Large Hyperedges
Pulak Purkait, Tat-Jun Chin, Hanno Ackermann, David Suter |
ECCV (4) | 4 |
| 2014 | Fast rotation search for real-time interactive point cloud registrationabstractOur goal is the registration of multiple 3D point clouds obtained from LIDAR scans of underground mines. Such a capability is crucial to the surveying and planning operations in mining. Often, the point clouds only partially overlap and initial alignment is unavailable. Here, we propose an interactive user-assisted point cloud registration system. Guided by the system, the user's role is simply to identify and search for overlapping regions across the point clouds. Specifically, given two point sets, the user clicks on a point in one set, then simply hovers the mouse on the other set to find a matching point. Each mouse position gives rise to a translation, and our system instantly optimises the rotation that aligns the point clouds. Tat-Jun Chin, Álvaro Parra Bustos, Michael S. Brown, David Suter |
I3D | 4 |
| 2014 | Sampling Minimal Subsets with Large Spans for Robust Estimation
Quoc-Huy Tran, Tat-Jun Chin, Wojciech Chojnacki, David Suter |
Int. J. Comput. Vis. | 4 |
| 2014 | The Random Cluster Model for Robust Geometric FittingabstractRandom hypothesis generation is central to robust geometric model fitting in computer vision. The predominant technique is to randomly sample minimal subsets of the data, and hypothesize the geometric models from the selected subsets. While taking minimal subsets increases the chance of successively "hitting" inliers in a sample, hypotheses fitted on minimal subsets may be severely biased due to the influence of measurement noise, even if the minimal subsets contain purely inliers. In this paper we propose Random Cluster Models, a technique used to simulate coupled spin systems, to conduct hypothesis generation using subsets larger than minimal. We show how large clusters of data from genuine instances of the model can be efficiently harvested to produce accurate hypotheses that are less affected by the vagaries of fitting on minimal subsets. A second aspect of the problem is the optimization of the set of structures that best fit the data. We show how our novel hypothesis sampler can be integrated seamlessly with graph cuts under a simple annealing framework to optimize the fitting efficiently. Unlike previous methods that conduct hypothesis sampling and fitting optimization in two disjoint stages, our algorithm performs the two subtasks alternatingly and in a mutually reinforcing manner. Experimental results show clear improvements in overall efficiency. Trung-Thanh Pham, Tat-Jun Chin, Jin Yu 0001, David Suter |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2014 | As-Projective-As-Possible Image Stitching with Moving DLTabstractThe success of commercial image stitching tools often leads to the impression that image stitching is a "solved problem". The reality, however, is that many tools give unconvincing results when the input photos violate fairly restrictive imaging assumptions; the main two being that the photos correspond to views that differ purely by rotation, or that the imaged scene is effectively planar. Such assumptions underpin the usage of 2D projective transforms or homographies to align photos. In the hands of the casual user, such conditions are often violated, yielding misalignment artifacts or "ghosting" in the results. Accordingly, many existing image stitching tools depend critically on post-processing routines to conceal ghosting. In this paper, we propose a novel estimation technique called Moving Direct Linear Transformation (Moving DLT) that is able to tweak or fine-tune the projective warp to accommodate the deviations of the input data from the idealized conditions. This produces as-projective-as-possible image alignment that significantly reduces ghosting without compromising the geometric realism of perspective image stitching. Our technique thus lessens the dependency on potentially expensive postprocessing algorithms. In addition, we describe how multiple as-projective-as-possible warps can be simultaneously refined via bundle adjustment to accurately align multiple images for large panorama creation. Julio Zaragoza, Tat-Jun Chin, Quoc-Huy Tran, Michael S. Brown, David Suter |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2014 | Multi-subregion based correlation filter bank for robust face recognition
Yan Yan 0001, Hanzi Wang, David Suter |
Pattern Recognit. | 3 |
| 2014 | Interacting Geometric Priors For Robust Multimodel FittingabstractRecent works on multimodel fitting are often formulated as an energy minimization task, where the energy function includes fitting error and regularization terms, such as low-level spatial smoothness and model complexity. In this paper, we introduce a novel energy with high-level geometric priors that consider interactions between geometric models, such that certain preferred model configurations may be induced.We argue that in many applications, such prior geometric properties are available and should be fruitfully exploited. For example, in surface fitting to point clouds, the building walls are usually either orthogonal or parallel to each other. Our proposed energy function is useful in dealing with unknown distributions of data errors and outliers, which are often the factors leading to biased estimation. Furthermore, the energy can be efficiently minimized using the expansion move method. We evaluate the performance on several vision applications using real data sets. Experimental results show that our method outperforms the state-of-the-art methods without significant increase in computation. Trung-Thanh Pham, Tat-Jun Chin, Konrad Schindler, David Suter |
IEEE Trans. Image Process. | 4 |
| 2013 | As-Projective-As-Possible Image Stitching with Moving DLTabstractWe investigate projective estimation under model inadequacies, i.e., when the underpinning assumptions of the projective model are not fully satisfied by the data. We focus on the task of image stitching which is customarily solved by estimating a projective warp - a model that is justified when the scene is planar or when the views differ purely by rotation. Such conditions are easily violated in practice, and this yields stitching results with ghosting artefacts that necessitate the usage of deghosting algorithms. To this end we propose as-projective-as-possible warps, i.e., warps that aim to be globally projective, yet allow local non-projective deviations to account for violations to the assumed imaging conditions. Based on a novel estimation technique called Moving Direct Linear Transformation (Moving DLT), our method seamlessly bridges image regions that are inconsistent with the projective model. The result is highly accurate image stitching, with significantly reduced ghosting effects, thus lowering the dependency on post hoc deghosting. Julio Zaragoza, Tat-Jun Chin, Michael S. Brown, David Suter |
CVPR | 4 |
| 2013 | Improved wireless tracking using radio frequency and video sensors
Thuraiappah Sathyan, Tat-Jun Chin, David Suter, Mark Hedley |
FUSION | 3 |
| 2013 | A General Two-Step Approach to Learning-Based HashingabstractMost existing approaches to hashing apply a single form of hash function, and an optimization process which is typically deeply coupled to this specific form. This tight coupling restricts the flexibility of the method to respond to the data, and can result in complex optimization problems that are difficult to solve. Here we propose a flexible yet simple framework that is able to accommodate different types of loss functions and hash functions. This framework allows a number of existing approaches to hashing to be placed in context, and simplifies the development of new problem-specific hashing methods. Our framework decomposes the hashing learning problem into two steps: hash bit learning and hash function learning based on the learned bits. The first step can typically be formulated as binary quadratic problems, and the second step can be accomplished by training standard binary classifiers. Both problems have been extensively studied in the literature. Our extensive experiments demonstrate that the proposed framework is effective, flexible and outperforms the state-of-the-art. Guosheng Lin, Chunhua Shen, David Suter, Anton van den Hengel |
ICCV | 3 |
| 2013 | A simultaneous sample-and-filter strategy for robust multi-structure model fitting
Hoi Sim Wong, Tat-Jun Chin, Jin Yu 0001, David Suter |
Comput. Vis. Image Underst. | 4 |
| 2013 | Mode seeking over permutations for rapid geometric model fitting
Hoi Sim Wong, Tat-Jun Chin, Jin Yu 0001, David Suter |
Pattern Recognit. | 4 |
| 2012 | Fast Training of Effective Multi-class Boosting Using Coordinate Descent Optimization
Guosheng Lin, Chunhua Shen, Anton van den Hengel, David Suter |
ACCV (2) | 4 |
| 2012 | The Random Cluster Model for robust geometric fittingabstractRandom hypothesis generation is central to robust geometric model fitting in computer vision. The predominant technique is to randomly sample minimal or elemental subsets of the data, and hypothesize the geometric model from the selected subsets. While taking minimal subsets increases the chance of simultaneously “hitting” inliers in a sample, it amplifies the noise of the underlying model, and hypotheses fitted on minimal subsets may be severely biased even if they contain purely inliers. In this paper we propose to use Random Cluster Models, a technique used to simulate coupled spin systems, to conduct hypothesis generation using subsets larger than minimal. We show how large clusters of data from genuine instances of the geometric model can be efficiently harvested to produce more accurate hypotheses. To take advantage of our hypothesis generator, we construct a simple annealing method based on graph cuts to fit multiple instances of the geometric model in the data. Experimental results show clear improvements in efficiency over other methods based on minimal subset samplers. Trung-Thanh Pham, Tat-Jun Chin, Jin Yu 0001, David Suter |
CVPR | 4 |
| 2012 | In Defence of RANSAC for Outlier Rejection in Deformable Registration
Quoc-Huy Tran, Tat-Jun Chin, Gustavo Carneiro 0001, Michael S. Brown, David Suter |
ECCV (4) | 5 |
| 2012 | Adaptive human silhouette reconstruction based on the exploration of temporal informationabstractHuman silhouette reconstruction has a wide range of applications in motion analysis, object segmentation and tracking, etc. In this paper, we propose a human silhouette reconstruction method based on the exploration of temporal information. Given a test silhouette, the proposed method aims to find its reliable templates for reconstruction by using the intrinsic temporal relationship among different frames. To effectively obtain such templates, we propose an adaptive criterion based on the non-negative least square optimization. Experimental results on two challenging datasets demonstrate the effectiveness of our method. Xi Li 0001, Tat-Jun Chin, David Suter |
ICASSP | 4 |
| 2012 | Superpixel-driven level set trackingabstractIn this paper, we propose a superpixel-driven method for level set tracking. In particular, by taking a superpixel-based speed function, the level set evolution is accelerated greatly. We define a mutual information based speed function using a superpixel-unit as the underlying representation, which captures the correlation of a superpixel with object/background. In order to enhance the robustness of our method, a shape prior is incorporated to constrain the contour evolution. Experimental results on a number of challenging sequences demonstrate the effectiveness and robustness of our method. Xi Li 0001, Tat-Jun Chin, David Suter |
ICIP | 4 |
| 2012 | Accelerated Hypothesis Generation for Multistructure Data via Preference AnalysisabstractRandom hypothesis generation is integral to many robust geometric model fitting techniques. Unfortunately, it is also computationally expensive, especially for higher order geometric models and heavily contaminated data. We propose a fundamentally new approach to accelerate hypothesis sampling by guiding it with information derived from residual sorting. We show that residual sorting innately encodes the probability of two points having arisen from the same model, and is obtained without recourse to domain knowledge (e.g., keypoint matching scores) typically used in previous sampling enhancement methods. More crucially, our approach encourages sampling within coherent structures and thus can very rapidly generate all-inlier minimal subsets that maximize the robust criterion. Sampling within coherent structures also affords a natural ability to handle multistructure data, a condition that is usually detrimental to other methods. The result is a sampling scheme that offers substantial speed-ups on common computer vision tasks such as homography and fundamental matrix estimation. We show on many computer vision data, especially those with multiple structures, that ours is the only method capable of retrieving satisfactory results within realistic time budgets. Tat-Jun Chin, Jin Yu 0001, David Suter |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Simultaneously Fitting and Segmenting Multiple-Structure Data with OutliersabstractWe propose a robust fitting framework, called Adaptive Kernel-Scale Weighted Hypotheses (AKSWH), to segment multiple-structure data even in the presence of a large number of outliers. Our framework contains a novel scale estimator called Iterative Kth Ordered Scale Estimator (IKOSE). IKOSE can accurately estimate the scale of inliers for heavily corrupted multiple-structure data and is of interest by itself since it can be used in other robust estimators. In addition to IKOSE, our framework includes several original elements based on the weighting, clustering, and fusing of hypotheses. AKSWH can provide accurate estimates of the number of model instances and the parameters and the scale of each model instance simultaneously. We demonstrate good performance in practical applications such as line fitting, circle fitting, range image segmentation, homography estimation, and two--view-based motion segmentation, using both synthetic data and real images. Hanzi Wang, Tat-Jun Chin, David Suter |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Visual tracking of numerous targets via multi-Bernoulli filtering of image data
Reza Hoseinnezhad, Ba-Ngu Vo, Ba-Tuong Vo, David Suter |
Pattern Recognit. | 4 |
| 2011 | A global optimization approach to robust multi-model fittingabstractWe present a novel Quadratic Program (QP) formulation for robust multi-model fitting of geometric structures in vision data. Our objective function enforces both the fidelity of a model to the data and the similarity between its associated inliers. Departing from most previous optimization-based approaches, the outcome of our method is a ranking of a given set of putative models, instead of a pre-specified number of “good” candidates (or an attempt to decide the right number of models). This is particularly useful when the number of structures in the data is a priori unascertainable due to unknown intent and purposes. Another key advantage of our approach is that it operates in a unified optimization framework, and the standard QP form of our problem formulation permits globally convergent optimization techniques. We tested our method on several geometric multi-model fitting problems on both synthetic and real data. Experiments show that our method consistently achieves state-of-the-art results. Jin Yu 0001, Tat-Jun Chin, David Suter |
CVPR | 3 |
| 2011 | Bayesian integration of audio and visual information for multi-target tracking using a CB-member filterabstractA new method is presented for integration of audio and visual information in multiple target tracking applications. The proposed approach uses a Bayesian filtering formulation and exploits multi-Bernoulli random finite set approximations. The work presented in this paper is the first principled Bayesian estimation approach to solve the sensor fusion problems that involve intermittent sensory data (e.g. audio data for a person who occasionally speaks.) We have examined our method with case studies from the SPEVI database. The results show nearly perfect tracking of people not only when they are silent but also when they are not visible to the camera (but speaking). Reza Hoseinnezhad, Ba-Ngu Vo, Ba-Tuong Vo, David Suter |
ICASSP | 4 |
| 2011 | Dynamic and hierarchical multi-structure geometric model fittingabstractThe ability to generate good model hypotheses is instrumental to accurate and robust geometric model fitting. We present a novel dynamic hypothesis generation algorithm for robust fitting of multiple structures. Underpinning our method is a fast guided sampling scheme enabled by analysing correlation of preferences induced by data and hypothesis residuals. Our method progressively accumulates evidence in the search space, and uses the information to dynamically (1) identify outliers, (2) filter unpromising hypotheses, and (3) bias the sampling for active discovery of multiple structures in the data-All achieved without sacrificing the speed associated with sampling-based methods. Our algorithm yields a disproportionately higher number of good hypotheses among the sampling outcomes, i.e., most hypotheses correspond to the genuine structures in the data. This directly supports a novel hierarchical model fitting algorithm that elicits the underlying stratified manner in which the structures are organized, allowing more meaningful results than traditional “flat” multi-structure fitting. Hoi Sim Wong, Tat-Jun Chin, Jin Yu 0001, David Suter |
ICCV | 4 |
| 2011 | An adversarial optimization approach to efficient outlier removalabstractThis paper proposes a novel adversarial optimization approach to efficient outlier removal in computer vision. We characterize the outlier removal problem as a game that involves two players of conflicting interests, namely, optimizer and outlier. Such an adversarial view not only brings new insights into various existing methods, but also gives rise to a general optimization framework that provably unifies them. Under the proposed framework, we develop a new outlier removal approach that is able to offer a much needed control over the trade-off between reliability and speed, which is otherwise not available in previous methods. The proposed approach is driven by a mixed-integer minmax (convex-concave) optimization process. Although a minmax problem is generally not amenable to efficient optimization, we show that for some commonly used vision objective functions, an equivalent Linear Program reformulation exists. We demonstrate our method on two representative multiview geometry problems. Experiments on real image data illustrate superior practical performance of our method over recent techniques. Jin Yu 0001, Anders P. Eriksson, Tat-Jun Chin, David Suter |
ICCV | 4 |
| 2011 | Simultaneous Sampling and Multi-Structure Fitting with Adaptive Reversible Jump MCMCabstractMulti-structure model fitting has traditionally taken a two-stage approach: First, sample a (large) number of model hypotheses, then select the subset of hypotheses that optimise a joint fitting and model selection criterion. This disjoint two-stage approach is arguably suboptimal and inefficient - if the random sampling did not retrieve a good set of hypotheses, the optimised outcome will not represent a good fit. To overcome this weakness we propose a new multi-structure fitting approach based on Reversible Jump MCMC. Instrumental in raising the effectiveness of our method is an adaptive hypothesis generator, whose proposal distribution is learned incrementally and online. We prove that this adaptive proposal satisfies the diminishing adaptation property crucial for ensuring ergodicity in MCMC. Our method effectively conducts hypothesis sampling and optimisation simultaneously, and gives superior computational efficiency over other methods. Trung-Thanh Pham, Tat-Jun Chin, Jin Yu 0001, David Suter |
NIPS | 4 |
| 2011 | Boosting histograms of descriptor distances for scalable multiclass specific scene recognition
Tat-Jun Chin, David Suter, Hanzi Wang |
Image Vis. Comput. | 2 |
| 2011 | Recognition of adult images, videos, and web page bagsabstractIn this article, we develop an integrated adult-content recognition system which can detect adult images, adult videos, and adult Web page bags, where a Web page bag consists of a Web page and a predefined number of Web pages linked to it through hyperlinks. In our adult image-recognition algorithm, we model skin patches rather than skin pixels, resulting in better results than state-of-the-art algorithms which model skin pixels. In our adult video-recognition algorithm, information from the accompanying audio section around an image in an adult video is used to obtain a prior classification of the image. The algorithm achieves a better performance than the ones which use image information alone or audio information alone. The adult Web page bag recognition is carried out using multi-instance learning based on the combination of classifying texts, images and videos in Web pages. Both the speed and the accuracy for recognizing the Web adult content are increased, in contrast to recognizing Web pages one-by-one. Weiming Hu 0004, Haiqiang Zuo, Ou Wu 0001, Yunfei Chen 0002, Zhongfei Zhang, David Suter |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2010 | Efficient Multi-structure Robust Fitting with Incremental Top-k Lists Comparison
Hoi Sim Wong, Tat-Jun Chin, Jin Yu 0001, David Suter |
ACCV (4) | 4 |
| 2010 | Multi-structure model selection via kernel optimisationabstractOur goal is to fit the multiple instances (or structures) of a generic model existing in data. Here we propose a novel model selection scheme to estimate the number of genuine structures present. In contrast to conventional model selection approaches, our method is driven by kernel-based learning. The input data is first clustered based on their potential to have emerged from the same structure. However the number of clusters is deliberately overestimated to obtain a set of initial model fits onto the data. We then resolve the oversegmentation via a series of kernel optimisation conducted through multiple kernel learning, and the concept of kernel-target alignment is used as a model selection criterion. Experiments on synthetic and real data show that our method outperforms previous model selection schemes. We also focus on the application of multi-body motion segmentation. In particular we demonstrate success on estimating the number of motions on sequences with more than 3 unique motions. Tat-Jun Chin, David Suter, Hanzi Wang |
CVPR | 2 |
| 2010 | Accelerated Hypothesis Generation for Multi-structure Robust Fitting
Tat-Jun Chin, Jin Yu 0001, David Suter |
ECCV (5) | 3 |
| 2010 | Multi-object filtering from image sequence without detectionabstractAlmost every single-view visual multi-target tracking method presented in the literature includes a detection routine that maps the image data to point measurements relevant to the target states. These measurements are commonly further processed by a filter to estimate the number of targets and their states. This paper presents a novel visual tracking technique based on a multi-object filtering algorithm that operates directly on the image observations without the need for any detection. Experimental results on tracking sport players show that our proposed method can automatically track numerous interacting targets and quickly finds players entering or leaving the scene. Reza Hoseinnezhad, Ba-Ngu Vo, David Suter, Ba-Tuong Vo |
ICASSP | 3 |
| 2010 | Visual localization and segmentation based on foreground/background modelingabstractIn this paper, we propose a novel method to localize (or track) a foreground object and segment the foreground object from the surrounding background with occlusions for a moving camera. We measure the likelihood of a target position by using a combination of a generative model and a discriminative model, considering not only the foreground similarity to the target model but also the dissimilarity between the foreground and the background appearances. Object segmentation is treated as a binary labeling problem. A Markov Random Field (MRF) is employed to add a spatial smooth prior on the foreground/background patterns. We demonstrate the advantages of the proposed method on several challenging videos and compare our results with the results of several other popular methods. The proposed method has achieved good results. Hanzi Wang, Tat-Jun Chin, David Suter |
ICASSP | 3 |
| 2010 | BoostML: An Adaptive Metric Learning for Nearest Neighbor Classification
Nayyar Abbas Zaidi, David McG. Squire, David Suter |
PAKDD (1) | 3 |
| 2009 | Keypoint induced distance profiles for visual recognitionabstractWe show that histograms of keypoint descriptor distances can make useful features for visual recognition. Descriptor distances are often exhaustively computed between sets of keypoints, but besides finding the k-smallest distances the structure of the distribution of these distances has been largely overlooked. We highlight the potential of such information in the task of particular scene recognition. Discriminative scene signatures in the form of histograms of keypoint descriptor distances are constructed in a supervised manner. The distances are computed between properly selected reference keypoints and the keypoints detected in the input image. The signature is low dimensional, computationally cheap to obtain, and can distinguish a large number of scenes. We introduce a scheme based on multiclass AdaBoost to select the appropriate reference keypoints. The resulting system is capable of handling a large number of scene classes at a fraction of the time required for exhaustively matching sets of keypoints. This supports supports a coarse-to-fine search strategy for approaches reliant on keypoint matching. We test the idea on 3 datasets for particular scene recognition and report the obtained results. Tat-Jun Chin, David Suter |
CVPR | 2 |
| 2009 | Bayesian multi-object estimation from image observations
Ba-Ngu Vo, Ba-Tuong Vo, Nam-Trung Pham, David Suter |
FUSION | 4 |
| 2009 | Robust fitting of multiple structures: The statistical learning approachabstractWe propose an unconventional but highly effective approach to robust fitting of multiple structures by using statistical learning concepts. We design a novel Mercer kernel for the robust estimation problem which elicits the potential of two points to have emerged from the same underlying structure. The Mercer kernel permits the application of well-grounded statistical learning methods, among which nonlinear dimensionality reduction, principal component analysis and spectral clustering are applied for robust fitting. Our method can remove gross outliers and in parallel discover the multiple structures present. It functions well under severe outliers (more than 90% of the data) and considerable inlier noise without requiring elaborate manual tuning or unrealistic prior information. Experiments on synthetic and real problems illustrate the superiority of the proposed idea over previous methods. Tat-Jun Chin, Hanzi Wang, David Suter |
ICCV | 3 |
| 2009 | The Ordered Residual Kernel for Robust Motion Subspace ClusteringabstractWe present a novel and highly effective approach for multi-body motion segmentation. Drawing inspiration from robust statistical model fitting, we estimate putative subspace hypotheses from the data. However, instead of ranking them we encapsulate the hypotheses in a novel Mercer kernel which elicits the potential of two point trajectories to have emerged from the same subspace. The kernel permits the application of well-established statistical learning methods for effective outlier rejection, automatic recovery of the number of motions and accurate segmentation of the point trajectories. The method operates well under severe outliers arising from spurious trajectories or mistracks. Detailed experiments on a recent benchmark dataset (Hopkins 155) show that our method is superior to other state-of-the-art approaches in terms of recovering the number of motions, segmentation accuracy, robustness against gross outliers and computational efficiency. Tat-Jun Chin, Hanzi Wang, David Suter |
NIPS | 3 |
| 2009 | 3D terrestrial LIDAR classifications with super-voxels and multi-scale Conditional Random Fields
Ee Hui Lim, David Suter |
Comput. Aided Des. | 2 |
| 2009 | Rank Constraints for Homographies over Two Views: Revisiting the Rank Four Constraint
Pei Chen 0001, David Suter |
Int. J. Comput. Vis. | 2 |
| 2009 | Human action recognition by feature-reduced Gaussian process classification
Hang Zhou 0005, Liang Wang 0001, David Suter |
Pattern Recognit. Lett. | 3 |
| 2009 | Simultaneously Estimating the Fundamental Matrix and HomographiesabstractThe estimation of the fundamental matrix (FM) and/or one or more homographies between two views is of great interest for a number of computer vision and robotics tasks. We consider the joint estimation of the FM and one or more homographies. Given point matches between two views (and assuming rigid geometry of the camera-scene displacement), it is well known that all of the matched points satisfy the epipolar constraint that is usually characterized by the FM. Subsets of these point matches may also obey a constraint characterized by a homography (all matches in the subset coming from three-dimensional (3-D) points lying on a 3-D plane). The estimations of homographies and the FM are well-studied problems, and therefore, the (separate) estimation of the FM, or the homography matrices, can be considered as effectively solved problems with mature algorithms. However, the homographies and FM are not independent of each other: therefore, separate estimation of each is likely to be suboptimal. In this paper, we propose to simultaneously estimate the FM and homographies by employing the compatibility constraint between them. This is done by first concentrating on a set of parameters that (jointly) parameterize the entire set of homographies and FM (simultaneously) and that also implicitly enforce the compatibility between the estimates of each set. We then derive a reduced form with the purpose of improving the speed. We propose a solution method in which the Sampson error for the FM and homographies is minimized by the Levenberg-Marquardt (LM) algorithm. Experiments show that the gains can be compared with separate estimates (the FM and/or the homographies). Pei Chen 0001, David Suter |
IEEE Trans. Robotics | 2 |
| 2008 | Improved building detection by Gaussian processes classification via feature space rescale and spectral kernel selectionabstractWe use spectral analysis to facilitate Gaussian processes (GP) classification. Our solution provides two improvements: scaling of the data to achieve a more isotropic nature, as well as a method to choose the kernel to match certain data characteristics. Given the dataset, from the Fourier transform of the training data we compare the frequency domain features of each dimension to estimate a rescaling (towards making the data isotropic). Also, the spectrum of the training data is compared with several candidate kernel spectrums. From this comparison the best matching kernel is chosen. In these ways, the training data matches better the GP classification kernel function (and hence the underlying assumed correlation characteristics), resulting in a better GP classification result. Test results on both non image and image data show the efficiency and effectiveness of our approach. Hang Zhou 0005, David Suter |
CVPR | 2 |
| 2008 | Manifold optimisation for motion factorisationabstractThis paper presents a novel formulation for the popular factorisation based solution for Structure from Motion. Since our measurement matrices are populated with incomplete and inaccurate data, SVD based total least squares solution are less than appropriate. Instead, we approach the problem as a non-linear unconstrained minimisation problem on the product manifold of the Special Euclidean Group (SE3). The restriction of the domain of optimisation to the SE3product manifold not only implies that each intermediate solution is a plausible object motion, but also ensures better intrinsic stability for the minimisation algorithm. We compare our method with existing state of art, and show that our algorithm exhibits superior performance. Appu Shaji, Sharat Chandran, David Suter |
ICPR | 3 |
| 2008 | Confidence rated boosting algorithm for generic object detectionabstractIn this paper we propose a confidence rated boosting algorithm based on Ada-boost for generic object detection. Confidence rated Ada-boost algorithm has not been applied to generic object detection problem; in that sense our work is novel. We represent images as bag of words, where the words are SIFT descriptors extracted over some interest points. We compare our boosting algorithm to another version of boosting algorithm called Gentle-boost. Our approach generalizes well and performs equal or better than Gentle-boost. We show our results on four categories from the Caltech data sets, in terms of ROC curves. Nayyar Abbas Zaidi, David Suter |
ICPR | 2 |
| 2008 | Improving Gaussian processes classification by spectral data reorganizingabstractWe improve Gaussian processes (GP) classification by reorganizing the (non-stationary and anisotropic) data to better fit to the isotropic GP kernel. First, the data is partitioned into two parts: along the feature with the highest frequency bandwidth. Secondly, for each part of the data, only the spectrally homogeneous features are chosen and used (the rest discarded) for GP classification. In this way, anisotropy of the data is lessened from the frequency point of view. Tests on synthetic data as well as real datasets show that our approach is effective and outperforms Automatic Relevance Determination (ARD). Hang Zhou 0005, David Suter |
ICPR | 2 |
| 2008 | Human motion recognition using Gaussian Processes classificationabstractThis paper investigates the applicability of Gaussian Processes (GP) classification for recognition of articulated and deformable human motions from image sequences. Using Tensor Subspace Analysis (TSA), space-time human silhouettes (extracted from motion videos) are transformed to low-dimensional multivariate time series, based on which structure-based statistical features are calculated to summarize the motion properties. GP classification is then used to learn and predict motion categories. Experimental results on two real-world state-of-the-art datasets show that the proposed approach is effective, and outperforms Support Vector Machine (SVM). Hang Zhou 0005, Liang Wang 0001, David Suter |
ICPR | 3 |
| 2008 | Visual learning and recognition of sequential data manifolds with applications to human movement analysis
Liang Wang 0001, David Suter |
Comput. Vis. Image Underst. | 2 |
| 2008 | A Model-Selection Framework for Multibody Structure-and-Motion of Image Sequences
Konrad Schindler, David Suter, Hanzi Wang |
Int. J. Comput. Vis. | 2 |
| 2008 | Out-of-Sample Extrapolation of Learned ManifoldsabstractWe investigate the problem of extrapolating the embedding of a manifold learned from finite samples to novel out-of-sample data. We concentrate on the manifold learning method called Maximum Variance Unfolding (MVU) for which the extrapolation problem is still largely unsolved. Taking the perspective of MVU learning being equivalent to Kernel PCA, our problem reduces to extending a kernel matrix generated from an unknown kernel function to novel points. Leveraging on previous developments, we propose a novel solution which involves approximating the kernel eigenfunction using Gaussian basis functions. We also show how the width of the Gaussian can be tuned to achieve extrapolation. Experimental results which demonstrate the effectiveness of the proposed approach are also included. Tat-Jun Chin, David Suter |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Object detection by global contour shape
Konrad Schindler, David Suter |
Pattern Recognit. | 2 |
| 2007 | Human Pose Extraction from Monocular Videos using Constrained Non-Rigid FactorizationabstractWe focus on the problem of automatically extracting the 3D configuration of human poses from 2D image features tracked over a finite interval of time . This problem is highly non-linear in nature and confounds standard regression techniques. Our approach effectively marries a non-rigid factorization algorithm with prior learned statistical models from archival motion capture database. We show that a stand alone non-rigid factorization algorithm is highly unsuitable for this problem. However, when coupled with the learned statistical model in the form of a constrained non- linear programming method, it yields a substantially better solution. Appu Shaji, Behjat Siddiquie, Sharat Chandran, David Suter |
BMVC | 4 |
| 2007 | Recognizing Human Activities from Silhouettes: Motion Subspace and Factorial Discriminative Graphical ModelabstractWe describe a probabilistic framework for recognizing human activities in monocular video based on simple silhouette observations in this paper. The methodology combines kernel principal component analysis (KPCA) based feature extraction and factorial conditional random field (FCRF) based motion modeling. Silhouette data is represented more compactly by nonlinear dimensionality reduction that explores the underlying structure of the articulated action space and preserves explicit temporal orders in projection trajectories of motions. FCRF models temporal sequences in multiple interacting ways, thus increasing joint accuracy by information sharing, with the ideal advantages of discriminative models over generative ones (e.g., relaxing independence assumption between observations and the ability to effectively incorporate both overlapping features and long-range dependencies). The experimental results on two recent datasets have shown that the proposed framework can not only accurately recognize human activities with temporal, intra-and inter-person variations, but also is considerably robust to noise and other factors such as partial occlusion and irregularities in motion styles. Liang Wang 0001, David Suter |
CVPR | 2 |
| 2007 | Fast Sparse Gaussian Processes Learning for Man-Made Structure ClassificationabstractInformative Vector Machine (IVM) is an efficient fast sparse Gaussian process's (GP) method previously suggested for active learning. It greatly reduces the computational cost of GP classification and makes the GP learning close to real time. We apply IVM for man-made structure classification (a two class problem). Our work includes the investigation of the performance of IVM with varied active data points as well as the effects of different choices of GP kernels. Satisfactory results have been obtained, showing that the approach keeps full GP classification performance and yet is significantly faster (by virtue if using a subset of the whole training data points). Hang Zhou 0005, David Suter |
CVPR | 2 |
| 2007 | Conditional Random Field for 3D Point Clouds with Adaptive Data ReductionabstractWe proposed using Conditional Random Fields with adaptive data reduction for the classification of 3D point clouds acquired from a Riegl Terrestrial laser scanner. The training and inference of the acquired large outdoor urban data can be time consuming. We approach the problem by computing an adaptive support region for each data point using 3D scale theory. For training and inference of the discriminative Conditional Random Fields, smaller set of data samples that contains relevant information within the support region is selected instead of using all point cloud data. We tested the algorithm on synthetically generated data and urban point clouds data acquired from the laser scanner. The computed support region is also used in feature extraction for urban point clouds data. The results showed improvement in the training and inference rate while maintaining comparable classification accuracy. Ee Hui Lim, David Suter |
CW | 2 |
| 2007 | Extrapolating Learned Manifolds for Human Activity RecognitionabstractThe problem of human activity recognition via visual stimuli can be approached using manifold learning, since the silhouette (binary) images of a person undergoing a smooth motion can be represented as a manifold in the image space. While manifold learning methods allow the characterization of the activity manifolds, performing activity recognition requires distinguishing between manifolds. This invariably involves the extrapolation of learned activity manifolds to new silhouettes -a task that is not fully addressed in the literature. This paper investigates and compares methods for the extrapolation of learned manifolds within the context of activity recognition. Also, the problem of obtaining dense samples for learning human silhouette manifolds is addressed. Tat-Jun Chin, Liang Wang 0001, Konrad Schindler, David Suter |
ICIP (1) | 4 |
| 2007 | Man-Made Structure Segmentation using Gaussian Processes and Wavelet FeaturesabstractWe apply Gaussian process classification (GPC) to man-made structure segmentation, treated as a two class problem. GPC is a discriminative approach, and thus focuses on modelling the posterior directly. It relaxes the strong assumption of conditional independence of the observed data (generally used in a generative model). In addition, wavelet transform features, which are effective in describing directional textures, are incorporated in the feature vector. Satisfactory results have been obtained which show the effectiveness of our approach. Hang Zhou 0005, David Suter |
ICIP (4) | 2 |
| 2007 | Adaptive Object Tracking Based on an Effective Appearance FilterabstractWe propose a similarity measure based on a Spatial-color Mixture of Gaussians (SMOG) appearance model for particle filters. This improves on the popular similarity measure based on color histograms because it considers not only the colors in a region but also the spatial layout of the colors. Hence, the SMOG-based similarity measure is more discriminative. To efficiently compute the parameters for SMOG, we propose a new technique, with which the computational time is greatly reduced. We also extend our method by integrating multiple cues to increase the reliability and robustness. Experiments show that our method can successfully track objects in many difficult situations. Hanzi Wang, David Suter, Konrad Schindler, Chunhua Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | A consensus-based method for tracking: Modelling background scenario and foreground appearance
Hanzi Wang, David Suter |
Pattern Recognit. | 2 |
| 2007 | Incremental Kernel Principal Component AnalysisabstractThe kernel principal component analysis (KPCA) has been applied in numerous image-related machine learning applications and it has exhibited superior performance over previous approaches, such as PCA. However, the standard implementation of KPCA scales badly with the problem size, making computations for large problems infeasible. Also, the "batch" nature of the standard KPCA computation method does not allow for applications that require online processing. This has somewhat restricted the domains in which KPCA can potentially be applied. This paper introduces an incremental computation algorithm for KPCA to address these two problems. The basis of the proposed solution lies in computing incremental linear PCA in the kernel induced feature space, and constructing reduced-set expansions to maintain constant update speed and memory usage. We also provide experimental results which demonstrate the effectiveness of the approach. Tat-Jun Chin, David Suter |
IEEE Trans. Image Process. | 2 |
| 2007 | Learning and Matching of Dynamic Shape Manifolds for Human Action RecognitionabstractIn this paper, we learn explicit representations for dynamic shape manifolds of moving humans for the task of action recognition. We exploit locality preserving projections (LPP) for dimensionality reduction, leading to a low-dimensional embedding of human movements. Given a sequence of moving silhouettes associated to an action video, by LPP, we project them into a low-dimensional space to characterize the spatiotemporal property of the action, as well as to preserve much of the geometric structure. To match the embedded action trajectories, the median Hausdorff distance or normalized spatiotemporal correlation is used for similarity measures. Action classification is then achieved in a nearest-neighbor framework. To evaluate the proposed method, extensive experiments have been carried out on a recent dataset including ten actions performed by nine different subjects. The experimental results show that the proposed method is able to not only recognize human actions effectively, but also considerably tolerate some challenging conditions, e.g., partial occlusion, low-quality videos, changes in viewpoints, scales, and clothes; within-class variations caused by different subjects with different physical build; styles of motion; etc. Liang Wang 0001, David Suter |
IEEE Trans. Image Process. | 2 |
| 2006 | A New Distance Criterion for Face Recognition Using Image Sets
Tat-Jun Chin, David Suter |
ACCV (1) | 2 |
| 2006 | Feature Detection with an Improved Anisotropic Filter
Mohamed Gobara, David Suter |
ACCV (2) | 2 |
| 2006 | A Novel Robust Statistical Method for Background Initialization and Visual Surveillance
Hanzi Wang, David Suter |
ACCV (1) | 2 |
| 2006 | Improving the Speed of Kernel PCA on Large Scale DatasetsabstractThis paper concerns making large scale Kernel Principal Component Analysis (KPCA) feasible on regular hardware. The KPCA has been proven a useful non-linear feature extractor in several computer vision applications. The standard computation method for KPCA, however, scales badly with the problem size, thus limiting the potential of the technique for large scale data. We propose a novel method to alleviate this problem. The essence of our solution lies in partitioning the data and greedily filtering each partition for a sparse representation. Incremental KPCA is then utilized to merge each partition to arrive at the overall KPCA. We also provide experimental results which demonstrate the effectiveness of the approach. Tat-Jun Chin, David Suter |
AVSS | 2 |
| 2006 | Analyzing Human Movements from Silhouettes Using Manifold LearningabstractA novel method for learning and recognizing sequential image data is proposed, and promising applications to vision-based human movement analysis are demonstrated. To find more compact representations of high-dimensional silhouette data, we exploit locality preserving projections (LPP) to achieve low-dimensional manifold embedding. Further, we present two kinds of methods to analyze and recognize learned motion manifolds. One is correlation matching based on the Hausdorrf distance, and the other is a probabilistic method using continuous hidden Markov models (HMM). Encouraging results are obtained in two representative experiments in the areas of human activity recognition and gait-based human identification. Liang Wang 0001, David Suter |
AVSS | 2 |
| 2006 | A Compact Architecture for Wireless Video Surveillance over CDMA NetworkabstractWireless video surveillance over CDMA network enables transfer of high quality live video via public wireless broadband, which is accessible anywhere with CDMA mobile phone network coverage. Real time video can be received on an internet connected PC, laptop or even mobile phone. This makes possible surveillance from moving vehicles, or quick deployment in almost unrestricted locations. While many general purpose solutions exist for sending video over CDMA, they have not taken into consideration certain specific features of the network nor the requirements of the related applications. This paper presents a new architecture for video transmission over CDMA with simple hardware complexity, high reliability and rich functions. The compact and flexible system has been under trial in China, Australia, Singapore, Italy, Egypt and Brazil, amongst others. Hang Zhou 0005, David Suter |
AVSS | 2 |
| 2006 | Incremental Kernel PCA for Efficient Non-linear Feature ExtractionabstractThe Kernel Principal Component Analysis (KPCA) has been effectively applied as an unsupervised non-linear feature extractor in many machine learning applications. However, with a time complexity of O(n3), the practicality of KPCA on large datasets is minimal. In this paper, we propose an approximate incremental KPCA algorithm which allows efficient processing of large datasets. We extend a linear PCA updating algorithm to the non-linear case by utilizing the kernel trick, and apply a reduced set construction method to compress expressions for the derived KPCA basis at each update. In addition, we show how multiple feature space vectors can be compressed efficiently, and how approximated KPCA bases can be re-orthogonalized using the kernel trick. The proposed method is justified through experimental validations. Tat-Jun Chin, David Suter |
BMVC | 2 |
| 2006 | 3D Object Pose Inference via Kernel Principal Component Analysis with Image Euclidian Distance (IMED)abstractKernel Principal Component Analysis (KPCA) is a powerful non-linear unsupervised learning technique for high dimensional pattern analysis. KPCA on images, however, usually considers each image pixel as an independent dimension and does not take into account the spatial relationship of nearby pixels. In this paper, we show how the Image Euclidian Distance (IMED), which takes into account local pixel intensities, can efficiently be embedded into KPCA via the Kronecker product and Eigenvector projections, whilst still retaining desirable properties of Euclidian distance (such as kernel positive definitiveness and effective image de-noising). We demonstrate that KPCA with embedded IMED is a more intuitive and accurate technique than standard KPCA through a 3D object pose estimation application. 1 Therdsak Tangkuampien, David Suter |
BMVC | 2 |
| 2006 | Real-Time Human Pose Inference using Kernel Principal Component Pre-image ApproximationsabstractWe present a real-time markerless human motion capture technique based on un-calibrated synchronized cameras. Training sets of real motions captured from marker based systems are used to learn an optimal pose manifold of human motion via Kernel Principal Component Analysis (KPCA). Similarly, a synthetic silhouette manifold is also learnt, and markerless motion capture can then be viewed as the problem of mapping from the silhouette manifold to the pose manifold. After training, novel silhouettes of previously unseen actors are projected through the two manifolds using Locally Linear Embedding (LLE) reconstruction. The output pose is generated by approximating the pre-image (inverse mapping) of the LLE reconstructed vector from the pose manifold. 1 Therdsak Tangkuampien, David Suter |
BMVC | 2 |
| 2006 | Effective Appearance Model and Similarity Measure for Particle Filtering and Visual Tracking
Hanzi Wang, David Suter, Konrad Schindler |
ECCV (3) | 2 |
| 2006 | Parametric model-based motion segmentation using surface selection criterion
Niloofar Gheissari, Alireza Bab-Hadiashar, David Suter |
Comput. Vis. Image Underst. | 3 |
| 2006 | An Analysis of Linear Subspace Approaches for Computer Vision and Pattern Recognition
Pei Chen 0001, David Suter |
Int. J. Comput. Vis. | 2 |
| 2006 | Two-View Multibody Structure-and-Motion with Outliers through Model SelectionabstractMultibody structure-and-motion (MSaM) is the problem to establish the multiple-view geometry of several views of a 3D scene taken at different times, where the scene consists of multiple rigid objects moving relative to each other. We examine the case of two views. The setting is the following: Given are a set of corresponding image points in two images, which originate from an unknown number of moving scene objects, each giving rise to a motion model. Furthermore, the measurement noise is unknown, and there are a number of gross errors, which are outliers to all models. The task is to find an optimal set of motion models for the measurements. It is solved through Monte-Carlo sampling, careful statistical analysis of the sampled set of motion models, and simultaneous selection of multiple motion models to best explain the measurements. The framework is not restricted to any particular model selection mechanism because it is developed from a Bayesian viewpoint: Different model selection criteria are seen as different priors for the set of moving objects, which allow one to bias the selection procedure for different purposes. Konrad Schindler, David Suter |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Two-View Multibody Structure-and-Motion with OutliersabstractMulti-body structure-and-motion (MSaM) is the problem to establish the multiple-view geometry of several views of a 3D scene taken at different times, where the scene consists of multiple rigid objects moving relative to each other. We examine the case of two views. The setting is the following: given are a set of corresponding image points in two images, which originate from an unknown number of moving scene objects, each giving rise to a motion model. Furthermore, the measurement noise is unknown, and there are a number of gross errors, which are outliers to all models. The task to find an optimal set of motion models for the measurements is solved through Monte-Carlo sampling, careful statistical analysis of the data and simultaneous selection of multiple motion models. Konrad Schindler, David Suter |
CVPR (2) | 2 |
| 2005 | A re-evaluation of mixture of Gaussian background modeling [video signal processing applications]abstractThe mixture of Gaussians (MOG) has been widely used for robustly modeling complicated backgrounds, especially those with small repetitive movements (such as leaves, bushes, rotating fan, ocean waves, rain). The performance of MOG can be greatly improved by tackling several practical issues. In this paper, we quantitatively evaluate (using the Wallflower benchmarks) the performance of the MOG with and without our modifications. The experimental results show that the MOG, with our modifications, can achieve much better results - even outperforming other state-of-the-art methods. Hanzi Wang, David Suter |
ICASSP (2) | 2 |
| 2005 | Tracking and segmenting people with occlusions by a sample consensus based methodabstractOne of the most difficult issues in visual tracking is to track people in groups, especially under occlusions. In this paper, we present a novel sample consensus based method, which utilizes both color and spatial information of human bodies, to model the appearance of people. We use this appearance model to segment and track people through occlusions. We show experimental results in several video sequences to validate the effectiveness of the proposed method. Hanzi Wang, David Suter |
ICIP (2) | 2 |
| 2005 | Subspace-based face recognition: outlier detection and a new distance criterionabstractIllumination effects, including shadows and varying lighting, make the problem of face recognition challenging. Experimental and theoretical results show that the face images under different illumination conditions approximately lie in a low-dimensional subspace, hence principal component analysis (PCA) or low-dimensional subspace techniques have been used. Following this spirit, we propose new techniques for the face recognition problem, including an outlier detection strategy (mainly for those points not following the Lambertian reflectance model), and a new error criterion for the recognition algorithm. Experiments using the Yale-B face database show the effectiveness of the new strategies. Pei Chen 0001, David Suter |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2005 | Object tracking in image sequences using point features
Prithiraj Tissainayagam, David Suter |
Pattern Recognit. | 2 |
| 2004 | Robust Fitting by Adaptive-Scale Residual Consensus
Hanzi Wang, David Suter |
ECCV (3) | 2 |
| 2004 | Shift-invariant wavelet denoising using interscale dependencyabstractUsing statistical modeling in the wavelet domain, we address the problem of image denoising. Despite being effective, the denoised images can suffer from the Gibbs-like artifacts, like ringing around the edges and speckles in the smooth regions. We employ shift-invariant (SI) wavelet denoising in order to reduce these unpleasant artifacts. Not only is the visual quality greatly improved but also a PSNR gain of about 0.7/spl sim/0.9 dB is obtained. The proposed approach, siPAB, outperforms siHMT, which is a competitive SI wavelet denoising approach, by 0.1/spl sim/0.5 dB. Pei Chen 0001, David Suter |
ICIP | 2 |
| 2004 | MDPE: A Very Robust Estimator for Model Fitting and Range Image Segmentation
Hanzi Wang, David Suter |
Int. J. Comput. Vis. | 2 |
| 2004 | Special issue on statistical methods in video processing
David Suter, Dorin Comaniciu, Kenichi Kanatani |
Image Vis. Comput. | 1 |
| 2004 | Assessing the performance of corner detectors for point feature tracking applications
Prithiraj Tissainayagam, David Suter |
Image Vis. Comput. | 2 |
| 2004 | Recovering the Missing Components in a Large Noisy Low-Rank Matrix: Application to SFMabstractIn computer vision, it is common to require operations on matrices with "missing data," for example, because of occlusion or tracking failures in the Structure from Motion (SFM) problem. Such a problem can be tackled, allowing the recovery of the missing values, if the matrix should be of low rank (when noise free). The filling in of missing values is known as imputation. Imputation can also be applied in the various subspace techniques for face and shape classification, online "recommender" systems, and a wide variety of other applications. However, iterative imputation can lead to the "recovery" of data that is seriously in error. In this paper, we provide a method to recover the most reliable imputation, in terms of deciding when the inclusion of extra rows or columns, containing significant numbers of missing entries, is likely to lead to poor recovery of the missing parts. Although the proposed approach can be equally applied to a wide range of imputation methods, this paper addresses only the SFM problem. The performance of the proposed method is compared with Jacobs' and Shum's methods for SFM. Pei Chen 0001, David Suter |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | Robust Adaptive-Scale Parametric Model Estimation for Computer VisionabstractRobust model fitting essentially requires the application of two estimators. The first is an estimator for the values of the model parameters. The second is an estimator for the scale of the noise in the (inlier) data. Indeed, we propose two novel robust techniques: the Two-Step Scale estimator (TSSE) and the Adaptive Scale Sample Consensus (ASSC) estimator. TSSE applies nonparametric density estimation and density gradient estimation techniques, to robustly estimate the scale of the inliers. The ASSC estimator combines Random Sample Consensus (RANSAC) and TSSE: using a modified objective function that depends upon both the number of inliers and the corresponding scale. ASSC is very robust to discontinuous signals and data with multiple structures, being able to tolerate more than 80 percent outliers. The main advantage of ASSC over RANSAC is that prior knowledge about the scale of inliers is not needed. ASSC can simultaneously estimate the parameters of a model and the scale of the inliers belonging to that model. Experiments on synthetic data show that ASSC has better robustness to heavily corrupted data than Least Median Squares (LMedS), Residual Consensus (RESC), and Adaptive Least Kth order Squares (ALKS). We also apply ASSC to two fundamental computer vision tasks: range image segmentation and robust fundamental matrix estimation. Experiments show very promising results. Hanzi Wang, David Suter |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | Variable Bandwidth QMDPE and Its Application in Robust Optical Flow EstimationabstractRobust estimators, such as least median of squared (LMedS) residuals, M-estimators, the least trimmed squares (LTS) etc., have been employed to estimate optical flow from image sequences in recent years. However, these robust estimators have a breakdown point of no more than 50%. We propose a novel robust estimator, called variable bandwidth quick maximum density power estimator (vbQMDPE), which can tolerate more than 50% outliers. We apply the novel proposed estimator to robust optical flow estimation. Our method yields better results than most other recently proposed methods, and it has the potential to better handle multiple motion effects. Hanzi Wang, David Suter |
ICCV | 2 |
| 2003 | Contour tracking with automatic motion model switching
Prithiraj Tissainayagam, David Suter |
Pattern Recognit. | 2 |
| 2003 | Using symmetry in robust model fitting
Hanzi Wang, David Suter |
Pattern Recognit. Lett. | 2 |
| 2002 | A novel robust method for large numbers of gross errorsabstractIn computer vision tasks, it frequently happens that gross noise occupies the absolute majority of the data. Most robust estimators can tolerate no more than 50% gross errors. In this article, we propose a highly robust estimator, called MDPE (maximum density power estimator), employing density estimation and density gradient estimation techniques in the residual space. This estimator can tolerate more than 85% outliers. Experiments illustrate that the MDPE has a higher breakdown point and less errors than other recently proposed similar estimators: least median of squares (LMedS), residual consensus (RESC), and adaptive least kth order squares (ALKS). Hanzi Wang, David Suter |
ICARCV | 2 |
| 2002 | LTSD: a highly efficient symmetry-based robust estimatorabstractAlthough the least median of squares (LMedS) method and the least trimmed squares (LTS) method are said to hive a high breakdown point (50%), they can break down at unexpectedly lower percentages of outliers when those outliers are clustered. In this paper, we investigate the breakdown of LMedS and the LTS when a large percentage of clustered outliers exist in the data. We introduce the concept of symmetry distance (SD) and propose an improved method, called the least trimmed symmetry distance (LTSD). The experimental results show the LTSD gives better results than the LMedS method and the LTS method particularly when there is a large percentage of clustered outliers and/or a large standard variance in the inlier population. Hanzi Wang, David Suter |
ICARCV | 2 |
| 2001 | Performance Prediction Analysis of a Point Feature Tracker Based on Different Motion Models
Prithiraj Tissainayagam, David Suter |
Comput. Vis. Image Underst. | 2 |
| 2001 | Visual tracking with automatic motion model switching
Prithiraj Tissainayagam, David Suter |
Pattern Recognit. | 2 |
| 2000 | Visual Tracking of Multiple Objects with Automatic Motion Model SwitchingabstractIn this paper we present an efficient contour tracking algorithm which can track 2D silhouettes of multiple objects in extended image sequences captured by a static camera. We represent contours using cubic B-splines, and our tracking algorithm is based on tracking a lower dimensional shape-space. The tracker is coupled with a multiple model filtering algorithm which caters for objects that move with variable motion. The model based technique we provide is capable of tracking rigid and nonrigid object contours with good accuracy. Prithiraj Tissainayagam, David Suter |
ICPR | 2 |
| 2000 | Left Ventricular Motion Reconstruction Based on Elastic Vector SplinesabstractIn medical imaging it is common to reconstruct dense motion estimates, from sparse measurements of that motion, using some form of elastic spline (thin-plate spline, snakes and other deformable models, etc.). Usually the elastic spline uses only bending energy (second-order smoothness constraint) or stretching energy (first-order smoothness constraint), or a combination of the two. These elastic splines belong to a family of elastic vector splines called the Laplacian splines. This spline family is derived from an energy minimization functional, which is composed of multiple-order smoothness constraints. These splines can be explicitly tuned to vary the smoothness of the solution according to the deformation in the modeled material/tissue. In this context, it is natural to question which members of the family will reconstruct the motion more accurately. We compare different members of this spline family to assess how well these splines reconstruct human cardiac motion. We find that the commonly used splines (containing first-order and/or second-order smoothness terms only) are not the most accurate for modeling human cardiac motion. David Suter |
IEEE Trans. Medical Imaging | 1 |
| 1998 | Robust Total least Squares Based Optic Flow Computation
Alireza Bab-Hadiashar, David Suter |
ACCV (1) | 2 |
| 1998 | Robust Motion Segmentation Using Rank Ordering Estimatiors
Alireza Bab-Hadiashar, David Suter |
ACCV (2) | 2 |
| 1998 | Multiscale Image Representation and Edge Detection
David Suter |
ACCV (2) | 2 |
| 1998 | Robust range segmentationabstractThis paper proposes a robust estimator which is capable of segmenting multi-structural data. The proposed estimation technique is used to segment range data into linear (planar) and quadratic surfaces. The performance of the proposed range segmentation algorithm is tested on a number of real data sets. Alireza Bab-Hadiashar, David Suter |
ICPR | 2 |
| 1998 | Image coordinate transformation based on DIV-CURL vector splinesabstractWe present a vector spline technique for vector field reconstruction. These vector splines are based on an energy minimization functional, which involves the divergence and the rotational fields of the approximated vector. This technique can be used to determine the underlying coordinate transformation in image mapping. David Suter |
ICPR | 2 |
| 1998 | Visual tracking and motion determination using the IMM algorithmabstractWe present a feature tracking system with automatic motion determination of features in an image sequence. The positions of features (corners) extracted in the first frame of a sequence are estimated and predicted in the subsequent frames by using an extension of Bayesian multiple hypothesis technique (MHT) based on different motion models. The tracking of features is based on the interacting multiple model (IMM). The paper shows how the IMM algorithm combined with a MHT framework can be used in a visual tracking scenario. We considered different order (types) velocity and acceleration models for the IMM algorithm and applied them to two image sequences, the PUMA sequence and toy car sequence. The study shows that the method proposed can distinguish between different motions depicted in an image sequence with very good tracking results. Prithiraj Tissainayagam, David Suter |
ICPR | 2 |
| 1998 | Robust Optic Flow Computation
Alireza Bab-Hadiashar, David Suter |
Int. J. Comput. Vis. | 2 |
| 1997 | Optic flow calculation using robust statisticsabstractA method for calculating optic flow, using robust statistics, is developed. The method generally out-performs all competing methods in terms of accuracy. One of the key features in the success of this method, is that we use least median of squares, which is known to be robust to outliers. The computational cost is kept very low by using an approximate solution to the least median of squares only in a first stage that detects outliers. The essential ingredients of our method should be applicable in a wide range of other computer vision problems. Alireza Bab-Hadiashar, David Suter |
CVPR | 2 |
| 1996 | Robust optic flow estimation using least median of squaresabstractA new approach to optic flow calculation, based on a highly robust statistical technique, is presented. In this algorithm, the optic flow problem is first formulated as a standard least squares problem. Then, its associated closest point problem is introduced and the transformation which takes this problem to a standard regression problem is provided. The least median of squares technique is used to solve the resulting regression problem. Some experimental results for both synthetic and real image sequences are also presented. Alireza Bab-Hadiashar, David Suter |
ICIP (1) | 2 |
| 1995 | Restoration of historic film for digital compression: a case studyabstractWith the advent of compressed video standards such as MPEG, film archives around the world are looking at using the medium as a new means of distributing historically significant film material. However, we find that a compression standard such as MPEG, which is optimised for the statistics of modern motion picture film and video, is poorly suited to coding motion pictures recorded with early technologies and suffering from severe age related degradation. Indeed restoration of some historically significant film is necessary, not just so that it may be returned to its former visual quality, but to prevent the film from looking significantly worse after compression. We describe the degradation artifacts encountered-in a historically significant film made in 1906, the strategies employed to improve its compressed quality and the results of these efforts. P. Richardson, David Suter |
ICIP | 2 |
| 1994 | Motion estimation and vector splinesabstractMany formulations of visual reconstruction problems (e.g. optic flow, shape from shading, biomedical motion estimation from CT data) involve the recovery of a vector field. Often the solution is characterized via a generalized spline or regularization formulation using a smoothness constraint. This paper introduces a decomposition of the smoothness constraint into two parts: one related to the divergence of the vector field and one related to the curl or vorticity. This allows one to "tune" the smoothness to the properties of the data. One can, for example, use a high weighting on the smoothness imposed upon the curl in order to preserve the divergent parts of the field. For a particular spline within the family introduced by this decomposition process, we derive an exact solution and demonstrate the approach on examples.> David Suter |
CVPR | 1 |
| 1992 | Mixed Finite Element Based Neural Networks in Visual ReconstructionabstractThis paper shows how visual reconstruction problems can be solved using analog networks resulting from the reformulation of the reconstruction problem into a series of coupled sub-problems. A novel analog implementation based upon Platt's constraint networks is suggested for neural network implementation and the effectiveness is demonstrated with examples. The complete approach bears a philosophical similarity to the Harris Coupled Depth-Slope model of visual reconstruction. Significant differences appear, though, in the Lagrangian formulation, the mixed finite element discretization, and the analog implementation. It is also shown that the resulting networks can be structurally different from those suggested by Harris. David Suter |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1991 | Constraint Networks in VisionabstractApplications in machine vision of constraint networks based on an augmented Lagrangian formulation are discussed. Only those applications that have a fundamental significance are addressed. The first of these provides a generalization of the Harris coupled depth-slope analog model of visual reconstruction. Because of the generality of the approach, one can derive many more alternative structures, and the mathematical setting places this approach within the bounds of mixed finite element theory. This offers many advantages in terms of the associated mathematical theory and implementation on digital machines. The second use is in data fusion, which is a crucial task for systems using multiple sensors or methods of analysis of data.> David Suter |
IEEE Trans. Computers | 1 |