EDBT 2026 Demo / reviewers in the wild / expert
Octavia I. Camps
dblp:69/6960
· DBLP profile ↗
70ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0003-1945-9172ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 64 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 56 · 5 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TailedCore: Few-Shot Sampling for Unsupervised Long-Tail Noisy Anomaly DetectionabstractWe aim to solve unsupervised anomaly detection in a practical challenging environment where the normal dataset is both contaminated with defective regions and its product class distribution is tailed but unknown. We observe that existing models suffer from tail-versus-noise trade-off where if a model is robust against pixel noise, then its performance deteriorates on tail class samples, and vice versa. To mitigate the issue, we handle the tail class and noise samples independently. To this end, we propose TailSampler, a novel class size predictor that estimates the class cardinality of samples based on a symmetric assumption on the class-wise distribution of embedding similarities. TailSampler can be utilized to sample the tail class samples exclusively, allowing to handle them separately. Based on these facets, we build a memory-based anomaly detection model TailedCore, whose memory both well captures tail class information and is noise-robust. We extensively validate the effectiveness of TailedCore on the unsupervised long-tail noisy anomaly detection setting, and show that TailedCore outperforms the state-of-the-art in most settings. Code is available in TailedCore Yoon Gyo Jung, Jaewoo Park 0001, Jaeho Yoon, Kuan-Chuan Peng, Wonchul Kim, Andrew Beng Jin Teoh, Octavia I. Camps |
CVPR | 7 |
| 2025 | 3D-HGS: 3D Half-Gaussian SplattingabstractPhoto-realistic image rendering from 3D scene reconstruction has advanced significantly with neural rendering techniques. Among these, 3D Gaussian Splatting (3D-GS) outperforms Neural Radiance Fields (NeRFs) in quality and speed but struggles with shape and color discontinuities. We propose 3D Half-Gaussian (3D-HGS) kernels as a plug-and-play solution to address these limitations. Our experiments show that 3D-HGS enhances existing 3D-GS methods, achieving state-of-the-art rendering quality without compromising speed. More demos and code are available at https://lihaolin88.github.io/CVPR-2025-3DHGS. Jinyang Liu 0005, Mario Sznaier, Octavia I. Camps |
CVPR | 4 |
| 2025 | InCoDe: Interpretable Compressed Descriptions For Image GenerationabstractGenerative models have been successfully applied in diverse domains, from natural language processing to image synthesis. However, despite this success, a key challenge that remains is the ability to control the semantic content of the scene being generated. We argue that adequate control of the generation process requires a data representation that allows users to access and efficiently manipulate the semantic factors shaping the data distribution. This work advocates for the adoption of succinct, informative, and interpretable representations, quantified using information-theoretic principles. Through extensive experiments, we demonstrate the efficacy of our proposed framework both qualitatively and quantitatively. Our work contributes to the ongoing quest to enhance both controllability and interpretability in the generation process. Code available at github.com/ArmandCom/InCoDe. Armand Comas Massague, Aditya Chattopadhyay, Feliu Formosa, Changyu Liu, Octavia I. Camps, René Vidal |
ICLR | 5 |
| 2025 | Generalization Error Analysis for Selective State-Space Models Through the Lens of AttentionabstractState-space models (SSMs) have recently emerged as a compelling alternative to Transformers for sequence modeling tasks. This paper presents a theoretical generalization analysis of selective SSMs, the core architectural component behind the Mamba model. We derive a novel covering number-based generalization bound for selective SSMs, building upon recent theoretical advances in the analysis of Transformer models. Using this result, we analyze how the spectral abscissa of the continuous-time state matrix influences the model’s stability during training and its ability to generalize across sequence lengths. We empirically validate our findings on a synthetic majority task, the IMDb sentiment classification benchmark, and the ListOps task, demonstrating how our theoretical insights translate into practical model behavior. Arya Honarpisheh, Mustafa Bozdag, Octavia I. Camps, Mario Sznaier |
NeurIPS | 3 |
| 2024 | Solving Masked Jigsaw Puzzles with Diffusion Vision TransformersabstractSolving image and video jigsaw puzzles poses the chal-lenging task of rearranging image fragments or video frames from unordered sequences to restore meaningful images and video sequences. Existing approaches often hinge on discriminative models tasked with predicting either the absolute positions of puzzle elements or the permutation actions applied to the original data. Unfortunately, these methods face limitations in effectively solving puzzles with a large number of elements. In this paper, we propose JPDVT, an innovative approach that harnesses diffusion transform-ers to address this challenge. Specifically, we generate positional information for image patches or video frames, conditioned on their underlying visual content. This infor-mation is then employed to accurately assemble the puzzle pieces in their correct positions, even in scenarios involving missing pieces. Our method achieves state-of-the-art performance on several datasets. Jinyang Liu 0005, Wondmgezahu Teshome, Sandesh Ghimire, Mario Sznaier, Octavia I. Camps |
CVPR | 5 |
| 2024 | Face Reconstruction Transfer Attack as Out-of-Distribution Generalization
Yoon Gyo Jung, Jaewoo Park 0001, Xingbo Dong, Hojin Park, Andrew Beng Jin Teoh, Octavia I. Camps |
ECCV (75) | 6 |
| 2024 | HAT: History-Augmented Anchor Transformer for Online Temporal Action Localization
Sakib Reza, Yuexi Zhang, Mohsen Moghaddam, Octavia I. Camps |
ECCV (21) | 4 |
| 2024 | FasterVD: On Acceleration of Video Diffusion Models
Pinrui Yu, Timothy Rupprecht, Zhenglun Kong, Pu Zhao 0001, Yanyu Li, Octavia I. Camps, Xue Lin 0001, Yanzhi Wang 0001 |
IJCAI | 8 |
| 2023 | Inferring Relational Potentials in Interacting SystemsabstractSystems consisting of interacting agents are prevalent in the world, ranging from dynamical systems in physics to complex biological networks. To build systems which can interact robustly in the real world, it is thus important to be able to infer the precise interactions governing such systems. Existing approaches typically discover such interactions by explicitly modeling the feed-forward dynamics of the trajectories. In this work, we propose Neural Interaction Inference with Potentials (NIIP) as an alternative approach to discover such interactions that enables greater flexibility in trajectory modeling: it discovers a set of relational potentials, represented as energy functions, which when minimized reconstruct the original trajectory. NIIP assigns low energy to the subset of trajectories which respect the relational constraints observed. We illustrate that with these representations NIIP displays unique capabilities in test-time. First, it allows trajectory manipulation, such as interchanging interaction types across separately trained models, as well as trajectory forecasting. Additionally, it allows adding external hand-crafted potentials at test-time. Finally, NIIP enables the detection of out-of-distribution samples and anomalies without explicit training. Armand Comas Massague, Yilun Du, Christian Fernandez Lopez, Sandesh Ghimire, Mario Sznaier, Josh Tenenbaum, Octavia I. Camps |
ICML | 7 |
| 2022 | Fast Two-View Motion Segmentation Using Christoffel Polynomials
Bengisu Özbay, Octavia I. Camps, Mario Sznaier |
ECCV (30) | 2 |
| 2022 | Dynamics-aware Representation Learning via Multivariate Time Series TransformersabstractWe propose a novel multivariate time series autoencoder, which produces interpretable linear-dynamical latent features that govern the predictions for several downstream tasks. To this end, we combine a transformer autoencoder with a dynamical atoms-based autoencoder to mimic Koopman operators in the latent space. We demonstrate that our approach significantly outperforms deep Koopman operator learning baselines for time series forecasting on chaotic systems such as the lorenz Attractor. Furthermore, the dynamics-aware representations, combined with a transformer classifier, lead to state-of-the-art classification accuracy on benchmark multivariate time series datasets. Our code is publicly available at https://github.com/mlpotter/T-DYAN-T. Michael Potter, Ilkay Yildiz, Octavia I. Camps, Mario Sznaier |
ESANN | 3 |
| 2020 | Towards Visually Explaining Variational AutoencodersabstractRecent advances in Convolutional Neural Network (CNN) model interpretability have led to impressive progress in visualizing and understanding model predictions. In particular, gradient-based visual attention methods have driven much recent effort in using visual attention maps as a means for visual explanations. A key problem, however, is these methods are designed for classification and categorization tasks, and their extension to explaining generative models, e.g., variational autoencoders (VAE) is not trivial. In this work, we take a step towards bridging this crucial gap, proposing the first technique to visually explain VAEs by means of gradient-based attention. We present methods to generate visual attention from the learned latent space, and also demonstrate such attention explanations serve more than just explaining VAE predictions. We show how these attention maps can be used to localize anomalies in images, demonstrating state-of-the-art performance on the MVTec-AD dataset. We also show how they can be infused into model training, helping bootstrap the VAE into learning improved latent space disentanglement, demonstrated on the Dsprites dataset. WenQian Liu, Runze Li 0003, Meng Zheng 0002, Srikrishna Karanam, Ziyan Wu 0001, Bir Bhanu, Richard J. Radke, Octavia I. Camps |
CVPR | 8 |
| 2020 | Key Frame Proposal Network for Efficient Pose Estimation in Videos
Yuexi Zhang, Octavia I. Camps, Mario Sznaier |
ECCV (17) | 3 |
| 2020 | Learning Disentangled Representations of Videos with Missing DataabstractMissing data poses significant challenges while learning representations of video sequences. We present Disentangled Imputed Video autoEncoder (DIVE), a deep generative model that imputes and predicts future video frames in the presence of missing data. Specifically, DIVE introduces a missingness latent variable, disentangles the hidden video representations into static and dynamic appearance, pose, and missingness factors for each object, while it imputes each object trajectory where data is missing. On a moving MNIST dataset with various missing scenarios, DIVE outperforms the state of the art baselines by a substantial margin. We also present comparisons on a real-world MOTSChallenge pedestrian dataset, which demonstrates the practical value of our method in a more realistic setting. Our code can be found in https://github.com/Rose-STL-Lab/DIVE. Armand Comas Massague, Zlatan Feric, Octavia I. Camps, Rose Yu |
NeurIPS | 4 |
| 2020 | Dynamic Motion Representation for Human Action RecognitionabstractDespite the advances in Human Activity Recognition, the ability to exploit the dynamics of human body motion in videos has yet to be achieved. In numerous recent works, researchers have used appearance and motion as independent inputs to infer the action that is taking place in a specific video. In this paper, we highlight that while using a novel representation of human body motion, we can benefit from appearance and motion simultaneously. As a result, better performance of action recognition can be achieved. We start with a pose estimator to extract the location and heat-map of body joints in each frame. We use a dynamic encoder to generate a fixed size representation from these body joint heat-maps. Our experimental results show that training a convolutional neural network with the dynamic motion representation outperforms state-of-the-art action recognition models. By modeling distinguishable activities as distinct dynamical systems and with the help of two stream networks, we obtain the best performance on HMDB, JHMDB, UCF-101, and AVA datasets. Sadjad Asghari-Esfeden, Mario Sznaier, Octavia I. Camps |
WACV | 3 |
| 2019 | A Systematic Evaluation and Benchmark for Person Re-Identification: Features, Metrics, and DatasetsabstractPerson re-identification (re-id) is a critical problem in video analytics applications such as security and surveillance. The public release of several datasets and code for vision algorithms has facilitated rapid progress in this area over the last few years. However, directly comparing re-id algorithms reported in the literature has become difficult since a wide variety of features, experimental protocols, and evaluation metrics are employed. In order to address this need, we present an extensive review and performance evaluation of single- and multi-shot re-id algorithms. The experimental protocol incorporates the most recent advances in both feature extraction and metric learning. To ensure a fair comparison, all of the approaches were implemented using a unified code library that includes 11 feature extraction algorithms and 22 metric learning and ranking techniques. All approaches were evaluated using a new large-scale dataset that closely mimics a real-world problem setting, in addition to 16 other publicly available datasets: VIPeR, GRID, CAVIAR, DukeMTMC4ReID, 3DPeS, PRID, V47, WARD, SAIVT-SoftBio, CUHK01, CHUK02, CUHK03, RAiD, iLIDSVID, HDA+, and Market1501. The evaluation codebase and results will be made publicly available for community use. Srikrishna Karanam, Mengran Gou, Ziyan Wu 0001, Angels Rates-Borras, Octavia I. Camps, Richard J. Radke |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2018 | MoNet: Moments Embedding NetworkabstractBilinear pooling has been recently proposed as a feature encoding layer, which can be used after the convolutional layers of a deep network, to improve performance in multiple vision tasks. Different from conventional global average pooling or fully connected layer, bilinear pooling gathers 2nd order information in a translation invariant fashion. However, a serious drawback of this family of pooling layers is their dimensionality explosion. Approximate pooling methods with compact properties have been explored towards resolving this weakness. Additionally, recent results have shown that significant performance gains can be achieved by adding 1st order information and applying matrix normalization to regularize unstable higher order information. However, combining compact pooling with matrix normalization and other order information has not been explored until now. In this paper, we unify bilinear pooling and the global Gaussian embedding layers through the empirical moment matrix. In addition, we propose a novel sub-matrix square-root layer, which can be used to normalize the output of the convolution layer directly and mitigate the dimensionality problem with off-the-shelf compact pooling methods. Our experiments on three widely used fine-grained classification datasets illustrate that our proposed architecture, MoNet, can achieve similar or better performance than with the state-of-art G2DeNet. Furthermore, when combined with compact pooling technique, MoNet obtains comparable performance with encoded features with 96% less dimensions. Mengran Gou, Octavia I. Camps, Mario Sznaier |
CVPR | 3 |
| 2018 | SoS-RSC: A Sum-of-Squares Polynomial Approach to Robustifying Subspace Clustering AlgorithmsabstractThis paper addresses the problem of subspace clustering in the presence of outliers. Typically, this scenario is handled through a regularized optimization, whose computational complexity scales polynomially with the size of the data. Further, the regularization terms need to be manually tuned to achieve optimal performance. To circumvent these difficulties, in this paper we propose an outlier removal algorithm based on evaluating a suitable sum-of-squares polynomial, computed directly from the data. This algorithm only requires performing two singular value decompositions of fixed size, and provides certificates on the probability of misclassifying outliers as inliers. Mario Sznaier, Octavia I. Camps |
CVPR | 2 |
| 2018 | DYAN: A Dynamical Atoms-Based Network for Video Prediction
WenQian Liu, Abhishek Sharma 0009, Octavia I. Camps, Mario Sznaier |
ECCV (12) | 3 |
| 2017 | Dynamics Enhanced Multi-camera Motion Segmentation from Unsynchronized VideosabstractThis paper considers the multi-camera motion segmentation problem using unsynchronized videos: given two video clips containing several moving objects, captured by unregistered, unsynchronized cameras with different viewpoints, our goal is to assign features to moving objects in the scene. This problem challenges existing methods, due to the lack of registration information and correspondences across cameras. To solve it, we propose a method that combines shape and dynamical information and does not require spatiotemporal registration or shared features. As shown in the paper, this combination results in improved performance even in the single camera case, and allows for solving the multi-camera segmentation problem with a computational cost similar to that of existing single-view techniques. Xikang Zhang, Bengisu Özbay, Mario Sznaier, Octavia I. Camps |
ICCV | 4 |
| 2017 | From the Lab to the Real World: Re-identification in an Airport Camera NetworkabstractOver the past ten years, human re-identification has received increased attention from the computer vision research community. However, for the most part, these research papers are divorced from the context of how such algorithms would be used in a real-world system. This paper describes the unique opportunity our group of academic researchers had to design and deploy a human re-identification system in a demanding real-world environment: a busy airport. The system had to be designed from the ground up, including robust modules for real-time human detection and tracking, a distributed, low-latency software architecture, and a front-end user interface designed for a specific scenario. None of these issues are typically addressed in re-identification research papers, but all are critical to an effective system that end users would actually be willing to adopt. We detail the challenges of the real-world airport environment, the computer vision algorithms underlying our human detection and re-identification algorithms, our robust software architecture, and the ground-truthing system required to provide the training and validation data for the algorithms. Our initial results show that despite the challenges and constraints of the airport environment, the proposed system achieves very good performance while operating in real time. Octavia I. Camps, Mengran Gou, Tom Hebble, Srikrishna Karanam, Oliver Lehmann, Yang Li 0056, Richard J. Radke, Ziyan Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Person Re-identification in Appearance Impaired Scenarios
Mengran Gou, Xikang Zhang, Angels Rates-Borras, Sadjad Asghari-Esfeden, Octavia I. Camps, Mario Sznaier |
BMVC | 5 |
| 2016 | Subspace Clustering with Priors via Sparse Quadratically Constrained Quadratic ProgrammingabstractThis paper considers the problem of recovering a subspace arrangement from noisy samples, potentially corrupted with outliers. Our main result shows that this problem can be formulated as a convex semi-definite optimization problem subject to an additional rank constrain that involves only a very small number of variables. This is established by first reducing the problem to a quadratically constrained quadratic problem and then using its special structure to find conditions guaranteeing that a suitably built convex relaxation is indeed exact. When combined with the standard nuclear norm relaxation for rank, the results above lead to computationally efficient algorithms with optimality guarantees. A salient feature of the proposed approach is its ability to incorporate existing a-priori information about the noise, co-ocurrences, and percentage of outliers. These results are illustrated with several examples. Yongfang Cheng, Mario Sznaier, Octavia I. Camps |
CVPR | 4 |
| 2016 | Solving Temporal PuzzlesabstractMany physical phenomena, within short time windows, can be explained by low order differential relations. In a discrete world, these relations can be described using low order difference equations or equivalently low order auto regressive (AR) models. In this paper, based on this intuition, we propose an algorithm for solving time-sort temporal puzzles, defined as scrambled time series that need to be sorted out. We frame this problem using a mixed-integer semi definite programming formulation and show how to turn it into a mixed-integer linear programming problem, which can be solved with off-the-shelf solvers, by using the recently introduced atomic norm framework. Our experiments show the effectiveness and generality of our approach in different scenarios. Caglayan Dicle, Burak Yilmaz, Octavia I. Camps, Mario Sznaier |
CVPR | 3 |
| 2016 | Efficient Temporal Sequence Comparison and Classification Using Gram Matrix Embeddings on a Riemannian ManifoldabstractIn this paper we propose a new framework to compare and classify temporal sequences. The proposed approach captures the underlying dynamics of the data while avoiding expensive estimation procedures, making it suitable to process large numbers of sequences. The main idea is to first embed the sequences into a Riemannian manifold by using positive definite regularized Gram matrices of their Hankelets. The advantages of the this approach are: 1) it allows for using non-Euclidean similarity functions on the Positive Definite matrix manifold, which capture better the underlying geometry than directly comparing the sequences or their Hankel matrices, and 2) Gram matrices inherit desirable properties from the underlying Hankel matrices: their rank measure the complexity of the underlying dynamics, and the order and coefficients of the associated regressive models are invariant to affine transformations and varying initial conditions. The benefits of this approach are illustrated with extensive experiments in 3D action recognition using 3D joints sequences. In spite of its simplicity, the performance of this approach is competitive or better than using state-of-art approaches for this problem. Further, these results hold across a variety of metrics, supporting the idea that the improvement stems from the embedding itself, rather than from using one of these metrics. Xikang Zhang, Mengran Gou, Mario Sznaier, Octavia I. Camps |
CVPR | 5 |
| 2016 | Jensen Bregman LogDet Divergence Optimal Filtering in the Manifold of Positive Definite Matrices
Octavia I. Camps, Mario Sznaier, Biel Roig-Solvas |
ECCV (7) | 2 |
| 2015 | A convex optimization approach to robust fundamental matrix estimationabstractThis paper considers the problem of estimating the fundamental matrix from corrupted point correspondences. A general nonconvex framework is proposed that explicitly takes into account the rank-2 constraint on the fundamental matrix and the presence of noise and outliers. The main result of the paper shows that this non-convex problem can be solved by solving a sequence of convex semi-definite programs, obtained by exploiting a combination of polynomial optimization tools and rank minimization techniques. Further, the algorithm can be easily extended to handle the case where only some of the correspondences are labeled, and, to exploit co-ocurrence information, if available. Consistent experiments show that the proposed method works well, even in scenarios characterized by a very high percentage of outliers. Yongfang Cheng, Jose A. Lopez, Octavia I. Camps, Mario Sznaier |
CVPR | 3 |
| 2015 | Self Scaled Regularized Robust RegressionabstractLinear Robust Regression (LRR) seeks to find the parameters of a linear mapping from noisy data corrupted from outliers, such that the number of inliers (i.e. pairs of points where the fitting error of the model is less than a given bound) is maximized. While this problem is known to be NP hard, several tractable relaxations have been recently proposed along with theoretical conditions guaranteeing exact recovery of the parameters of the model. However, these relaxations may perform poorly in cases where the fitting error for the outliers is large. In addition, these approaches cannot exploit available a-priori information, such as cooccurrences. To circumvent these difficulties, in this paper we present an alternative approach to robust regression. Our main result shows that this approach is equivalent to a “self-scaled” ℓ1regularized robust regression problem, where the cost function is automatically scaled, with scalings that depend on the a-priori information. Thus, the proposed approach achieves substantially better performance than traditional regularized approaches in cases where the outliers are far from the linear manifold spanned by the inliers, while at the same time exhibits the same theoretical recovery properties. These results are illustrated with several application examples using both synthetic and real data. Caglayan Dicle, Mario Sznaier, Octavia I. Camps |
CVPR | 4 |
| 2015 | Hankelet-based dynamical systems modeling for 3D action recognitionabstractThis paper proposes to model an action as the output of a sequence of atomic Linear Time Invariant (LTI) systems. The sequence of LTI systems generating the action is modeled as a Markov chain , where a Hidden Markov Model (HMM) is used to model the transition from one atomic LTI system to another. In turn, the LTI systems are represented in terms of their Hankel matrices. For classification purposes, the parameters of a set of HMMs (one for each action class) are learned via a discriminative approach. This work proposes a novel method to learn the atomic LTI systems from training data , and analyzes in detail the action representation in terms of a sequence of Hankel matrices. Extensive evaluation of the proposed approach on two publicly available datasets demonstrates that the proposed method attains state-of-the-art accuracy in action classification from the 3D locations of body joints (skeleton). Liliana Lo Presti, Marco La Cascia, Stan Sclaroff, Octavia I. Camps |
Image Vis. Comput. | 4 |
| 2014 | Gesture Modeling by Hanklet-Based Hidden Markov Model
Liliana Lo Presti, Marco La Cascia, Stan Sclaroff, Octavia I. Camps |
ACCV (3) | 4 |
| 2014 | Person Re-Identification Using Kernel-Based Metric Learning Methods
Mengran Gou, Octavia I. Camps, Mario Sznaier |
ECCV (7) | 3 |
| 2013 | Finding Causal Interactions in Video SequencesabstractThis paper considers the problem of detecting causal interactions in video clips. Specifically, the goal is to detect whether the actions of a given target can be explained in terms of the past actions of a collection of other agents. We propose to solve this problem by recasting it into a directed graph topology identification, where each node corresponds to the observed motion of a given target, and each link indicates the presence of a causal correlation. As shown in the paper, this leads to a block-sparsification problem that can be efficiently solved using a modified Group-Lasso type approach, capable of handling missing data and outliers (due for instance to occlusion and mis-identified correspondences). Moreover, this approach also identifies time instants where the interactions between agents change, thus providing event detection capabilities. These results are illustrated with several examples involving non-trivial interactions amongst several human subjects. Mustafa Ayazoglu, Burak Yilmaz, Mario Sznaier, Octavia I. Camps |
ICCV | 4 |
| 2013 | The Way They Move: Tracking Multiple Targets with Similar AppearanceabstractWe introduce a computationally efficient algorithm for multi-object tracking by detection that addresses four main challenges: appearance similarity among targets, missing data due to targets being out of the field of view or occluded behind other objects, crossing trajectories, and camera motion. The proposed method uses motion dynamics as a cue to distinguish targets with similar appearance, minimize target mis-identification and recover missing data. Computational efficiency is achieved by using a Generalized Linear Assignment (GLA) coupled with efficient procedures to recover missing data and estimate the complexity of the underlying dynamics. The proposed approach works with track lets of arbitrary length and does not assume a dynamical model a priori, yet it captures the overall motion dynamics of the targets. Experiments using challenging videos show that this framework can handle complex target motions, non-stationary cameras and long occlusions, on scenarios where appearance cues are not available or poor. Caglayan Dicle, Octavia I. Camps, Mario Sznaier |
ICCV | 2 |
| 2012 | Fast algorithms for structured robust principal component analysisabstractA large number of problems arising in computer vision can be reduced to the problem of minimizing the nuclear norm of a matrix, subject to additional structural and sparsity constraints on its elements. Examples of relevant applications include, among others, robust tracking in the presence of outliers, manifold embedding, event detection, in-painting and tracklet matching across occlusion. In principle, these problems can be reduced to a convex semi-definite optimization form and solved using interior point methods. However, the poor scaling properties of these methods limit the use of this approach to relatively small sized problems. The main result of this paper shows that structured nuclear norm minimization problems can be efficiently solved by using an iterative Augmented Lagrangian Type (ALM) method that only requires performing at each iteration a combination of matrix thresholding and matrix inversion steps. As we illustrate in the paper with several examples, the proposed algorithm results in a substantial reduction of computational time and memory requirements when compared against interior-point methods, opening up the possibility of solving realistic, large sized problems. Mustafa Ayazoglu, Mario Sznaier, Octavia I. Camps |
CVPR | 3 |
| 2012 | Cross-view activity recognition using HankeletsabstractHuman activity recognition is central to many practical applications, ranging from visual surveillance to gaming interfacing. Most approaches addressing this problem are based on localized spatio-temporal features that can vary significantly when the viewpoint changes. As a result, their performances rapidly deteriorate as the difference between the viewpoints of the training and testing data increases. In this paper, we introduce a new type of feature, the “Hankelet” that captures dynamic properties of short tracklets. While Hankelets do not carry any spatial information, they bring invariant properties to changes in viewpoint that allow for robust cross-view activity recognition, i.e. when actions are recognized using a classifier trained on data from a different viewpoint. Our experiments on the IXMAS dataset show that using Hanklets improves the state of the art performance by over 20%. Binlong Li, Octavia I. Camps, Mario Sznaier |
CVPR | 2 |
| 2012 | Dynamic Context for Tracking behind Occlusions
Octavia I. Camps, Mario Sznaier |
ECCV (5) | 2 |
| 2011 | Activity recognition using dynamic subspace anglesabstractCameras are ubiquitous everywhere and hold the promise of significantly changing the way we live and interact with our environment. Human activity recognition is central to understanding dynamic scenes for applications ranging from security surveillance, to assisted living for the elderly, to video gaming without controllers. Most current approaches to solve this problem are based in the use of local temporal-spatial features that limit their ability to recognize long and complex actions. In this paper, we propose a new approach to exploit the temporal information encoded in the data. The main idea is to model activities as the output of unknown dynamic systems evolving from unknown initial conditions. Under this framework, we show that activity videos can be compared by computing the principal angles between subspaces representing activity types which are found by a simple SVD of the experimental data. The proposed approach outperforms state-of-the-art methods classifying activities in the KTH dataset as well as in much more complex scenarios involving interacting actors. Binlong Li, Mustafa Ayazoglu, Teresa Mao, Octavia I. Camps, Mario Sznaier |
CVPR | 4 |
| 2011 | Dynamic subspace-based coordinated multicamera trackingabstractThis paper considers the problem of sustained multicamera tracking in the presence of occlusion and changes in the target motion model. The key insight of the proposed method is the fact that, under mild conditions, the 2D trajectories of the target in the image planes of each of the cameras are constrained to evolve in the same subspace. This observation allows for identifying, at each time instant, a single (piecewise) linear model that explains all the available 2D measurements. In turn, this model can be used in the context of a modified particle filter to predict future target locations. In the case where the target is occluded to some of the cameras, the missing measurements can be estimated using the facts that they must lie both in the subspace spanned by previous measurements and satisfy epipolar constraints. Hence, by exploiting both dynamical and geometrical constraints the proposed method can robustly handle substantial occlusion, without the need for performing 3D reconstruction, calibrated cameras or constraints on sensor separation. The performance of the proposed tracker is illustrated with several challenging examples involving targets that substantially change appearance and motion models while occluded to some of the cameras. Mustafa Ayazoglu, Binlong Li, Caglayan Dicle, Mario Sznaier, Octavia I. Camps |
ICCV | 5 |
| 2011 | Low order dynamics embedding for high dimensional time seriesabstractThis paper considers the problem of finding low order nonlinear embeddings of dynamic data, that is, data characterized by a temporal ordering. Examples where this problem arises include, among others, appearance-based multi-frame tracking, activity recognition from video and dynamic texture analysis/synthesis. Our main result is a semi-definite programming based manifold embedding algorithm that exploits the dynamical information encapsulated in the temporal ordering of the sequence to obtain embeddings with lower complexity that those produced by existing algorithms, at a comparable computational cost. In addition, the use of spatio-temporal information allows for minimizing the effects of outliers on the manifold structure and for handling fragmented sequences, where some of the data is missing, for instance due to occlusion. Octavia I. Camps, Mario Sznaier |
ICCV | 2 |
| 2010 | GPCA with denoising: A moments-based convex approachabstractThis paper addresses the problem of segmenting a combination of linear subspaces and quadratic surfaces from sample data points corrupted by (not necessarily small) noise. Our main result shows that this problem can be reduced to minimizing the rank of a matrix whose entries are affine in the optimization variables, subject to a convex constraint imposing that these variables are the moments of an (unknown) probability distribution function with finite support. Exploiting the linear matrix inequality based characterization of the moments problem and appealing to well known convex relaxations of rank leads to an overall semi-definite optimization problem. We apply our method to problems such as simultaneous 2D motion segmentation and motion segmentation from two perspective views and illustrate that our formulation substantially reduces the noise sensitivity of existing approaches. Necmiye Ozay, Mario Sznaier, Constantino M. Lagoa, Octavia I. Camps |
CVPR | 4 |
| 2010 | Euclidean Structure Recovery from Motion in Perspective Image Sequences via Hankel Rank Minimization
Mustafa Ayazoglu, Mario Sznaier, Octavia I. Camps |
ECCV (2) | 3 |
| 2008 | Fast track matching and event detectionabstractThis paper addresses the problems of track stitching and dynamic event detection in a sequence of frames. The input data consists of tracks, possibly fragmented due to occlusion, belonging to multiple targets. The goals are to (i) establish track identity across occlusion, and (ii) detect points where the motion of these targets undergo substantial changes. The main result of the paper is a simple, computationally inexpensive approach that achieves these goals in a unified way. Given a continuous track, the main idea is to detect changes in the dynamics by parsing it into segments according to the complexity of the model required to explain the observed data. Intuitively, changes in this complexity correspond to points where the dynamics change. Since the problem of estimating the complexity of the underlying model can be reduced to estimating the rank of a matrix constructed from the observed data, these changes can be found with a simple algorithm, computationally no more expensive that a sequence of SVDs. Proceeding along the same lines, fragmented tracks corresponding to multiple targets can be linked by searching for sets corresponding to minimal complexity joint models. As we show in the paper, this problem can be reduced to a semi-definite optimization and efficiently solved with commonly available software. Mario Sznaier, Octavia I. Camps |
CVPR | 3 |
| 2008 | Sequential sparsification for change detectionabstractThis paper presents a general method for segmenting a vector valued sequence into an unknown number of subsequences where all data points from a subsequence can be represented with the same affine parametric model. The idea is to cluster the data into the minimum number of such subsequences which, as we show, can be cast as a sparse signal recovery problem by exploiting the temporal correlation between consecutive data points. We try to maximize the sparsity (i.e. the number of zero elements) of the first order differences of the sequence of parameter vectors. Each non-zero element in the first order difference sequence corresponds to a change. A weighted l1norm based convex approximation is adopted to solve the change detection problem. We apply the proposed method to video segmentation and temporal segmentation of dynamic textures. Necmiye Ozay, Mario Sznaier, Octavia I. Camps |
CVPR | 3 |
| 2008 | Extraction and temporal segmentation of multiple motion trajectories in human motion
Junghye Min, Rangachar Kasturi, Octavia I. Camps |
Image Vis. Comput. | 3 |
| 2007 | A Rank Minimization Approach to Video InpaintingabstractThis paper addresses the problem of video inpainting, that is seamlessly reconstructing missing portions in a set of video frames. We propose to solve this problem proceeding as follows: (i) finding a set of descriptors that encapsulate the information necessary to reconstruct a frame, (ii) finding an optimal estimate of the value of these descriptors for the missing/corrupted frames, and (iii) using the estimated values to reconstruct the frames. The main result of the paper shows that the optimal descriptor estimates can be efficiently obtained by minimizing the rank of a matrix directly constructed from the available data, leading to a simple, computationally attractive, dynamic inpainting algorithm that optimizes the use of spatio/temporal information. Moreover, contrary to most currently available techniques, the method can handle non-periodic target motions, non-stationary backgrounds and moving cameras. These results are illustrated with several examples, including reconstructing dynamic textures and object disocclusion in cases involving both moving targets and camera. Mario Sznaier, Octavia I. Camps |
ICCV | 3 |
| 2006 | Dynamic Appearance Modeling for Human TrackingabstractDynamic appearance is one of the most important cues for tracking and identifying moving people. However, direct modeling spatio-temporal variations of such appearance is often a difficult problem due to their high dimensionality and nonlinearities. In this paper we present a human tracking system that uses a dynamic appearance and motion modeling framework based on the use of robust system dynamics identification and nonlinear dimensionality reduction techniques. The proposed system learns dynamic appearance and motion models from a small set of initial frames and does not require prior knowledge such as gender or type of activity. The advantages of the proposed tracking system are illustrated with several examples where the learned dynamics accurately predict the location and appearance of the targets in future frames, preventing tracking failures due to model drifting, target occlusion and scene clutter. Hwasup Lim, Octavia I. Camps, Mario Sznaier, Vlad I. Morariu |
CVPR (1) | 2 |
| 2006 | Dynamics Based Robust Motion SegmentationabstractIn this paper we consider the problem of segmenting multiple rigid motions using multi-frame point correspondence data. The main idea of the method is to group points according to the complexity of the model required to explain their relative motion. Intuitively, this formalizes the idea that points on the same rigid share more modes of motion (for instance a common translation or rotation) than points on different objects, leading to less complex models. By exploiting results from systems theory, the problem of estimating the complexity of the underlying model is reduced to simply computing the rank of a matrix constructed from the correspondence data. This leads to a simple segmentation algorithm, computationally no more expensive than a sequence of SVDs. Since the proposed method exploits both spatial and temporal constraints, is less sensitive to the effect of noise or outliers than approaches that rely solely on factorizations of the measurements matrix. In addition, the method can also naturally handle "degenerate cases", e.g. cases where the objects partially share motion modes. These results are illustrated using several examples involving both degenerate and non-degenerate cases. Roberto Lublinerman, Mario Sznaier, Octavia I. Camps |
CVPR (1) | 3 |
| 2006 | Modeling Correspondences for Multi-Camera Tracking Using Nonlinear Manifold Learning and Target DynamicsabstractMulti-camera tracking systems often must maintain consistent identity labels of the targets across views to recover 3D trajectories and fully take advantage of the additional information available from the multiple sensors. Previous approaches to the "correspondence across views" problem include matching features, using camera calibration information, and computing homographies between views under the assumption that the world is planar. However, it can be difficult to match features across significantly different views. Furthermore, calibration information is not always available and planar world hypothesis can be too restrictive. In this paper, a new approach is presented for matching correspondences based on the use of nonlinear manifold learning and system dynamics identification. The proposed approach does not require similar views, calibration nor geometric assumptions of the 3D environment, and is robust to noise and occlusion. Experimental results demonstrate the use of this approach to generate and predict views in cases where identity labels become ambiguous. Vlad I. Morariu, Octavia I. Camps |
CVPR (1) | 2 |
| 2006 | Performance characterization of the dynamic programming obstacle detection algorithmabstractA computer vision-based system using images from an airborne aircraft can increase flight safety by aiding the pilot to detect obstacles in the flight path so as to avoid mid-air collisions. Such a system fits naturally with the development of an external vision system proposed by NASA for use in high-speed civil transport aircraft with limited cockpit visibility. The detection techniques should provide high detection probability for obstacles that can vary from subpixels to a few pixels in size, while maintaining a low false alarm probability in the presence of noise and severe background clutter. Furthermore, the detection algorithms must be able to report such obstacles in a timely fashion, imposing severe constraints on their execution time. For this purpose, we have implemented a number of algorithms to detect airborne obstacles using image sequences obtained from a camera mounted on an aircraft. This paper describes the methodology used for characterizing the performance of the dynamic programming obstacle detection algorithm and its special cases. The experimental results were obtained using several types of image sequences, with simulated and real backgrounds. The approximate performance of the algorithm is also theoretically derived using principles of statistical analysis in terms of the signal-to-noise ration (SNR) required for the probabilities of false alarms and misdetections to be lower than prespecified values. The theoretical and experimental performance are compared in terms of the required SNR. Tarak Gandhi, Mau-Tsuen Yang, Rangachar Kasturi, Octavia I. Camps, Lee D. Coraor, Jeffrey McCandless |
IEEE Trans. Image Process. | 4 |
| 2005 | A Caratheodory-Fejer Approach to Dynamic Appearance ModelingabstractThis paper presents a technique to learn dynamic appearance models from a small number of training frames. Under this framework, dynamic appearance is modelled as an unknown operator that satisfies certain interpolation conditions and that can be efficiently identified using very little a priori information with off the shelf software. The advantages of the proposed method are illustrated with several examples where the leaned dynamics accurately predict the appearance of the targets preventing tracking failures due to occlusion or clutter. Hwasup Lim, Octavia I. Camps, Mario Sznaier |
CVPR (1) | 2 |
| 2005 | Robust Structure from Motion and Identified DynamicsabstractThis paper addresses robust recovery of structure and motion for rigid bodies in video sequences. For small inter-frame motion, feature appearance is commonly estimated according to the previous frame. However, in the case of occlusion or corrupt frames, the small interframe motion model fails. In this paper, we propose to use robust identification techniques to estimate the motion dynamics based on a set of previous frames. These dynamics are then recursively used to robustly estimate object structure and motion and to predict the object appearance in future frames. Results are tested for 3D reconstructions of rigid bodies in real and synthetic image sequences. Benjamin R. Fransen, Octavia I. Camps, Mario Sznaier |
ICCV | 2 |
| 2005 | A Model (In)Validation Approach to Gait ClassificationabstractThis paper addresses the problem of human gait classification from a robust model (in)validation perspective. The main idea is to associate to each class of gaits a nominal model, subject to bounded uncertainty and measurement noise. In this context, the problem of recognizing an activity from a sequence of frames can be formulated as the problem of determining whether this sequence could have been generated by a given (model, uncertainty, and noise) triple. By exploiting interpolation theory, this problem can be recast into a nonconvex optimization. In order to efficiently solve it, we propose two convex relaxations, one deterministic and one stochastic. As we illustrate experimentally, these relaxations achieve over 83 percent and 86 percent success rates, respectively, even in the face of noisy data. Maria Cecilia Mazzaro, Mario Sznaier, Octavia I. Camps |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2004 | Segmentation for robust tracking in the presence of severe occlusionabstractTracking an object in a sequence of images can fail due to partial occlusion or clutter. Robustness to occlusion can be increased by tracking the object as a set of "parts" such that not all of these are occluded at the same time. However, successful implementation of this idea hinges upon finding a suitable set of parts. In this paper we propose a novel segmentation, specifically designed to improve robustness against occlusion in the context of tracking. The main result shows that tracking the parts resulting from this segmentation outperforms both tracking parts obtained through traditional segmentations, and tracking the entire target. Additional results include a statistical analysis of the correlation between features of a part and tracking error, and identifying a cost function that exhibits a high degree of correlation with the tracking error. Camillo Gentile, Octavia I. Camps, Mario Sznaier |
IEEE Trans. Image Process. | 2 |
| 2003 | A Caratheodory-Fejer Approach to Robust Multiframe TrackingabstractA requirement common to most dynamic vision applications is the ability to track objects in a sequence of frames. This problem has been extensively studied in the past few years, leading to several techniques, such as unscented particle filter based trackers, that exploit a combination of the (assumed) target dynamics, empirically learned noise distributions and past position observations. While successful in many scenarios, these trackers remain fragile to occlusion and model uncertainty in the target dynamics. As we show in this paper, these difficulties can be addressed by modeling the dynamics of the target as an unknown operator that satisfies certain interpolation conditions. Results from interpolation theory can then be used to find this operator by solving a convex optimization problem. As illustrated with several examples, combining this operator with Kalman and UPF techniques leads to both robustness improvement and computational complexity reduction. Octavia I. Camps, Hwasup Lim, Maria Cecilia Mazzaro, Mario Sznaier |
ICCV | 1 |
| 2001 | Segmentation for Robust Tracking in the Presence of Severe OcclusionabstractTracking an object in a sequence of images can fail due to partial occlusion or clutter Robustness can be increased by tracking a set of "parts", provided that a suitable set can be identified. In this paper we propose a novel segmentation, specifically designed to improve robustness against occlusion in the context of tracking. The main result shows that tracking the parts resulting from this segmentation outperforms both tracking parts obtained through traditional segmentations, and tracking the entire target. Additional results include a statistical analysis of the correlation between features of a part and tracking error, and identifying a cost function highly correlated with the tracking error. Camillo Gentile, Octavia I. Camps, Mario Sznaier |
CVPR (2) | 2 |
| 2000 | Detection of Obstacles in the Flight Path of an AircraftabstractThe National Aeronautics and Space Administration (NASA), along with members of the aircraft industry, recently developed technologies for a new supersonic aircraft. One of the technological areas considered for this aircraft is the use of video cameras and image processing equipment to aid the pilot in detecting other aircraft in the sky. The detection techniques should provide high detection probability for obstacles that can vary from sub-pixel to a few pixels in size, while maintaining a low false alarm probability in the presence of noise and severe background clutter Furthermore, the detection algorithms must be able to report such obstacles in a timely fashion, imposing severe constraints on their execution time. This paper describes approaches to detect airborne obstacles on collision course and crossing trajectories in video images captured from an airborne aircraft. In both cases the approaches consist of an image processing stage to identify possible obstacles followed by a tracking stage to distinguish between true obstacles and image clutter, based on their behavior. The crossing target detection algorithm was also implemented on a pipelined architecture from DataCube and runs in real time. Both algorithms have been successfully tested on flight tests conducted by NASA. Tarak Gandhi, Mau-Tsuen Yang, Rangachar Kasturi, Octavia I. Camps, Lee D. Coraor, Jeffrey McCandless |
CVPR | 4 |
| 2000 | Detection of obstacles on runways using ego-motion compensation and tracking of significant features
Tarak Gandhi, Sadashiva Devadiga, Rangachar Kasturi, Octavia I. Camps |
Image Vis. Comput. | 4 |
| 1999 | Line-Based Recognition Using A Multidimensional Hausdorff DistanceabstractA line-feature-based approach for model based recognition using a four-dimensional Hausdorff distance is proposed. This approach reduces the problem of finding the rotation, scaling, and translation transformations between a model and an image to the problem of finding a single translation minimizing the Hausdorff distance between two sets of points in a four-dimensional space. The implementation of the proposed algorithm can be naturally extended to higher dimensional spaces to efficiently find correspondences between n-dimensional patterns. The method performance and sensitivity to segmentation problems are quantitatively characterized using an experimental protocol with simulated data. It is shown that the algorithm performs well, is robust to occlusion and outliers, and that it degrades nicely as the segmentation problems increase. Experiments with real images are also presented. Xilin Yi, Octavia I. Camps |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1998 | Hierarchical Organization of Appearance-Based Parts and Relations for Object RecognitionabstractPreviously a new object representation using appearance-based parts and relations to recognize 3D objects from 2D images, in the presence of occlusion and background clutter, was introduced. Appearance-based parts and relations are defined in terms of closed regions and the union of these regions, respectively. The regions are segmented using the MDL principle, and their appearance is obtained from collection of images and compactly represented by parametric manifolds in the eigenspaces spanned by the parts and the relations. In this paper we introduce the discriminatory power of the proposed features and describe how to use it to organize large databases of objects. Octavia I. Camps, Chien-Yuan Huang, Tapas Kanungo |
CVPR | 1 |
| 1998 | 3D Object Depth Recovery from Highlights Using Active Sensor and Illumination ControlabstractTwo approaches for 3D curved object reconstruction using active sensor and illumination control are proposed and compared to each other. In both cases, the highlight information is fully utilized rather than discarded, and knowledge of the object surface is not required. The first approach requires camera control only and recovers shape (depth) from highlights and occluding contours. The second approach requires both camera and illumination control and recovers 3D depth from highlights only. Xilin Yi, Octavia I. Camps |
CVPR | 2 |
| 1998 | Recent Progress in CAD-Based Computer Vision: An Introduction to the Special Issue
Octavia I. Camps, Patrick J. Flynn, George C. Stockman |
Comput. Vis. Image Underst. | 1 |
| 1997 | Object Recognition Using Appearance-Based Parts and RelationsabstractThe recognition of general three-dimensional objects in cluttered scenes is a challenging problem. In particular, the design of a good representation suitable to model large numbers of generic objects that is also robust to occlusion has been a stumbling block in achieving success. In this paper, we propose a representation using appearance-based parts and relations to overcome these problems. Appearance-based parts and relations are defined in terms of closed regions and the union of these regions, respectively. The regions are segmented using the MDL principle, and their appearance is obtained from collection of images and compactly represented by parametric manifolds in the two eigenspaces spanned by the parts and the relations. Chien-Yuan Huang, Octavia I. Camps, Tapas Kanungo |
CVPR | 2 |
| 1997 | Robust Occluding Contour Detection Using the Hausdorff DistanceabstractIn this paper, a correlational approach for distinguishing occluding contours from object markings for 3D object modeling is presented. The proposed method is valid under weak perspective projection, does not require to search for correspondences between frames, can handle scaling between consecutive images. Thus can estimate the full Euclidean surface structure, and does not require camera calibration or camera motion measurement. Extensive experimental results show that the method is robust to the occlusion of feature points and image noise unlike previous affine-based approaches. Qualitative and quantitative results for the relation between the required minimum viewing angle change for the detection and the surface curvature are also presented. Xilin Yi, Octavia I. Camps |
CVPR | 2 |
| 1996 | Detection of obstacles on runway using ego-motion compensation and tracking of significant featuresabstractThe paper proposes a method for obstacle detection on a runway for autonomous navigation and landing of an aircraft. Detection is done in the presence of extraneous features such as tire marks. Suitable features are extracted from the image and warping using approximately known camera and plane parameters is performed in order to compensate ego-motion as far as possible. Residual disparity after warping is estimated using an optical flow algorithm. Features are tracked from frame to frame so as to obtain more reliable estimates of their motion. Corrections are made to motion parameters with the residual disparities using a robust method, and features having large residual disparities are signalled as obstacles. Sensitivity analysis of the procedure is also studied. A Bayesian framework is used at every stage so that the confidence in the estimates can be determined. Tarak Gandhi, Sadashiva Devadiga, Rangachar Kasturi, Octavia I. Camps |
WACV | 4 |
| 1996 | Weighted Parzen Windows for Pattern ClassificationabstractThis paper introduces the weighted-Parzen-window classifier. The proposed technique uses a clustering procedure to find a set of reference vectors and weights which are used to approximate the Parzen-window (kernel-estimator) classifier. The weighted-Parzen-window classifier requires less computation and storage than the full Parzen-window classifier. Experimental results showed that significant savings could be achieved with only minimal, if any, error rate degradation for synthetic and real data sets. Gregory A. Babich, Octavia I. Camps |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1996 | Gray-scale structuring element decompositionabstractEfficient implementation of morphological operations requires the decomposition of structuring elements into the dilation of smaller structuring elements. Zhuang and Haralick (1986) presented a search algorithm to find optimal decompositions of structuring elements in binary morphology. We use the concepts of Top of a set and Umbra of a surface to extend this algorithm to find an optimal decomposition of any arbitrary gray-scale structuring element. Octavia I. Camps, Tapas Kanungo, Robert M. Haralick |
IEEE Trans. Image Process. | 1 |
| 1994 | Robust feature selection for object recognition using uncertain 2D image dataabstractThe use of a small set of features is recurrent in the object recognition literature. If the image data is perfect with no sensor uncertainty and there are not incorrect feature correspondences between the model and the image, then the pose of the object can be computed with no error using these few correspondences. However, in most real cases the noise in the data will propagate into the pose. Moreover, the extent of the effect of the uncertainty will depend on the selection of the correspondences used to compute it. In this paper we address the problem of how to select these correspondences so that the effect of the data uncertainty on the pose estimation is minimized.> Tarak Gandhi, Octavia I. Camps |
CVPR | 2 |
| 1993 | Bayesian view class determinationabstractA Bayesian approach to the view class determination problem is presented. The view classes used contained probabilistic information that takes into account both geometrical and illumination characteristics. The test images match best or second best to the correct view class in approximately 80% of the cases and above 90% of the cases, respectively. The images that fail to be correctly classified correspond to views near the boundaries of the clusters. Even though these views have the same segments as the rest of the views in their class, they look significantly different. This suggests that different definitions of clustering should be studied.> Anjali Pathak, Octavia I. Camps |
CVPR | 2 |
| 1992 | Gray scale structuring element decompositionabstractEfficient implementation of morphological operations requires the decomposition of structuring elements into the dilation of smaller structuring elements. Zhuang and Haralick (1986) presented an algorithm to find optimal decompositions of structuring elements in binary morphology. In this paper the authors extend the algorithm to find optimal structuring element decomposition for gray scale morphology.> Octavia I. Camps, Robert M. Haralick, Tapas Kanungo |
ICPR (3) | 1 |
| 1992 | Object Recognition Using Prediction And Probabilistic MatchabstractPREMIO is a CAD-based object recognition and localization system that uses CAD models of 3D objects and knowledge of lighting and sensors to predict the detectability of features in various views of the object. The predictions that PREMIO produces are powerful new tools in recognizing and determining the pose of a 3D object. In order to take advantage of these tools, we have developed a new matching algorithm: an iterative-deepening-A* search that explicitly takes advantage of the predictions to guide the search and reduce the search space. The purpose of this paper is to describe the matching algorithm and illustrative results. I Introduction Most feature-based matching schemes assume that all the features that are potentially visible in a view of an object will appear with equal probability. The resultant matching algorithms have to allow for “errors” without really understanding what the errors mean. PREMIO [2] is an object recognition/localization system that attempts to model some of the physical processes that can cause these “errors”. It uses CAD models of 3D objects and knowledge of lighting and sensors to predict the detectability of features in various views of the object. From these predictions, PREMIO calculates probabilities for each feature of being detected as a whole, being missed entirely, or breaking into pieces and conditional probabilities of the detection of one feature given the detection or nondetection of other features. The predictions that PREMIO produces are powerful new tools in recognizing and determining the pose of a 3D object. In order to take advantage of these tools, we have developed a new matching algorithm: an iterative-deepening-A* search that explicitly takes advantage of the probabilities to guide the search and prune the tree. The matching algorithm represents a large theoretical effort that is actually independent of the PREMIO system. The algorithm has been implemented as a C program and teated on data specifically generated to fit the abstract paradigm for the probabilistic search. The purpose of this paper is to describe the theory, the algorithm, and illustrative results. Octavia I. Camps, Linda G. Shapiro, Robert M. Haralick |
IROS | 1 |