Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Mario Sznaier

dblp:14/1686 · DBLP profile ↗
← Back
46ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0003-4439-3988ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 43 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
33 papers
Video understanding and tracking · 34% Deep learning architectures and training · 16% Representation and self-supervised learning · 12%
Theoretical computer science
11 papers
Mathematical optimization · 76% Computational complexity · 17% Algorithms and data structures · 7%
Computer graphics and multimedia
7 papers
Rendering · 53% Image and video processing · 44% Geometric modeling and processing · 3%
Databases, data mining, and information retrieval
2 papers
Data mining · 100%

Topics — the 30 heaviest of 93, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
motion segmentation
1.042022
Fast Two-View Motion Segmentation Using Christoffel Polynomials · ECCV (30) 2022
Dynamics Enhanced Multi-camera Motion Segmentation from Unsynchronized Videos · ICCV 2017
GPCA with denoising: A moments-based convex approach · CVPR 2010
Machine learning › Learning theory › generalization bounds
covering number
0.912025
Generalization Error Analysis for Selective State-Space Models Through the Lens of Attention · NeurIPS 2025
Machine learning › Learning theory
generalization bounds
0.912025
Generalization Error Analysis for Selective State-Space Models Through the Lens of Attention · NeurIPS 2025
Machine learning › Deep learning architectures and training
sequence modeling
0.912025
Generalization Error Analysis for Selective State-Space Models Through the Lens of Attention · NeurIPS 2025
Machine learning › Deep learning architectures and training
state space model
0.912025
Generalization Error Analysis for Selective State-Space Models Through the Lens of Attention · NeurIPS 2025
Rendering › gaussian splatting
3d gaussian splatting
0.912025
3D-HGS: 3D Half-Gaussian Splatting · CVPR 2025
Rendering
neural rendering
0.912025
3D-HGS: 3D Half-Gaussian Splatting · CVPR 2025
Image and video processing
jigsaw puzzle solving
0.812024
Solving Masked Jigsaw Puzzles with Diffusion Vision Transformers · CVPR 2024
Data mining
clustering
0.622018
SoS-RSC: A Sum-of-Squares Polynomial Approach to Robustifying Subspace Clustering Algorithms · CVPR 2018
Subspace Clustering with Priors via Sparse Quadratically Constrained Quadratic Programming · CVPR 2016
Data mining › clustering › high-dimensional clustering
subspace clustering
0.622018
SoS-RSC: A Sum-of-Squares Polynomial Approach to Robustifying Subspace Clustering Algorithms · CVPR 2018
Subspace Clustering with Priors via Sparse Quadratically Constrained Quadratic Programming · CVPR 2016
Computer vision › Video understanding and tracking › motion segmentation
two-view motion segmentation
0.612022
Fast Two-View Motion Segmentation Using Christoffel Polynomials · ECCV (30) 2022
Computational complexity
polynomial method
0.612022
Fast Two-View Motion Segmentation Using Christoffel Polynomials · ECCV (30) 2022
Computer vision › Face, body and person analysis
human pose estimation
0.412020
Key Frame Proposal Network for Efficient Pose Estimation in Videos · ECCV (17) 2020
Computer vision › Video understanding and tracking › video summarization
keyframe selection
0.412020
Key Frame Proposal Network for Efficient Pose Estimation in Videos · ECCV (17) 2020
Computer vision › Face, body and person analysis › human pose estimation
video pose estimation
0.412020
Key Frame Proposal Network for Efficient Pose Estimation in Videos · ECCV (17) 2020
Computer vision › Video understanding and tracking
activity recognition
0.432012
Cross-view activity recognition using Hankelets · CVPR 2012
Low order dynamics embedding for high dimensional time series · ICCV 2011
Activity recognition using dynamic subspace angles · CVPR 2011
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.412019
Solving Interpretable Kernel Dimensionality Reduction · NeurIPS 2019
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
kernel dimension reduction
0.412019
Solving Interpretable Kernel Dimensionality Reduction · NeurIPS 2019
Mathematical optimization › continuous optimization
convex optimization
0.422015
A convex optimization approach to robust fundamental matrix estimation · CVPR 2015
Fast algorithms for structured robust principal component analysis · CVPR 2012
Mathematical optimization
convex relaxation
0.332016
Subspace Clustering with Priors via Sparse Quadratically Constrained Quadratic Programming · CVPR 2016
A Model (In)Validation Approach to Gait Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2005
GPCA with denoising: A moments-based convex approach · CVPR 2010
Machine learning › Deep learning architectures and training › neural network layer design › pooling
bilinear pooling
0.312018
MoNet: Moments Embedding Network · CVPR 2018
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.312018
MoNet: Moments Embedding Network · CVPR 2018
Computer vision › Video understanding and tracking
video prediction
0.312018
DYAN: A Dynamical Atoms-Based Network for Video Prediction · ECCV (12) 2018
Data mining
anomaly detection
0.312018
SoS-RSC: A Sum-of-Squares Polynomial Approach to Robustifying Subspace Clustering Algorithms · CVPR 2018
Data mining › anomaly detection
outlier removal
0.312018
SoS-RSC: A Sum-of-Squares Polynomial Approach to Robustifying Subspace Clustering Algorithms · CVPR 2018
Computer vision › Video understanding and tracking › object tracking
occlusion handling
0.342013
Dynamic Context for Tracking behind Occlusions · ECCV (5) 2012
Dynamic subspace-based coordinated multicamera tracking · ICCV 2011
The Way They Move: Tracking Multiple Targets with Similar Appearance · ICCV 2013
Computer vision › 3D vision › multi-view geometry
multi-camera geometry
0.312017
Dynamics Enhanced Multi-camera Motion Segmentation from Unsynchronized Videos · ICCV 2017
Mathematical optimization
semidefinite programming
0.332015
A convex optimization approach to robust fundamental matrix estimation · CVPR 2015
GPCA with denoising: A moments-based convex approach · CVPR 2010
Fast track matching and event detection · CVPR 2008
Computer vision › 3D vision
3d scene reconstruction
0.312025
3D-HGS: 3D Half-Gaussian Splatting · CVPR 2025
Computer vision › Video understanding and tracking › action recognition
3d action recognition
0.212016
Efficient Temporal Sequence Comparison and Classification Using Gram Matrix Embeddings on a Riemannian Manifold · CVPR 2016

Methods — techniques the papers use, named apart from their topics

neural rendering · 1.7gaussian splatting · 1.7motion segmentation · 1.7christoffel polynomials · 1.7spectral abscissa analysis · 0.9covering number analysis · 0.9vision transformer · 0.8diffusion model · 0.8neural interaction inference with potentials · 0.7energy-based modeling · 0.7quadratically constrained quadratic programming · 0.5nuclear norm relaxation · 0.5temporal modeling · 0.4key frame proposal · 0.4sum-of-squares polynomial · 0.3singular value decomposition · 0.3rank minimization · 0.3convex relaxation · 0.3
YearPublicationVenuePosition
2025 3D-HGS: 3D Half-Gaussian Splatting
abstract
Photo-realistic image rendering from 3D scene reconstruction has advanced significantly with neural rendering techniques. Among these, 3D Gaussian Splatting (3D-GS) outperforms Neural Radiance Fields (NeRFs) in quality and speed but struggles with shape and color discontinuities. We propose 3D Half-Gaussian (3D-HGS) kernels as a plug-and-play solution to address these limitations. Our experiments show that 3D-HGS enhances existing 3D-GS methods, achieving state-of-the-art rendering quality without compromising speed. More demos and code are available at https://lihaolin88.github.io/CVPR-2025-3DHGS.
Jinyang Liu 0005, Mario Sznaier, Octavia I. Camps
CVPR3
2025 Generalization Error Analysis for Selective State-Space Models Through the Lens of Attention
abstract
State-space models (SSMs) have recently emerged as a compelling alternative to Transformers for sequence modeling tasks. This paper presents a theoretical generalization analysis of selective SSMs, the core architectural component behind the Mamba model. We derive a novel covering number-based generalization bound for selective SSMs, building upon recent theoretical advances in the analysis of Transformer models. Using this result, we analyze how the spectral abscissa of the continuous-time state matrix influences the model’s stability during training and its ability to generalize across sequence lengths. We empirically validate our findings on a synthetic majority task, the IMDb sentiment classification benchmark, and the ListOps task, demonstrating how our theoretical insights translate into practical model behavior.
Arya Honarpisheh, Mustafa Bozdag, Octavia I. Camps, Mario Sznaier
NeurIPS4
2024 Solving Masked Jigsaw Puzzles with Diffusion Vision Transformers
abstract
Solving image and video jigsaw puzzles poses the chal-lenging task of rearranging image fragments or video frames from unordered sequences to restore meaningful images and video sequences. Existing approaches often hinge on discriminative models tasked with predicting either the absolute positions of puzzle elements or the permutation actions applied to the original data. Unfortunately, these methods face limitations in effectively solving puzzles with a large number of elements. In this paper, we propose JPDVT, an innovative approach that harnesses diffusion transform-ers to address this challenge. Specifically, we generate positional information for image patches or video frames, conditioned on their underlying visual content. This infor-mation is then employed to accurately assemble the puzzle pieces in their correct positions, even in scenarios involving missing pieces. Our method achieves state-of-the-art performance on several datasets.
Jinyang Liu 0005, Wondmgezahu Teshome, Sandesh Ghimire, Mario Sznaier, Octavia I. Camps
CVPR4
2023 Inferring Relational Potentials in Interacting Systems
abstract
Systems consisting of interacting agents are prevalent in the world, ranging from dynamical systems in physics to complex biological networks. To build systems which can interact robustly in the real world, it is thus important to be able to infer the precise interactions governing such systems. Existing approaches typically discover such interactions by explicitly modeling the feed-forward dynamics of the trajectories. In this work, we propose Neural Interaction Inference with Potentials (NIIP) as an alternative approach to discover such interactions that enables greater flexibility in trajectory modeling: it discovers a set of relational potentials, represented as energy functions, which when minimized reconstruct the original trajectory. NIIP assigns low energy to the subset of trajectories which respect the relational constraints observed. We illustrate that with these representations NIIP displays unique capabilities in test-time. First, it allows trajectory manipulation, such as interchanging interaction types across separately trained models, as well as trajectory forecasting. Additionally, it allows adding external hand-crafted potentials at test-time. Finally, NIIP enables the detection of out-of-distribution samples and anomalies without explicit training.
Armand Comas Massague, Yilun Du, Christian Fernandez Lopez, Sandesh Ghimire, Mario Sznaier, Josh Tenenbaum, Octavia I. Camps
ICML5
2022 Fast Two-View Motion Segmentation Using Christoffel Polynomials
Bengisu Özbay, Octavia I. Camps, Mario Sznaier
ECCV (30)3
2022 Dynamics-aware Representation Learning via Multivariate Time Series Transformers
abstract
We propose a novel multivariate time series autoencoder, which produces interpretable linear-dynamical latent features that govern the predictions for several downstream tasks. To this end, we combine a transformer autoencoder with a dynamical atoms-based autoencoder to mimic Koopman operators in the latent space. We demonstrate that our approach significantly outperforms deep Koopman operator learning baselines for time series forecasting on chaotic systems such as the lorenz Attractor. Furthermore, the dynamics-aware representations, combined with a transformer classifier, lead to state-of-the-art classification accuracy on benchmark multivariate time series datasets. Our code is publicly available at https://github.com/mlpotter/T-DYAN-T.
Michael Potter, Ilkay Yildiz, Octavia I. Camps, Mario Sznaier
ESANN4
2020 Key Frame Proposal Network for Efficient Pose Estimation in Videos
Yuexi Zhang, Octavia I. Camps, Mario Sznaier
ECCV (17)4
2020 Dynamic Motion Representation for Human Action Recognition
abstract
Despite the advances in Human Activity Recognition, the ability to exploit the dynamics of human body motion in videos has yet to be achieved. In numerous recent works, researchers have used appearance and motion as independent inputs to infer the action that is taking place in a specific video. In this paper, we highlight that while using a novel representation of human body motion, we can benefit from appearance and motion simultaneously. As a result, better performance of action recognition can be achieved. We start with a pose estimator to extract the location and heat-map of body joints in each frame. We use a dynamic encoder to generate a fixed size representation from these body joint heat-maps. Our experimental results show that training a convolutional neural network with the dynamic motion representation outperforms state-of-the-art action recognition models. By modeling distinguishable activities as distinct dynamical systems and with the help of two stream networks, we obtain the best performance on HMDB, JHMDB, UCF-101, and AVA datasets.
Sadjad Asghari-Esfeden, Mario Sznaier, Octavia I. Camps
WACV2
2019 Solving Interpretable Kernel Dimensionality Reduction
abstract
Kernel dimensionality reduction (KDR) algorithms find a low dimensional representation of the original data by optimizing kernel dependency measures that are capable of capturing nonlinear relationships. The standard strategy is to first map the data into a high dimensional feature space using kernels prior to a projection onto a low dimensional space. While KDR methods can be easily solved by keeping the most dominant eigenvectors of the kernel matrix, its features are no longer easy to interpret. Alternatively, Interpretable KDR (IKDR) is different in that it projects onto a subspace \textit{before} the kernel feature mapping, therefore, the projection matrix can indicate how the original features linearly combine to form the new features. Unfortunately, the IKDR objective requires a non-convex manifold optimization that is difficult to solve and can no longer be solved by eigendecomposition. Recently, an efficient iterative spectral (eigendecomposition) method (ISM) has been proposed for this objective in the context of alternative clustering. However, ISM only provides theoretical guarantees for the Gaussian kernel. This greatly constrains ISM's usage since any kernel method using ISM is now limited to a single kernel. This work extends the theoretical guarantees of ISM to an entire family of kernels, thereby empowering ISM to solve any kernel method of the same objective. In identifying this family, we prove that each kernel within the family has a surrogate $\Phi$ matrix and the optimal projection is formed by its most dominant eigenvectors. With this extension, we establish how a wide range of IKDR applications across different learning paradigms can be solved by ISM. To support reproducible results, the source code is made publicly available on \url{https://github.com/ANONYMIZED}.
Chieh Wu, Jared Miller, Yale Chang, Mario Sznaier, Jennifer G. Dy
NeurIPS4
2018 Iterative Spectral Method for Alternative Clustering
abstract
Given a dataset and an existing clustering as input, alternative clustering aims to find an alternative partition. One of the state-of-the-art approaches is Kernel Dimension Alternative Clustering (KDAC). We propose a novel Iterative Spectral Method (ISM) that greatly improves the scalability of KDAC. Our algorithm is intuitive, relies on easily implementable spectral decompositions, and comes with theoretical guarantees. Its computation time improves upon existing implementations of KDAC by as much as 5 orders of magnitude.
Chieh Wu, Stratis Ioannidis, Mario Sznaier, Xiangyu Li 0006, David R. Kaeli, Jennifer G. Dy
AISTATS3
2018 MoNet: Moments Embedding Network
abstract
Bilinear pooling has been recently proposed as a feature encoding layer, which can be used after the convolutional layers of a deep network, to improve performance in multiple vision tasks. Different from conventional global average pooling or fully connected layer, bilinear pooling gathers 2nd order information in a translation invariant fashion. However, a serious drawback of this family of pooling layers is their dimensionality explosion. Approximate pooling methods with compact properties have been explored towards resolving this weakness. Additionally, recent results have shown that significant performance gains can be achieved by adding 1st order information and applying matrix normalization to regularize unstable higher order information. However, combining compact pooling with matrix normalization and other order information has not been explored until now. In this paper, we unify bilinear pooling and the global Gaussian embedding layers through the empirical moment matrix. In addition, we propose a novel sub-matrix square-root layer, which can be used to normalize the output of the convolution layer directly and mitigate the dimensionality problem with off-the-shelf compact pooling methods. Our experiments on three widely used fine-grained classification datasets illustrate that our proposed architecture, MoNet, can achieve similar or better performance than with the state-of-art G2DeNet. Furthermore, when combined with compact pooling technique, MoNet obtains comparable performance with encoded features with 96% less dimensions.
Mengran Gou, Octavia I. Camps, Mario Sznaier
CVPR4
2018 SoS-RSC: A Sum-of-Squares Polynomial Approach to Robustifying Subspace Clustering Algorithms
abstract
This paper addresses the problem of subspace clustering in the presence of outliers. Typically, this scenario is handled through a regularized optimization, whose computational complexity scales polynomially with the size of the data. Further, the regularization terms need to be manually tuned to achieve optimal performance. To circumvent these difficulties, in this paper we propose an outlier removal algorithm based on evaluating a suitable sum-of-squares polynomial, computed directly from the data. This algorithm only requires performing two singular value decompositions of fixed size, and provides certificates on the probability of misclassifying outliers as inliers.
Mario Sznaier, Octavia I. Camps
CVPR1
2018 DYAN: A Dynamical Atoms-Based Network for Video Prediction
WenQian Liu, Abhishek Sharma 0009, Octavia I. Camps, Mario Sznaier
ECCV (12)4
2017 Dynamics Enhanced Multi-camera Motion Segmentation from Unsynchronized Videos
abstract
This paper considers the multi-camera motion segmentation problem using unsynchronized videos: given two video clips containing several moving objects, captured by unregistered, unsynchronized cameras with different viewpoints, our goal is to assign features to moving objects in the scene. This problem challenges existing methods, due to the lack of registration information and correspondences across cameras. To solve it, we propose a method that combines shape and dynamical information and does not require spatiotemporal registration or shared features. As shown in the paper, this combination results in improved performance even in the single camera case, and allows for solving the multi-camera segmentation problem with a computational cost similar to that of existing single-view techniques.
Xikang Zhang, Bengisu Özbay, Mario Sznaier, Octavia I. Camps
ICCV3
2016 Person Re-identification in Appearance Impaired Scenarios
Mengran Gou, Xikang Zhang, Angels Rates-Borras, Sadjad Asghari-Esfeden, Octavia I. Camps, Mario Sznaier
BMVC6
2016 Subspace Clustering with Priors via Sparse Quadratically Constrained Quadratic Programming
abstract
This paper considers the problem of recovering a subspace arrangement from noisy samples, potentially corrupted with outliers. Our main result shows that this problem can be formulated as a convex semi-definite optimization problem subject to an additional rank constrain that involves only a very small number of variables. This is established by first reducing the problem to a quadratically constrained quadratic problem and then using its special structure to find conditions guaranteeing that a suitably built convex relaxation is indeed exact. When combined with the standard nuclear norm relaxation for rank, the results above lead to computationally efficient algorithms with optimality guarantees. A salient feature of the proposed approach is its ability to incorporate existing a-priori information about the noise, co-ocurrences, and percentage of outliers. These results are illustrated with several examples.
Yongfang Cheng, Mario Sznaier, Octavia I. Camps
CVPR3
2016 Solving Temporal Puzzles
abstract
Many physical phenomena, within short time windows, can be explained by low order differential relations. In a discrete world, these relations can be described using low order difference equations or equivalently low order auto regressive (AR) models. In this paper, based on this intuition, we propose an algorithm for solving time-sort temporal puzzles, defined as scrambled time series that need to be sorted out. We frame this problem using a mixed-integer semi definite programming formulation and show how to turn it into a mixed-integer linear programming problem, which can be solved with off-the-shelf solvers, by using the recently introduced atomic norm framework. Our experiments show the effectiveness and generality of our approach in different scenarios.
Caglayan Dicle, Burak Yilmaz, Octavia I. Camps, Mario Sznaier
CVPR4
2016 Efficient Temporal Sequence Comparison and Classification Using Gram Matrix Embeddings on a Riemannian Manifold
abstract
In this paper we propose a new framework to compare and classify temporal sequences. The proposed approach captures the underlying dynamics of the data while avoiding expensive estimation procedures, making it suitable to process large numbers of sequences. The main idea is to first embed the sequences into a Riemannian manifold by using positive definite regularized Gram matrices of their Hankelets. The advantages of the this approach are: 1) it allows for using non-Euclidean similarity functions on the Positive Definite matrix manifold, which capture better the underlying geometry than directly comparing the sequences or their Hankel matrices, and 2) Gram matrices inherit desirable properties from the underlying Hankel matrices: their rank measure the complexity of the underlying dynamics, and the order and coefficients of the associated regressive models are invariant to affine transformations and varying initial conditions. The benefits of this approach are illustrated with extensive experiments in 3D action recognition using 3D joints sequences. In spite of its simplicity, the performance of this approach is competitive or better than using state-of-art approaches for this problem. Further, these results hold across a variety of metrics, supporting the idea that the improvement stems from the embedding itself, rather than from using one of these metrics.
Xikang Zhang, Mengran Gou, Mario Sznaier, Octavia I. Camps
CVPR4
2016 Jensen Bregman LogDet Divergence Optimal Filtering in the Manifold of Positive Definite Matrices
Octavia I. Camps, Mario Sznaier, Biel Roig-Solvas
ECCV (7)3
2015 A convex optimization approach to robust fundamental matrix estimation
abstract
This paper considers the problem of estimating the fundamental matrix from corrupted point correspondences. A general nonconvex framework is proposed that explicitly takes into account the rank-2 constraint on the fundamental matrix and the presence of noise and outliers. The main result of the paper shows that this non-convex problem can be solved by solving a sequence of convex semi-definite programs, obtained by exploiting a combination of polynomial optimization tools and rank minimization techniques. Further, the algorithm can be easily extended to handle the case where only some of the correspondences are labeled, and, to exploit co-ocurrence information, if available. Consistent experiments show that the proposed method works well, even in scenarios characterized by a very high percentage of outliers.
Yongfang Cheng, Jose A. Lopez, Octavia I. Camps, Mario Sznaier
CVPR4
2015 Self Scaled Regularized Robust Regression
abstract
Linear Robust Regression (LRR) seeks to find the parameters of a linear mapping from noisy data corrupted from outliers, such that the number of inliers (i.e. pairs of points where the fitting error of the model is less than a given bound) is maximized. While this problem is known to be NP hard, several tractable relaxations have been recently proposed along with theoretical conditions guaranteeing exact recovery of the parameters of the model. However, these relaxations may perform poorly in cases where the fitting error for the outliers is large. In addition, these approaches cannot exploit available a-priori information, such as cooccurrences. To circumvent these difficulties, in this paper we present an alternative approach to robust regression. Our main result shows that this approach is equivalent to a “self-scaled” ℓ1regularized robust regression problem, where the cost function is automatically scaled, with scalings that depend on the a-priori information. Thus, the proposed approach achieves substantially better performance than traditional regularized approaches in cases where the outliers are far from the linear manifold spanned by the inliers, while at the same time exhibits the same theoretical recovery properties. These results are illustrated with several application examples using both synthetic and real data.
Caglayan Dicle, Mario Sznaier, Octavia I. Camps
CVPR3
2014 Person Re-Identification Using Kernel-Based Metric Learning Methods
Mengran Gou, Octavia I. Camps, Mario Sznaier
ECCV (7)4
2013 Finding Causal Interactions in Video Sequences
abstract
This paper considers the problem of detecting causal interactions in video clips. Specifically, the goal is to detect whether the actions of a given target can be explained in terms of the past actions of a collection of other agents. We propose to solve this problem by recasting it into a directed graph topology identification, where each node corresponds to the observed motion of a given target, and each link indicates the presence of a causal correlation. As shown in the paper, this leads to a block-sparsification problem that can be efficiently solved using a modified Group-Lasso type approach, capable of handling missing data and outliers (due for instance to occlusion and mis-identified correspondences). Moreover, this approach also identifies time instants where the interactions between agents change, thus providing event detection capabilities. These results are illustrated with several examples involving non-trivial interactions amongst several human subjects.
Mustafa Ayazoglu, Burak Yilmaz, Mario Sznaier, Octavia I. Camps
ICCV3
2013 The Way They Move: Tracking Multiple Targets with Similar Appearance
abstract
We introduce a computationally efficient algorithm for multi-object tracking by detection that addresses four main challenges: appearance similarity among targets, missing data due to targets being out of the field of view or occluded behind other objects, crossing trajectories, and camera motion. The proposed method uses motion dynamics as a cue to distinguish targets with similar appearance, minimize target mis-identification and recover missing data. Computational efficiency is achieved by using a Generalized Linear Assignment (GLA) coupled with efficient procedures to recover missing data and estimate the complexity of the underlying dynamics. The proposed approach works with track lets of arbitrary length and does not assume a dynamical model a priori, yet it captures the overall motion dynamics of the targets. Experiments using challenging videos show that this framework can handle complex target motions, non-stationary cameras and long occlusions, on scenarios where appearance cues are not available or poor.
Caglayan Dicle, Octavia I. Camps, Mario Sznaier
ICCV3
2012 Fast algorithms for structured robust principal component analysis
abstract
A large number of problems arising in computer vision can be reduced to the problem of minimizing the nuclear norm of a matrix, subject to additional structural and sparsity constraints on its elements. Examples of relevant applications include, among others, robust tracking in the presence of outliers, manifold embedding, event detection, in-painting and tracklet matching across occlusion. In principle, these problems can be reduced to a convex semi-definite optimization form and solved using interior point methods. However, the poor scaling properties of these methods limit the use of this approach to relatively small sized problems. The main result of this paper shows that structured nuclear norm minimization problems can be efficiently solved by using an iterative Augmented Lagrangian Type (ALM) method that only requires performing at each iteration a combination of matrix thresholding and matrix inversion steps. As we illustrate in the paper with several examples, the proposed algorithm results in a substantial reduction of computational time and memory requirements when compared against interior-point methods, opening up the possibility of solving realistic, large sized problems.
Mustafa Ayazoglu, Mario Sznaier, Octavia I. Camps
CVPR2
2012 Cross-view activity recognition using Hankelets
abstract
Human activity recognition is central to many practical applications, ranging from visual surveillance to gaming interfacing. Most approaches addressing this problem are based on localized spatio-temporal features that can vary significantly when the viewpoint changes. As a result, their performances rapidly deteriorate as the difference between the viewpoints of the training and testing data increases. In this paper, we introduce a new type of feature, the “Hankelet” that captures dynamic properties of short tracklets. While Hankelets do not carry any spatial information, they bring invariant properties to changes in viewpoint that allow for robust cross-view activity recognition, i.e. when actions are recognized using a classifier trained on data from a different viewpoint. Our experiments on the IXMAS dataset show that using Hanklets improves the state of the art performance by over 20%.
Binlong Li, Octavia I. Camps, Mario Sznaier
CVPR3
2012 Dynamic Context for Tracking behind Occlusions
Octavia I. Camps, Mario Sznaier
ECCV (5)3
2011 Activity recognition using dynamic subspace angles
abstract
Cameras are ubiquitous everywhere and hold the promise of significantly changing the way we live and interact with our environment. Human activity recognition is central to understanding dynamic scenes for applications ranging from security surveillance, to assisted living for the elderly, to video gaming without controllers. Most current approaches to solve this problem are based in the use of local temporal-spatial features that limit their ability to recognize long and complex actions. In this paper, we propose a new approach to exploit the temporal information encoded in the data. The main idea is to model activities as the output of unknown dynamic systems evolving from unknown initial conditions. Under this framework, we show that activity videos can be compared by computing the principal angles between subspaces representing activity types which are found by a simple SVD of the experimental data. The proposed approach outperforms state-of-the-art methods classifying activities in the KTH dataset as well as in much more complex scenarios involving interacting actors.
Binlong Li, Mustafa Ayazoglu, Teresa Mao, Octavia I. Camps, Mario Sznaier
CVPR5
2011 Dynamic subspace-based coordinated multicamera tracking
abstract
This paper considers the problem of sustained multicamera tracking in the presence of occlusion and changes in the target motion model. The key insight of the proposed method is the fact that, under mild conditions, the 2D trajectories of the target in the image planes of each of the cameras are constrained to evolve in the same subspace. This observation allows for identifying, at each time instant, a single (piecewise) linear model that explains all the available 2D measurements. In turn, this model can be used in the context of a modified particle filter to predict future target locations. In the case where the target is occluded to some of the cameras, the missing measurements can be estimated using the facts that they must lie both in the subspace spanned by previous measurements and satisfy epipolar constraints. Hence, by exploiting both dynamical and geometrical constraints the proposed method can robustly handle substantial occlusion, without the need for performing 3D reconstruction, calibrated cameras or constraints on sensor separation. The performance of the proposed tracker is illustrated with several challenging examples involving targets that substantially change appearance and motion models while occluded to some of the cameras.
Mustafa Ayazoglu, Binlong Li, Caglayan Dicle, Mario Sznaier, Octavia I. Camps
ICCV4
2011 Low order dynamics embedding for high dimensional time series
abstract
This paper considers the problem of finding low order nonlinear embeddings of dynamic data, that is, data characterized by a temporal ordering. Examples where this problem arises include, among others, appearance-based multi-frame tracking, activity recognition from video and dynamic texture analysis/synthesis. Our main result is a semi-definite programming based manifold embedding algorithm that exploits the dynamical information encapsulated in the temporal ordering of the sequence to obtain embeddings with lower complexity that those produced by existing algorithms, at a comparable computational cost. In addition, the use of spatio-temporal information allows for minimizing the effects of outliers on the manifold structure and for handling fragmented sequences, where some of the data is missing, for instance due to occlusion.
Octavia I. Camps, Mario Sznaier
ICCV3
2010 GPCA with denoising: A moments-based convex approach
abstract
This paper addresses the problem of segmenting a combination of linear subspaces and quadratic surfaces from sample data points corrupted by (not necessarily small) noise. Our main result shows that this problem can be reduced to minimizing the rank of a matrix whose entries are affine in the optimization variables, subject to a convex constraint imposing that these variables are the moments of an (unknown) probability distribution function with finite support. Exploiting the linear matrix inequality based characterization of the moments problem and appealing to well known convex relaxations of rank leads to an overall semi-definite optimization problem. We apply our method to problems such as simultaneous 2D motion segmentation and motion segmentation from two perspective views and illustrate that our formulation substantially reduces the noise sensitivity of existing approaches.
Necmiye Ozay, Mario Sznaier, Constantino M. Lagoa, Octavia I. Camps
CVPR2
2010 Euclidean Structure Recovery from Motion in Perspective Image Sequences via Hankel Rank Minimization
Mustafa Ayazoglu, Mario Sznaier, Octavia I. Camps
ECCV (2)2
2008 Fast track matching and event detection
abstract
This paper addresses the problems of track stitching and dynamic event detection in a sequence of frames. The input data consists of tracks, possibly fragmented due to occlusion, belonging to multiple targets. The goals are to (i) establish track identity across occlusion, and (ii) detect points where the motion of these targets undergo substantial changes. The main result of the paper is a simple, computationally inexpensive approach that achieves these goals in a unified way. Given a continuous track, the main idea is to detect changes in the dynamics by parsing it into segments according to the complexity of the model required to explain the observed data. Intuitively, changes in this complexity correspond to points where the dynamics change. Since the problem of estimating the complexity of the underlying model can be reduced to estimating the rank of a matrix constructed from the observed data, these changes can be found with a simple algorithm, computationally no more expensive that a sequence of SVDs. Proceeding along the same lines, fragmented tracks corresponding to multiple targets can be linked by searching for sets corresponding to minimal complexity joint models. As we show in the paper, this problem can be reduced to a semi-definite optimization and efficiently solved with commonly available software.
Mario Sznaier, Octavia I. Camps
CVPR2
2008 Sequential sparsification for change detection
abstract
This paper presents a general method for segmenting a vector valued sequence into an unknown number of subsequences where all data points from a subsequence can be represented with the same affine parametric model. The idea is to cluster the data into the minimum number of such subsequences which, as we show, can be cast as a sparse signal recovery problem by exploiting the temporal correlation between consecutive data points. We try to maximize the sparsity (i.e. the number of zero elements) of the first order differences of the sequence of parameter vectors. Each non-zero element in the first order difference sequence corresponds to a change. A weighted l1norm based convex approximation is adopted to solve the change detection problem. We apply the proposed method to video segmentation and temporal segmentation of dynamic textures.
Necmiye Ozay, Mario Sznaier, Octavia I. Camps
CVPR2
2007 A Rank Minimization Approach to Video Inpainting
abstract
This paper addresses the problem of video inpainting, that is seamlessly reconstructing missing portions in a set of video frames. We propose to solve this problem proceeding as follows: (i) finding a set of descriptors that encapsulate the information necessary to reconstruct a frame, (ii) finding an optimal estimate of the value of these descriptors for the missing/corrupted frames, and (iii) using the estimated values to reconstruct the frames. The main result of the paper shows that the optimal descriptor estimates can be efficiently obtained by minimizing the rank of a matrix directly constructed from the available data, leading to a simple, computationally attractive, dynamic inpainting algorithm that optimizes the use of spatio/temporal information. Moreover, contrary to most currently available techniques, the method can handle non-periodic target motions, non-stationary backgrounds and moving cameras. These results are illustrated with several examples, including reconstructing dynamic textures and object disocclusion in cases involving both moving targets and camera.
Mario Sznaier, Octavia I. Camps
ICCV2
2006 Dynamic Appearance Modeling for Human Tracking
abstract
Dynamic appearance is one of the most important cues for tracking and identifying moving people. However, direct modeling spatio-temporal variations of such appearance is often a difficult problem due to their high dimensionality and nonlinearities. In this paper we present a human tracking system that uses a dynamic appearance and motion modeling framework based on the use of robust system dynamics identification and nonlinear dimensionality reduction techniques. The proposed system learns dynamic appearance and motion models from a small set of initial frames and does not require prior knowledge such as gender or type of activity. The advantages of the proposed tracking system are illustrated with several examples where the learned dynamics accurately predict the location and appearance of the targets in future frames, preventing tracking failures due to model drifting, target occlusion and scene clutter.
Hwasup Lim, Octavia I. Camps, Mario Sznaier, Vlad I. Morariu
CVPR (1)3
2006 Dynamics Based Robust Motion Segmentation
abstract
In this paper we consider the problem of segmenting multiple rigid motions using multi-frame point correspondence data. The main idea of the method is to group points according to the complexity of the model required to explain their relative motion. Intuitively, this formalizes the idea that points on the same rigid share more modes of motion (for instance a common translation or rotation) than points on different objects, leading to less complex models. By exploiting results from systems theory, the problem of estimating the complexity of the underlying model is reduced to simply computing the rank of a matrix constructed from the correspondence data. This leads to a simple segmentation algorithm, computationally no more expensive than a sequence of SVDs. Since the proposed method exploits both spatial and temporal constraints, is less sensitive to the effect of noise or outliers than approaches that rely solely on factorizations of the measurements matrix. In addition, the method can also naturally handle "degenerate cases", e.g. cases where the objects partially share motion modes. These results are illustrated using several examples involving both degenerate and non-degenerate cases.
Roberto Lublinerman, Mario Sznaier, Octavia I. Camps
CVPR (1)2
2005 A Caratheodory-Fejer Approach to Dynamic Appearance Modeling
abstract
This paper presents a technique to learn dynamic appearance models from a small number of training frames. Under this framework, dynamic appearance is modelled as an unknown operator that satisfies certain interpolation conditions and that can be efficiently identified using very little a priori information with off the shelf software. The advantages of the proposed method are illustrated with several examples where the leaned dynamics accurately predict the appearance of the targets preventing tracking failures due to occlusion or clutter.
Hwasup Lim, Octavia I. Camps, Mario Sznaier
CVPR (1)3
2005 Robust Structure from Motion and Identified Dynamics
abstract
This paper addresses robust recovery of structure and motion for rigid bodies in video sequences. For small inter-frame motion, feature appearance is commonly estimated according to the previous frame. However, in the case of occlusion or corrupt frames, the small interframe motion model fails. In this paper, we propose to use robust identification techniques to estimate the motion dynamics based on a set of previous frames. These dynamics are then recursively used to robustly estimate object structure and motion and to predict the object appearance in future frames. Results are tested for 3D reconstructions of rigid bodies in real and synthetic image sequences.
Benjamin R. Fransen, Octavia I. Camps, Mario Sznaier
ICCV3
2005 A Model (In)Validation Approach to Gait Classification
abstract
This paper addresses the problem of human gait classification from a robust model (in)validation perspective. The main idea is to associate to each class of gaits a nominal model, subject to bounded uncertainty and measurement noise. In this context, the problem of recognizing an activity from a sequence of frames can be formulated as the problem of determining whether this sequence could have been generated by a given (model, uncertainty, and noise) triple. By exploiting interpolation theory, this problem can be recast into a nonconvex optimization. In order to efficiently solve it, we propose two convex relaxations, one deterministic and one stochastic. As we illustrate experimentally, these relaxations achieve over 83 percent and 86 percent success rates, respectively, even in the face of noisy data.
Maria Cecilia Mazzaro, Mario Sznaier, Octavia I. Camps
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Segmentation for robust tracking in the presence of severe occlusion
abstract
Tracking an object in a sequence of images can fail due to partial occlusion or clutter. Robustness to occlusion can be increased by tracking the object as a set of "parts" such that not all of these are occluded at the same time. However, successful implementation of this idea hinges upon finding a suitable set of parts. In this paper we propose a novel segmentation, specifically designed to improve robustness against occlusion in the context of tracking. The main result shows that tracking the parts resulting from this segmentation outperforms both tracking parts obtained through traditional segmentations, and tracking the entire target. Additional results include a statistical analysis of the correlation between features of a part and tracking error, and identifying a cost function that exhibits a high degree of correlation with the tracking error.
Camillo Gentile, Octavia I. Camps, Mario Sznaier
IEEE Trans. Image Process.3
2003 A Caratheodory-Fejer Approach to Robust Multiframe Tracking
abstract
A requirement common to most dynamic vision applications is the ability to track objects in a sequence of frames. This problem has been extensively studied in the past few years, leading to several techniques, such as unscented particle filter based trackers, that exploit a combination of the (assumed) target dynamics, empirically learned noise distributions and past position observations. While successful in many scenarios, these trackers remain fragile to occlusion and model uncertainty in the target dynamics. As we show in this paper, these difficulties can be addressed by modeling the dynamics of the target as an unknown operator that satisfies certain interpolation conditions. Results from interpolation theory can then be used to find this operator by solving a convex optimization problem. As illustrated with several examples, combining this operator with Kalman and UPF techniques leads to both robustness improvement and computational complexity reduction.
Octavia I. Camps, Hwasup Lim, Maria Cecilia Mazzaro, Mario Sznaier
ICCV4
2001 Segmentation for Robust Tracking in the Presence of Severe Occlusion
abstract
Tracking an object in a sequence of images can fail due to partial occlusion or clutter Robustness can be increased by tracking a set of "parts", provided that a suitable set can be identified. In this paper we propose a novel segmentation, specifically designed to improve robustness against occlusion in the context of tracking. The main result shows that tracking the parts resulting from this segmentation outperforms both tracking parts obtained through traditional segmentations, and tracking the entire target. Additional results include a statistical analysis of the correlation between features of a part and tracking error, and identifying a cost function highly correlated with the tracking error.
Camillo Gentile, Octavia I. Camps, Mario Sznaier
CVPR (2)3
2001 An improved Voronoi-diagram-based neural net for pattern classification
abstract
We propose a novel two-layer neural network to answer a point query in R(n) which is partitioned into polyhedral regions; such a task solves among others nearest neighbor clustering. As in previous approaches to the problem, our design is based on the use of Voronoi diagrams. However, our approach results in substantial reduction of the number of neurons, completely eliminating the second layer, at the price of requiring only two additional clock steps. In addition, the design process is also simplified while retaining the main advantage of the approach, namely its ability to furnish precise values for the number of neurons and the connection weights necessitating neither trial and error type iterations nor ad hoc parameters.
Camillo Gentile, Mario Sznaier
IEEE Trans. Neural Networks2
1999 An improved Voronoi-diagram based neural net for pattern classification
abstract
We propose a novel two-layer neural network to answer a point query in R/sup n/ which is partitioned into polyhedral regions. Such a task solves amongst others nearest neighbor clustering. As in previous approaches to the problem, our design is based on the use of Voronoi diagrams. However, our approach results in substantial reduction of the the number of neurons, completely eliminating the middle layer at the price of requiring only one additional clock step. In addition, the design process is also simplified while retaining the main advantage of the approach, namely its ability to furnish precise values for the number of neurons and the connection weights requiring neither trial and error type iterations nor ad-hoc parameters.
Camillo Gentile, Mario Sznaier
IJCNN2
1989 An adaptive controller for a one-legged mobile robot
abstract
An adaptive controller based upon the online minimization of a performance criterion is described. The adaptive controller is used to improve the performance of a one-legged mobile robot, removing problems experienced with previous controllers. Specifically, this controller eliminates the problem of a bias in the forward velocity experienced with the nonadaptive controller and, at the same time, provides the flexibility required to allow a dynamically stabilized legged machine to perform satisfactorily under the widely varying conditions that exist in the real world. The performance of several minimization algorithms is analyzed, and the adaptive step size random search algorithm is selected. Finally, a series of experiments illustrating the ability of the adaptive controller to handle a changing environment is presented.>
Mario Sznaier, Mark J. Damborg
IEEE Trans. Robotics Autom.1