Jeffrey Ho

dblp:23/1866 · DBLP profile ↗
← Back
44ranked-venue papers
9as first author
0since 2021 · last 2017
0000-0001-8038-1605ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 7 first-authorGraphics, computer vision, multimedia, augmented reality and games · 30 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 4Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
17 papers
Representation and self-supervised learning · 32% 3D vision · 26% Face, body and person analysis · 17%
Computer graphics and multimedia
12 papers
Visual content generation and editing · 36% Geometric modeling and processing · 26% Image and video processing · 18%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%
Theoretical computer science
2 papers
Algorithms and data structures · 85% Computational geometry · 15%

Topics — the 30 heaviest of 53, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding
0.532013
On A Nonlinear Generalization of Sparse Coding and Dictionary Learning · ICML (3) 2013
Affine-Constrained Group Sparse Coding and Its Application to Image-Based Classifications · ICCV 2013
Block and Group Regularized Sparse Modeling for Dictionary Learning · CVPR 2013
Visual content generation and editing
image retargeting
0.422016
CASAIR: Content and Shape-Aware Image Retargeting and Its Applications · IEEE Trans. Image Process. 2016
Seam Segment Carving: Retargeting Images to Irregularly-Shaped Image Domains · ECCV (6) 2012
Visual content generation and editing › image retargeting
seam carving
0.422016
CASAIR: Content and Shape-Aware Image Retargeting and Its Applications · IEEE Trans. Image Process. 2016
Seam Segment Carving: Retargeting Images to Irregularly-Shaped Image Domains · ECCV (6) 2012
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning
0.322013
On A Nonlinear Generalization of Sparse Coding and Dictionary Learning · ICML (3) 2013
Block and Group Regularized Sparse Modeling for Dictionary Learning · CVPR 2013
Computer vision › Image recognition and object detection › medical image analysis
medical image classification
0.212016
Covariant Image Representation with Applications to Classification Problems in Medical Imaging · Int. J. Comput. Vis. 2016
Computer vision › 3D vision › image registration
affine registration
0.222011
Higher-Dimensional Affine Registration and Vision Applications · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Higher Dimensional Affine Registration and Vision Applications · ECCV (4) 2008
Computer vision › Face, body and person analysis
face recognition
0.242006
Face Recognition Using 3-D Models: Pose and Illumination · Proc. IEEE 2006
Acquiring Linear Subspaces for Face Recognition under Variable Lighting · IEEE Trans. Pattern Anal. Mach. Intell. 2005
Video-Based Face Recognition Using Probabilistic Appearance Manifolds · CVPR (1) 2003
Computer vision › Video understanding and tracking
object tracking
0.232011
Visual tracking with histograms and articulating blocks · CVPR 2008
Visual Tracking Using Learned Linear Subspaces · CVPR (1) 2004
Online Sparse Gaussian Process Regression and Its Applications · IEEE Trans. Image Process. 2011
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
group sparse coding
0.212013
Affine-Constrained Group Sparse Coding and Its Application to Image-Based Classifications · ICCV 2013
Computer vision › Face, body and person analysis › face recognition
sparse representation-based classification
0.212013
Affine-Constrained Group Sparse Coding and Its Application to Image-Based Classifications · ICCV 2013
Information retrieval › image retrieval
hashing-based image retrieval
0.212013
Recursive Estimation of the Stein Center of SPD Matrices and Its Applications · ICCV 2013
Information retrieval › image retrieval
image indexing
0.212013
Recursive Estimation of the Stein Center of SPD Matrices and Its Applications · ICCV 2013
Computer animation and physical simulation › motion synthesis
motion interpolation
0.212013
Computing Diffeomorphic Paths for Large Motion Interpolation · CVPR 2013
Geometric modeling and processing
shape analysis
0.212013
Computing Diffeomorphic Paths for Large Motion Interpolation · CVPR 2013
Algorithms and data structures
clustering
0.212013
Recursive Estimation of the Stein Center of SPD Matrices and Its Applications · ICCV 2013
Image and video processing
image registration
0.222008
Higher Dimensional Affine Registration and Vision Applications · ECCV (4) 2008
USSR: A Unified Framework for Simultaneous Smoothing, Segmentation, and Registration of Multiple Images · ICCV 2007
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
gaussian process regression
0.112011
Online Sparse Gaussian Process Regression and Its Applications · IEEE Trans. Image Process. 2011
Computer vision › 3D vision › point cloud registration
point set registration
0.112011
Higher-Dimensional Affine Registration and Vision Applications · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Image and video processing › image registration
affine transform estimation
0.112009
An algebraic approach to affine registration of point sets · ICCV 2009
Geometric modeling and processing
point set registration
0.112009
An algebraic approach to affine registration of point sets · ICCV 2009
Computer vision › Face, body and person analysis › face recognition › robust face recognition
illumination-invariant face recognition
0.122005
Acquiring Linear Subspaces for Face Recognition under Variable Lighting · IEEE Trans. Pattern Anal. Mach. Intell. 2005
Nine Points of Light: Acquiring Subspaces for Face Recognition under Variable Lighting · CVPR (1) 2001
Computer vision › Video understanding and tracking › object tracking
appearance modeling
0.112008
Visual tracking with histograms and articulating blocks · CVPR 2008
Computer vision › Video understanding and tracking › object tracking
articulated object tracking
0.112008
Visual tracking with histograms and articulating blocks · CVPR 2008
Computer vision › 3D vision › stereo vision
stereo matching
0.122011
Binocular Helmholtz Stereopsis · ICCV 2003
Higher-Dimensional Affine Registration and Vision Applications · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Medical and health informatics
medical imaging
0.112016
Covariant Image Representation with Applications to Classification Problems in Medical Imaging · Int. J. Comput. Vis. 2016
Image and video processing
image segmentation
0.112007
USSR: A Unified Framework for Simultaneous Smoothing, Segmentation, and Registration of Multiple Images · ICCV 2007
Image and video processing › image filtering
image smoothing
0.112007
USSR: A Unified Framework for Simultaneous Smoothing, Segmentation, and Registration of Multiple Images · ICCV 2007
Computer vision › 3D vision › 3d face modeling
3d morphable model
0.112006
Face Recognition Using 3-D Models: Pose and Illumination · Proc. IEEE 2006
Geometric modeling and processing › surface reconstruction
normal integration
0.112006
Integrating Surface Normal Vectors Using Fast Marching Method · ECCV (3) 2006
Computational photography and imaging
photometric stereo
0.112005
Passive Photometric Stereo from Motion · ICCV 2005

Methods — techniques the papers use, named apart from their topics

seam segment carving · 0.5covariant representation · 0.5cost function optimization · 0.5stein distance · 0.5recursive estimation · 0.5logdet divergence · 0.5algebraic geometry · 0.2riemannian geometry · 0.2quotient space projection · 0.2quadratic programming · 0.2intra-block coherence suppression · 0.2dictionary learning · 0.2block-gradient descent · 0.2block coordinate descent · 0.2augmented lagrangian · 0.2affine constraint · 0.2seam carving · 0.1karcher mean · 0.1
YearPublicationVenuePosition
2017 High-Level Concepts for Affective Understanding of Images
abstract
This paper aims to bridge the affective gap between image content and the emotional response of the viewer it elicits by using High-Level Concepts (HLCs). In contrast to previous work that relied solely on low-level features or used convolutional neural network (CNN) as a blackbox, we use HLCs generated by pretrained CNNs in an explicit way to investigate the relations/associations between these HLCs and a (small) set of Ekman's emotional classes. As a proof-of-concept, we first propose a linear admixture model for modeling these relations, and the resulting computational framework allows us to determine the associations between each emotion class and certain HLCs (objects and places). This linear model is further extended to a nonlinear model using support vector regression (SVR) that aims to predict the viewer's emotional response using both low-level image features and HLCs extracted from images. These class-specific regressors are then assembled into a regressor ensemble that provide a flexible and effective predictor for predicting viewer's emotional responses from images. Experimental results have demonstrated that our results are comparable to existing methods, with a clear view of the association between HLCs and emotional classes that is ostensibly missing in most existing work.
Afsheen Rafaqat Ali, Usman Shahid, Mohsen Ali, Jeffrey Ho
WACV4
2016 Covariant Image Representation with Applications to Classification Problems in Medical Imaging
Dohyung Seo, Jeffrey Ho, Baba C. Vemuri
Int. J. Comput. Vis.2
2016 CASAIR: Content and Shape-Aware Image Retargeting and Its Applications
abstract
This paper proposes a novel image-retargeting algorithm that can retarget images to a large family of non-rectangular shapes. Specifically, we study image retargeting from a broader perspective that includes the content as well as the shape of an image, and the proposed content and shape-aware image-retargeting (CASAIR) algorithm is driven by the dual objectives of image content preservation and image domain transformation, with the latter defined by an application-specific target shape. The algorithm is based on the idea of seam segment carving that successively removes low-cost seam segments from the image to simultaneously achieve the two objectives, with the selection of seam segments determined by a cost function incorporating inputs from image content and target shape. To provide a complete characterization of shapes that can be obtained using CASAIR, we introduce the notion of bhv-convex shapes, and we show that bhv-convex shapes are precisely the family of shapes that can be retargeted to by CASAIR. The proposed algorithm is simple in both its design and implementation, and in practice, it offers an efficient and effective retargeting platform that provides its users with considerable flexibility in choosing target shapes. To demonstrate the potential of CASAIR for broadening the application scope of image retargeting, this paper also proposes a smart camera-projector system that incorporates CASAIR. In the context of ubiquitous display, CASAIR equips the camera-projector system with the capability of retargeting images online in order to maximize the quality and fidelity of the displayed images whenever the situation demands.
Shaoyu Qi, Yu-Tseh Chi, Adrian M. Peter, Jeffrey Ho
IEEE Trans. Image Process.4
2014 Deconstructing Binary Classifiers in Computer Vision
Mohsen Ali, Jeffrey Ho
ACCV (3)2
2014 Deconstructing Kernel Machines
Mohsen Ali, Muhammad Ali Rushdi 0001, Jeffrey Ho
ECML/PKDD (1)3
2013 Recursive Karcher Expectation Estimators And Geometric Law of Large Numbers
abstract
This paper studies a form of law of large numbers on Pn, the space of nxn symmetric positive-definite matrices equipped with Fisher-Rao metric. Specifically, we propose a recursive algorithm for estimating the Karcher expectation of an arbitrary distribution defined on Pn, and we show that the estimates computed by the recursive algorithm asymptotically converge in probability to the correct Karcher expectation. The steps in the recursive algorithm mainly consist of making appropriate moves on geodesics in Pn, and the algorithm is simple to implement and it offers a tremendous gain in computation time of several orders in magnitude over existing non-recursive algorithms. We elucidate the connection between the more familiar law of large numbers for real-valued random variables and the asymptotic convergence of the proposed recursive algorithm, and our result provides an example of a new form of law of large numbers for random variables taking values in a Riemannian manifold. From the practical side, the computation of the mean of a collection of symmetric positive-definite (SPD) matrices is a fundamental ingredient in many algorithms in machine learning, computer vision and medical imaging applications. We report an experiment using the proposed recursive algorithm for K-means clustering, demonstrating the algorithm’s efficiency, accuracy and stability.
Jeffrey Ho, Guang Cheng 0002, Hesamoddin Salehian, Baba C. Vemuri
AISTATS1
2013 Block and Group Regularized Sparse Modeling for Dictionary Learning
abstract
This paper proposes a dictionary learning framework that combines the proposed block/group (BGSC) or reconstructed block/group (R-BGSC) sparse coding schemes with the novel Intra-block Coherence Suppression Dictionary Learning algorithm. An important and distinguishing feature of the proposed framework is that all dictionary blocks are trained simultaneously with respect to each data group while the intra-block coherence being explicitly minimized as an important objective. We provide both empirical evidence and heuristic support for this feature that can be considered as a direct consequence of incorporating both the group structure for the input data and the block structure for the dictionary in the learning process. The optimization problems for both the dictionary learning and sparse coding can be solved efficiently using block-gradient descent, and the details of the optimization algorithms are presented. We evaluate the proposed methods using well-known datasets, and favorable comparisons with state-of-the-art dictionary learning methods demonstrate the viability and validity of the proposed framework.
Yu-Tseh Chi, Mohsen Ali, Jeffrey Ho
CVPR4
2013 Computing Diffeomorphic Paths for Large Motion Interpolation
abstract
In this paper, we introduce a novel framework for computing a path of diffeomorphisms between a pair of input diffeomorphisms. Direct computation of a geodesic path on the space of diffeomorphisms Diff(Ω) is difficult, and it can be attributed mainly to the infinite dimensionality of Diff(Ω). Our proposed framework, to some degree, bypasses this difficulty using the quotient map of Diff(Ω) to the quotient space Diff(M)/Diff(M)μobtained by quotienting out the subgroup of volume-preserving diffeomorphisms Diff(M)μ. This quotient space was recently identified as the unit sphere in a Hilbert space in mathematics literature, a space with well-known geometric properties. Our framework leverages this recent result by computing the diffeomorphic path in two stages. First, we project the given diffeomorphism pair onto this sphere and then compute the geodesic path between these projected points. Second, we lift the geodesic on the sphere back to the space of diffeomerphisms, by solving a quadratic programming problem with bilinear constraints using the augmented Lagrangian technique with penalty terms. In this way, we can estimate the path of diffeomorphisms, first, staying in the space of diffeomorphisms, and second, preserving shapes/volumes in the deformed images along the path as much as possible. We have applied our framework to interpolate intermediate frames of frame-sub-sampled video sequences. In the reported experiments, our approach compares favorably with the popular Large Deformation Diffeomorphic Metric Mapping framework (LDDMM).
Dohyung Seo, Jeffrey Ho, Baba C. Vemuri
CVPR2
2013 Affine-Constrained Group Sparse Coding and Its Application to Image-Based Classifications
abstract
This paper proposes a novel approach for sparse coding that further improves upon the sparse representation-based classification (SRC) framework. The proposed framework, Affine-Constrained Group Sparse Coding (ACGSC), extends the current SRC framework to classification problems with multiple input samples. Geometrically, the affineconstrained group sparse coding essentially searches for the vector in the convex hull spanned by the input vectors that can best be sparse coded using the given dictionary. The resulting objective function is still convex and can be efficiently optimized using iterative block-coordinate descent scheme that is guaranteed to converge. Furthermore, we provide a form of sparse recovery result that guarantees, at least theoretically, that the classification performance of the constrained group sparse coding should be at least as good as the group sparse coding. We have evaluated the proposed approach using three different recognition experiments that involve illumination variation of faces and textures, and face recognition under occlusions. Preliminary experiments have demonstrated the effectiveness of the proposed approach, and in particular, the results from the recognition/occlusion experiment are surprisingly accurate and robust.
Yu-Tseh Chi, Mohsen Ali, Muhammad Ali Rushdi 0001, Jeffrey Ho
ICCV4
2013 Recursive Estimation of the Stein Center of SPD Matrices and Its Applications
abstract
Symmetric positive-definite (SPD) matrices are ubiquitous in Computer Vision, Machine Learning and Medical Image Analysis. Finding the center/average of a population of such matrices is a common theme in many algorithms such as clustering, segmentation, principal geodesic analysis, etc. The center of a population of such matrices can be defined using a variety of distance/divergence measures as the minimizer of the sum of squared distances/divergences from the unknown center to the members of the population. It is well known that the computation of the Karcher mean for the space of SPD matrices which is a negatively-curved Riemannian manifold is computationally expensive. Recently, the LogDet divergence-based center was shown to be a computationally attractive alternative. However, the LogDet-based mean of more than two matrices can not be computed in closed form, which makes it computationally less attractive for large populations. In this paper we present a novel recursive estimator for center based on the Stein distance - which is the square root of the LogDet divergence - that is significantly faster than the batch mode computation of this center. The key theoretical contribution is a closed-form solution for the weighted Stein center of two SPD matrices, which is used in the recursive computation of the Stein center for a population of SPD matrices. Additionally, we show experimental evidence of the convergence of our recursive Stein center estimator to the batch mode Stein center. We present applications of our recursive estimator to K-means clustering and image indexing depicting significant time gains over corresponding algorithms that use the batch mode computations. For the latter application, we develop novel hashing functions using the Stein distance and apply it to publicly available data sets, and experimental results have shown favorable comparisons to other competing methods.
Hesamoddin Salehian, Guang Cheng 0002, Baba C. Vemuri, Jeffrey Ho
ICCV4
2013 Color de-rendering using coupled dictionary learning
abstract
Consumer-level digital cameras typically post-process raw captured image data to produce enhanced visually appealing output RGB images. Post-processing operations include color gamut compression, tone mapping and other non-linear color corrections. However, raw image data is needed for many computer vision applications such as photometric stereo, shape from shading, and color constancy. Recovering raw image data from RGB images is complicated by the high non-linearity of the post-processing operations. In this paper, we propose a coupled dictionary scheme to model the relationship between the raw and RGB color image spaces of consumer cameras. Dictionary learning is regularized by sparsity constraints on feature representation. As well, we explore a more elaborate variant of coupled dictionary schemes that models the feature coupling more accurately. We test the proposed dictionary learning schemes on many commercial camera datasets. Our experimental results show accurate recovery of raw image data that looks visually indistinguishable from the ground truth.
Muhammad Ali Rushdi 0001, Mohsen Ali, Jeffrey Ho
ICIP3
2013 On A Nonlinear Generalization of Sparse Coding and Dictionary Learning
abstract
Existing dictionary learning algorithms are based on the assumption that the data are vectors in an Euclidean vector space, and the dictionary is learned from the training data using the vector space structure and its Euclidean metric. However, in many applications, features and data often originated from a Riemannian manifold that does not support a global linear (vector space) structure. Furthermore, the extrinsic viewpoint of existing dictionary learning algorithms becomes inappropriate for modeling and incorporating the intrinsic geometry of the manifold that is potentially important and critical to the application. This paper proposes a novel framework for sparse coding and dictionary learning for data on a Riemannian manifold, and it shows that the existing sparse coding and dictionary learning methods can be considered as special (Euclidean) cases of the more general framework proposed here. We show that both the dictionary and sparse coding can be effectively computed for several important classes of Riemannian manifolds, and we validate the proposed method using two well-known classification problems in computer vision and medical imaging analysis.
Jeffrey Ho, Baba C. Vemuri
ICML (3)1
2013 Visual Tracking Using Superpixel-Based Appearance Model
S. M. Nejhum Shahed, Muhammad Ali Rushdi 0001, Jeffrey Ho
ICVS3
2013 Multiple Atlas Construction From A Heterogeneous Brain MR Image Collection
abstract
In this paper, we propose a novel framework for computing single or multiple atlases (templates) from a large population of images. Unlike many existing methods, our proposed approach is distinguished by its emphasis on the sharpness of the computed atlases and the requirement of rotational invariance. In particular, we argue that sharp atlas images that retain crucial and important anatomical features with high fidelity are more useful for many medical imaging applications when compared with the blurry and fuzzy atlas images computed by most existing methods. The geometric notion that underlies our approach is the idea of manifold learning in a quotient space, the quotient space of the image space by the rotations. We present an extension of the existing manifold learning approach to quotient spaces by using invariant metrics, and utilizing the manifold structure for partitioning the images into more homogeneous sub-collections, each of which can be represented by a single atlas image. Specifically, we propose a three-step algorithm. First, we partition the input images into subgroups using unsupervised or semi-supervised learning methods on manifolds. Then we formulate a convex optimization problem in each subgroup to locate the atlases and determine the crucial neighbors that are used in the realization step to form the template images. We have evaluated our algorithm using whole brain MR volumes from OASIS database. Experimental results demonstrate that the atlases computed using the proposed algorithm not only discover the brain structural changes in different age groups but also preserve important structural details and generally enjoy better image quality.
Jeffrey Ho, Baba C. Vemuri
IEEE Trans. Medical Imaging2
2012 Shift-Map Based Stereo Image Retargeting with Disparity Adjustment
Shaoyu Qi, Jeffrey Ho
ACCV (4)2
2012 Seam Segment Carving: Retargeting Images to Irregularly-Shaped Image Domains
Shaoyu Qi, Jeffrey Ho
ECCV (6)2
2012 Augmented Coupled Dictionary Learning for Image Super-Resolution
abstract
Recent approaches in image super-resolution suggest learning dictionary pairs to model the relationship between low-resolution and high-resolution image patches with sparsity constraints on the patch representation. Most of the previous approaches in this direction assume for simplicity that the sparse codes for a low-resolution patch are equal to those of the corresponding high-resolution patch. However, this invariance assumption is not quite accurate especially for large scaling factors where the optimal weights and indices of representative features are not fixed across the scaling transformation. In this paper, we propose an augmented coupled dictionary learning scheme that compensates for the inaccuracy of the invariance assumption. First, we learn a dictionary for the low-resolution image space. Then, we compute an augmented dictionary in the high-resolution image space where novel augmented dictionary atoms are inferred from the training error of the low-resolution dictionary. For a low-resolution test image, the sparse codes of the low-resolution patches and the lowresolution dictionary training error are combined with the trained high-resolution dictionary to produce a high-resolution image. Our experimental results compare favourably with the non-augmented scheme.
Muhammad Ali Rushdi 0001, Jeffrey Ho
ICMLA (1)2
2012 Approximating Symmetric Positive Semidefinite Tensors of Even Order
abstract
Tensors of various orders can be used for modeling physical quantities such as strain and diffusion as well as curvature and other quantities of geometric origin. Depending on the physical properties of the modeled quantity, the estimated tensors are often required to satisfy the positivity constraint, which can be satisfied only with tensors of even order. Although the space [Formula: see text] of 2m(th)-order symmetric positive semi-definite tensors is known to be a convex cone, enforcing positivity constraint directly on [Formula: see text] is usually not straightforward computationally because there is no known analytic description of [Formula: see text] for m > 1. In this paper, we propose a novel approach for enforcing the positivity constraint on even-order tensors by approximating the cone [Formula: see text] for the cases 0 < m < 3, and presenting an explicit characterization of the approximation Σ(2) (m) ⊂ Ω(2) (m) for m ≥ 1, using the subset [Formula: see text] of semi-definite tensors that can be written as a sum of squares of tensors of order m. Furthermore, we show that this approximation leads to a non-negative linear least-squares (NNLS) optimization problem with the complexity that equals the number of generators in Σ(2) (m). Finally, we experimentally validate the proposed approach and we present an application for computing 2m(th)-order diffusion tensors from Diffusion Weighted Magnetic Resonance Images.
Angelos Barmpoutis, Jeffrey Ho, Baba C. Vemuri
SIAM J. Imaging Sci.2
2011 On affine registration of planar point sets using complex numbers
Jeffrey Ho, Ming-Hsuan Yang 0001
Comput. Vis. Image Underst.1
2011 Higher-Dimensional Affine Registration and Vision Applications
abstract
Affine registration has a long and venerable history in computer vision literature, and in particular, extensive work has been done for affine registration in R(2) and R(3). This paper studies affine registration in R(m) with m typically ranging from 4 to 12. To justify breaking of this dimension barrier, the first part of the paper describes three novel matching problems that can be formulated and solved as affine point-set registration problems in dimensions greater than three: stereo correspondence under motion, image set matching, and covariant point-set matching, problems that are not only interesting in their own right but also have potential for important vision applications. Unfortunately, most of the existing affine registration algorithms do not generalize easily to higher dimensions due to their inefficiency. Therefore, the second part of this paper develops a novel algorithm for estimating the affine transform between two point sets in R(m). Specifically, the algorithm follows the common approach of iteratively solving the correspondences and transform. The initial correspondences are determined using the novel notion of local spectral features, features constructed from local distance matrices. Unlike many correspondence-based methods, the proposed algorithm is capable of registering point sets of different size, and the use of local features provides some degree of robustness against noise and outliers. The proposed algorithm is validated on a variety of synthetic point sets in different dimensions with varying degrees of deformation and noise, and the paper also shows experimentally that several instances of the aforementioned three matching problems can indeed be solved satisfactorily using the proposed affine registration algorithm.
S. M. Nejhum Shahed, Yu-Tseh Chi, Jeffrey Ho, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2011 Online Sparse Gaussian Process Regression and Its Applications
abstract
We present a new Gaussian process (GP) inference algorithm, called online sparse matrix Gaussian processes (OSMGP), and demonstrate its merits by applying it to the problems of head pose estimation and visual tracking. The OSMGP is based upon the observation that for kernels with local support, the Gram matrix is typically sparse. Maintaining and updating the sparse Cholesky factor of the Gram matrix can be done efficiently using Givens rotations. This leads to an exact, online algorithm whose update time scales linearly with the size of the Gram matrix. Further, we provide a method for constant time operation of the OSMGP using matrix downdates. The downdates maintain the Cholesky factor at a constant size by removing certain rows and columns corresponding to discarded training examples. We demonstrate that, using these matrix downdates, online hyperparameter estimation can be included at cost linear in the number of total training examples. We describe a robust appearance-based head pose estimation system based upon the OSMGP. Numerous experiments and comparisons with existing methods using a large dataset system demonstrate the efficiency and accuracy of our system. Further, to showcase the applicability of OSMGP to a wide variety of problems, we also describe a regression-based visual tracking method. Experiments show that our OSMGP algorithm generalizes well using online learning.
Ananth Ranganathan, Ming-Hsuan Yang 0001, Jeffrey Ho
IEEE Trans. Image Process.3
2010 A Direct Method for Estimating Planar Projective Transform
Yu-Tseh Chi, Jeffrey Ho, Ming-Hsuan Yang 0001
ACCV (2)2
2010 Image atlas construction via intrinsic averaging on the manifold of images
abstract
In this paper, we propose a novel algorithm for computing an atlas from a collection of images. In the literature, atlases have almost always been computed as some types of means such as the straightforward Euclidean means or the more general Karcher means on Riemannian manifolds. In the context of images, the paper's main contribution is a geometric framework for computing image atlases through a two-step process: the localization of mean and the realization of it as an image. In the localization step, a few nearest neighbors of the mean among the input images are determined, and the realization step then proceeds to reconstruct the atlas image using these neighbors. Decoupling the localization step from the realization step provides the flexibility that allows us to formulate a general algorithm for computing image atlas. More specifically, we assume the input images belong to some smooth manifold M modulo image rotations. We use a graph structure to represent the manifold, and for the localization step, we formulate a convex optimization problem in ℝ(k) (k the number of input images) to determine the crucial neighbors that are used in the realization step to form the atlas image. The algorithm is both unbiased and rotation-invariant. We have evaluated the algorithm using synthetic and real images. In particular, experimental results demonstrate that the atlases computed using the proposed algorithm preserve important image features and generally enjoy better image quality in comparison with atlases computed using existing methods.
Jeffrey Ho, Baba C. Vemuri
CVPR2
2010 Statistical Analysis of Tensor Fields
abstract
In this paper, we propose a Riemannian framework for statistical analysis of tensor fields. Existing approaches to this problem have been mainly voxel-based that overlook the correlation between tensors at different voxels. In our approach, the tensor fields are considered as points in a high-dimensional Riemannian product space and accordingly, we extend Principal Geodesic Analysis (PGA) to the product space. This provides us with a principled method for linearizing the problem, and coupled with the usual log-exp maps that relate points on manifold to tangent vectors, the global correlation of the tensor field can be captured using Principal Component Analysis in a tangent space. Using the proposed method, the modes of variation of tensor fields can be efficiently determined, and dimension reduction of the data is also easily implemented. Experimental results on characterizing the variation of a large set of tensor fields are presented in the paper, and results on classifying tensor fields using the proposed method are also reported. These preliminary experimental results demonstrate the advantages of our method over the voxel-based approach.
Baba C. Vemuri, Jeffrey Ho
MICCAI (1)3
2010 Online visual tracking with histograms and articulating blocks
S. M. Nejhum Shahed, Jeffrey Ho, Ming-Hsuan Yang 0001
Comput. Vis. Image Underst.2
2009 An algebraic approach to affine registration of point sets
abstract
This paper proposes a new affine registration algorithm for matching two point sets in IR2or IR3. The input point sets are represented as probability density functions, using either Gaussian mixture models or discrete density models, and the problem of registering the point sets is treated as aligning the two distributions. Since polynomials transform as symmetric tensors under an affine transformation, the distributions' moments, which are the expected values of polynomials, also transform accordingly. Therefore, instead of solving the harder problem of aligning the two distributions directly, we solve the softer problem of matching the distributions' moments. By formulating a least-squares problem for matching moments of the two distributions up to degree three, the resulting cost function is a polynomial that can be efficiently optimized using techniques originated from algebraic geometry: the global minimum of this polynomial can be determined by solving a system of polynomial equations. The algorithm is robust in the presence of noises and outliers, and we validate the proposed algorithm on a variety of point sets with varying degrees of deformation and noise.
Jeffrey Ho, Adrian M. Peter, Anand Rangarajan 0001, Ming-Hsuan Yang 0001
ICCV1
2008 Shape L'Âne rouge: Sliding wavelets for indexing and retrieval
abstract
Shape representation and retrieval of stored shape models are becoming increasingly more prominent in fields such as medical imaging, molecular biology and remote sensing. We present a novel framework that directly addresses the necessity for a rich and compressible shape representation, while simultaneously providing an accurate method to index stored shapes. The core idea is to represent point-set shapes as the square root of probability densities expanded in a wavelet basis. We then use this representation to develop a natural similarity metric that respects the geometry of these probability distributions, i.e. under the wavelet expansion, densities are points on a unit hypersphere and the distance between densities is given by the separating arc length. The process uses a linear assignment solver for non-rigid alignment between densities prior to matching; this has the connotation of "sliding" wavelet coefficients akin to the sliding block puzzle L'Âne Rouge. We illustrate the utility of this framework by matching shapes from the MPEG-7 data set and provide comparisons to other similarity measures, such as Euclidean distance shape distributions.
Adrian M. Peter, Anand Rangarajan 0001, Jeffrey Ho
CVPR3
2008 Visual tracking with histograms and articulating blocks
abstract
We propose an algorithm for accurate tracking of (articulated) objects using online update of appearance and shape. The challenge here is to model foreground appearance with histograms in a way that is both efficient and accurate. In this algorithm, the constantly changing foreground shape is modeled as a small number of rectangular blocks, whose positions within the tracking window are adaptively determined. Under the general assumption of stationary foreground appearance, we show that robust object tracking is possible by adaptively adjusting the locations of these blocks. Implemented in MATLAB without substantial optimization, our tracker runs already at 3.7 frames per second on a 3GHz machine. Experimental results have demonstrated that the algorithm is able to efficiently track articulated objects undergoing large variation in appearance and shape.
S. M. Nejhum Shahed, Jeffrey Ho, Ming-Hsuan Yang 0001
CVPR2
2008 Higher Dimensional Affine Registration and Vision Applications
Yu-Tseh Chi, S. M. Nejhum Shahed, Jeffrey Ho, Ming-Hsuan Yang 0001
ECCV (4)3
2007 USSR: A Unified Framework for Simultaneous Smoothing, Segmentation, and Registration of Multiple Images
abstract
Image smoothing, segmentation and registration are three key processing steps in many computer vision applications. In this paper, we present a novel framework for achieving all three seemingly disparate goals simultaneously across multiple images in a unified framework via a single variational principle. The proposed method ensures that the estimated registration is unbiased and all compositions of registration maps are compatible. The solution to the variational problem is achieved efficiently by solving a coupled system of partial differential equations over the common domain on which the registration maps are defined. The effectiveness of the proposed framework is demonstrated on sets of real images.
Nicholas A. Lord, Jeffrey Ho, Baba C. Vemuri
ICCV2
2007 A New Affine Registration Algorithm for Matching 2D Point Sets
abstract
We propose a novel affine registration algorithm for matching 2D feature points. Unlike many previously published work on affine point matching, the proposed algorithm does not require any optimization and in the absence of data noise, the algorithm will recover the exact affine transformation and the unknown correspondence. The two-step algorithm first reduces the general affine case to the orthogonal case, and the unknown rotation is computed as the roots of a low-degree polynomial with complex coefficients. The algebraic and geometric ideas behind the proposed method are both clear and transparent, and its implementation is straightforward. We validate the algorithm on a variety of synthetic 2D point sets as well as feature points on images of real-world objects
Jeffrey Ho, Ming-Hsuan Yang 0001, Anand Rangarajan 0001, Baba C. Vemuri
WACV1
2007 Simultaneous Registration and Parcellation of Bilateral Hippocampal Surface Pairs for Local Asymmetry Quantification
abstract
In clinical applications where structural asymmetries between homologous shapes have been correlated with pathology, the questions of definition and quantification of "asymmetry" arise naturally. When not only the degree but the position of deformity is thought relevant, asymmetry localization must also be addressed. Asymmetries between paired shapes have already been formulated in terms of (nonrigid) diffeomorphisms between the shapes. For the infinity of such maps possible for a given pair, we define optimality as the minimization of deviation from isometry under the constraint of piecewise deformation homogeneity. We propose a novel variational formulation for segmenting asymmetric regions from surface pairs based on the minimization of a functional of both the deformation map and the segmentation boundary, which defines the regions within which the homogeneity constraint is to be enforced. The functional minimization is achieved via a quasi-simultaneous evolution of the map and the segmenting curve, conducted on and between two-dimensional surface parametric domains. We present examples using both synthetic data and pairs of left and right hippocampal structures and demonstrate the relevance of the extracted features through a clinical epilepsy classification analysis
Nicholas A. Lord, Jeffrey Ho, Baba C. Vemuri, Stephan J. Eisenschenk
IEEE Trans. Medical Imaging2
2006 Integrating Surface Normal Vectors Using Fast Marching Method
Jeffrey Ho, Jongwoo Lim, Ming-Hsuan Yang 0001, David J. Kriegman
ECCV (3)1
2006 Face Recognition Using 3-D Models: Pose and Illumination
abstract
Unconstrained illumination and pose variation lead to significant variation in the photographs of faces and constitute a major hurdle preventing the widespread use of face recognition systems. The challenge is to generalize from a limited number of images of an individual to a broad range of conditions. Recently, advances in modeling the effects of illumination and pose have been accomplished using three-dimensional (3-D) shape information coupled with reflectance models. Notable developments in understanding the effects of illumination include the nonexistence of illumination invariants, a characterization of the set of images of objects in fixed pose under variable illumination (the illumination cone), and the introduction of spherical harmonics and low-dimensional linear subspaces for modeling illumination. To generalize to novel conditions, either multiple images must be available to reconstruct 3-D shape or, if only a single image is accessible, prior information about the 3-D shape and appearance of faces in general must be used. The 3-D Morphable Model was introduced as a generative model to predict the appearances of an individual while using a statistical prior on shape and texture allowing its parameters to be estimated from single image. Based on these new understandings, face recognition algorithms have been developed to address the joint challenges of pose and lighting. In this paper, we review these developments and provide a brief survey of the resulting face recognition algorithms and their performance
Sami Romdhani, Jeffrey Ho, Thomas Vetter, David J. Kriegman
Proc. IEEE2
2005 Passive Photometric Stereo from Motion
abstract
We introduce an iterative algorithm for shape reconstruction from multiple images of a moving (Lambertian) object illuminated by distant (and possibly time varying) lighting. Starting with an initial piecewise linear surface, the algorithm iteratively estimates a new surface based on the previous surface estimate and the photometric information available from the input image sequence. During each iteration, standard photometric stereo techniques are applied to estimate the surface normals up to an unknown generalized bas-relief transform, and a new surface is computed by integrating the estimated normals. The algorithm essentially consists of a sequence of matrix factorizations (of intensity values) followed by minimization using gradient descent (integration of the normals). Conceptually, the algorithm admits a clear geometric interpretation, which is used to provide a qualitative analysis of the algorithm's convergence. Implementation-wise, it is straightforward being based on several established photometric stereo and structure from motion algorithms. We demonstrate experimentally the effectiveness of our algorithm using several videos of hand-held objects moving in front of a fixed light and camera.
Jongwoo Lim, Jeffrey Ho, Ming-Hsuan Yang 0001, David J. Kriegman
ICCV2
2005 Visual tracking and recognition using probabilistic appearance manifolds
Kuang-chih Lee, Jeffrey Ho, Ming-Hsuan Yang 0001, David J. Kriegman
Comput. Vis. Image Underst.2
2005 Acquiring Linear Subspaces for Face Recognition under Variable Lighting
abstract
Previous work has demonstrated that the image variation of many objects (human faces in particular) under variable lighting can be effectively modeled by low-dimensional linear spaces, even when there are multiple light sources and shadowing. Basis images spanning this space are usually obtained in one of three ways: A large set of images of the object under different lighting conditions is acquired, and principal component analysis (PCA) is used to estimate a subspace. Alternatively, synthetic images are rendered from a 3D model (perhaps reconstructed from images) under point sources and, again, PCA is used to estimate a subspace. Finally, images rendered from a 3D model under diffuse lighting based on spherical harmonics are directly used as basis images. In this paper, we show how to arrange physical lighting so that the acquired images of each object can be directly used as the basis vectors of a low-dimensional linear space and that this subspace is close to those acquired by the other methods. More specifically, there exist configurations of k point light source directions, with k typically ranging from 5 to 9, such that, by taking k images of an object under these single sources, the resulting subspace is an effective representation for recognition under a wide range of lighting conditions. Since the subspace is generated directly from real images, potentially complex and/or brittle intermediate steps such as 3D reconstruction can be completely avoided; nor is it necessary to acquire large numbers of training images or to physically construct complex diffuse (harmonic) light fields. We validate the use of subspaces constructed in this fashion within the context of face recognition.
Kuang-chih Lee, Jeffrey Ho, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Visual Tracking Using Learned Linear Subspaces
Jeffrey Ho, Kuang-chih Lee, Ming-Hsuan Yang 0001, David J. Kriegman
CVPR (1)1
2004 Image Clustering with Metric, Local Linear Structure, and Affine Symmetry
Jongwoo Lim, Jeffrey Ho, Ming-Hsuan Yang 0001, Kuang-chih Lee, David J. Kriegman
ECCV (1)2
2003 Clustering Appearances of Objects Under Varying Illumination Conditions
abstract
We introduce two appearance-based methods for clustering a set of images of 3D (three-dimensional) objects, acquired under varying illumination conditions, into disjoint subsets corresponding to individual objects. The first algorithm is based on the concept of illumination cones. According to the theory, the clustering problem is equivalent to finding convex polyhedral cones in the high-dimensional image space. To efficiently determine the conic structures hidden in the image data, we introduce the concept of conic affinity, which measures the likelihood of a pair of images belonging to the same underlying polyhedral cone. For the second method, we introduce another affinity measure based on image gradient comparisons. The algorithm operates directly on the image gradients by comparing the magnitudes and orientations of the image gradient at each pixel. Both methods have clear geometric motivations, and they operate directly on the images without the need for feature extraction or computation of pixel statistics. We demonstrate experimentally that both algorithms are surprisingly effective in clustering images acquired under varying illumination conditions with two large, well-known image data sets.
Jeffrey Ho, Ming-Hsuan Yang 0001, Jongwoo Lim, Kuang-chih Lee, David J. Kriegman
CVPR (1)1
2003 Video-Based Face Recognition Using Probabilistic Appearance Manifolds
abstract
This paper presents a method to model and recognize human faces in video sequences. Each registered person is represented by a low-dimensional appearance manifold in the ambient image space, the complex nonlinear appearance manifold expressed as a collection of subsets (named pose manifolds), and the connectivity among them. Each pose manifold is approximated by an affine plane. To construct this representation, exemplars are sampled from videos, and these exemplars are clustered with a K-means algorithm; each cluster is represented as a plane computed through principal component analysis (PCA). The connectivity between the pose manifolds encodes the transition probability between images in each of the pose manifold and is learned from a training video sequences. A maximum a posteriori formulation is presented for face recognition in test video sequences by integrating the likelihood that the input image comes from a particular pose manifold and the transition probability to this pose manifold from the previous frame. To recognize faces with partial occlusion, we introduce a weight mask into the process. Extensive experiments demonstrate that the proposed algorithm outperforms existing frame-based face recognition methods with temporal voting schemes.
Kuang-chih Lee, Jeffrey Ho, Ming-Hsuan Yang 0001, David J. Kriegman
CVPR (1)2
2003 Binocular Helmholtz Stereopsis
abstract
Helmholtz stereopsis has been introduced recently as a surface reconstruction technique that does not assume a model of surface reflectance. In the reported formulation, correspondence was established using a rank constraint, necessitating at least three viewpoints and three pairs of images. Here, it is revealed that the fundamental Helmholtz stereopsis constraint defines a nonlinear partial differential equation, which can be solved using only two images. It is shown that, unlike conventional stereo, binocular Helmholtz stereopsis is able to establish correspondence (and thereby recover surface depth) for objects having an arbitrary and unknown BRDF and in textureless regions (i.e., regions of constant or slowly varying BRDF). An implementation and experimental results validate the method for specular surfaces with and without texture.
Todd E. Zickler, Jeffrey Ho, David J. Kriegman, Jean Ponce, Peter N. Belhumeur
ICCV2
2001 Nine Points of Light: Acquiring Subspaces for Face Recognition under Variable Lighting
abstract
Previous work has demonstrated that the image variations of many objects (human faces in particular) under variable lighting can be effectively modeled by low dimensional linear spaces. Basis images spanning this space are usually obtained in one of two ways: A large number of images of the object under different conditions is acquired, and principal component analysis (PCA) is used to estimate a subspace. Alternatively, a 3D model (perhaps reconstructed from images) is used to render virtual images under either point sources from which a subspace is derived using PCA or more recently under diffuse synthetic lighting based on spherical harmonics. In this paper we show that there exists a configuration of nine point light source directions such that by taking nine images of each individual under these single sources, the resulting subspace is effective at recognition under a wide range of lighting conditions. Since the subspace is generated directly from real images, potentially complex intermediate steps such as PCA and 3D reconstruction can be completely avoided; nor is it necessary to acquire large numbers of training images or physically construct complex diffuse (harmonic) light fields. We provide both theoretical and empirical results to explain why these linear spaces should be good for recognition.
Kuang-chih Lee, Jeffrey Ho, David J. Kriegman
CVPR (1)2
2001 Compressing Large Polygonal Models
abstract
Presents an algorithm that uses partitioning and gluing to compress large triangular meshes which are too complex to fit in main memory. The algorithm is based largely on the existing mesh compression algorithms, most of which require an 'in-core' representation of the input mesh. Our solution is to partition the mesh into smaller submeshes and compress these submeshes separately using existing mesh compression techniques. Since a direct partition of the input mesh is out of question, instead we partition a simplified mesh and use the partition on the simplified model to obtain a partition on the original model. In order to recover the full connectivity, we present a simple scheme for encoding/decoding the resulting boundary structure from the mesh partition. When compressing large models with few singular vertices, a negligible portion of the compressed output is devoted to gluing information. On desktop computers, we have run experiments on models with millions of vertices, which could not be compressed using standard compression software packages, and have observed compression ratios as high as 17 to 1 using our technique.
Jeffrey Ho, Kuang-chih Lee, David J. Kriegman
IEEE Visualization1