Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ruei-Sung Lin

dblp:89/3312 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Generative modeling · 42% Video understanding and tracking · 23% Autonomous driving · 18%
Computer graphics and multimedia
1 paper
Computer animation and physical simulation · 100%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%

Topics — the 27 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.022024
Bidirectional Autoregressive Diffusion Model for Dance Generation · CVPR 2024
DiffMOT: A Real-time Diffusion-based Multiple Object Tracker with Non-linear Prediction · CVPR 2024
Machine learning › Generative modeling › diffusion model › diffusion model architecture
autoregressive diffusion models
0.812024
Bidirectional Autoregressive Diffusion Model for Dance Generation · CVPR 2024
Computer vision › Video understanding and tracking
multi-object tracking
0.812024
DiffMOT: A Real-time Diffusion-based Multiple Object Tracker with Non-linear Prediction · CVPR 2024
Robotics › Autonomous driving
trajectory prediction
0.812024
DiffMOT: A Real-time Diffusion-based Multiple Object Tracker with Non-linear Prediction · CVPR 2024
Computer animation and physical simulation › motion synthesis
human motion synthesis
0.812024
Bidirectional Autoregressive Diffusion Model for Dance Generation · CVPR 2024
Computer animation and physical simulation
music-driven dance generation
0.812024
Bidirectional Autoregressive Diffusion Model for Dance Generation · CVPR 2024
Computer vision › Video understanding and tracking
object tracking
0.242008
Incremental Learning for Robust Visual Tracking · Int. J. Comput. Vis. 2008
Adaptive Discriminative Generative Model and Its Applications · NIPS 2004
Incremental Learning for Visual Tracking · NIPS 2004
Machine learning › Representation and self-supervised learning › representation learning › metric learning
ordinal embedding
0.112011
The power of comparative reasoning · ICCV 2011
Information retrieval › similarity search
fast similarity search
0.112011
The power of comparative reasoning · ICCV 2011
Information retrieval
similarity search
0.112011
The power of comparative reasoning · ICCV 2011
Information retrieval
hashing
0.112010
SPEC hashing: Similarity preserving algorithm for entropy-based coding · CVPR 2010
Information retrieval › similarity search
nearest neighbor search
0.112010
SPEC hashing: Similarity preserving algorithm for entropy-based coding · CVPR 2010
Information retrieval › hashing
similarity-preserving hashing
0.112010
SPEC hashing: Similarity preserving algorithm for entropy-based coding · CVPR 2010
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
belief propagation
0.112007
Efficient Message Representations for Belief Propagation · ICCV 2007
Computer vision › 3D vision › stereo vision › stereo matching
dense stereo matching
0.112007
Efficient Message Representations for Belief Propagation · ICCV 2007
Computer vision › 3D vision › 3d reconstruction › multi-view stereo
stereo reconstruction
0.112007
Efficient Message Representations for Belief Propagation · ICCV 2007
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning
0.112006
Learning Nonlinear Manifolds from Time Series · ECCV (2) 2006
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning
nonlinear manifold learning
0.112006
Learning Nonlinear Manifolds from Time Series · ECCV (2) 2006
Machine learning › Time series and sequential data › time series analysis
time series learning
0.112006
Learning Nonlinear Manifolds from Time Series · ECCV (2) 2006
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › subspace learning
online subspace learning
0.012004
Incremental Learning for Visual Tracking · NIPS 2004
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
subspace learning
0.012004
Incremental Learning for Visual Tracking · NIPS 2004
Algorithms and data structures › data structure design › search structures
hashing
0.012011
The power of comparative reasoning · ICCV 2011
Algorithms and data structures › data structure design › search structures › hashing › locality-sensitive hashing
minhash
0.012011
The power of comparative reasoning · ICCV 2011
Computer vision › Face, body and person analysis
face recognition
0.012010
SPEC hashing: Similarity preserving algorithm for entropy-based coding · CVPR 2010
Computer vision › 3D vision
camera pose estimation
0.012001
Tracking of Object with SVM Regression · CVPR (2) 2001
Computer vision › 3D vision › multi-view geometry
epipolar geometry estimation
0.012001
Tracking of Object with SVM Regression · CVPR (2) 2001
Computer vision › Video understanding and tracking
feature tracking
0.012001
Tracking of Object with SVM Regression · CVPR (2) 2001

Methods — techniques the papers use, named apart from their topics

local information decoding · 1.5bidirectional autoregressive diffusion · 1.5beat conditioning · 1.5kalman filter · 0.8diffusion probabilistic model · 0.8decoupled diffusion · 0.8rank correlation · 0.4polynomial kernel · 0.4partial order statistics · 0.4binary hash function learning · 0.2conditional entropy maximization · 0.1
YearPublicationVenuePosition
2024 DiffMOT: A Real-time Diffusion-based Multiple Object Tracker with Non-linear Prediction
abstract
In Multiple Object Tracking, objects often exhibit nonlinear motion of acceleration and deceleration, with irregular direction changes. Tacking-by-detection (TBD) trackers with Kalman Filter motion prediction work well in pedestrian-dominant scenarios but fall short in complex situations when multiple objects perform non-linear and diverse motion simultaneously. To tackle the complex non-linear motion, we propose a real-time diffusion-based MOT approach named DiffMOT. Specifically, for the motion predictor component, we propose a novel Decoupled Diffusion based Motion Predictor (D2MP). It models the entire distribution of various motion presented by the data as a whole. It also predicts an individual object's motion conditioning on an individual's historical motion information. Furthermore, it optimizes the diffusion process with much fewer sampling steps. As a MOT tracker, the DiffMOT is real-time at 22.7FPS, and also outperforms the state-of-the-art on DanceTrack[30] and SportsMOT[6] datasets with 62.3% and 76.2% in HOTA metrics, respectively. To the best of our knowledge, DiffMOT is the first to introduce a diffusion probabilistic model into the MOT to tackle non-linear motion prediction.
Weiyi Lv, Yuhang Huang 0006, Ning Zhang 0023, Ruei-Sung Lin, Dan Zeng 0001
CVPR4
2024 Bidirectional Autoregressive Diffusion Model for Dance Generation
abstract
Dance serves as a powerful medium for expressing human emotions, but the lifelike generation of dance is still a considerable challenge. Recently, diffusion models have showcased remarkable generative abilities across various domains. They hold promise for human motion generation due to their adaptable many-to-many nature. Nonetheless, current diffusion-based motion generation models often create entire motion sequences directly and unidirectionally, lacking focus on the motion with local and bidirectional enhancement. When choreographing high-quality dance movements, people need to take into account not only the musical context but also the nearby music-aligned dance motions. To authentically capture human behavior, we propose a Bidirectional Autoregressive Diffusion Model (BADM) for music-to-dance generation, where a bidirectional encoder is built to enforce that the generated dance is harmonious in both the forward and backward directions. To make the generated dance motion smoother, a local information decoder is built for local motion enhancement. The proposed framework is able to generate new motions based on the input conditions and nearby motions, which foresees individual motion slices iteratively and con-solidates all predictions. To further refine the synchronicity between the generated dance and the beat, the beat information is incorporated as an input to generate better music-aligned dance movements. Experimental results demonstrate that the proposed model achieves state-of-the-art performance compared to existing unidirectional approaches on the prominent benchmark for music-to-dance generation. The code and models are available: https://github.com/czzhang179/BADM.
Canyu Zhang 0002, Youbao Tang, Ruei-Sung Lin, Jing Xiao 0006, Song Wang 0002
CVPR4
2023 Prior-Enhanced Temporal Action Localization Using Subject-Aware Spatial Attention
abstract
Temporal action localization (TAL) aims to detect the boundary and identify the class of each action instance in a long untrimmed video. Current approaches treat video frames homogeneously, and tend to give background and key objects excessive attention. This limits their sensitivity to localize action boundaries. To this end, we propose a prior-enhanced temporal action localization method (PETAL), which only takes in RGB input and incorporates action subjects as priors. This proposal leverages action subjects’ information with a plug-and-play subject-aware spatial attention module (SA-SAM) to generate an aggregated and subject-prioritized representation. Experimental results on THUMOS-14 and ActivityNet-1.3 datasets demonstrate that the proposed PETAL achieves competitive performance using only RGB features, e.g., boosting mAP by 2.41% or 0.25% over the state-of-the-art approach that uses RGB features or with additional optical flow features on the THUMOS-14 dataset.
Youbao Tang, Ruei-Sung Lin, Haoqian Wang
ICASSP4
2022 Accurate and Robust Lesion RECIST Diameter Prediction and Segmentation with Transformers
Youbao Tang, Yirui Wang 0002, Shenghua He, Jing Xiao 0006, Ruei-Sung Lin
MICCAI (4)7
2011 The power of comparative reasoning
abstract
Rank correlation measures are known for their resilience to perturbations in numeric values and are widely used in many evaluation metrics. Such ordinal measures have rarely been applied in treatment of numeric features as a representational transformation. We emphasize the benefits of ordinal representations of input features both theoretically and empirically. We present a family of algorithms for computing ordinal embeddings based on partial order statistics. Apart from having the stability benefits of ordinal measures, these embeddings are highly nonlinear, giving rise to sparse feature spaces highly favored by several machine learning methods. These embeddings are deterministic, data independent and by virtue of being based on partial order statistics, add another degree of resilience to noise. These machine-learning-free methods when applied to the task of fast similarity search outperform state-of-the-art machine learning methods with complex optimization setups. For solving classification problems, the embeddings provide a nonlinear transformation resulting in sparse binary codes that are well-suited for a large class of machine learning algorithms. These methods show significant improvement on VOC 2010 using simple linear classifiers which can be trained quickly. Our method can be extended to the case of polynomial kernels, while permitting very efficient computation. Further, since the popular Min Hash algorithm is a special case of our method, we demonstrate an efficient scheme for computing Min Hash on conjunctions of binary features. The actual method can be implemented in about 10 lines of code in most languages (2 lines in MATLAB), and does not require any data-driven optimization.
Jay Yagnik, Dennis Strelow, David A. Ross, Ruei-Sung Lin
ICCV4
2010 SPEC hashing: Similarity preserving algorithm for entropy-based coding
abstract
Searching approximate nearest neighbors in large scale high dimensional data set has been a challenging problem. This paper presents a novel and fast algorithm for learning binary hash functions for fast nearest neighbor retrieval. The nearest neighbors are defined according to the semantic similarity between the objects. Our method uses the information of these semantic similarities and learns a hash function with binary code such that only objects with high similarity have small Hamming distance. The hash function is incrementally trained one bit at a time, and as bits are added to the hash code Hamming distances between dissimilar objects increase. We further link our method to the idea of maximizing conditional entropy among pair of bits and derive an extremely efficient linear time hash learning algorithm. Experiments on similar image retrieval and celebrity face recognition show that our method produces apparent improvement in performance over some state-of-the-art methods.
Ruei-Sung Lin, David A. Ross, Jay Yagnik
CVPR1
2008 Incremental Learning for Robust Visual Tracking
David A. Ross, Jongwoo Lim, Ruei-Sung Lin, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.3
2007 Efficient Message Representations for Belief Propagation
abstract
Belief propagation (BP) has been successfully used to approximate the solutions of various Markov random field (MRF) formulated energy minimization problems. However, large MRFs require a significant amount of memory to store the intermediate belief messages. We observe that these messages have redundant information due to the imposed smoothness prior. In this paper, we study the feasibility of applying compression techniques to the messages in the min-sum/max-product BP algorithm with 1D labels to improve the memory efficiency and reduce the read/write bandwidth. We articulate properties that an efficient message representation should satisfy. We investigate two common compression schemes, predictive coding and linear transform coding (PCA), and then propose a novel envelope point transform (EPT) method. Predictive coding is efficient and supports linear operations directly in the compressed domain, but it is only compatible with the L1smoothness function. PCA has the disadvantage that it does not guarantee the preservation of the minimal label. EPT is not limited to L1smoothness cost and allows a flexible quality vs. compression ratio tradeoff compared with predictive coding. Experiments on dense stereo reconstruction have shown that the predictive scheme and EPT can achieve 8times or more compression without significant loss of depth accuracy.
Tian-Li Yu 0002, Ruei-Sung Lin, Boaz J. Super, Bei Tang
ICCV2
2006 Dynamic Textures Synthesis as Nonlinear Manifold Learning and Traversing
abstract
We formulate the problem of dynamic texture synthesis as a nonlinear manifold learning and traversing problem. We characterize dynamic textures as the temporal changes in spectral parameters of image sequences. For continuous changes of such parameters, it is commonly assumed that all these parameters lie on or close to a low-dimensional manifold embedded in the original configuration space. For complex dynamic data, the manifolds are usually nonlinear and we propose to use a mixture of linear subspaces to model a nonlinear manifold. These locally linear subspaces are further aligned within a global coordinate system. With the nonlinear manifold being globally parameterized, we overcome motion discontinuity problems encountered in switching linear models and dynamics. We present a nonparametric method to describe the complex dynamics of data sequences on the manifold. We also apply such approach to dynamic spatial parameters such as motion capture data. The experimental results suggest that our approach is able to synthesize smooth, complex dynamic textures and human motions, and has potential applications to other dynamic data synthesis problems. 1
Che-Bin Liu, Ruei-Sung Lin, Narendra Ahuja, Ming-Hsuan Yang 0001
BMVC2
2006 Learning Nonlinear Manifolds from Time Series
Ruei-Sung Lin, Che-Bin Liu, Ming-Hsuan Yang 0001, Narendra Ahuja, Stephen E. Levinson
ECCV (2)1
2005 Modeling Dynamic Textures Using Subspace Mixtures
abstract
In this paper, we aim at modeling video sequences that exhibit temporal appearance variation. The dynamic texture model proposed in [6] is effective to model simple dynamic scenes. However, because of its over-simplified appearance model and under-constrained dynamics model, the visual quality of its synthesized video sequences is often not satisfactory. This leads to our new model. We parameterize the nonlinear image manifold using mixtures of probabilistic principal component analyzers. We then align coefficients from different mixture components in a global coordinate system, and model the image dynamics in the global coordinate using an autoregressive process. The experimental results show that our method is capable of capturing complex temporal appearance variation and offers improved synthesis results over previous works.
Che-Bin Liu, Ruei-Sung Lin, Narendra Ahuja
ICME2
2004 Adaptive Discriminative Generative Model for Object Tracking
Ruei-Sung Lin, Ming-Hsuan Yang 0001, Stephen E. Levinson
ECAI1
2004 Incremental Learning for Visual Tracking
abstract
Most existing tracking algorithms construct a representation of a target object prior to the tracking task starts, and utilize invariant features to handle appearance variation of the target caused by lighting, pose, and view angle change. In this paper, we present an efficient and effec- tive online algorithm that incrementally learns and adapts a low dimen- sional eigenspace representation to reflect appearance changes of the tar- get, thereby facilitating the tracking task. Furthermore, our incremental method correctly updates the sample mean and the eigenbasis, whereas existing incremental subspace update methods ignore the fact the sample mean varies over time. The tracking problem is formulated as a state inference problem within a Markov Chain Monte Carlo framework and a particle filter is incorporated for propagating sample distributions over time. Numerous experiments demonstrate the effectiveness of the pro- posed tracking algorithm in indoor and outdoor environments where the target objects undergo large pose and lighting changes. 1 Introduction The main challenges of visual tracking can be attributed to the difficulty in handling appear- ance variability of a target object. Intrinsic appearance variabilities include pose variation and shape deformation of a target object, whereas extrinsic illumination change, camera motion, camera viewpoint, and occlusions inevitably cause large appearance variation. Due to the nature of the tracking problem, it is imperative for a tracking algorithm to model such appearance variation. Here we developed a method that, during visual tracking, constantly and efficiently up- dates a low dimensional eigenspace representation of the appearance of the target object. The advantages of this adaptive subspace representation are several folds. The eigenspace representation provides a compact notion of the "thing" being tracked rather than treating the target as a set of independent pixels, i.e., "stuff" [1]. The use of an incremental method continually updates the eigenspace to reflect the appearance change caused by intrinsic and extrinsic factors, thereby facilitating the tracking process. To estimate the locations of the target objects in consecutive frames, we used a sampling algorithm with likelihood estimates, which is in direct contrast to other tracking methods that usually solve complex optimization problems using gradient-descent approach. The proposed method differs from our prior work [14] in several aspects. First, the pro- posed algorithm does not require any training images of the target object before the tracking task starts. That is, our tracker learns a low dimensional eigenspace representation on-line and incrementally updates it as time progresses (We assume, like most tracking algorithms, that the target region has been initialized in the first frame). Second, we extend our sam- pling method to incorporate a particle filter so that the sample distributions are propagated over time. Based on the eigenspace model with updates, an effective likelihood estimation function is developed. Third, we extend the R-SVD algorithm [6] so that both the sample mean and eigenbasis are correctly updated as new data arrive. Though there are numerous subspace update algorithms in the literature, only the method by Hall et al. [8] is also able to update the sample mean. However, their method is based on the addition of a single col- umn (single observation) rather than blocks (a number of observations in our case) and thus is less efficient than ours. While our formulation provides an exact solution, their algorithm gives only approximate updates and thus it may suffer from numerical instability. Finally, the proposed tracker is extended to use a robust error norm for likelihood estimation in the presence of noisy data or partial occlusions, thereby rendering more accurate and robust tracking results. 2 Previous Work and Motivation Black et al. [4] proposed a tracking algorithm using a pre-trained view-based eigenbasis representation and a robust error norm. Instead of relying on the popular brightness con- stancy working principal, they advocated the use of subspace constancy assumption for visual tracking. Although their algorithm demonstrated excellent empirical results, it re- quires to build a set of view-based eigenbases before the tracking task starts. Furthermore, their method assumes that certain factors, such as illumination conditions, do not change significantly as the eigenbasis, once constructed, is not updated. Hager and Belhumeur [7] presented a tracking algorithm to handle the geometry and illu- mination variations of target objects. Their method extends a gradient-based optical flow algorithm to incorporate research findings in [2] for object tracking under varying illumi- nation conditions. Prior to the tracking task starts, a set of illumination basis needs to be constructed at a fixed pose in order to account for appearance variation of the target due to lighting changes. Consequently, it is not clear whether this method is effective if a target object undergoes changes in illumination with arbitrary pose. In [9] Isard and Blake developed the Condensation algorithm for contour tracking in which multiple plausible interpretations are propagated over time. Though their probabilistic ap- proach has demonstrated success in tracking contours in clutter, the representation scheme is rather primitive, i.e., curves or splines, and is not updated as the appearance of a target varies due to pose or illumination change. Mixture models have been used to describe appearance change for motion estimation [3] [10]. In Black et al. [3] four possible causes are identified in a mixture model for estimating appearance change in consecutive frames, and thereby more reliable image motion can be obtained. A more elaborate mixture model with an online EM algorithm was recently proposed by Jepson et al. [10] in which they use three components and wavelet filters to account for appearance changes during tracking. Their method is able to handle variations in pose, illumination and expression. However, their WSL appearance model treats pixels within the target region independently, and therefore does not have notion of the "thing" being tracked. This may result in modeling background rather than the foreground, and fail to track the target. In contrast to the eigentracking algorithm [4], our algorithm does not require a training phase but learns the eigenbases on-line during the object tracking process, and constantly updates this representation as the appearance changes due to pose, view angle, and illumi- nation variation. Further, our method uses a particle filter for motion parameter estimation rather than the Gauss-Newton method which often gets stuck in local minima or is dis- tracted by outliers [4]. Our appearance-based model provides a richer description than simple curves or splines as used in [9], and has notion of the "thing" being tracked. In addition, the learned representation can be utilized for other tasks such as object recog- nition. In this work, an eigenspace representation is learned directly from pixel values within a target object in the image space. Experiments show that good tracking results can be obtained with this representation without resorting to wavelets as used in [10], and better performance can potentially be achieved using wavelet filters. Note also that the view-based eigenspace representation has demonstrated its ability to model appearance of objects at different pose [13], and under different lighting conditions [2]. 3 Incremental Learning for Tracking We present the details of the proposed incremental learning algorithm for object tracking in this section. 3.1 Incremental Update of Eigenbasis and Mean The appearance of a target object may change drastically due to intrinsic and extrinsic factors as discussed earlier. Therefore it is important to develop an efficient algorithm to update the eigenspace as the tracking task progresses. Numerous algorithms have been developed to update eigenbasis from a time-varying covariance matrix as more data arrive [6] [8] [11] [5]. However, most methods assume zero mean in updating the eigenbasis except the method by Hall et al. [8] in which they consider the change of the mean when updating eigenbasis as each new datum arrives. Their update algorithm only handles one datum per update and gives approximate results, while our formulation handles multiple data at the same time and renders exact solutions. We extend the work of the classic R-SVD method [6] in which we update the eigenbasis while taking the shift of the sample mean into account. To the best of our knowledge, this formulation with mean update is new in the literature. Given a d n data matrix A = {I1, . . . , In} where each column Ii is an observation (a d- dimensional image vector in this paper), we can compute the singular value decomposition (SVD) of A, i.e., A = U V . When a dm matrix E of new observations is available, the R-SVD algorithm efficiently computes the SVD of the matrix A = (A|E) = U V based on the SVD of A as follows: 1. Apply QR decomposition to and get orthonormal basis ~ E of E, and U = (U | ~ E). 2. Let V = V 0 0 I where Im is an m m identity matrix. It follows then, m = U A V = U (A|E) V 0 = U AV U E = U E . ~ E 0 Im ~ E AV ~ E E 0 ~ E E 3. Compute the SVD of = ~ U ~ ~ V and the SVD of A is A = U ( ~ U ~ ~ V )V = (U ~ U ) ~ ( ~ V V ). Exploiting the properties of orthonormal bases and block structures, the R-SVD algorithm computes the new eigenbasis efficiently. The computational complexity analysis and more details are described in [6]. One problem with the R-SVD algorithm is that the eigenbasis U is computed from AA with the zero mean assumption. We modify the R-SVD algorithm and compute the eigen- basis with mean update. The following derivation is based on scatter matrix, which is same as covariance matrix except a scalar factor. Proposition 1 Let Ip = {I1, I2, . . . , In}, Iq= {In+1, In+2, . . . , In+m}, and Ir = (Ip|Iq). Denote the means and scatter matrices of Ip, Iq, Ir as Ip, Iq, Ir, and Sp, Sq, Sr respec- tively, then Sr = Sp + Sq + nm (I n+m q - Ip)(Iq - Ip) . Proof: By definition, I r = n I I (I n+m p + m n+m q , Ip - Ir = m n+m p - Iq); Iq - Ir = n (I n+m q - Ip) and, Sr = n ( ( i=1 Ii - Ir)(Ii - Ir) + n+m i=n+1 Ii - Ir)(Ii - Ir) = n ( i=1 Ii - Ip + Ip - Ir)(Ii - Ip + Ip - Ir) + n+m ( i=m+1 Ii - Iq + Iq - Ir)(Ii - Iq + Iq - Ir) = Sp + n(Ip - Ir)(Ip - Ir) + Sq + m(Iq - Ir)(Iq - Ir) = Sp + nm2 ( ( ( I I n+m)2 p - Iq)(Ip - Iq) + Sq + n2m (n+m)2 p - Iq)(Ip - Iq) = Sp + Sq + nm (I n+m p - Iq)(Ip - Iq) Let ^ Ip = {I1 - Ip, . . . , In - Ip}, ^ Iq = {In+1 - Iq, . . . , In+m - Iq}, and ^ Ir = {I1 - Ir, . . . , In+m - Ir}, and the SVD of ^Ir = UrrVr . Let ~ E = ^ Iq| nm (I n+m p - Iq) , and use Proposition 1, Sr = (^ Ip| ~ E)(^ Ip| ~ E) . Therefore, we compute SVD on ( ^ Ip| ~ E) to get Ur. This can be done efficiently by the R-SVD algorithm as described above. In summary, given the mean Ip and the SVD of existing data Ip, i.e., UppVp and new data Iq, we can compute the the mean Ir and the SVD of Ir, i.e., UrrVr easily: 1. Compute I r = n I I (I n+m p + m n+m q , and ~ E = Iq - Ir 1(1m) | nm n+m p - Iq) . 2. Compute R-SVD with (UppVp ) and ~ E to obtain (UrrVr ). In numerous vision problems, we can further exploit the low dimensional approximation of image data and put larger weights on the recent observations, or equivalently downweight the contributions of previous observations. For example as the appearance of a target object gradually changes, we may want to put more weights on recent observations in updating the eigenbasis since they are more likely to be similar to the current appearance of the target. The forgetting factor f can be used under this premise as suggested in [11] , i.e., A = (f A |E) = (U (f )V |E) where A and A are original and weighted data matrices, respectively. 3.2 Sequential Inference Model The visual tracking problem is cast as an inference problem with a Markov model and hidden state variable, where a state variable Xt describes the affine motion parameters (and thereby the location) of the target at time t. Given a set of observed images It = {I1, . . . , It}. we aim to estimate the value of the hidden state variable Xt. Using Bayes' theorem, we have p(Xt| It) p(It|Xt) p(Xt|Xt-1) p(Xt-1| It-1) dXt-1 The tracking process is governed by the observation model p(It|Xt) where we estimate the likelihood of Xt observing It, and the dynamical model between two states p(Xt|Xt-1). The Condensation algorithm [9], based on factored sampling, approximates an arbitrary distribution of observations with a stochastically generated set of weighted samples. We use a variant of the Condensation algorithm to model the distribution over the object's location, as it evolves over time. 3.3 Dynamical and Observation Models The motion of a target object between two consecutive frames can be approximated by an affine image warping. In this work, we use the six parameters of affine transform to model the state transition from Xt-1 to Xt of a target object being tracked. Let Xt = (xt, yt, t, st, t, t) where xt, yt, t, st, t, t, denote x, y translation, rotation angle, scale, aspect ratio, and skew direction at time t. Each parameter in Xt is modeled independently by a Gaussian distribution around its counterpart in Xt-1. That is, p(Xt|Xt-1) = N (Xt; Xt-1, ) where is a diagonal covariance matrix whose elements are the corresponding variances of affine parameters, i.e., 2x, 2y, 2, 2 . s , 2 , 2 Since our goal is to use a representation to model the "thing" that we are tracking, we model the image observations using a probabilistic interpretation of principal component analysis [16]. Given an image patch predicated by Xt, we assume the observed image It was generated from a subspace spanned by U centered at . The probability that a sample being generated from the subspace is inversely proportional to the distance d from the sample to the reference point (i.e., center) of the subspace, which can be decomposed into the distance-to-subspace, dt, and the distance-within-subspace from the projected sample to the subspace center, dw. This distance formulation, based on a orthonormal subspace and its complement space, is similar to [12] in spirit. The probability of a sample generated from a subspace, pd (I t t|Xt), is governed by a Gaus- sian distribution: pd (I t t | Xt) = N (It ; , U U + I ) where I is an identity matrix, is the mean, and I term corresponds to the additive Gaus- sian noise in the observation process. It can be shown [15] that the negative exponential distance from It to the subspace spanned by U , i.e., exp(-||(It - ) - U U (It - )||2), is proportional to N (It; , U U + I) as 0. Within a subspace, the likelihood of the projected sample can be modeled by the Maha- lanobis distance from the mean as follows: pd (I w t | Xt) = N (It ; , U -2U ) where is the center of the subspace and is the matrix of singular values corresponding to the columns of U . Put together, the likelihood of a sample being generated from the subspace is governed by p(It|Xt) = pd (I (I t t|Xt) pdw t|Xt) = N (It; , U U + I) N (It; , U-2U ) (1) Given a drawn sample Xt and the corresponding image region It, we aim to compute p(It|Xt) using (1). To minimize the effects of noisy pixels, we utilize a robust error norm [4], (x, ) = x2 instead of the Euclidean norm d(x) = ||x||2, to ignore the "outlier" 2+x2 pixels (i.e., the pixels that are not likely to appear inside the target region given the current eigenspace). We use a method similar to that used in [4] in order to compute dt and dw. This robust error norm is helpful especially when we use a rectangular region to enclose the target (which inevitably contains some noisy background pixels). 4 Experiments To test the performance of our proposed tracker, we collected a number of videos recorded in indoor and outdoor environments where the targets change pose in different lighting con- ditions. Each video consists of 320 240 gray scale images and are recorded at 15 frames per second unless specified otherwise. For the eigenspace representation, each target image region is resized to 32 32 patch, and the number of eigenvectors used in all experiments is set to 16 though fewer eigenvectors may also work well. Implemented in MATLAB with MEX, our algorithm runs at 4 frames per second on a standard computer with 200 particles. We present some tracking results in this section and more tracking results as well as videos can be found at http://vision.ucsd.edu/~jwlim/ilt/. 4.1 Experimental Results Figure 1 shows the tracking results using a challenging sequence recorded with a mov- ing digital camera in which a person moves from a dark room toward a bright area while changing his pose, moving underneath spot lights, changing facial expressions and taking off glasses. All the eigenbases are constructed automatically from scratch and constantly updated to model the appearance of the target object while undergoing appearance changes. Even with the significant camera motion and low frame rate (which makes the motions be- tween frames more significant, or equivalently to tracking fast moving objects), our tracker stays stably on the target throughout the sequence. The second sequence contains an animal doll moving in different pose, scale, and lighting conditions as shown in Figure 2. Experimental results demonstrate that our tracker is able to follow the target as it undergoes large pose change, cluttered background, and lighting variation. Notice that the non-convex target object is localized with an enclosing rectan- gular window, and thus it inevitably contains some background pixels in its appearance representation. The robust error norm enables the tracker to ignore background pixels and estimate the target location correctly. The results also show that our algorithm faithfully Figure 1: A person moves from dark toward bright area with large lighting and pose changes. The images in the second row shows the current sample mean, tracked region, reconstructed image, and the reconstruction error respectively. The third and forth rows shows 10 largest eigenbases. Figure 2: An animal doll moving with large pose, lighting variation in a cluttered background. models the appearance of the target, as shown in eigenbases and reconstructed images, in the presence of noisy background pixels. We recorded a sequence to demonstrate that our tracker performs well in outdoor environ- ment where lighting conditions change drastically. The video was acquired when a person walking underneath a trellis covered by vines. As shown in Figure 3, the cast shadow changes the appearance of the target face drastically. Furthermore, the combined pose and lighting variation with low frame rate makes the tracking task extremely difficult. Nev- ertheless, the results show that our tracker successfully follows the target accurately and robustly. Due to heavy shadows and drastic lighting change, other tracking methods based on gradient, contour, or color information are unlikely to perform well in this case.
Jongwoo Lim, David A. Ross, Ruei-Sung Lin, Ming-Hsuan Yang 0001
NIPS3
2004 Adaptive Discriminative Generative Model and Its Applications
abstract
This paper presents an adaptive discriminative generative model that gen- eralizes the conventional Fisher Linear Discriminant algorithm and ren- ders a proper probabilistic interpretation. Within the context of object tracking, we aim to find a discriminative generative model that best sep- arates the target from the background. We present a computationally efficient algorithm to constantly update this discriminative model as time progresses. While most tracking algorithms operate on the premise that the object appearance or ambient lighting condition does not significantly change as time progresses, our method adapts a discriminative genera- tive model to reflect appearance variation of the target and background, thereby facilitating the tracking task in ever-changing environments. Nu- merous experiments show that our method is able to learn a discrimina- tive generative model for tracking target objects undergoing large pose and lighting changes.
Ruei-Sung Lin, David A. Ross, Jongwoo Lim, Ming-Hsuan Yang 0001
NIPS1
2003 Automatic language acquisition by an autonomous robot
abstract
There is no such thing as a disembodied mind. We posit that cognitive development can only occur through interaction with the physical world. To this end, we are developing a robotic platform for the purpose of studying cognition. We suggest that the central component of cognition is a memory which is primarily associative, one where learning occurs as the correlation of events from diverse inputs. We also posit that human-like cognition requires a well-integrated sensory-motor system, to provide these diverse inputs. As implemented in our robot, this system includes binaural hearing, stereo vision, tactile sense, and basic proprioceptive control. On top of these abilities, we are implementing and studying various models of processing, learning and decision making. Our goal is to produce a robot that will learn to carry out simple tasks in response to natural language requests. The robot's understanding of language will be learned concurrently with its other cognitive abilities. We have already developed a robust system and conducted a number or experiments on the way to this goal, some details of which appear in this paper. This is a first progress report of what we believe will be a long term project with significant implications.
Stephen E. Levinson, Weiyu Zhu, Danfeng Li, Kevin Squire, Ruei-Sung Lin, Matthew Kleffner, Matthew McClain, Johnny Lee
IJCNN5
2001 Tracking of Object with SVM Regression
abstract
This paper presents a novel feature-matching based approach for rigid object tracking. The proposed method models the tracking problem as discovering the affine transforms of object images between frames according to the extracted feature correspondences. False feature matches (outliers) are automatically detected and removed with a new SVM regression technique, where outliers are iteratively identified as support vectors with the gradually decreased insensitive margin /spl epsi/. This method, in addition to object tracking, can also be used for general feature-based epipolar constraint estimation, in which it can quickly detect outliers even if they make up, in theory, over 50% of the whole data. We have applied the proposed method to track real objects under cluttering backgrounds with very encouraging results.
Weiyu Zhu, Ruei-Sung Lin, Stephen E. Levinson
CVPR (2)3