Zhuolin Jiang

dblp:82/1160 · DBLP profile ↗
← Back
31ranked-venue papers
10as first author
1since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 8 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-authorSystems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
16 papers
Representation and self-supervised learning · 31% Trustworthy machine learning · 19% Video understanding and tracking · 16%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%
Theoretical computer science
3 papers
Mathematical optimization · 100%

Topics — the 30 heaviest of 35, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
action recognition
1.062016
Cross-View Action Recognition via Transferable Dictionary Learning · IEEE Trans. Image Process. 2016
Submodular Attribute Selection for Action Recognition in Video · NIPS 2014
Learning View-Invariant Sparse Representations for Cross-View Action Recognition · ICCV 2013
Machine learning › Representation and self-supervised learning
information bottleneck
0.912025
A Variational Information Theoretic Approach to Out-of-Distribution Detection · ICML 2025
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.912025
A Variational Information Theoretic Approach to Out-of-Distribution Detection · ICML 2025
Machine learning › Trustworthy machine learning
uncertainty estimation
0.912025
A Variational Information Theoretic Approach to Out-of-Distribution Detection · ICML 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning
0.742016
Cross-View Action Recognition via Transferable Dictionary Learning · IEEE Trans. Image Process. 2016
Label Consistent K-SVD: Learning a Discriminative Dictionary for Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Learning Structured Low-Rank Representations for Image Classification · CVPR 2013
Mathematical optimization
submodular optimization
0.532017
Submodular Attribute Selection for Visual Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Submodular dictionary learning for sparse coding · CVPR 2012
Submodular Attribute Selection for Action Recognition in Video · NIPS 2014
Computer vision › Video understanding and tracking › action recognition › robust action recognition
view-invariant action recognition
0.422016
Cross-View Action Recognition via Transferable Dictionary Learning · IEEE Trans. Image Process. 2016
Learning View-Invariant Sparse Representations for Cross-View Action Recognition · ICCV 2013
Computer vision › Image recognition and object detection
image classification
0.422015
Modeling Neuron Selectivity Over Simple Midlevel Features for Image Classification · IEEE Trans. Image Process. 2015
Learning Structured Low-Rank Representations for Image Classification · CVPR 2013
Information retrieval
cross-language information retrieval
0.412019
Neural-Network Lexical Translation for Cross-lingual IR from Text and Speech · SIGIR 2019
Computer vision › Face, body and person analysis
face recognition
0.432015
Label Consistent K-SVD: Learning a Discriminative Dictionary for Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Learning a discriminative dictionary for sparse coding via label consistent K-SVD · CVPR 2011
Modeling Neuron Selectivity Over Simple Midlevel Features for Image Classification · IEEE Trans. Image Process. 2015
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding
0.332013
Learning View-Invariant Sparse Representations for Cross-View Action Recognition · ICCV 2013
Submodular dictionary learning for sparse coding · CVPR 2012
Learning a discriminative dictionary for sparse coding via label consistent K-SVD · CVPR 2011
Computer vision › Image recognition and object detection
attribute-based recognition
0.312017
Submodular Attribute Selection for Visual Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Machine learning › Learning theory
information-theoretic learning
0.312025
A Variational Information Theoretic Approach to Out-of-Distribution Detection · ICML 2025
Machine learning › Transfer learning and domain adaptation › feature-based transfer learning
transfer dictionary learning
0.212016
Cross-View Action Recognition via Transferable Dictionary Learning · IEEE Trans. Image Process. 2016
Machine learning › Learning paradigms › unsupervised learning
clustering-based feature learning
0.212015
Modeling Neuron Selectivity Over Simple Midlevel Features for Image Classification · IEEE Trans. Image Process. 2015
Machine learning › Representation and self-supervised learning › visual representation › image representation › mid-level representation
mid-level representation learning
0.212015
Modeling Neuron Selectivity Over Simple Midlevel Features for Image Classification · IEEE Trans. Image Process. 2015
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.212015
Modeling Neuron Selectivity Over Simple Midlevel Features for Image Classification · IEEE Trans. Image Process. 2015
Machine learning › Kernel, tree and ensemble methods
attribute selection
0.212014
Submodular Attribute Selection for Action Recognition in Video · NIPS 2014
Computer vision › Segmentation and scene understanding › image segmentation › binary segmentation
foreground-background segmentation
0.212014
Submodular Object Recognition · CVPR 2014
Computer vision › Image recognition and object detection
object recognition
0.212014
Submodular Object Recognition · CVPR 2014
Machine learning › Optimization for machine learning › combinatorial optimization
submodular optimization
0.212014
Submodular Object Recognition · CVPR 2014
Computer vision › Image recognition and object detection › image classification
object classification
0.222012
Learning a discriminative dictionary for sparse coding via label consistent K-SVD · CVPR 2011
Submodular dictionary learning for sparse coding · CVPR 2012
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
discriminative sparse coding
0.212013
Label Consistent K-SVD: Learning a Discriminative Dictionary for Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › subspace learning
low-rank representation
0.212013
Learning Structured Low-Rank Representations for Image Classification · CVPR 2013
Computer vision › Segmentation and scene understanding
saliency detection
0.212013
Submodular Salient Region Detection · CVPR 2013
Information retrieval › multimedia analysis and retrieval
image annotation
0.212013
Tag Taxonomy Aware Dictionary Learning for Region Tagging · CVPR 2013
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding › dictionary learning
sparse dictionary learning
0.112011
Sparse dictionary-based representation and recognition of action attributes · ICCV 2011
Information retrieval › document retrieval
spoken document retrieval
0.112019
Neural-Network Lexical Translation for Cross-lingual IR from Text and Speech · SIGIR 2019
Computer vision › Face, body and person analysis › facial attribute analysis
gender recognition
0.112015
Modeling Neuron Selectivity Over Simple Midlevel Features for Image Classification · IEEE Trans. Image Process. 2015
Mathematical optimization
discrete optimization
0.112014
Submodular Attribute Selection for Action Recognition in Video · NIPS 2014

Methods — techniques the papers use, named apart from their topics

dictionary learning · 1.3variational inference · 0.9information bottleneck · 0.9KL divergence · 0.9sparse coding · 0.8greedy optimization · 0.6submodular optimization · 0.5linear classifier · 0.5greedy algorithm · 0.5sparse representation · 0.4word alignment · 0.4neural network · 0.4character sequence encoding · 0.4attribute selection · 0.2submodular maximization · 0.1matroid constraint · 0.1
YearPublicationVenuePosition
2025 A Variational Information Theoretic Approach to Out-of-Distribution Detection
abstract
We present a theory for the construction of out-of-distribution (OOD) detection features for neural networks. We introduce random features for OOD through a novel information-theoretic loss functional consisting of two terms, the first based on the KL divergence separates resulting in-distribution (ID) and OOD feature distributions and the second term is the Information Bottleneck, which favors compressed features that retain the OOD information. We formulate a variational procedure to optimize the loss and obtain OOD features. Based on assumptions on OOD distributions, one can recover properties of existing OOD features, i.e., shaping functions. Furthermore, we show that our theory can predict a new shaping function that out-performs existing ones on OOD benchmarks. Our theory provides a general framework for constructing a variety of new features with clear explainability.
Sudeepta Mondal, Zhuolin Jiang, Ganesh Sundaramoorthi
ICML2
2020 Towards a New Understanding of the Training of Neural Networks with Mislabeled Training Data
Herbert Gish, Jan Silovský, Man-Ling Sung, Man-Hung Siu, William Hartmann, Zhuolin Jiang
ICASSP6
2020 Analysis of the Average Neutral-Point Current Limits of the Neutral-Point-Clamped Converter Under Three-Level Modulation
abstract
The neutral-point-clamped (NPC) converter employed in applications like bipolar dc-bus, back-to-back converters, power compensator, etc., requires a nonzero average neutral-point (NP) current injection under certain operating conditions. In order to better implement the average NP current injection, this paper studies the maximum and minimum average NP current that can be injected with an NPC converter. For this study, a three-level modulation technique is employed in which one of the three phases is selected to switch between the positive dc-rail and the negative dc-rail while the other two phases switch in two consecutive voltage levels. The maximum and minimum average NP current injection limits are analyzed under different loading conditions. The presented study also evaluates the effects of modulation index and operating power factor on the limits.
Zhuolin Jiang, Neha Beniwal, Salvador Ceballos, Josep Pou, Hossein Dehghani Tafti, Glen Farivar
IECON1
2019 Neural-Network Lexical Translation for Cross-lingual IR from Text and Speech
abstract
We propose a neural network model to estimate word translation probabilities for Cross-Lingual Information Retrieval (CLIR). The model estimates better probabilities for word translations than automatic word alignments alone, and generalizes to unseen source-target word pairs. We further improve the lexical neural translation model (and subsequently CLIR), by incorporating source word context, and by encoding the character sequences of input source words to generate translations of out-of-vocabulary words. To be effective, neural network models typically need training on large amounts of data labeled directly on the final task, in this case relevance to queries. In contrast, our approach only requires parallel data to train the translation model, and uses an unsupervised model to compute CLIR relevance scores.
Rabih Zbib, Lingjun Zhao, Damianos Karakos, William Hartmann, Jay DeYoung, Zhongqiang Huang, Zhuolin Jiang, Noah Rivkin, Le Zhang 0002, Richard M. Schwartz, John Makhoul
SIGIR7
2017 Learning Transferable Representation for Bilingual Relation Extraction via Convolutional Neural Networks
abstract
Typically, relation extraction models are trained to extract instances of a relation ontology using only training data from a single language. However, the concepts represented by the relation ontology (e.g. ResidesIn, EmployeeOf) are language independent. The numbers of annotated examples available for a given ontology vary between languages. For example, there are far fewer annotated examples in Spanish and Japanese than English and Chinese. Furthermore, using only language-specific training data results in the need to manually annotate equivalently large amounts of training for each new language a system encounters. We propose a deep neural network to learn transferable, discriminative bilingual representation. Experiments on the ACE 2005 multilingual training corpus demonstrate that the joint training process results in significant improvement in relation classification performance over the monolingual counterparts. The learnt representation is discriminative and transferable between languages. When using 10% (25K English words, or 30K Chinese characters) of the training data, our approach results in doubling F1 compared to a monolingual baseline. We achieve comparable performance to the monolingual system trained with 250K English words (or 300K Chinese characters) With 50% of training data.
Bonan Min, Zhuolin Jiang, Marjorie Freedman, Ralph M. Weischedel
IJCNLP(1)2
2017 Learning Discriminative Features via Label Consistent Neural Network
abstract
Deep Convolutional Neural Networks (CNN) enforce supervised information only at the output layer, and hidden layers are trained by back propagating the prediction error from the output layer without explicit supervision. We propose a supervised feature learning approach, Label Consistent Neural Network, which enforces direct supervision in late hidden layers in a novel way. We associate each neuron in a hidden layer with a particular class label and encourage it to be activated for input signals from the same class. More specifically, we introduce a label consistency regularization called "discriminative representation error" loss for late hidden layers and combine it with classification error loss to build our overall objective function. This label consistency constraint alleviates the common problem of gradient vanishing and tends to faster convergence, it also makes the features derived from late hidden layers discriminative enough for classification even using a simple k-NN classifier. Experimental results demonstrate that our approach achieves state-of-the-art performances on several public datasets for action and object category recognition.
Zhuolin Jiang, Yaming Wang, Larry Davis 0001, Walter Andrews, Viktor Rozgic
WACV1
2017 Submodular Attribute Selection for Visual Recognition
abstract
In real-world visual recognition problems, low-level features cannot adequately characterize the semantic content in images, or the spatio-temporal structure in videos. In this work, we encode objects or actions based on attributes that describe them as high-level concepts. We consider two types of attributes. One type of attributes is generated by humans, while the second type is data-driven attributes extracted from data using dictionary learning methods. Attribute-based representation may exhibit variations due to noisy and redundant attributes. We propose a discriminative and compact attribute-based representation by selecting a subset of discriminative attributes from a large attribute set. Three attribute selection criteria are proposed and formulated as a submodular optimization problem. A greedy optimization algorithm is presented and its solution is guaranteed to be at least (1-1/e)-approximation to the optimum. Experimental results on four public datasets demonstrate that the proposed attribute-based representation significantly boosts the performance of visual recognition and outperforms most recently proposed recognition approaches.
Zhuolin Jiang, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 Cross-View Action Recognition via Transferable Dictionary Learning
abstract
Discriminative appearance features are effective for recognizing actions in a fixed view, but may not generalize well to a new view. In this paper, we present two effective approaches to learn dictionaries for robust action recognition across views. In the first approach, we learn a set of view-specific dictionaries where each dictionary corresponds to one camera view. These dictionaries are learned simultaneously from the sets of correspondence videos taken at different views with the aim of encouraging each video in the set to have the same sparse representation. In the second approach, we additionally learn a common dictionary shared by different views to model view-shared features. This approach represents the videos in each view using a view-specific dictionary and the common dictionary. More importantly, it encourages the set of videos taken from the different views of the same action to have the similar sparse representations. The learned common dictionary not only has the capability to represent actions from unseen views, but also makes our approach effective in a semi-supervised setting where no correspondence videos exist and only a few labeled videos exist in the target view. The extensive experiments using three public datasets demonstrate that the proposed approach outperforms recently developed approaches for cross-view action recognition.
Zhuolin Jiang, Rama Chellappa
IEEE Trans. Image Process.2
2015 Discriminative feature learning from big data for visual recognition
Zhuolin Jiang, Zhe Lin 0001, Haibin Ling, Fatih Porikli, Ling Shao 0001, Pavan Turaga
Pattern Recognit.1
2015 Modeling Neuron Selectivity Over Simple Midlevel Features for Image Classification
abstract
We now know that good mid-level features can greatly enhance the performance of image classification, but how to efficiently learn the image features is still an open question. In this paper, we present an efficient unsupervised midlevel feature learning approach (MidFea), which only involves simple operations, such as k-means clustering, convolution, pooling, vector quantization, and random projection. We show this simple feature can also achieve good performance in traditional classification task. To further boost the performance, we model the neuron selectivity (NS) principle by building an additional layer over the midlevel features prior to the classifier. The NS-layer learns category-specific neurons in a supervised manner with both bottom-up inference and top-down analysis, and thus supports fast inference for a query image. Through extensive experiments, we demonstrate that this higher level NS-layer notably improves the classification accuracy with our simple MidFea, achieving comparable performances for face recognition, gender classification, age estimation, and object categorization. In particular, our approach runs faster in inference by an order of magnitude than sparse coding-based feature learning methods. As a conclusion, we argue that not only do carefully learned features (MidFea) bring improved performance, but also a sophisticated mechanism (NS-layer) at higher level boosts the performance further.
Shu Kong, Zhuolin Jiang, Qiang Yang 0001
IEEE Trans. Image Process.2
2014 Submodular Reranking with Multiple Feature Modalities for Image Retrieval
Fan Yang 0016, Zhuolin Jiang, Larry Davis 0001
ACCV (1)2
2014 Submodular Object Recognition
abstract
We present a novel object recognition framework based on multiple figure-ground hypotheses with a large object spatial support, generated by bottom-up processes and mid-level cues in an unsupervised manner. We exploit the benefit of regression for discriminating segments' categories and qualities, where a regressor is trained to each category using the overlapping observations between each figure-ground segment hypothesis and the ground-truth of the target category in an image. Object recognition is achieved by maximizing a submodular objective function, which maximizes the similarities between the selected segments (i.e., facility locations) and their group elements (i.e., clients), penalizes the number of selected segments, and more importantly, encourages the consistency of object categories corresponding to maximum regression values from different category-specific regressors for the selected segments. The proposed framework achieves impressive recognition results on three benchmark datasets, including PASCAL VOC 2007, Caltech-101 and ETHZ-shape.
Fan Zhu 0001, Zhuolin Jiang, Ling Shao 0001
CVPR2
2014 Submodular Attribute Selection for Action Recognition in Video
Zhuolin Jiang, Rama Chellappa, P. Jonathon Phillips
NIPS2
2014 Online discriminative dictionary learning for visual tracking
abstract
Dictionary learning has been applied to various computer vision problems, such as image restoration, object classification and face recognition. In this work, we propose a tracking framework based on sparse representation and online discriminative dictionary learning. By associating dictionary items with label information, the learned dictionary is both reconstructive and discriminative, which better distinguishes target objects from the background. During tracking, the best target candidate is selected by a joint decision measure. Reliable tracking results and augmented training samples are accumulated into two sets to update the dictionary. Both online dictionary learning and the proposed joint decision measure are important for the final tracking performance. Experiments show that our approach outperforms several recently proposed trackers.
Fan Yang 0016, Zhuolin Jiang, Larry Davis 0001
WACV2
2013 System and algorithms on detection of objects embedded in perspective geometry using monocular cameras
abstract
In this work, we present a framework to detect objects embedded in complex perspective geometry. Our goal is to accurately identify objects such as people standing in balconies or windows on building facades of surrounding buildings. Compared to traditional computer vision work focused on activity analysis from a horizontal view, our framework provides a solution for the application domain of mobile surveillance in urban areas. A novel solution for a monocular camera is formulated by tightly coupling various computational modules including geometric analysis, segmentation, scale estimation, and object detection. In particular, our proposed approach alleviates the effect of the perspective geometry and corresponding distortion in object appearance effectively, and provides accurate scale priors to eliminate unlikely object detection hypotheses. The experimental results on collected video dataset show that the proposed approach is more accurate than traditional detection approaches based on brute-force scanning windows.
Yiliang Xu, Sangmin Oh, Fan Yang 0016, Zhuolin Jiang, Naresh P. Cuntoor, Anthony Hoogs, Larry Davis 0001
AVSS4
2013 Discriminative Tensor Sparse Coding for Image Classification
abstract
A novel approach to learn a discriminative dictionary over a tensor sparse model is presented. A structural incoherence constraint between dictionary atoms from different classes is introduced to promote discriminating information into the dictionary. The incoherence term encourages dictionary atoms to be as independent as possible. In addition, we incorporate classification error into the objective function of dictionary learning. The dictionary is learned in a supervised setting to make it useful for classification. A linear multi-class classifier and the dictionary are learned simultaneously during the training phase. Our approach is evaluated on three types of public databases, including texture, digit, and face databases. Experimental results demonstrate the effectiveness of our approach. 1
Yangmuzi Zhang, Zhuolin Jiang, Larry Davis 0001
BMVC2
2013 Submodular Salient Region Detection
abstract
The problem of salient region detection is formulated as the well-studied facility location problem from operations research. High-level priors are combined with low-level features to detect salient regions. Salient region detection is achieved by maximizing a sub modular objective function, which maximizes the total similarities (i.e., total profits) between the hypothesized salient region centers (i.e., facility locations) and their region elements (i.e., clients), and penalizes the number of potential salient regions (i.e., the number of open facilities). The similarities are efficiently computed by finding a closed-form harmonic solution on the constructed graph for an input image. The saliency of a selected region is modeled in terms of appearance and spatial location. By exploiting the sub modularity properties of the objective function, a highly efficient greedy-based optimization algorithm can be employed. This algorithm is guaranteed to be at least a (e - 1)/e 0.632-approximation to the optimum. Experimental results demonstrate that our approach outperforms several recently proposed saliency detection approaches.
Zhuolin Jiang, Larry Davis 0001
CVPR1
2013 Learning Structured Low-Rank Representations for Image Classification
abstract
An approach to learn a structured low-rank representation for image classification is presented. We use a supervised learning method to construct a discriminative and reconstructive dictionary. By introducing an ideal regularization term, we perform low-rank matrix recovery for contaminated training data from all categories simultaneously without losing structural information. A discriminative low-rank representation for images with respect to the constructed dictionary is obtained. With semantic structure information and strong identification capability, this representation is good for classification tasks even using a simple linear multi-classifier. Experimental results demonstrate the effectiveness of our approach.
Yangmuzi Zhang, Zhuolin Jiang, Larry Davis 0001
CVPR2
2013 Tag Taxonomy Aware Dictionary Learning for Region Tagging
abstract
Tags of image regions are often arranged in a hierarchical taxonomy based on their semantic meanings. In this paper, using the given tag taxonomy, we propose to jointly learn multi-layer hierarchical dictionaries and corresponding linear classifiers for region tagging. Specifically, we generate a node-specific dictionary for each tag node in the taxonomy, and then concatenate the node-specific dictionaries from each level to construct a level-specific dictionary. The hierarchical semantic structure among tags is preserved in the relationship among node-dictionaries. Simultaneously, the sparse codes obtained using the level-specific dictionaries are summed up as the final feature representation to design a linear classifier. Our approach not only makes use of sparse codes obtained from higher levels to help learn the classifiers for lower levels, but also encourages the tag nodes from lower levels that have the same parent tag node to implicitly share sparse codes obtained from higher levels. Experimental results using three benchmark datasets show that the proposed approach yields the best performance over recently proposed methods.
Zhuolin Jiang
CVPR2
2013 Learning View-Invariant Sparse Representations for Cross-View Action Recognition
abstract
We present an approach to jointly learn a set of view-specific dictionaries and a common dictionary for cross-view action recognition. The set of view-specific dictionaries is learned for specific views while the common dictionary is shared across different views. Our approach represents videos in each view using both the corresponding view-specific dictionary and the common dictionary. More importantly, it encourages the set of videos taken from different views of the same action to have similar sparse representations. In this way, we can align view-specific features in the sparse feature spaces spanned by the view-specific dictionary set and transfer the view-shared features in the sparse feature space spanned by the common dictionary. Meanwhile, the incoherence between the common dictionary and the view-specific dictionary set enables us to exploit the discrimination information encoded in view-specific features and view-shared features separately. In addition, the learned common dictionary not only has the capability to represent actions from unseen views, but also makes our approach effective in a semi-supervised setting where no correspondence videos exist and only a few labels exist in the target view. Extensive experiments using the multi-view IXMAS dataset demonstrate that our approach outperforms many recent approaches for cross-view action recognition.
Zhuolin Jiang
ICCV2
2013 A unified tree-based framework for joint action localization, recognition and segmentation
Zhuolin Jiang, Zhe Lin 0001, Larry Davis 0001
Comput. Vis. Image Underst.1
2013 Label Consistent K-SVD: Learning a Discriminative Dictionary for Recognition
abstract
A label consistent K-SVD (LC-KSVD) algorithm to learn a discriminative dictionary for sparse coding is presented. In addition to using class labels of training data, we also associate label information with each dictionary item (columns of the dictionary matrix) to enforce discriminability in sparse codes during the dictionary learning process. More specifically, we introduce a new label consistency constraint called "discriminative sparse-code error" and combine it with the reconstruction error and the classification error to form a unified objective function. The optimal solution is efficiently obtained using the K-SVD algorithm. Our algorithm learns a single overcomplete dictionary and an optimal linear classifier jointly. The incremental dictionary learning algorithm is presented for the situation of limited memory resources. It yields dictionaries so that feature points with the same class labels have similar sparse codes. Experimental results demonstrate that our algorithm outperforms many recently proposed sparse-coding techniques for face, action, scene, and object category recognition under the same learning conditions.
Zhuolin Jiang, Zhe Lin 0001, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 Discriminative Dictionary Learning with Pairwise Constraints
Huimin Guo, Zhuolin Jiang, Larry Davis 0001
ACCV (1)2
2012 Online Semi-Supervised Discriminative Dictionary Learning for Sparse Representation
Guangxiao Zhang, Zhuolin Jiang, Larry Davis 0001
ACCV (1)2
2012 Cross-View Action Recognition via a Transferable Dictionary Pair
abstract
Discriminative appearance features are effective for recognizing actions in a fixed view, but generalize poorly to changes in viewpoint. We present a method for viewinvariant action recognition based on sparse representations using a transferable dictionary pair. A transferable dictionary pair consists of two dictionaries that correspond to the source and target views respectively. The two dictionaries are learned simultaneously from pairs of videos taken at different views and aim to encourage each video in the pair to have the same sparse representation. Thus, the transferable dictionary pair links features between the two views that are useful for action recognition. Both unsupervised and supervised algorithms are presented for learning transferable dictionary pairs. Using the sparse representation as features, a classifier built in the source view can be directly transferred to the target view. We extend our approach to transferring an action model learned from multiple source views to one target view. We demonstrate the effectiveness of our approach on the multi-view IXMAS data set. Our results compare favorably to the the state of the art.
Zhuolin Jiang, P. Jonathon Phillips, Rama Chellappa
BMVC2
2012 Submodular dictionary learning for sparse coding
abstract
A greedy-based approach to learn a compact and discriminative dictionary for sparse representation is presented. We propose an objective function consisting of two components: entropy rate of a random walk on a graph and a discriminative term. Dictionary learning is achieved by finding a graph topology which maximizes the objective function. By exploiting the monotonicity and submodularity properties of the objective function and the matroid constraint, we present a highly efficient greedy-based optimization algorithm. It is more than an order of magnitude faster than several recently proposed dictionary learning approaches. Moreover, the greedy algorithm gives a near-optimal solution with a (1/2)-approximation bound. Our approach yields dictionaries having the property that feature points from the same class have very similar sparse codes. Experimental results demonstrate that our approach outperforms several recently proposed dictionary learning techniques for face, action and object category recognition.
Zhuolin Jiang, Guangxiao Zhang, Larry Davis 0001
CVPR1
2012 Class consistent k-means: Application to face and action recognition
Zhuolin Jiang, Zhe Lin 0001, Larry Davis 0001
Comput. Vis. Image Underst.1
2012 Recognizing Human Actions by Learning and Matching Shape-Motion Prototype Trees
abstract
A shape-motion prototype-based approach is introduced for action recognition. The approach represents an action as a sequence of prototypes for efficient and flexible action matching in long video sequences. During training, an action prototype tree is learned in a joint shape and motion space via hierarchical K-means clustering and each training sequence is represented as a labeled prototype sequence; then a look-up table of prototype-to-prototype distances is generated. During testing, based on a joint probability model of the actor location and action prototype, the actor is tracked while a frame-to-prototype correspondence is established by maximizing the joint probability, which is efficiently performed by searching the learned prototype tree; then actions are recognized using dynamic prototype sequence matching. Distance measures used for sequence matching are rapidly obtained by look-up table indexing, which is an order of magnitude faster than brute-force computation of frame-to-frame distances. Our approach enables robust action matching in challenging situations (such as moving cameras, dynamic backgrounds) and allows automatic alignment of action sequences. Experimental results demonstrate that our approach achieves recognition rates of 92.86 percent on a large gesture data set (with dynamic backgrounds), 100 percent on the Weizmann action data set, 95.77 percent on the KTH action data set, 88 percent on the UCF sports data set, and 87.27 percent on the CMU action data set.
Zhuolin Jiang, Zhe Lin 0001, Larry Davis 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2011 Learning a discriminative dictionary for sparse coding via label consistent K-SVD
abstract
A label consistent K-SVD (LC-KSVD) algorithm to learn a discriminative dictionary for sparse coding is presented. In addition to using class labels of training data, we also associate label information with each dictionary item (columns of the dictionary matrix) to enforce discriminability in sparse codes during the dictionary learning process. More specifically, we introduce a new label consistent constraint called `discriminative sparse-code error' and combine it with the reconstruction error and the classification error to form a unified objective function. The optimal solution is efficiently obtained using the K-SVD algorithm. Our algorithm learns a single over-complete dictionary and an optimal linear classifier jointly. It yields dictionaries so that feature points with the same class labels have similar sparse codes. Experimental results demonstrate that our algorithm outperforms many recently proposed sparse coding techniques for face and object category recognition under the same learning conditions.
Zhuolin Jiang, Zhe Lin 0001, Larry Davis 0001
CVPR1
2011 Sparse dictionary-based representation and recognition of action attributes
abstract
We present an approach for dictionary learning of action attributes via information maximization. We unify the class distribution and appearance information into an objective function for learning a sparse dictionary of action attributes. The objective function maximizes the mutual information between what has been learned and what remains to be learned in terms of appearance information and class distribution for each dictionary item. We propose a Gaussian Process (GP) model for sparse representation to optimize the dictionary objective function. The sparse coding property allows a kernel with a compact support in GP to realize a very efficient dictionary learning process. Hence we can describe an action video by a set of compact and discriminative action attributes. More importantly, we can recognize modeled action categories in a sparse feature space, which can be generalized to unseen and unmodeled action categories. Experimental results demonstrate the effectiveness of our approach in action recognition applications.
Qiang Qiu 0002, Zhuolin Jiang, Rama Chellappa
ICCV2
2009 Recognizing actions by shape-motion prototype trees
abstract
A prototype-based approach is introduced for action recognition. The approach represents an action as a sequence of prototypes for efficient and flexible action matching in long video sequences. During training, first, an action prototype tree is learned in a joint shape and motion space via hierarchical k-means clustering; then a lookup table of prototype-to-prototype distances is generated. During testing, based on a joint likelihood model of the actor location and action prototype, the actor is tracked while a frame-to-prototype correspondence is established by maximizing the joint likelihood, which is efficiently performed by searching the learned prototype tree; then actions are recognized using dynamic prototype sequence matching. Distance matrices used for sequence matching are rapidly obtained by look-up table indexing, which is an order of magnitude faster than brute-force computation of frame-to-frame distances. Our approach enables robust action matching in very challenging situations (such as moving cameras, dynamic backgrounds) and allows automatic alignment of action sequences. Experimental results demonstrate that our approach achieves recognition rates of 91.07% on a large gesture dataset (with dynamic backgrounds), 100% on the Weizmann action dataset and 95.77% on the KTH action dataset.
Zhe Lin 0001, Zhuolin Jiang, Larry Davis 0001
ICCV2