Lei Zhang 0093

dblp:97/8704-93 · DBLP profile ↗
← Back
40ranked-venue papers
21as first author
21since 2021 · last 2026
0000-0002-8494-0504ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 19 first-author · 10 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 : Localized text prompt refinement for zero-shot referring image segmentation
Lei Zhang 0093, Yongqiu Huang, Yingjun Du, Fang Lei, Zhiying Yang, Cees Snoek, Yehui Wang
Comput. Vis. Image Underst.1
2026 Instance-aware visual-semantic interaction for zero-shot learning
Yongcai Chen, Xinfa Shi, Yingjun Du, Lei Zhang 0093
Expert Syst. Appl.5
2026 UniAD: Unified cross-modal prompt regularization for zero-shot anomaly detection across domains
Xuezhi Xiang, Songran Luo, Yingjun Du, Lei Zhang 0093, Xiantong Zhen
Neurocomputing4
2025 LKA-ReID: Vehicle Re-Identification with Large Kernel Attention
abstract
With the rapid development of intelligent transportation systems and the popularity of smart city infrastructure, Vehicle Re-ID technology has become an important research field. The vehicle Re-ID task faces an important challenge, which is the high similarity between different vehicles. Existing methods use additional detection or segmentation models to extract differentiated local features. However, these methods either rely on additional annotations or greatly increase the computational cost. Using attention mechanism to capture global and local features is crucial to solve the challenge of high similarity between classes in vehicle Re-ID tasks. In this paper, we propose LKA-ReID with large kernel attention. Specifically, the large kernel attention (LKA) utilizes the advantages of self-attention and also benefits from the advantages of convolution, which can extract the global and local features of the vehicle more comprehensively. We also introduce hybrid channel attention (HCA), which combines channel attention with spatial information, so that the model can better focus on channels and feature regions, and ignore background and other disturbing information. Experiments on VeRi-776 dataset demonstrated the effectiveness of LKA-ReID, with mAP reaches 86.65% and Rank-1 reaches 98.03%.
Xuezhi Xiang, Zhushan Ma, Lei Zhang 0093, Denis Ombati, Himaloy Himu, Xiantong Zhen
ICASSP3
2025 Mamba-SF: Monocular Scene Flow Learning with State Space Models
abstract
Monocular scene flow estimation has been a long-standing problem in computer vision. Methods based on the RAFT architecture are currently the mainstream approaches, while often overlooking the long-range dependencies in motion and texture features and fail to fully utilize the spatial information in texture features. In this paper, we consider that using Transformers introduces high computational complexity. Therefore, we propose the Mamba Motion Module based on State Space Models design, which first models long-range dependencies in motion and texture features and then fully leverages the spatial information in texture features to enhance the motion features, generating global motion features while maintaining low computational complexity. Additionally, texture features play a significant role in constraining motion boundaries. Therefore, we propose the Enhanced Texture Module, which predicts a set of channel weights from the global motion features to enrich the channel properties of the texture features and concatenates texture features with the global motion features along the channels, thereby improving scene flow accuracy. Experimental results show that our method achieves highly competitive results on the KITTI 2015 and Eigen Split datasets, increasing by 18.82% and 2.15% compared to the baseline, respectively.
Xuezhi Xiang, Xianye Ben, Insha Hassan, Mingliang Zhai, Lei Zhang 0093, Xiantong Zhen
ICIP6
2025 Refining and adaption: multi-modal learning in wear debris analysis
abstract
Wear debris analysis(WDA) is a critical predictive maintenance technique that provides key data for diagnosing wear faults in mechanical equipment. This technique helps ensure long-term stable operation and fault prediction for machinery. However, recent approaches in WDA rely heavily on data-driven methods and often neglect the most important expert knowledge, such as the morphology, size, and color of debris. This knowledge is essential for fault diagnosis, but traditional networks struggle to represent and integrate it effectively, in contrast, large language models (LLMs) and vision-language models (VLMs), with their exceptional generalization capabilities, are rapidly evolving into ideal tools for representing and integrating expert knowledge. Inspired by this, we propose a novel multimodal learning method that combines VLMs and LLMs to integrate expert knowledge into image information, extending general models to specialized industrial domains. Specifically, we use an LLMs to regenerate textual descriptions from expert knowledge and refine CLIP’s Layer Normalization parameters to enhance the model’s transferability to specialized datasets, such as ferrography images, for zero-shot learning. Additionally, we introduce an adaptive meta-learning strategy that allows the model to quickly adapt to ferrography images with limited samples, further improving performance. Experimental results on industrial datasets show that the proposed method effectively addresses the challenge of applying VLMs in data-limited specialized domains, significantly improving the accuracy and robustness of WDA.
Yongcai Chen, Fang Lei, Xiantong Zhen, Xin Li 0100, Lei Zhang 0093
IJCNN5
2025 Deep scene flow learning from point cloud with Transformer
Xuezhi Xiang, Rokia Abdein, Lei Zhang 0093, Xiantong Zhen
Neurocomputing4
2025 PVFT-Net: A point-voxel fusion method for self-supervised scene flow estimation with transformer
Xuezhi Xiang, Xiaoheng Li, Xiankun Zhou, Lei Zhang 0093, Xiantong Zhen
Neurocomputing5
2025 De-noising mask transformer for referring image segmentation
Yehui Wang, Fang Lei, Baoyan Wang, Xiantong Zhen, Lei Zhang 0093
Image Vis. Comput.6
2025 Vehicle re-identification with large separable kernel attention and hybrid channel attention
Xuezhi Xiang, Zhushan Ma, Xiaoheng Li, Lei Zhang 0093, Xiantong Zhen
Image Vis. Comput.4
2025 Self-supervised monocular depth estimation with large kernel attention and dynamic scene perception
Xuezhi Xiang, Xiaoheng Li, Lei Zhang 0093, Xiantong Zhen
J. Vis. Commun. Image Represent.4
2024 Deep Optical Flow Learning With Deformable Large-Kernel Cross-Attention
abstract
Optical flow estimation from image sequences is a fundamental problem in computer vision. In recent years, some methods have utilized Transformer to model global dependencies and improve optical flow, achieving impressive performance. However, in these methods, Transformers typically treat two-dimensional image features as one-dimensional sequences. While position encoding partially mitigates the loss of position information between different feature patches, Transformer still lacks inherent biases for modeling local visual patterns and tend to overlook channel characteristics in image features. Therefore, this paper introduces a deformable large kernel attention module, combining the strengths of convolution and attention mechanisms, which can preserve feature channel adaptability while modeling global dependencies without compromising the two-dimensional structure of features, significantly enhancing optical flow estimation. Additionally, the introduced deformable mechanism allows the model to adapt appropriately to different data patterns. Experimental results demonstrate that our optical flow estimation method achieves competitive results on publicly available benchmarks such as Sintel and KITTI.
Xuezhi Xiang, Denis Ombati, Lei Zhang 0093, Xiantong Zhen
ICIP4
2024 Lightweight improved residual network for efficient inverse tone mapping
Liqi Xue, Yongbao Song, Yan Liu 0004, Lei Zhang 0093, Xiantong Zhen, Jun Xu 0019
Multim. Tools Appl.5
2024 Variational Neuron Shifting for Few-Shot Image Classification Across Domains
abstract
Few-shot image classification aims to recognize unseen classes with few labeled samples. Existing meta-learning models learn the ability of learning good representation or model parameters, in order to adapt to new tasks with a few training samples. However, when there exists a domain gap between training and test tasks, the learned ability often does not generalize well across domains, resulting in degraded performance on new tasks. In this article, we propose variational neuron shifting to generate adapted feature representations for few-shot learning. To do so, we introduce a working memory module to store the shifted neurons from the support set, which will be accessed to generate adapted feature representations of query samples. Under the meta-learning paradigm, the model is learned to acquire the ability of adaptation with single sample at meta-training time so as to further adapt itself to each single test sample at meta-test time. We formulate the adaptation process as a variational Bayesian inference problem, which incorporates the test sample as the condition into the generation of the model neuron shifting. We conduct extensive experiments on both within and across domain few-shot classification tasks. The new state-of-the-art performance substantiates the effectiveness of our variational neuron shifting. The thorough ablation studies further demonstrate the benefit of each component in our model.
Liyun Zuo, Baoyan Wang, Lei Zhang 0093, Jun Xu 0019, Xiantong Zhen
IEEE Trans. Multim.3
2023 Learning to Learn With Variational Inference for Cross-Domain Image Classification
abstract
Learning models that can generalize to previously unseen domains to which we have no access is a fundamental yet challenging problem in machine learning. In this paper, we propose meta variational inference (MetaVI), a variational Bayesian framework of meta-learning for cross domain image classification. Within the meta learning setting, MetaVI is derived to learn a probabilistic latent variable model by maximizing a meta evidence lower bound (Meta ELBO) for knowledge transfer across domains. To enhance the discriminative ability of the model, we further introduce a Wasserstein distance based constraint to the variational objective, leading to the Wasserstein MetaVI, which largely improves classification performance. By casting into a probabilistic inference problem, MetaVI offers the first, principled variational meta-learning framework for cross domain learning. In addition, we collect a new visual recognition dataset to contribute a more challenging benchmark for cross domain learning, which will be released to the public. Extensive experimental evaluation and ablation studies on four benchmarks show that our Wasserstein MetaVI achieves new state-of-the-art performance and surpasses previous methods, demonstrating its great effectiveness.
Lei Zhang 0093, Yingjun Du, Xiantong Zhen
IEEE Trans. Multim.1
2022 Meta conditional variational auto-encoder for domain generalization
Zhiqiang Ge, Xin Li 0100, Lei Zhang 0093
Comput. Vis. Image Underst.4
2022 Spherical Zero-Shot Learning
abstract
Zero-shot Learning (ZSL) is a highly non-trivial task to generalize from seen to unseen classes. In this paper, we propose spherical zero-shot learning (SZSL) to address the major challenges in ZSL. By decoupling the similarity metric in the spherical embedding space into radius and angle, our SZSL can map classes to hyperspherical surfaces of different radiuses, which greatly increases its flexibility. Specifically, we introduce the spherical alignment on angles to spread classes as uniformly as possible to alleviate the hubness problem and simultaneously preserve the inter-class semantic structure to make the alignment more reasonable. We also introduce the spherical calibration with a minimum entropy based regularizer by adopting a larger radius for unseen classes than seen classes to reduce the prediction bias. Extensive experiments on five middle-scale benchmarks and large-scale ImageNet dataset demonstrate that the proposed approach consistently achieves superior performance for the traditional and generalized settings of ZSL.
Zehao Xiao, Xiantong Zhen, Lei Zhang 0093
IEEE Trans. Circuits Syst. Video Technol.4
2022 Variational Hyperparameter Inference for Few-Shot Learning Across Domains
abstract
The focus of few shot learning research has been on the development of meta-learning recently, where a meta-learner is trained on a variety of tasks in hopes of being generalizable to new tasks. Tasks in meta training and meta test are usually assumed to be from the same domain, which would not necessarily hold in real world scenarios. In this paper, we propose variational hyperparameter inference for few-shot learning across domains. Based on an especially successful algorithm named model agnostic meta learning, the proposed variational hyperparameter inference integrates meta learning and variational inference into the optimization of hyperparameters, which enables the meta-learner with adaptivity for generalization across domains. In particular, we choose to learn adaptive hyperparameters including the learning rate and weight decay to avoid the failure in the face of few labeled examples across domain. Moreover, we model hyperparameters as distributions instead of fixed values, which will further enhance the generalization ability by capturing the uncertainty. Extensive experiments are conducted on two benchmark datasets including few shot learning dataset within-domain and across-domain. The results demonstrate that our methods outperforms previous approaches consistently, and comprehensive ablation studies further validate its effectiveness on few shot learning both within domains and across domains.
Lei Zhang 0093, Liyun Zuo, Baoyan Wang, Xin Li 0100, Xiantong Zhen
IEEE Trans. Circuits Syst. Video Technol.1
2021 Unbalanced data processing using deep sparse learning technique
Xin Li 0100, Lei Zhang 0093
Future Gener. Comput. Syst.2
2021 A Comprehensive Study on VLAD
Xin Li 0100, Lei Zhang 0093, Zhiping Jian, Liyun Zuo
Neural Process. Lett.2
2021 Learning to Adapt With Memory for Probabilistic Few-Shot Learning
abstract
Few-shot learning has recently generated increasing popularity in machine learning, which addresses the fundamental yet challenging problem of learning to adapt to new tasks with the limited data. In this paper, we propose a new probabilistic framework that learns to fast adapt with external memory. We model the classifier parameters as distributions that are inferred from the support set and directly applied to the query set for prediction. The model is optimized by formulating as a variational inference problem. The probabilistic modeling enables better handling prediction uncertainty due to the limited data. We impose a discriminative constraint on the feature representations by exploring the class structure, which can improve the classification performance. We further introduce a memory unit to store task-specific information extracted from the support set and used for the query set to achieve explicit adaption to individual tasks. By episodic training, the model learns to acquire the capability of adapting to specific tasks, which guarantees its performance on new related tasks. We conduct extensive experiments on widely-used benchmarks for few-shot recognition. Our method achieves new state-of-the-art performance and largely surpassing previous methods by large margins. The ablation study further demonstrates the effectiveness of the proposed discriminative learning and memory unit.
Lei Zhang 0093, Liyun Zuo, Yingjun Du, Xiantong Zhen
IEEE Trans. Circuits Syst. Video Technol.1
2020 Variational Image Deraining
abstract
Images captured in severe weather such as rain and snow significantly degrade the accuracy of vision systems, e.g., for outdoor video surveillance or autonomous driving. Image deraining is a critical yet highly challenging task, due to the fact that rain density varies across spatial locations, while the distribution patterns simultaneously vary across color channels. In this paper, we propose a variational image deraining (VID) method by formulating image deraining in a conditional variational auto-encoder framework. To achieve adaptive deraining to spatial rain density, we generate a density estimation map for each color channel, which can largely avoid over and under deraining. In addition, to address cross-channel variations, we conduct channel-wise deraining, motivated by our observation that bright pixels do not tend to remain bright after deraining unless their color channels are handled separately. Experimental results show that the proposed deraining method achieves superior performance on both synthesized and real rainy images, surpassing previous state-of-the-art methods by large margins.
Yingjun Du, Jun Xu 0019, Qiang Qiu 0001, Xiantong Zhen, Lei Zhang 0093
WACV5
2020 Video event classification based on two-stage neural network
Lei Zhang 0093, Xuezhi Xiang
Multim. Tools Appl.1
2020 Heterogenous output regression network for direct face alignment
Xiantong Zhen, Mengyang Yu, Zehao Xiao, Lei Zhang 0093, Ling Shao 0001
Pattern Recognit.4
2020 Calibrated Multivariate Regression Networks
abstract
In this paper, we propose a new multi-layer learning architecture, the calibrated multivariate regression network (CMRN). Compared to previous multivariate models, the CMRN is able to simultaneously handle major challenges in multivariate regression including highly nonlinear input-output relationships, underlying inter-output correlations and calibration of multiple outputs within one single framework. The CMRN is comprised of a nonlinear module with cosine activations and a linear module with the low-rank expansion, which establishes a compact multivariate regression network. By seamlessly working with the ℓ2,1loss, the CMRN automatically calibrates multiple outputs with distinct noise levels to achieve improved performance. Being succinctly formulated but theoretically well-founded, the CMRN offers a compact multi-layer learning architecture that can be efficiently trained to scale up with massive datasets. We conduct extensive experimental evaluation on two representative large multivariate regression tasks for both machine learning and computer vision. The proposed CMRN can produce high performance on all tasks, which is better or competitive to state-of-the-art models. Extensive ablation studies offer deep insights into the effectiveness of the proposed CMRN.
Lei Zhang 0093, Yingjun Du, Xin Li 0100, Xiantong Zhen
IEEE Trans. Circuits Syst. Video Technol.1
2019 Learning Match Kernels on Grassmann Manifolds for Action Recognition
abstract
Action recognition has been extensively researched in computer vision due to its potential applications in a broad range of areas. The key to action recognition lies in modeling actions and measuring their similarity, which however poses great challenges. In this paper, we propose learning match kernels between actions on Grassmann manifold for action recognition. Specifically, we propose modeling actions as a linear subspace on the Grassmann manifold; the subspace is a set of convolutional neural network (CNN) feature vectors pooled temporally over frames in semantic video clips, which simultaneously captures local discriminant patterns and temporal dynamics of motion. To measure the similarity between actions, we propose Grassmann match kernels (GMK) based on canonical correlations of linear subspaces to directly match videos for action recognition; GMK is learned in a supervised way via kernel target alignment, which is endowed with a great discriminative ability to distinguish actions from different classes. The proposed approach leverages the strengths of CNNs for feature extraction and kernels for measuring similarity, which accomplishes a general learning framework of match kernels for action recognition. We have conducted extensive experiments on five challenging realistic data sets including Youtube, UCF50, UCF101, Penn action, and HMDB51. The proposed approach achieves high performance and substantially surpasses the state-of-the-art algorithms by large margins, which demonstrates the great effectiveness of proposed approach for action recognition.
Lei Zhang 0093, Xiantong Zhen, Ling Shao 0001, Jingkuan Song
IEEE Trans. Image Process.1
2018 Spatial Ensemble Kernel Learning for Scene Classification
abstract
Scene recognition is one of the most important tasks in computer vision. Apart from appearance, spatial layout carries the crucial cue for discriminative representation. In this paper, we propose spatial ensemble kernel (SEK) learning, which enables fusion of multi-scale spatial information to achieve compact while discriminative representation of scenes. Based on the spatial pyramid, SEK combines the CNN features in each level of the pyramid in an ensemble and fuse them by kernels. By kernel approximation, we achieve Fourier feature embedding of CNN features in each scale, which establishes a nonlinear layer of the neural network with a cosine activation function. The parameters of the nonlinear layer can be learned jointly in one single optimization framework by supervised learning, which enables compact and discriminative feature representations. We show the effectiveness of the proposed SEK on two recent scene benchmark datasets, i.e., MIT indoor and SUN 397. The propose SEK produces high performance on two datasets which are competitive to state-of-the-art algorithms.
Lei Zhang 0093, Xiantong Zhen, Qiujing Zhang
ICASSP1
2017 Realistic human action recognition: When CNNS meet LDS
abstract
In this paper, we proposed new framework for human action representation, which leverages the strengths of convolutional neural networks (CNNs) and the linear dynamical system (LDS) to represent both spatial and temporal structures of actions in videos. We make two principal contributions: first, we incorporate image-trained CNNs to detect action clip concepts, which takes advantage of different levels of information by combining the two layers in CNNs trained from images; Second, we further propose adopting a linear dynamical system (LDS) to model the relationships between these clip concepts, which captures temporal structures of actions. We have applied the proposed method on two challenging realistic benchmark datasets, and our method achieves high performance up to 86.16% on the YouTube and 82.76% UCF50 datasets, which largely outperforms most of the state-of-the-art algorithms with more sophisticated techniques.
Lei Zhang 0093, Yangyang Feng, Xuezhi Xiang, Xiantong Zhen
ICASSP1
2017 Learning discriminant grassmann kernels for image-set classification
abstract
Image-set classification has recently generated great popularity due to widespread application to challenging tasks in computer vision. The great challenges arise from measuring the similarity between image sets which usually exhibit huge inter-class ambiguity and intra-class variation. In this paper, based on the assumption that each image set as a linear subspace can be treated as a point on a Grassmann manifold, we propose discriminant Grassmann kernels (DGK) of principal angles between subspaces. To tackle the ambiguity and variation, we propose learning the DGK via kernel target alignment, which achieves kernels of great discrimination by maximizing correlations with class labels. The proposed DGK has been evaluated on two challenging datasets including the ETH-80 and UCSD datasets for object recognition and video-based traffic congestion recognition, respectively. Extensive experiments have shown that the proposed DGKs achieves state-of-the-art performance and surpasses most of previous methods, which demonstrates the great effectiveness of the DGKs for image-set classification.
Lei Zhang 0093, Xuezhi Xiang, Xiantong Zhen
ICIP1
2016 Realistic human action recognition: When deep learning meets VLAD
abstract
Human action recognition from realistic scenarios is extremely challenging due to large intra-class variation and complex background clutters. In this paper, by leveraging the strength of deep learning and vector of locally aggregated descriptors (VLAD), we propose a new methods for human action recognition from realistic datsets. We adopt stack convolu-tional independent subspace analysis (ISA) networks to learn 3D cuboid representation directly from spatio-temporal video data; we propose an improved VLAD by incorporating the spatio-temporal geometrical information to encode the deep learned local features. On two challenging realistic datasets: the YouTube action and HMDB51 datasets, the proposed method achieves state-of-the-art performance with an efficient linear SVM classifier, which is competitive with and even better than existing sophisticated algorithms.
Lei Zhang 0093, Yangyang Feng, Jiqing Han 0001, Xiantong Zhen
ICASSP1
2016 Towards optimal vlad for human action recognition from still images
abstract
Human action recognition from still image has recently drawn increasing attention in human behavior analysis vision and also poses great challenges due to the huge inter ambiguity and intra variability. Vector of locally aggregated descriptors (VLAD) has achieved state-of-the-art performance in many image classification tasks based on local features. The great success of VLAD is largely due to its high descriptive ability and computational efficiency. In this paper, towards optimal VLAD representations for human action recognition from still images, we improve VLAD by tackling two important issues in VLAD including empty cavity and assignment ambiguity. The empty cavity issue severely compromises the performance of VLAD and has long been overlooked. We investigate the empty cavity and provide an effective solution to deal with it, which largely improves the performance of VLAD; we propose middle level assignments to conquer the assignment ambiguity, which are more reliable and can provide more useful information for realistic activity. We have conducted extensive experiments on two widely-used benchmarks to validate the proposed method for human action recognition from still images. Our method produces competitive performance with state-of-the-art algorithms.
Lei Zhang 0093, Xiantong Zhen, Jiqing Han 0001
ICASSP1
2016 Towards optimal VLAD for human action recognition from still images
Lei Zhang 0093, Changxi Li, Peipei Peng, Xuezhi Xiang, Jingkuan Song
Image Vis. Comput.1
2015 Dimensionality reduction by supervised locality analysis
abstract
High-dimensional feature representations have recently been widely used for image classification, which not only induce large storage requirement and high computational complexity, but also tend to be lack of discrimination due to redundant and noisy features. In this paper, we propose a novel algorithm named supervised locality analysis (SLA) for dimensionality reduction. In contrast to conventional dimensionality reduction methods, the proposed SLA incorporates supervision into locality analysis by fully exploring multi-class distributions, which can handle the non-linear data structure while preserving intrinsic discriminative information. The obtained compact and highly discriminative features by the SLA is enables more accurate and efficient classification. Moreover, the SLA can be used for supervised dimensionality reduction of both handcrafted and deep learning based features. We have conduced experiments to evaluate the proposed SLA on three datasets for image classification. The SLA has produced state-of-the-art performance and largely outperformed widely-used dimensionality reduction methods.
Lei Zhang 0093, Peipei Peng, Xuezhi Xiang, Xiantong Zhen
ICIP1
2014 Learning semantic kernels for scene classification
abstract
In this paper we propose to learn semantic kernels for scene classification. We first decompose the Object Bank representation into subspaces associated with each object, Anchor Objects are then created by clustering for each scene class separately. The Anchor Distances are computed to measure the distance between objects to scene classes. In order to take the advantage of the discriminative information from different scene classes, we propose semantic kernels based on the anchor distances to different classes for scene classification. Through extensive experiments on two benchmark datasets: UIUC-Sports dataset and 15-Scene dataset, we prove that the proposed Semantic Kernels can significantly improve the original Object Bank and achieve state-of-the-art performance.
Lei Zhang 0093, Xiantong Zhen, Jiqing Han 0001, Xuezhi Xiang
ICASSP1
2014 Learning Object-to-Class Kernels for Scene Classification
abstract
High-level image representations have drawn increasing attention in visual recognition, e.g., scene classification, since the invention of the object bank. The object bank represents an image as a response map of a large number of pretrained object detectors and has achieved superior performance for visual recognition. In this paper, based on the object bank representation, we propose the object-to-class (O2C) distances to model scene images. In particular, four variants of O2C distances are presented, and with the O2C distances, we can represent the images using the object bank by lower-dimensional but more discriminative spaces, called distance spaces, which are spanned by the O2C distances. Due to the explicit computation of O2C distances based on the object bank, the obtained representations can possess more semantic meanings. To combine the discriminant ability of the O2C distances to all scene classes, we further propose to kernalize the distance representation for the final classification. We have conducted extensive experiments on four benchmark data sets, UIUC-Sports, Scene-15, MIT Indoor, and Caltech-101, which demonstrate that the proposed approaches can significantly improve the original object bank approach and achieve the state-of-the-art performance.
Lei Zhang 0093, Xiantong Zhen, Ling Shao 0001
IEEE Trans. Image Process.1
2013 Towards optimal object bank for scene classification
abstract
High-level image representations have drawn increasing attention in visual recognition, e.g., scene classification, since the invention of the object bank (OB). The object bank represents an image as a response map of a large number of pre-trained object detectors and has achieved superior performances for visual recognition. However, the object bank representation can be further improved by considering the distributions of the object across categories and the discriminative contributions to the image representation. In this paper, we propose an optimal object bank (OOB) by imposing weights on the detectors according to their discriminative abilities. Through extensive experiments on two benchmark datasets: UIUC-Sports dataset and 15-Scene dataset, we prove that the proposed OOB can significantly improve the original object bank and achieves state-of-the-art performances.
Lei Zhang 0093, Shouzhi Xie, Xiantong Zhen
ICASSP1
2013 Recognizing actions via sparse coding on structure projection
abstract
In this paper, we propose a novel method for human action recognition based on sparse coding with a pyramid matching. Spatio-temporal interest points (STIPs) are firstly detected by a newly developed detector named spatio-temporal steerable detector (STSD). To effectively capture the distribution of STIPs in the video sequence, we propose to project the STIPs onto the three orthogonal planes (TOP), and we employ a sparse coding algorithm combined with the spatial pyramid matching to encode the layout of STIPs. Therefore the structure of an action are sufficiently encoded, obtaining a informative holistic descriptor for action representation. Extensive experiments have been conducted on KTH and HMDB51 datasets. Our method achieves the state-of-the-art performance for action recognition showing the effectiveness of the proposed methods for human action representation.
Lei Zhang 0093, Xiantong Zhen
ICIP1
2013 Discriminative high-level representations for scene classification
abstract
High-level image representations, e.g, Object Bank, have drawn increasing attention in visual recognition. In this paper, we propose a discriminative high-level representation based on object bank for scene classification. By projecting the high-level features from the object bank into discriminative subspaces, which are obtained by clustering the features in a supervised way, the final representations are more compact and discriminative. We have conducted extensive experiments on two benchmark datasets: UIUC-Sports dataset and 15-Scene dataset, which demonstrates that the proposed approach can significantly improve the original object bank and achieves state-of-the-art performances.
Lei Zhang 0093, Shouzhi Xie, Xiantong Zhen
ICIP1
2012 High order co-occurrence of visualwords for action recognition
abstract
This paper exploits the high order co-occurrence information for human action representation. Based on the bag-of-words (BoW) model, visual words are mapped into a co-occurrence space through latent semantic analysis (LSA). High order co-occurrence of the visual words is well captured and therefore the representation of actions in the co-occurrence space becomes more informative and compact. Since the representation is effective and efficient, and is less affected by the sizes of the codebook, it can be easily integrated into models based on BoW. Evaluations on the benchmark KTH dataset and the realistic HMDB51 dataset demonstrates that the proposed approach significantly improves the baseline BoW model and therefore is promising for human action recognition.
Lei Zhang 0093, Xiantong Zhen, Ling Shao 0001
ICIP1
2000 An environment model-based robust speech recognition
Lei Zhang 0093, Jiqing Han 0001, Chengguo Lv, Chengfa Wang
INTERSPEECH1