Chao Zhu 0003

dblp:76/445-3 · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0001-5486-7492ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 13 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 OSS-OCL: Occlusion Scenario Simulation and Occluded-edge Concentrated Learning for pedestrian detection
Keqi Lu, Chao Zhu 0003, Mengyin Liu, Xu-Cheng Yin
Pattern Recognit. Lett.2
2025 HA-FGOVD: Highlighting Fine-Grained Attributes via Explicit Linear Composition for Open-Vocabulary Object Detection
abstract
Open-vocabulary object detection (OVD) models are considered to be Large Multi-modal Models (LMM), due to their extensive training data and a large number of parameters. Mainstream OVD models prioritize object coarse-grained category rather than focus on their fine-grained attributes, e.g., colors or materials, thus failed to identify objects specified with certain attributes. Despite being pretrained on large-scale image-text pairs with rich attribute information, their latent feature space does not highlight these fine-grained attributes. In this paper, we introduce HA-FGOVD, a universal and explicit method that enhances the attribute-level detection capabilities of frozen OVD models by highlighting fine-grained attributes in explicit linear space. Our approach uses a LLM to extract attribute words in input text as a zero-shot task. Then, token attention masks are adjusted to guide text encoders in extracting both global and attribute-specific features, which are explicitly composited as two vectors in linear space to form a new attribute-highlighted feature for detection tasks. The composition weight scalars can be learned or transferred across different OVD models, showcasing the universality of our method. Experimental results show that HA-FGOVD achieves state-of-the-art performance on the FG-OVD benchmark and demonstrates promising generalization on the OVDEval benchmark, suggesting that our method addresses significant limitations in fine-grained attribute detection and has potential for broader fine-grained detection applications.
Mengyin Liu, Chao Zhu 0003, Xu-Cheng Yin
IEEE Trans. Multim.3
2024 Unsupervised Multi-view Pedestrian Detection
abstract
With the prosperity of the intelligent surveillance, multiple cameras have been applied to localize pedestrians more accurately. However, previous methods rely on laborious annotations of pedestrians in every frame and camera view. Therefore, we propose in this paper an Unsupervised Multi-view Pedestrian Detection approach (UMPD) to learn an annotation-free detector via vision-language models and 2D-3D cross-modal mapping: 1) Firstly, Semantic-aware Iterative Segmentation (SIS) is proposed to extract unsupervised representations of multi-view images, which are converted into 2D masks as pseudo labels, via our proposed iterative PCA and zero-shot semantic classes from vision-language models; 2) Secondly, we propose Geometry-aware Volume-based Detector (GVD) to end-to-end encode multi-view 2D images into a 3D volume to predict voxel-wise density and color via 2D-to-3D geometric projection, trained by 3D-to-2D rendering losses with SIS pseudo labels; 3) Thirdly, for better detection results, i.e., the 3D density projected on Birds-Eye-View, we propose Vertical-aware BEV Regularization (VBR) to constrain pedestrians to be vertical like the natural poses. Extensive experiments on popular multi-view pedestrian detection benchmarks Wildtrack, Terrace, and MultiviewX, show that our proposed UMPD, as the first fully-unsupervised method to our best knowledge, performs competitively to the previous state-of-the-art supervised methods. Code is available at https://github.com/lmy98129/UMPD.
Mengyin Liu, Chao Zhu 0003, Shiqi Ren, Xu-Cheng Yin
ACM Multimedia2
2024 Improving Small License Plate Detection with Bidirectional Vehicle-Plate Relation
Songkang Dai, Song-Lu Chen, Qi Liu 0041, Chao Zhu 0003, Feng Chen 0040, Xu-Cheng Yin
MMM (2)4
2023 VLPD: Context-Aware Pedestrian Detection via Vision-Language Semantic Self-Supervision
abstract
Detecting pedestrians accurately in urban scenes is significant for realistic applications like autonomous driving or video surveillance. However, confusing human-like objects often lead to wrong detections, and small scale or heavily occluded pedestrians are easily missed due to their unusual appearances. To address these challenges, only object regions are inadequate, thus how to fully utilize more explicit and semantic contexts becomes a key problem. Meanwhile, previous context-aware pedestrian detectors either only learn latent contexts with visual clues, or need laborious annotations to obtain explicit and semantic contexts. Therefore, we propose in this paper a novel approach via Vision-Language semantic self-supervision for context-aware Pedestrian Detection (VLPD) to model explicitly semantic contexts without any extra annotations. Firstly, we propose a self-supervised Vision-Language Semantic (VLS) segmentation method, which learns both fully-supervised pedestrian detection and contextual segmentation via self-generated explicit labels of semantic classes by vision-language models. Furthermore, a self-supervised Prototypical Semantic Contrastive (PSC) learning method is proposed to better discriminate pedestrians and other classes, based on more explicit and semantic contexts obtained from VLS. Extensive experiments on popular benchmarks show that our proposed VLPD achieves superior performances over the previous state-of-the-arts, particularly under challenging circumstances like small scale and heavy occlusion. Code is available at https://github.com/lmy98129/VLPD.
Mengyin Liu, Jie Jiang 0015, Chao Zhu 0003, Xu-Cheng Yin
CVPR3
2023 Towards Discriminative Semantic Relationship for Fine-grained Crowd Counting
abstract
As an extended task of crowd counting, fine-grained crowd counting aims to estimate the number of people in each semantic category instead of the whole in an image, and faces challenges including 1) inter-category crowd appearance similarity, 2) intra-category crowd appearance variations, and 3) frequent scene changes. In this paper, we propose a new fine-grained crowd counting approach named DSR to tackle these challenges by modeling Discriminative Semantic Relationship, which consists of two key components: Word Vector Module (WVM) and Adaptive Kernel Module (AKM). The WVM introduces more explicit semantic relationship information to better distinguish people of different semantic groups with similar appearance. The AKM dynamically adjusts kernel weights according to the features from different crowd appearance and scenes. The proposed DSR achieves superior results over state-of-the-art on the standard dataset. Our approach can serve as a new solid baseline and facilitate future research for the task of fine-grained crowd counting.
Shiqi Ren, Chao Zhu 0003, Mengyin Liu, Xu-Cheng Yin
ICME2
2023 Feature Implicit Enhancement via Super-Resolution for Small Object Detection
Zhehao Xu, Mengyin Liu, Chao Zhu 0003, Xu-Cheng Yin
PRCV (12)3
2022 Efficient Text-based Person Search via Single-stage Identity-guided Attribute Parsing and Alignment
abstract
Cross-modal text-based person search aims at retrieving target person in a large image gallery by natural language description. This task is quite challenging due to the complex environment of person image acquisition and the semantic gap between different modalities. The existing popular deep learning based models rely heavily on a large amount of labeled data to obtain good performance, which is labor-consuming and not always available in real applications. In order to achieve an effective alignment within and between modalities, additional semantic information or pre-trained network is often introduced to assist human body positioning, which will further shape the increase in training parameters and the decrease in efficiency. To address these problems, we propose a single-stage Identity-guided image-text Attribute Parsing and Alignment network (IAPA). IAPA realizes cross-modal alignment of human body parts in an unsupervised manner through image itself, resulting in great efficiency improvement while maintaining promising accuracy. This is also the first attempt to apply pixel-level supervision to cross-modal person retrieval task. The experiments on the CUHK-PEDES data-set validate the effectiveness and the efficiency of IAPA compared to the other state-of-the-art methods.
Chao Zhu 0003
ICPR2
2022 CAliC: Accurate and Efficient Image-Text Retrieval via Contrastive Alignment and Visual Contexts Modeling
abstract
Image-text retrieval is an essential task of information retrieval, in which the models with the Vision-and-Language Pretraining(VLP) are able to achieve ideal accuracy compared with the ones without VLP. Among different VLP approaches, the single-stream models achieve the overall best retrieval accuracy, but slower inference speed. Recently, researchers have introduced the two-stage retrieval setting commonly used in the information retrieval field to the single-stream VLP model for a better accuracy/efficiency trade-off. However, the retrieval accuracy and efficiency are still unsatisfactory mainly due to the limitations of the patch-based visual unimodal encoder in these VLP models. The unimodal encoders are trained on pure visual data, so the visual features extracted by them are difficult to align with the textual features and it is also difficult for the multi-modal encoder to understand visual information. Under these circumstances, we propose an accurate and efficient two-stage image-text retrieval model via Contrastive Alignment and visual Contexts modeling(CAliC). In the first stage of the proposed model, the visual unimodal encoder is pretrained with cross-modal contrastive learning to extract easily aligned visual features, which improves the retrieval accuracy and the inference speed. In the second stage of the proposed model, we introduce a new visual contexts modeling task during pretraining to help the multi-modal encoder better understand the visual information and get more accurate predictions. Extensive experimental evaluation validates the effectiveness of our proposed approach, which achieves a higher retrieval accuracy while keeping a faster inference speed, and outperforms existing state-of-the-art retrieval methods on image-text retrieval tasks over Flickr30K and COCO benchmarks.
Chao Zhu 0003, Mengyin Liu, Weibo Gu, Hongfa Wang, Wei Liu 0005, Xu-Cheng Yin
ACM Multimedia2
2022 Occluded Pedestrian Detection via Distribution-Based Mutual-Supervised Feature Learning
abstract
Pedestrian detection is a very important task in intelligent transportation system. State-of-the-art detectors work well on non-occluded pedestrians, but they are still far from satisfactory for heavily occluded ones. Recently, to deal with occlusion problems, the popular two-stage approaches are to build a two-branch architecture with the help of additional visible body annotations. However, these methods still have disadvantages. Either the two branches only use score-level fusion, which cannot guarantee the detectors to learn more robust pedestrian features. Or they only focus on the features of visible part via the attention mechanisms. However, the visible body features of heavily occluded pedestrians are only concentrated in a relatively small area, which may easily lead to missed detections. To alleviate the above issues, we propose a novel Distribution-based Mutual-Supervised Feature Learning Network (DMSFLN), to better deal with occluded pedestrian detection. The key DMSFL module in our network is to learn more discriminative feature representations of pedestrians by minimizing the similarity loss between feature distributions of full body and visible body, which has two advantages: enhancing the feature representations of occluded pedestrians and reducing the intra-class variance in pedestrians. To facilitate the DMSFL module, we also propose a novel two-branch network architecture, which is trained in a mutual-supervised way with both full body and visible body annotations respectively. Extensive experiments are conducted on four challenging pedestrian datasets: Caltech, CityPersons, CrowdHuman and CUHK occlusion. Our approach achieves superior performance compared to other state-of-the-art methods, especially on heavy occlusion subsets.
Ye He 0004, Chao Zhu 0003, Xu-Cheng Yin
IEEE Trans. Intell. Transp. Syst.2
2021 Adaptive Pattern-Parameter Matching for Robust Pedestrian Detection
abstract
Pedestrians with challenging patterns, e.g. small scale or heavy occlusion, appear frequently in practical applications like autonomous driving, which remains tremendous obstacle to higher robustness of detectors. Although plenty of previous works have been dedicated to these problems, properly matching patterns of pedestrian and parameters of detector, i.e., constructing a detector with proper parameter sizes for certain pedestrian patterns of different complexity, has been seldom investigated intensively. Pedestrian instances are usually handled equally with the same amount of parameters, which in our opinion is inadequate for those with more difficult patterns and leads to unsatisfactory performance. Thus, we propose in this paper a novel detection approach via adaptive pattern-parameter matching. The input pedestrian patterns, especially the complex ones, are first disentangled into simpler patterns for detection head by Pattern Disentangling Module (PDM) with various receptive fields. Then, Gating Feature Filtering Module (GFFM) dynamically decides the spatial positions where the patterns are still not simple enough and need further disentanglement by the next-level PDM. Cooperating with these two key components, our approach can adaptively select the best matched parameter size for the input patterns according to their complexity. Moreover, to further explore the relationship between parameter sizes and their performance on the corresponding patterns, two parameter selection policies are designed: 1) extending parameter size to maximum, aiming at more difficult patterns for different occlusion types; 2) specializing parameter size by group division, aiming at complex patterns for scale variations. Extensive experiments on two popular benchmarks, Caltech and CityPersons, show that our proposed method achieves superior performance compared with other state-of-the-art methods on subsets of different scales and occlusion types.
Mengyin Liu, Chao Zhu 0003, Xu-Cheng Yin
AAAI2
2020 A Hybrid Self-Attention Model for Pedestrians Detection
Chao Zhu 0003, Xu-Cheng Yin
ICONIP (1)2
2020 Mutual-Supervised Feature Modulation Network for Occluded Pedestrian Detection
abstract
State-of-the-art pedestrian detectors have achieved significant progress on non-occluded pedestrians, yet they are still struggling under heavy occlusions. The recent occlusion handling strategy of popular two-stage approaches is to build a two-branch architecture with the help of additional visible body annotations. Nonetheless, these methods still have some weaknesses. Either the two branches are trained independently with only score-level fusion, which cannot guarantee the detectors to learn robust enough pedestrian features. Or the attention mechanisms are exploited to only emphasize on the visible body features. However, the visible body features of heavily occluded pedestrians are concentrated on a relatively small area, which will easily cause missing detections. To address the above issues, we propose in this paper a novel Mutual-Supervised Feature Modulation (MSFM) network, to better handle occluded pedestrian detection. The key MSFM module in our network calculates the similarity loss of full body boxes and visible body boxes corresponding to the same pedestrian so that the full-body detector could learn more complete and robust pedestrian features with the assist of contextual features from the occluding parts. To facilitate the MSFM module, we also propose a novel two-branch architecture, consisting of a standard full body detection branch and an extra visible body classification branch. These two branches are trained in a mutual-supervised way with full body annotations and visible body annotations, respectively. To verify the effectiveness of our proposed method, extensive experiments are conducted on two challenging pedestrian datasets: Caltech and CityPersons, and our approach achieves superior performance compared to other state-of-the-art methods on both datasets, especially in heavy occlusion cases.
Ye He 0004, Chao Zhu 0003, Xu-Cheng Yin
ICPR2
2020 Semantic Bilinear Pooling for Fine-Grained Recognition
abstract
Naturally, fine-grained recognition, e.g., vehicle identification or bird classification, has specific hierarchical labels, where fine categories are always harder to be classified than coarse categories. However, most of the recent deep learning based methods neglect the semantic structure of fine-grained objects and do not take advantage of the traditional fine-grained recognition techniques (e.g. coarse-to-fine classification). In this paper, we propose a novel framework with a two-branch network (coarse branch and fine branch), i.e., semantic bilinear pooling, for fine-grained recognition with a hierarchical label tree. This framework can adaptively learn the semantic information from the hierarchical levels. Specifically, we design a generalized cross-entropy loss for the training of the proposed framework to fully exploit the semantic priors via considering the relevance between adjacent levels and enlarge the distance between samples of different coarse classes. Furthermore, our method leverages only the fine branch when testing so that it adds no overhead to the testing time. Experimental results show that our proposed method achieves state-of-the-art performance on four public datasets.
Xinjie Li 0002, Song-Lu Chen, Chao Zhu 0003, Xu-Cheng Yin
ICPR4
2017 Discriminative latent semantic feature learning for pedestrian detection
Chao Zhu 0003, Yuxin Peng 0001
Neurocomputing1
2017 Tracking Based Multi-Orientation Scene Text Detection: A Unified Framework With Dynamic Programming
abstract
There are a variety of grand challenges for multi-orientation text detection in scene videos, where the typical issues include skew distortion, low contrast, and arbitrary motion. Most conventional video text detection methods using individual frames have limited performance. In this paper, we propose a novel tracking based multi-orientation scene text detection method using multiple frames within a unified framework via dynamic programming. First, a multi-information fusion-based multi-orientation text detection method in each frame is proposed to extensively locate possible character candidates and extract text regions with multiple channels and scales. Second, an optimal tracking trajectory is learned and linked globally over consecutive frames by dynamic programming to finally refine the detection results with all detection, recognition, and prediction information. Moreover, the effectiveness of our proposed system is evaluated with the state-of-the-art performances on several public data sets of multi-orientation scene text images and videos, including MSRA-TD500, USTB-SV1K, and ICDAR 2015 Scene Videos.
Xu-Cheng Yin, Wei-Yi Pei, Shu Tian, Ze-Yu Zuo, Chao Zhu 0003, Junchi Yan
IEEE Trans. Image Process.6
2016 Group Cost-Sensitive Boosting for Multi-Resolution Pedestrian Detection
abstract
As an important yet challenging problem in computer vision, pedestrian detection has achieved impressive progress in recent years. However, the significant performance decline with decreasing resolution is a major bottleneck of current state-of-the-art methods. For the popular boosting-based detectors, one of the main reasons is that low resolution samples, which are usually more difficult to detect than high resolution ones, are treated by equal costs in the boosting process, leading to the consequence that they are more easily being rejected in early stages and can hardly be recovered in late stages as false negatives. To address this problem, we propose in this paper a new multi-resolution detection approach based on a novel group cost-sensitive boosting algorithm, which extends the popular AdaBoost by exploring different costs for different resolution groups in the boosting process, and places more emphases on low resolution group in order to better handle detection of hard samples. The proposed approach is evaluated on the challenging Caltech pedestrian benchmark, and outperforms other state-of-the-art on different resolution-specific test sets.
Chao Zhu 0003, Yuxin Peng 0001
AAAI1
2015 A Boosted Multi-Task Model for Pedestrian Detection with Occlusion Handling
abstract
Pedestrian detection is a challenging problem in computer vision. Especially, a major bottleneck for current state-of-the-art methods is the significant performance decline with increasing occlusion. A common technique for occlusion handling is to train a set of occlusion-specific detectors and merge their results directly. These detectors are trained independently and the relationship among them is ignored. In this paper, we consider pedestrian detection in different occlusion levels as different but related problems, and propose a multi-task model to jointly consider their relatedness and differences. The proposed model adopts multi-task learning algorithm to map pedestrians in different occlusion levels to a common space, where all models corresponding to different occlusion levels are constrained to share a common set of features, and a boosted detector is then constructed to distinguish pedestrians from background. The proposed approach is evaluated on the challenging Caltech pedestrian detection benchmark, and achieves state-of-the-art results on different occlusion-specific test sets.
Chao Zhu 0003, Yuxin Peng 0001
AAAI1
2015 A Boosted Multi-Task Model for Pedestrian Detection With Occlusion Handling
abstract
Pedestrian detection is a challenging problem in computer vision, and has achieved impressive progress in recent years. However, the current state-of-the-art methods suffer from significant performance decline with increasing occlusion level of pedestrians. A common approach for occlusion handling is to train a set of occlusion-specific detectors and merge their results directly, but these detectors are trained independently and the relationship among them is ignored. In this paper, we consider pedestrian detection in different occlusion levels as different but related problems, and propose a boosted multi-task model to jointly consider their relatedness and differences. The proposed model adopts multi-task learning algorithm to map pedestrians in different occlusion levels to a common space, where all models corresponding to different occlusion levels are constrained to share a common set of features, and a boosted detector is then constructed to distinguish pedestrians from background. The proposed approach is evaluated on three challenging pedestrian detection data sets, including Caltech, TUD-Brussels, and INRIA, and achieves superior performances against state of the art in the literature on different occlusion-specific test sets.
Chao Zhu 0003, Yuxin Peng 0001
IEEE Trans. Image Process.1
2014 HSOG: A Novel Local Image Descriptor Based on Histograms of the Second-Order Gradients
abstract
Recent investigations on human vision discover that the retinal image is a landscape or a geometric surface, consisting of features such as ridges and summits. However, most of existing popular local image descriptors in the literature, e.g., scale invariant feature transform (SIFT), histogram of oriented gradient (HOG), DAISY, local binary Patterns (LBP), and gradient location and orientation histogram, only employ the first-order gradient information related to the slope and the elasticity, i.e., length, area, and so on of a surface, and thereby partially characterize the geometric properties of a landscape. In this paper, we introduce a novel and powerful local image descriptor that extracts the histograms of second-order gradients (HSOGs) to capture the curvature related geometric properties of the neural landscape, i.e., cliffs, ridges, summits, valleys, basins, and so on. We conduct comprehensive experiments on three different applications, including the problem of local image matching, visual object categorization, and scene classification. The experimental results clearly evidence the discriminative power of HSOG as compared with its first-order gradient-based counterparts, e.g., SIFT, HOG, DAISY, and center-symmetric LBP, and the complementarity in terms of image representation, demonstrating the effectiveness of the proposed local descriptor.
Di Huang 0001, Chao Zhu 0003, Yunhong Wang 0001, Liming Chen 0002
IEEE Trans. Image Process.2
2013 Encoding Local Binary Descriptors by Bag-of-Features with Hamming Distance for Visual Object Categorization
Yu Zhang 0052, Chao Zhu 0003, Stéphane Bres, Liming Chen 0002
ECIR2
2013 Graph-based image segmentation using weighted color patch
abstract
Constructing a discriminative affinity graph plays an essential role in graph-based image segmentation, and feature directly influences the discriminative power of the affinity graph. In this paper, we propose a new method based on the weighted color patch to compute the weight of edges in an affinity graph. The proposed method intends to incorporate both color and neighborhood information by representing pixels with color patches. Furthermore, we assign both local and global weights adaptively for each pixel in a patch in order to alleviate the over-smooth effect of using patches. The normalized cut (NCut) algorithm is then applied on the resulting affinity graph to find partitions. We evaluate the proposed method on the Prague color texture image benchmark and the Berkeley image segmentation database. The extensive experiments show that our method is competitive compared to the other standard methods with multiple evaluation metrics.
Chao Zhu 0003, Charles-Edmond Bichot, Simon Masnou
ICIP2
2013 HSOG: a novel local descriptor based on histograms of second order gradients for object categorization
abstract
This paper presents a novel local image descriptor for object categorization that extracts the Histograms of the Second Order Gradients and is thereby named as HSOG. The HSOG descriptor is in contrast to the widely used ones in the literature, e.g. SIFT, DAISY, HOG, LBP, etc., which are based on the first order gradient information. The contributions of this work can be summarized as: (1) the design of HSOG; (2) the prove of its discriminative power and its complementation to the first order gradient based descriptors; (3) the analysis of performance variation caused by different parameter settings; and (4) the multi-scale extension which further improves the categorization accuracy. The experimental results achieved on the Caltech 101 and Caltech 256 databases clearly highlight the effectiveness of the proposed approach.
Di Huang 0001, Chao Zhu 0003, Charles-Edmond Bichot, Yunhong Wang 0001, Liming Chen 0002
ICMR2
2013 Multimodal recognition of visual concepts using histograms of textual concepts and selective weighted late fusion scheme
Ningning Liu, Emmanuel Dellandréa, Liming Chen 0002, Chao Zhu 0003, Yu Zhang 0052, Charles-Edmond Bichot, Stéphane Bres, Bruno Tellez
Comput. Vis. Image Underst.4
2013 Image region description using orthogonal combination of local binary patterns enhanced with color information
Chao Zhu 0003, Charles-Edmond Bichot, Liming Chen 0002
Pattern Recognit.1
2013 Pixel to Patch Sampling Structure and Local Neighboring Intensity Relationship Patterns for Texture Classification
abstract
In this letter, we explore local image descriptors for texture classification. We mainly propose two novel contributions: an effective sampling structure based on Pixel To Patch (PTP) to mimic the retinal sampling pattern; and a novel Local Neighboring Intensity Relationship Pattern (LNIRP) descriptor to extract texture feature by exploring neighboring gray-scale properties. The LNIRP descriptor is extended by using the PTP sampling structure which aims to capture not only micro-patterns but also macro-patterns, while reducing feature dimensionality and improving computational efficiency. The proposed descriptor has advantages of computational simplicity, no texton dictionary learning step and training-free. Moreover, the LNIRP descriptor is complementary to the Local Binary Pattern (LBP) descriptor. Extensive experiments were conducted on Outex database to evaluate the proposed descriptor and sampling structure. The proposed descriptor can achieve superior classification performance compared to most of the state-of-the-art methods, including what we believe to be the best results reported for Outex, while offering a smallest feature dimension.
Kai Wang 0021, Charles-Edmond Bichot, Chao Zhu 0003, Bailin Li
IEEE Signal Process. Lett.3
2011 Visual object recognition using DAISY descriptor
abstract
Visual content description is a key issue for the task of machine-based visual object categorization (VOC). A good visual descriptor should be both discriminative enough and computationally efficient while possessing some properties of robustness to viewpoint changes and lighting condition variations. The recent literature has featured local image descriptors, e.g. SIFT, as the main trend in VOC. However, it is well known that SIFT is computationally expensive, especially when the number of objects/concepts and learning data increase significantly. In this paper, we investigate the DAISY, which is a new fast local descriptor introduced for wide baseline matching problem, in the context of VOC. We carefully evaluate and compare the DAISY descriptor with SIFT both in terms of recognition accuracy and computation complexity on two standard image benchmarks - Caltech 101 and PASCAL VOC 2007. The experimental results show that DAISY outperforms the state-of-the-art SIFT while using shorter descriptor length and operating 3 times faster. When displaying a similar recognition accuracy to SIFT, DAISY can operate 12 times faster.
Chao Zhu 0003, Charles-Edmond Bichot, Liming Chen 0002
ICME1
2010 Multi-scale Color Local Binary Patterns for Visual Object Classes Recognition
abstract
The Local Binary Pattern (LBP) operator is a computationally efficient yet powerful feature for analyzing local texture structures. While the LBP operator has been successfully applied to tasks as diverse as texture classification, texture segmentation, face recognition and facial expression recognition, etc., it has been rarely used in the domain of Visual Object Classes (VOC) recognition mainly due to its deficiency of power for dealing with various changes in lighting and viewing conditions in real-world scenes. In this paper, we propose six novel multi-scale color LBP operators in order to increase photometric invariance property and discriminative power of the original LBP operator. The experimental results on the PASCAL VOC 2007 image benchmark show significant accuracy improvement by the proposed operators as compared with both the original LBP and other popular texture descriptors such as Gabor filter.
Chao Zhu 0003, Charles-Edmond Bichot, Liming Chen 0002
ICPR1
2009 Visual Object Categorization via Sparse Representation
abstract
In this paper, we consider the problem of classifying a real world image to the corresponding object class based on its visual content via sparse representation, which is originally used as a powerful tool for acquiring, representing and compressing high-dimensional signals. Assuming the intuitive hypothesis that an image could be represented by a linear combination of the training images from the same class, we propose a novel approach for visual object categorization in which a sparse representation of the image is first of all obtained by solving a L1 (or L0)-minimization problem and then fed into a traditional classifier such as Support Vector Machine (SVM) to finally perform the specified task. Experimental results obtained on the SIMPLIcity database have shown that this new approach can improve the classification performance compared to standard SVM using directly features extracted from the image.
Huanzhang Fu, Chao Zhu 0003, Emmanuel Dellandréa, Charles-Edmond Bichot, Liming Chen 0002
ICIG2