EDBT 2026 Demo / reviewers in the wild / expert
Zhenhua Chai
dblp:14/1604
· DBLP profile ↗
29ranked-venue papers
3as first author
18since 2021 · last 2025
0000-0002-1169-6133ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 1 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 2 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unpaired Image-Text Matching via Multimodal Aligned Conceptual KnowledgeabstractRecently, the accuracy of image-text matching has been greatly improved by multimodal pretrained models, all of which use millions or billions of paired images and texts for supervised model learning. Different from them, human brains can well match images with texts using their stored multimodal knowledge. Inspired by that, this paper studies a new scenario as unpaired image-text matching, in which paired images and texts are assumed to be unavailable during model learning. To deal with it, we accordingly propose a simple yet effective method namely Multimodal Aligned Conceptual Knowledge (MACK). First, we collect a set of words and their related image regions from publicly available datasets, and compute prototypical region representations to obtain pretrained general knowledge. To make the obtained knowledge better suit for certain datasets, we refine it using unpaired images and texts in a self-supervised learning manner to obtain fine-tuned domain knowledge. Then, to match given images with texts based on the knowledge, we represent parsed words in the texts by prototypical region representations, and compute region-word similarity scores. At last, the scores are aggregated based on bidirectional similarity pooling into an image-text similarity score, which can be directly used for unpaired image-text matching. The proposed MACK is complementary with existing models, which can be easily extended as a re-ranking method to substantially improve their performance of zero-shot and cross-dataset image-text matching. Yan Huang 0008, Yunan Zeng, Junshi Huang, Zhenhua Chai, Liang Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language NavigationabstractEmbodied agents equipped with GPT as their brains have exhibited extraordinary decisionmaking and generalization abilities across various tasks.However, existing zero-shot agents for vision-and-language navigation (VLN) only prompt GPT-4 to select potential locations within localized environments, without constructing an effective "global-view" for the agent to understand the overall environment.In this work, we present a novel map-guided GPT-based agent, dubbed MapGPT, which introduces an online linguistic-formed map to encourage global exploration.Specifically, we build an online map and incorporate it into the prompts that include node information and topological relationships, to help GPT understand the spatial environment.Benefiting from this design, we further propose an adaptive planning mechanism to assist the agent in performing multi-step path planning based on a map, systematically exploring multiple candidate nodes or sub-goals step by step.Extensive experiments demonstrate that our MapGPT is applicable to both GPT-4 and GPT-4V, achieving state-of-the-art zero-shot performance on R2R and REVERIE simultaneously (∼10% and ∼12% improvements in SR), and showcasing the newly emergent global thinking and path planning abilities of the GPT. Bingqian Lin, Zhenhua Chai, Xiaodan Liang, Kwan-Yee Kenneth Wong |
ACL (1) | 4 |
| 2024 | Investigating Compositional Challenges in Vision-Language Models for Visual GroundingabstractPre-trained vision-language models (VLMs) have achieved high performance on various downstream tasks, which have been widely used for visual grounding tasks in a weakly supervised manner. However, despite the per-formance gains contributed by large vision and language pre-training, we find that state-of-the-art VLMs struggle with compositional reasoning on grounding tasks. To demonstrate this, we propose Attribute, Relation, and Pri-ority grounding (ARPGrounding) benchmark to test VLMs' compositional reasoning ability on visual grounding tasks. ARPGrounding contains 11,425 samples and evaluates the compositional understanding of VLMs in three dimensions: 1) attribute, denoting comprehension of objects' properties; 2) relation, indicating an understanding of relation between objects; 3) priority, reflecting an awareness of the part of speech associated with nouns. Using the ARPGrounding benchmark, we evaluate several mainstream VLMs. We empirically find that these models perform quite well on conventional visual grounding datasets, achieving performance comparable to or surpassing state-of-the-art methods but showing strong deficiencies in compositional reasoning. Furthermore, we propose a composition-aware fine-tuning pipeline, demonstrating the potential to lever-age cost-effective image-text annotations for enhancing the compositional understanding of VLMs in grounding tasks. Code is available at link. Yunan Zeng, Yan Huang 0008, Zequn Jie, Zhenhua Chai, Liang Wang 0001 |
CVPR | 5 |
| 2024 | Exploration and Exploitation of Unlabeled Data for Open-Set Semi-supervised Learning
Ganlong Zhao, Guanbin Li, Yipeng Qin, Zhenhua Chai, Xiaolin Wei, Liang Lin 0004, Yizhou Yu |
Int. J. Comput. Vis. | 5 |
| 2024 | Contrastive Open-Set Active Learning-Based Sample Selection for Image ClassificationabstractIn this paper, we address a complex but practical scenario in Active Learning (AL) known as open-set AL, where the unlabeled data consists of both in-distribution (ID) and out-of-distribution (OOD) samples. Standard AL methods will fail in this scenario as OOD samples are highly likely to be regarded as uncertain samples, leading to their selection and wasting of the budget. Existing methods focus on selecting the highly likely ID samples, which tend to be easy and less informative. To this end, we introduce two criteria, namely contrastive confidence and historical divergence, which measure the possibility of being ID and the hardness of a sample, respectively. By balancing the two proposed criteria, highly informative ID samples can be selected as much as possible. Furthermore, unlike previous methods that require additional neural networks to detect the OOD samples, we propose a contrastive clustering framework that endows the classifier with the ability to identify the OOD samples and further enhances the network's representation learning. The experimental results demonstrate that the proposed method achieves state-of-the-art performance on several benchmark datasets. Zizheng Yan, Delian Ruan, Yushuang Wu, Junshi Huang, Zhenhua Chai, Xiaoguang Han 0001, Shuguang Cui, Guanbin Li |
IEEE Trans. Image Process. | 5 |
| 2023 | Divide and Adapt: Active Domain Adaptation via Customized LearningabstractActive domain adaptation (ADA) aims to improve the model adaptation performance by incorporating active learning (AL) techniques to label a maximally-informative subset of target samples. Conventional AL methods do not consider the existence of domain shift, and hence, fail to identify the truly valuable samples in the context of domain adaptation. To accommodate active learning and domain adaption, the two naturally different tasks, in a collaborative framework, we advocate that a customized learning strategy for the target data is the key to the success of ADA solutions. We present Divide-and-Adapt (DiaNA), a new ADA framework that partitions the target instances into four categories with stratified transferable properties. With a novel data subdivision protocol based on uncertainty and domainness, DiaNA can accurately recognize the most gainful samples. While sending the informative instances for annotation, DiaNA employs tailored learning strategies for the remaining categories. Furthermore, we propose an informativeness score that unifies the data partitioning criteria. This enables the use of a Gaussian mixture model (GMM) to automatically sample unlabeled data into the proposed four categories. Thanks to the “divide-and-adapt” spirit, DiaNA can handle data with large variations of domain gap. In addition, we show that DiaNA can generalize to different domain adaptation settings, such as unsupervised domain adaptation (UDA), semi-supervised domain adaptation (SSDA), source-free domain adaptation (SFDA), etc. Duojun Huang, Jichang Li, Weikai Chen 0001, Junshi Huang, Zhenhua Chai, Guanbin Li |
CVPR | 5 |
| 2023 | Progressive Temporal Transformer for Bird's-Eye-View Camera Pose Estimation
Zhuoyuan Wu, Jiancheng Cai, Ranran Huang 0001, Xinmin Liu, Zhenhua Chai |
ICONIP (6) | 5 |
| 2023 | DRKF: Distilled Rotated Kernel Fusion for Efficient Rotation Invariant Descriptors in Local Feature MatchingabstractThe performance of local feature descriptors degrades in the presence of large rotation variations. To address this issue, we present an efficient approach to learning rotation invariant descriptors. Specifically, we propose Rotated Kernel Fusion (RKF) which imposes rotations on the convolution kernel to improve the inherent nature of CNN. Since RKF can be processed by the subsequent re-parameterization, no extra computational costs will be introduced in the inference stage. Moreover, we present Multi-oriented Feature Aggregation (MOFA) which aggregates features extracted from multiple rotated versions of the input image and can provide auxiliary knowledge for the training of RKF by leveraging the distillation strategy. We refer to the distilled RKF model as DRKF. Besides the evaluation on a rotation-augmented version of the public dataset HPatches, we also contribute a new dataset named DiverseBEV which is collected during the drone's flight and consists of bird's eye view images with large viewpoint changes and camera rotations. Extensive experiments show that our method can outperform other state-of-the-art techniques when exposed to large rotation variations. Ranran Huang 0001, Jiancheng Cai, Zhuoyuan Wu, Xinmin Liu, Zhenhua Chai |
IROS | 5 |
| 2022 | Compressing Models with Few Samples: Mimicking then ReplacingabstractFew-sample compression aims to compress a big redundant model into a small compact one with only few samples. If we fine-tune models with these limited few samples directly, models will be vulnerable to overfit and learn almost nothing. Hence, previous methods optimize the compressed model layer-by-layer and try to make every layer have the same outputs as the corresponding layer in the teacher model, which is cumbersome. In this paper, we propose a new framework named Mimicking then Replacing (MiR) for few-sample compression, which firstly urges the pruned model to output the same features as the teacher's in the penultimate layer, and then replaces teacher's layers before penultimate with a well-tuned compact one. Unlike previous layer-wise reconstruction methods, our MiR optimizes the entire network holistically, which is not only simple and effective, but also unsupervised and general. MiR outperforms previous methods with large margins. Codes is available at https://github.com/cjnjuwhy/MiR. Junjie Liu 0003, Xin Ma 0031, Yang Yong, Zhenhua Chai, Jianxin Wu 0001 |
CVPR | 5 |
| 2022 | Towards Accurate Post-Training Quantization for Vision TransformerabstractVision transformer emerges as a potential architecture for vision tasks. However, the intense computation and non-negligible delay hinder its application in the real world. As a widespread model compression technique, existing post-training quantization methods still cause severe performance drops. We find the main reasons lie in (1) the existing calibration metric is inaccurate in measuring the quantization influence for extremely low-bit representation, and (2) the existing quantization paradigm is unfriendly to the power-law distribution of Softmax. Based on these observations, we propose a novel Accurate Post-training Quantization framework for Vision Transformer, namely APQ-ViT. We first present a unified Bottom-elimination Blockwise Calibration scheme to optimize the calibration metric to perceive the overall quantization disturbance in a blockwise manner and prioritize the crucial quantization errors that influence more on the final output. Then, we design a Matthew-effect Preserving Quantization for Softmax to maintain the power-law character and keep the function of the attention mechanism. Comprehensive experiments on large-scale classification and detection datasets demonstrate that our APQ-ViT surpasses the existing post-training quantization methods by convincing margins, especially in lower bit-width settings (e.g., averagely up to 5.17% improvement for classification and 24.43% for detection on W4A4). We also highlight that APQ-ViT enjoys versatility and works well on diverse transformer variants. Yifu Ding 0001, Haotong Qin, Qinghua Yan, Zhenhua Chai, Junjie Liu 0003, Xiaolin Wei, Xianglong Liu 0001 |
ACM Multimedia | 4 |
| 2022 | Contrastive attention network with dense field estimation for face completion
Xin Ma 0031, Xiaoqiang Zhou, Huaibo Huang, Gengyun Jia, Zhenhua Chai, Xiaolin Wei |
Pattern Recognit. | 5 |
| 2022 | Dimension-aware attention for efficient mobile networks
Rongyun Mo, Shenqi Lai, Yan Yan 0001, Zhenhua Chai, Xiaolin Wei |
Pattern Recognit. | 4 |
| 2022 | SPGNet: Serial and Parallel Group NetworkabstractNeural-network Processing Units (NPU), which specializes in the acceleration of deep neural networks (DNN), is of great significance to latency-sensitive areas like robotics or edge computing. However, there are few works focusing on the network design for NPU in recent studies. Most of the popular lightweight structures (e.g. MobileNet) are designed with depthwise convolution, which has less computation in theory but is not friendly to existing hardwares, and the speed tested on NPU is not always satisfactory. Even under similar FLOPs (the number of multiply-accumulates), vanilla convolution operation is always faster than depthwise one. In this paper, we will propose a novel architecture named Serial and Parallel Group Network (SPGNet), which can capture discriminative multi-scale information and at the same time keep the structure compact. Extensive evaluations have been conducted on different computer vision tasks, e.g. image classification (CIFAR and ImageNet), object detection (PASCAL VOC and MS COCO) and person re-identification (Market-1501 and DukeMTMC-ReID). The experimental results show that our proposed SPGNet can achieve comparable performance with the state-of-the-art networks while the speed is 120% faster than MobileNetV2 under similar FLOPS and over 300% faster than GhostNet with similar accuracy on NPU. Xuan Wang 0018, Shenqi Lai, Zhenhua Chai, Xingjun Zhang, Xueming Qian |
IEEE Trans. Multim. | 3 |
| 2021 | Rethinking BiSeNet for Real-Time Semantic SegmentationabstractBiSeNet [28], [27] has been proved to be a popular two-stream network for real-time segmentation. However, its principle of adding an extra path to encode spatial information is time-consuming, and the backbones borrowed from pretrained tasks, e.g., image classification, may be inefficient for image segmentation due to the deficiency of task-specific design. To handle these problems, we propose a novel and efficient structure named Short-Term Dense Concatenate network (STDC network) by removing structure redundancy. Specifically, we gradually reduce the dimension of feature maps and use the aggregation of them for image representation, which forms the basic module of STDC network. In the decoder, we propose a Detail Aggregation module by integrating the learning of spatial information into low-level layers in single-stream manner. Finally, the low-level features and deep features are fused to predict the final segmentation results. Extensive experiments on Cityscapes and CamVid dataset demonstrate the effectiveness of our method by achieving promising trade-off between segmentation accuracy and inference speed. On Cityscapes, we achieve 71.9% mIoU on the test set with a speed of 250.4 FPS on NVIDIA GTX 1080Ti, which is 45.2% faster than the latest methods, and achieve 76.8% mIoU with 97.0 FPS while inferring on higher resolution images. Code is available at https://github.com/MichaelFan01/STDC-Seg. Mingyuan Fan 0002, Shenqi Lai, Junshi Huang, Xiaoming Wei, Zhenhua Chai, Junfeng Luo, Xiaolin Wei |
CVPR | 5 |
| 2021 | Feature Decomposition and Reconstruction Learning for Effective Facial Expression RecognitionabstractIn this paper, we propose a novel Feature Decomposition and Reconstruction Learning (FDRL) method for effective facial expression recognition. We view the expression information as the combination of the shared information (expression similarities) across different expressions and the unique information (expression-specific variations) for each expression. More specifically, FDRL mainly consists of two crucial networks: a Feature Decomposition Network (FDN) and a Feature Reconstruction Network (FRN). In particular, FDN first decomposes the basic features extracted from a backbone network into a set of facial action-aware latent features to model expression similarities. Then, FRN captures the intra-feature and inter-feature relationships for la-tent features to characterize expression-specific variations, and reconstructs the expression feature. To this end, two modules including an intra-feature relation modeling module and an inter-feature relation modeling module are developed in FRN. Experimental results on both the in-the-lab databases (including CK+, MMI, and Oulu-CASIA) and the in-the-wild databases (including RAF-DB and SFEW) show that the proposed FDRL method consistently achieves higher recognition accuracy than several state-of-the-art methods. This clearly highlights the benefit of feature decomposition and reconstruction for classifying expressions. Delian Ruan, Yan Yan 0001, Shenqi Lai, Zhenhua Chai, Chunhua Shen, Hanzi Wang |
CVPR | 4 |
| 2021 | Trash to Treasure: Harvesting OOD Data with Cross-Modal Matching for Open-Set Semi-Supervised LearningabstractOpen-set semi-supervised learning (open-set SSL) investigates a challenging but practical scenario where out-of-distribution (OOD) samples are contained in the unlabeled data. While the mainstream technique seeks to completely filter out the OOD samples for semi-supervised learning (SSL), we propose a novel training mechanism that could effectively exploit the presence of OOD data for enhanced feature learning while avoiding its adverse impact on the SSL. We achieve this goal by first introducing a warm-up training that leverages all the unlabeled data, including both the in-distribution (ID) and OOD samples. Specifically, we perform a pretext task that enforces our feature extractor to obtain a high-level semantic understanding of the training images, leading to more discriminative features that can benefit the downstream tasks. Since the OOD samples are inevitably detrimental to SSL, we propose a novel cross-modal matching strategy to detect OOD samples. Instead of directly applying binary classification [39], we train the network to predict whether the data sample is matched to an assigned one-hot class label. The appeal of the proposed cross-modal matching over binary classification is the ability to generate a compatible feature space that aligns with the core classification task. Extensive experiments show that our approach substantially lifts the performance on open-set SSL and outperforms the state-of-the-art by a large margin. Chaowei Fang, Weikai Chen 0001, Zhenhua Chai, Xiaolin Wei, Pengxu Wei, Liang Lin 0004, Guanbin Li |
ICCV | 4 |
| 2021 | Selective Wavelet Attention Learning for Single Image Deraining
Huaibo Huang, Aijing Yu, Zhenhua Chai, Ran He 0001, Tieniu Tan |
Int. J. Comput. Vis. | 3 |
| 2021 | Coupled adversarial learning for semi-supervised heterogeneous face recognition
Ran He 0001, Yi Li 0018, Xiang Wu 0001, Lingxiao Song, Zhenhua Chai, Xiaolin Wei |
Pattern Recognit. | 5 |
| 2020 | Free-Form Image Inpainting via Contrastive Attention NetworkabstractMost deep learning based image inpainting approaches adopt autoencoder or its variants to fill missing regions in images. Encoders are usually utilized to learn powerful representational spaces, which are important for dealing with sophisticated learning tasks. Specifically, in image inpainting tasks, masks with any shapes can appear anywhere in images (i.e., free-form masks) which form complex patterns. It is difficult for encoders to capture such powerful representations under this complex situation. To tackle this problem, we propose a self-supervised Siamese inference network to improve the robustness and generalization. It can encode contextual semantics from full resolution images and obtain more discriminative representations. we further propose a multi-scale decoder with a novel dual attention fusion module (DAF), which can combine both the restored and known regions in a smooth way. This multi-scale architecture is benefit for decoding discriminative representations learned by encoders into images layer by layer. In this way, unknown regions will be filled naturally from outside to inside. Qualitative and quantitative experiments on multiple datasets, including facial and natural datasets (i.e., Celeb-HQ, Pairs Street View, Places2 and ImageNet), demonstrate that our proposed method outperforms state-of-the-art methods in generating high-quality inpainting results. Xin Ma 0031, Xiaoqiang Zhou, Huaibo Huang, Zhenhua Chai, Xiaolin Wei, Ran He 0001 |
ICPR | 4 |
| 2020 | Reference Guided Face Component EditingabstractFace portrait editing has achieved great progress in recent years. However, previous methods either 1) operate on pre-defined face attributes, lacking the flexibility of controlling shapes of high-level semantic facial components (e.g., eyes, nose, mouth), or 2) take manually edited mask or sketch as an intermediate representation for observable changes, but such additional input usually requires extra efforts to obtain. To break the limitations (e.g. shape, mask or sketch) of the existing methods, we propose a novel framework termed r FACE (Reference Guided FAce Component Editing) for diverse and controllable face component editing with geometric changes. Specifically, r-FACE takes an image inpainting model as the backbone, utilizing reference images as conditions for controlling the shape of face components. In order to encourage the framework to concentrate on the target face components, an example-guided attention module is designed to fuse attention features and the target face component features extracted from the reference image. Through extensive experimental validation and comparisons, we verify the effectiveness of the proposed framework. Qiyao Deng, Jie Cao 0002, Yunfan Liu 0001, Zhenhua Chai, Qi Li 0005, Zhenan Sun |
IJCAI | 4 |
| 2020 | Query Twice: Dual Mixture Attention Meta Learning for Video SummarizationabstractVideo summarization aims to select representative frames to retain high-level information, which is usually solved by predicting the segment-wise importance score via a softmax function. However, softmax function suffers in retaining high-rank representations for complex visual or sequential information, which is known as the Softmax Bottleneck problem. In this paper, we propose a novel framework named Dual Mixture Attention (DMASum) model with Meta Learning for video summarization that tackles the softmax bottleneck problem, where the Mixture of Attention layer (MoA) effectively increases the model capacity by employing twice self-query attention that can capture the second-order changes in addition to the initial query-key attention, and a novel Single Frame Meta Learning rule is then introduced to achieve more generalization to small datasets with limited training sources. Furthermore, the DMASum significantly exploits both visual and sequential attention that connects local key-frame and global attention in an accumulative way. We adopt the new evaluation protocol on two public datasets, SumMe, and TVSum. Both qualitative and quantitative experiments manifest significant improvements over the state-of-the-art methods. Junyan Wang 0001, Yang Bai 0011, Yang Long 0001, Bingzhang Hu, Zhenhua Chai, Yu Guan 0001, Xiaolin Wei |
ACM Multimedia | 5 |
| 2019 | Enhanced Normalized Mean Error loss for Robust Facial Landmark detection
Shenqi Lai, Zhenhua Chai, Shengxi Li, Huanhuan Meng, Mengzhao Yang, Xiaoming Wei |
BMVC | 2 |
| 2019 | Facial Expression Recognition: Disentangling Expression Based on Self-attention Conditional Generative Adversarial Nets
Haohao Li, Xiaoming Wei, Zhenhua Chai |
PRCV (2) | 4 |
| 2015 | Online Multi-Target Tracking With Unified Handling of Complex ScenariosabstractComplex scenarios, including miss detections, occlusions, false detections, and trajectory terminations, make the data association challenging. In this paper, we propose an online tracking-by-detection method to track multiple targets with unified handling of aforementioned complex scenarios, where current detection responses are linked to the previous trajectories. We introduce a dummy node to each trajectory to allow it to temporally disappear. If a trajectory fails to find its matching detection, it will be linked to its corresponding dummy node until the emergence of its matching detection. Source nodes are also incorporated to account for the entrance of new targets. The standard Hungarian algorithm, extended by the dummy nodes, can be exploited to solve the online data association implicitly in a global manner, although it is formulated between two consecutive frames. Moreover, as dummy nodes tend to accumulate in a fake or disappeared trajectory while they only occasionally appear in a real trajectory, we can deal with false detections and trajectory terminations by simply checking the number of consecutive dummy nodes. Our approach works on a single, uncalibrated camera, and requires neither scene prior knowledge nor explicit occlusion reasoning, running at 132 frames/s on the PETS09-S2L1 benchmark sequence. The experimental results validate the effectiveness of the dummy nodes in complex scenarios and show that our proposed approach is robust against false detections and miss detections. Quantitative comparisons with other methods on five benchmark sequences demonstrate that we can achieve comparable results with the most existing offline methods and better results than other online algorithms. Huaizu Jiang, Jinjun Wang, Yihong Gong, Na Rong, Zhenhua Chai, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 5 |
| 2014 | Learning Flexible Block based Local Binary Patterns for unconstrained face detectionabstractFace detection has been a very active research topic in recent years. However, when applied to uncontrolled environments, some systems exhibit poor generalization ability. Even though few of existing methods can achieve promising results in some challenging situations, they usually have the requirement of high computational cost. This will definitely limit the use of those methods in some mobile platforms which have limited computational resources and strict power-consumption control. In this paper, a novel facial representation method for multi-view face detection in uncontrolled environment is presented. The proposed method, named Flexible Block based Local Binary Patterns (FBLBP), has low storage requirements and it is fast to compute; while its performance is comparable with the state of the art methods, demonstrated on the challenging Face Detection Data set and Benchmark (FDDB). Zhenhua Chai, Zhijun Du, Heydi Mendez Vazquez |
ICME | 1 |
| 2014 | Real-time summarization of user-generated videos based on semantic recognitionabstractUser-generated contents play an important role in the Internet video-sharing activities. Techniques for summarizing the user-generated videos (UGVs) into short representative clips are useful in many applications. This paper introduces an approach for UGV summarization based on semantic recognition. Different from other types of videos like movies or broadcasting news, where the semantic contents may vary greatly across different shots, most UGVs have only a single long shot with relatively consistent high-level semantics. Therefore, a few semantically representative segments are generally sufficient for a UGV summary, which can be selected based on the distribution of semantic recognition scores. In addition, due to the poor shooting quality of many UGVs, factors such as camera shaking and lighting condition are also considered to achieve more pleasant summaries. Experiments on over 100 UGVs with both subjective and objective evaluations show that our approach clearly outperforms several alternative methods and is highly efficient. Using a regular laptop, it can produce a summary for a 2-minute video in just 10 seconds. Xi Wang 0008, Yu-Gang Jiang 0001, Zhenhua Chai, Zichen Gu |
ACM Multimedia | 3 |
| 2014 | Gabor Ordinal Measures for Face RecognitionabstractGreat progress has been achieved in face recognition in the last three decades. However, it is still challenging to characterize the identity related features in face images. This paper proposes a novel facial feature extraction method named Gabor ordinal measures (GOM), which integrates the distinctiveness of Gabor features and the robustness of ordinal measures as a promising solution to jointly handle inter-person similarity and intra-person variations in face images. In the proposal, different kinds of ordinal measures are derived from magnitude, phase, real, and imaginary components of Gabor images, respectively, and then are jointly encoded as visual primitives in local regions. The statistical distributions of these visual primitives in face image blocks are concatenated into a feature vector and linear discriminant analysis is further used to obtain a compact and discriminative feature representation. Finally, a two-stage cascade learning method and a greedy block selection method are used to train a strong classifier for face recognition. Extensive experiments on publicly available face image databases, such as FERET, AR, and large scale FRGC v2.0, demonstrate state-of-the-art face recognition performance of GOM. Zhenhua Chai, Zhenan Sun, Heydi Mendez Vazquez, Ran He 0001, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2013 | Face recognition using Histogram of co-occurrence Gabor phase patternsabstractThe fusion of Local Binary Patterns (LBP) and Gabor magnitude features has been demonstrated to be one of the most successful descriptors for face recognition. Recently, several Gabor phase based features like Histogram of Gabor Phase Patterns (HGPP) and Local Gabor XOR Patterns (LGXP) also show competitive results and complementary attributes to Gabor magnitude based features. However, in these two typical Gabor phase based approaches only the binary relationship between neighboring Gabor phases is used, which may lose some discriminative information. To investigate the potential of Gabor phase features for robust face recognition, this paper proposes a novel local descriptor, named Histogram of Co-occurrence Gabor Phase Patterns (HCGPP). In HCGPP, Gabor Phase features are first extracted and quantized into different ranges. Second we estimate the histograms of cooccurrence Gabor phase patterns in each face region. Finally, a nearest-neighbor classifier with the dissimilarity measure χ2is used for classification. Extensive experimental results on FERET and AR databases show the significant advantages of the proposed method over the state-of-the art ones in terms of recognition rate. Zhenhua Chai, Zhenan Sun |
ICIP | 2 |
| 2012 | Semantic Pixel Sets Based Local Binary Patterns for Face Recognition
Zhenhua Chai, Heydi Mendez Vazquez, Ran He 0001, Zhenan Sun, Tieniu Tan |
ACCV (2) | 1 |