EDBT 2026 Demo / reviewers in the wild / expert
Chengju Zhou
dblp:150/6745
· DBLP profile ↗
27ranked-venue papers
12as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Watch Where You Move: Region-Aware Dynamic Aggregation and Excitation for Gait RecognitionabstractDeep learning-based gait recognition has achieved great success in various applications. The key to accurate gait recognition lies in considering the unique and diverse behavior patterns indifferent motion regions, especially when covariates affect visual appearance. However, existing methods typically use predefined regions for temporal modeling, with fixed or equivalent temporal scales assigned to different types of regions, which makes it difficult to model motion regions that change dynamically over time and adapt to their specific patterns. To tackle this problem, we introduce a Region-aware Dynamic Aggregation and Excitation framework (GaitRDAE) that automatically searches for motion regions, assigns adaptive temporal scales and applies corresponding attention. Specifically, the framework includes two core modules: the Region-aware Dynamic Aggregation (RDA) module, which dynamically searches the optimal temporal receptive field for each region, and the Region-aware Dynamic Excitation (RDE) module, which emphasizes the learning of motion regions containing more stable behavior patterns while suppressing attention to static regions that are more susceptible to covariates. Experimental results show that GaitRDAE achieves state-of-the-art performance on several benchmark datasets. The source code will be published athttps://github.com/HUAFOR/GaitRDAE. Binyuan Huang, Yongdong Luo, Xianda Guo, Xiawu Zheng, Jiahui Pan 0003, Chengju Zhou |
IEEE Trans. Multim. | 7 |
| 2025 | Prior-Prompt-Based GCN for Depression Recognition Through Gait Observation
Chengju Zhou, Yutao Xu |
CogSci | 1 |
| 2025 | Non-invasive Emotion Perception from Gait by Sparse and Spatial-Temporal Excitation Based Graph Convolutional Network
Liangyu Lu, Chengju Zhou |
ICIC (21) | 2 |
| 2025 | Multi-soft-label Guided Supervised Contrastive Learning for Gait Emotion RecognitionabstractGait-based emotion recognition has received considerable attention due to its non-invasive capturing manner. However, most existing works learn the gait representations by treating different classes independently, which ignores the inherent class ambiguity in this field. Therefore, we propose a Multi-soft-label Guided Supervised Contrastive Learning (MSL-SCL) framework, which leverages the class correlation information in soft labels to explicitly guide the SCL, thereby alleviating the ambiguous gait representation. Specifically, a Soft-Label SCL (Sof-SCL) module is designed to select positive and negative samples based on their soft-label similarity to the anchors, and the similarity is further incorporated into a novel contrastive loss function for the refinement. Moreover, a Prior-Guided SCL (Prior-SCL) is introduced to capture the subtle changes in gait and employed as soft labels to provide an adaptive supervision for SCL. Extensive experimental results on the Emotion-Gait dataset demonstrate that our method outperforms SOTAs with a mean average precision of 89.8%. Chengju Zhou, Mengxin Xu, Xiaotong Fan, Liangyu Lu, Jiahui Pan 0003, Lewei He |
ICME | 1 |
| 2025 | PerioDet: Large-Scale Panoramic Radiograph Benchmark for Clinical-Oriented Apical Periodontitis Detection
Xiaocheng Fang, Jieyi Cai, Chengju Zhou, Minhua Lu, Bingzhi Chen |
MICCAI (16) | 4 |
| 2025 | Med-BiasX: Robust Medical Visual Question Answering with Language Biases
Huanjia Zhu, Yishu Liu 0001, Chengju Zhou, Guangming Lu 0002, Bingzhi Chen |
MICCAI (14) | 3 |
| 2025 | Attention-aware spatio-temporal learning for multi-view gait-based age estimation and gender classificationabstractAbstract Recently, gait‐based age and gender recognition have attracted considerable attention in the fields of advertisement marketing and surveillance retrieval due to the unique advantage that gaits can be perceived at a long distance. Intuitively, age and gender can be recognised by observing people's static shape (e.g. different hairstyles between males and females) and dynamic motion (e.g. different walking velocities between the elderly and youth). However, most of the existing gait‐based age and gender recognition methods are based on Gait Energy Image (GEI), which loses the capability of explicitly modelling temporal dynamic information and is not robust to the multi‐view recognition that inevitably happens in a real application. Therefore, in this study, an Attention‐aware Spatio‐Temporal Learning (ASTL) framework is proposed, which employs a silhouette sequence as input to learn essential and invariable spatial‐temporal gait representations. More specifically, a Multi‐Scale Temporal Aggregation (MSTA) module provides an effective scheme for dynamic gait description by exploring and aggregating multi‐scale temporal interval information, which is a core supplement to spatial representation. Then, a Multiple Attention Aggregation (MAA) module is designed to help the network focus on the most discriminatory information along temporal, spatial and channel dimensions. Finally, a Multimodal Collaborative Learning (MCL) block gives full play to the advantages of different modal features through a multimodal cooperative learning strategy. The mean absolute error (MAE) for the age estimation and the correct classification rate (CCR) for the gender classification on OU‐MVLP achieve 6.68 years and 97%, respectively, demonstrating the superiority of the method. Ablation experiments and visualisation results also prove the effectiveness of the three individual modules in their framework. Binyuan Huang, Yongdong Luo, Jiahui Xie, Jiahui Pan 0003, Chengju Zhou |
IET Comput. Vis. | 5 |
| 2025 | MTADA: A Multi-Task Adversarial Domain Adaptation Network for EEG-Based Cross-Subject Emotion RecognitionabstractIn electroencephalogram (EEG)-based emotion recognition, the applicability of most current models is limited by inter-subject variability and emotion complexity. This study proposes a multi-task adversarial domain adaptation (MTADA) network to enhance cross-subject emotion recognition performance. The model first employs a domain matching strategy to select the source domain that best matches the target domain. Then, adversarial domain adaptation is used to learn the difference between source and target domains, and a fine-grained joint domain discriminator is constructed to align them by incorporating category information. At the same time, a multi-task learning mechanism is utilized to learn the intrinsic relationships between different emotions and predict multiple emotions simultaneously. We conducted comprehensive experiments on two public datasets, DEAP and FACED. On DEAP, the average accuracies for valence, arousal and dominance are 76.39%, 69.74% and 68.26%, respectively. On FACED, the average accuracies for valence and arousal are 78.90% and 77.95%. When using the subject from DEAP as the source domain to predict the subjects in FACED, the accuracies for valence and arousal are 61.07% and 60.82%. These results show that our MTADA model improves cross-subject emotion recognition and outperforms most state-of-the-art methods, which may provide new approach for EEG-based emotion brain-computer interface systems. Lina Qiu, Zuorui Ying, Xianyue Song, Weisen Feng, Chengju Zhou, Jiahui Pan 0003 |
IEEE Trans. Affect. Comput. | 5 |
| 2024 | GaitCTCG: cross-view gait recognition via cascaded residual temporal shift and comprehensive multi-granularity learning
Binyuan Huang, Chengju Zhou, Lewei He, Jiahui Pan 0003 |
Appl. Intell. | 2 |
| 2024 | An attention-based adaptive spatial-temporal graph convolutional network for long-video ergonomic risk assessmentabstractErgonomic risk assessment (ERA) is commonly used to identify and analyze postures that are detrimental to the health of workers in industrial workplaces, which is vital to prevent work-related musculoskeletal disorders (WMSDs). Among the automatic approaches, algorithms based on graph convolutional networks (GCNs) have shown promising results in ERA using skeleton sequence as input. However, previous GCN-based methods still have certain limitations. First, the separated modeling of spatial and temporal information and the manually pre-defined topology of graph may restrict the representation diversity of the networks. Additionally, RNN-based temporal modeling often incurs high computational costs and fails to capture long-range temporal dependencies, thereby reducing flexibility in describing long videos. To overcome these challenges, in this study, we propose an attention-based adaptive spatial–temporal graph convolutional network (AAST-GCN), aiming to achieve effective and efficient action representation for ERA in long video. First, we employ an alternate modeling strategy to effectively capture the spatial–temporal information, and propose an improved adaptive adjacency matrix scheme to learn various coordination and relations of body-joints, thus enhancing the flexibility to model diverse postures. Furthermore, we introduce an efficient multi-scale temporal convolutional network as a replacement for RNN-based algorithms, enabling the network to extract various granularities of temporal features. Moreover, to make the network focuses on more valuable information, we employ a spatial–temporal interaction attention (STIA) module. Finally, the aforementioned modules are aggregated within a multi-task learning framework, with the action segmentation serving as the auxiliary task to further improve the accuracy of ERA. We conducted the ergonomic risk assessment on the UW-IOM and TUM Kitchen datasets using our network. Extensive experiments conducted on the most popular datasets UW-IOM and TUM Kitchen demonstrated that our proposed AAST-GCN outperforms other GCN-based methods. Ablation studies and visualization also prove the effectiveness of the individual sub-modules. Chengju Zhou, Jiayu Zeng, Lina Qiu, Shuxi Wang, Pingzhi Liu, Jiahui Pan 0003 |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Portable vision-based gait assessment for post-stroke rehabilitation using an attention-based lightweight CNN
Chengju Zhou, Daqin Feng, Shuyu Chen 0002, Nianming Ban, Jiahui Pan 0003 |
Expert Syst. Appl. | 1 |
| 2024 | SFT-SGAT: A semi-supervised fine-tuning self-supervised graph attention network for emotion recognition and consciousness detection
Lina Qiu, Liangquan Zhong, Weisen Feng, Chengju Zhou, Jiahui Pan 0003 |
Neural Networks | 5 |
| 2023 | Automatic Hemiplegia Gait Assessment for Post-Stroke by an Efficient Hybrid Attention-Based GhostNetabstractVision-based gait analysis provides the possibility to automatically and unobtrusively detect walking pattern alterations caused by stoke. Therefore, it can be used to determine the severity of stroke during stroke rehabilitation outside the hospital, which greatly releases the economic and labor burden on patients and their families. However, state-of-the-art deep learning algorithms for gait analysis usually suffer from high computational complexity and can even lead to overfitting problems on small-scale pathological gait datasets. To realize an efficient and effective system, we constructed a specially designed dataset and proposed a novel lightweight network to lean discriminative gait representation to map the input into one of the stroke severity levels. More specifically, a simulated hemiplegia gait dataset with multiple severity levels is first constructed, including sufficient 2D image sequences collected from 14 subjects. Different from the existing pathological datasets used for coarse classification, which only distinguish different pathological gait types, our proposed dataset is specifically designed for fine classification to assess the severity of hemiplegia that is defined according to medical prior. Second, considering that pathological datasets are usually small-scale, an attention-based lightweight network is proposed. In detail, a lightweight hybrid attention module (LHAM) based on the 1D adaptive convolution for channel attention interaction was developed to enhance the network's ability to integrate and focus on meaningful spatial and channel features. To further lighten the networks, a proposed efficient ghost module (EGM) is used in the bottleneck structure instead of the normal convolutional layer. Extensive experiments on both self-constructed and publicly available datasets demonstrate that the proposed efficient hybrid attention-based GhostNet realizes an effective and efficient gait analysis for stroke rehabilitation. Chengju Zhou, Daqin Feng, Lewei He, Nianming Ban, Shuxi Wang, Jiahui Pan 0003 |
IJCNN | 1 |
| 2023 | ICE-GCN: An interactional channel excitation-enhanced graph convolutional network for skeleton-based action recognitionabstractAbstract Thanks to the development of depth sensors and pose estimation algorithms, skeleton-based action recognition has become prevalent in the computer vision community. Most of the existing works are based on spatio-temporal graph convolutional network frameworks, which learn and treat all spatial or temporal features equally, ignoring the interaction with channel dimension to explore different contributions of different spatio-temporal patterns along the channel direction and thus losing the ability to distinguish confusing actions with subtle differences. In this paper, an interactional channel excitation (ICE) module is proposed to explore discriminative spatio-temporal features of actions by adaptively recalibrating channel-wise pattern maps. More specifically, a channel-wise spatial excitation (CSE) is incorporated to capture the crucial body global structure patterns to excite the spatial-sensitive channels. A channel-wise temporal excitation (CTE) is designed to learn temporal inter-frame dynamics information to excite the temporal-sensitive channels. ICE enhances different backbones as a plug-and-play module. Furthermore, we systematically investigate the strategies of graph topology and argue that complementary information is necessary for sophisticated action description. Finally, together equipped with ICE, an interactional channel excited graph convolutional network with complementary topology (ICE-GCN) is proposed and evaluated on three large-scale datasets, NTU RGB+D 60, NTU RGB+D 120, and Kinetics-Skeleton. Extensive experimental results and ablation studies demonstrate that our method outperforms other SOTAs and proves the effectiveness of individual sub-modules. The code will be published at https://github.com/shuxiwang/ICE-GCN . Shuxi Wang, Jiahui Pan 0003, Binyuan Huang, Pingzhi Liu, Zina Li, Chengju Zhou |
Mach. Vis. Appl. | 6 |
| 2022 | GaitMSTP: Multi-Granularity Spatio-Temporal Pyramid for Gait Recognition Under Complex Covariation ConditionsabstractGait, with its unique advantage of remote perception without any cooperation from the perceived subject, has become a popular biometric modality for human identity authentication. Diverse spatial representation and temporal modeling are crucial information for gait recognition, especially under covariation conditions. However, most existing algorithms do not fully and explicitly exploit the rich spatial-temporal clues in the gait sequences, leading to a decline in the discriminative ability of gait feature representations. In this paper, we propose a GaitMSTP network for gait recognition under complex covariation conditions, which explicitly models spatio-temporal representations at multi-granularity and multi-sematic levels. More specifically, a Multi-Granularity Temporal Pyramid (MGTP) module is incorporated to extract features at different temporal granularity, which simulates coarse- and fine-grained motion patterns at diverse temporal scales. A Multi-Granularity Spatial Pyramid (MGSP) is designed to capture global and local features at multiple spatial locations and scales. In addition, the multi-granularity spatial-temporal features extracted from shallow to deep semantic levels are further used for supervised learning, aiming to exploit both high-level and low-level of multi-sematic gait characteristics. Extensive experiments on the CASIA-B dataset show that our method outperforms the state-of-the-art algorithms for gait recognition in all scenarios. Binyuan Huang, Chengju Zhou, Jiahui Pan 0003 |
IJCB | 2 |
| 2022 | A Unified Multi-Task Learning Architecture for Fast and Accurate Pedestrian DetectionabstractWe present a unified multi-task learning architecture for fast and accurate pedestrian detection. Different from existing methods which often focus on either a new loss function or architecture, we propose an improved multi-task convolutional neural network learning architecture to effectively and efficiently interfuse the task of pedestrian detection and semantic segmentation. To achieve this, we integrate a lightweight semantic segmentation branch to Faster R-CNN detection framework that enables end-to-end hard parameter sharing in order to boost the detection performance and maintain computational efficiency as follows. Firstly, a Semantic Segmentation to Feature Module (SS2FM) refines the convolutional features in RPN stage by integrating the features generated from the semantic segmentation branch. Secondly, a Semantic Segmentation to Confidence Module (SS2CM) refines the classification confidence in RPN stage by fusing it with the semantic segmentation confidence. We also introduce an effective anchor matching point transform to alleviate the problem of feature misalignment for heavily occluded pedestrians. The proposed unified multi-task learning architecture lends itself well to more robust pedestrian detection in diverse scenarios with negligible computation overhead. In addition, the proposed architecture can achieve high detection performance with low resolution input images, which significantly reduces the computational complexity. Experiment results on CityPersons and Caltech datasets show that our method is the fastest among all state-of-the-art pedestrian detection methods while exhibiting competitive detection performance. Chengju Zhou, Meiqing Wu, Siew-Kei Lam |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Enhanced Multi-Task Learning Architecture for Detecting Pedestrian at Far DistanceabstractExisting pedestrian detection methods suffer from performance degradation in the presence of small-scale pedestrians who are positioned at far distance from the camera. We present a pedestrian detection framework that is not only robust to small- and large-scale pedestrians, but is also significantly faster than state-of-the-art methods. The proposed framework incorporates semantic segmentation to confidence modules for RPN (Region Proposal Network) head and R-FCN (Region-based Fully Convolutional Networks) head, and a cascaded R-FCN head. The semantic segmentation confidence is extracted and utilized as auxiliary classification prior knowledge for RPN proposal selection and R-FCN head prediction. Finally, the cascaded R-FCN head progressively refine the pedestrian prediction accuracy with negligible computation overhead. The proposed framework is also capable of maintaining high detection performance on down-sampled input images, which leads to further reduction in overall computational complexity. Experiment results on CityPersons and MOT17Det datasets show that the proposed framework achieves competitive detection performance with about$3\times $speedup over state-of-the-art methods. Chengju Zhou, Meiqing Wu, Siew-Kei Lam |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Information Entropy-Based Leakage ProfilingabstractAn accurate leakage model is critical to side-channel attacks and evaluations. Leakage certification plays an important role to address the following question: “how good is my leakage model?” Moreover, most of the current leakage model profiling only exploit the information from lower orders of moments. They still need to tolerate assumption error and estimation error from unknown leakage models. There are many probability density functions (PDFs) satisfying given moment constraints. As such, finding an unbiased, objective, and reasonable model still remains an unresolved problem. In this article, we address a more fundamental question: “which model can approach the leakage infinitely and is the optimal in theory?” In particular, we extract information from higher order moments and propose maximum entropy distribution (MED) to estimate the leakage model as MED is an unbiased, objective, and theoretically the most reasonable PDF conditioned upon the available information. MED is a moment-based statistical PDF model in side-channel attacks. It can theoretically use information on arbitrary higher order moments to infinitely approximate the leakage distribution, and well compensates the theory vacancy of model profiling and evaluation. Experimental results demonstrate the superiority of our proposed method for approximating the leakage model using MED estimation. Changhai Ou, Xinping Zhou, Siew-Kei Lam, Chengju Zhou, Fangxin Ning |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Multiple-Differential Mechanism for Collision-Optimized Divide-and-Conquer AttacksabstractSeveral combined attacks have shown promising results in recovering cryptographic keys by introducing collision information into divide-and-conquer attacks to transform a part of the best key candidates within given thresholds into a much smaller collision space. However, these Collision-Optimized Divide-and-Conquer Attacks (CODCAs) uniformly demarcate the thresholds for all sub-keys, which is unreasonable. Moreover, the inadequate exploitation of collision information and backward fault tolerance mechanisms of CODCAs also lead to low attack efficiency. Finally, existing CODCAs mainly focus on improving collision detection algorithms but lack theoretical basis. We exploit Correlation-Enhanced Collision Attack (CECA) to optimize Template Attack (TA). To overcome the above-mentioned problems, we first introduce guessing theory into TA to enable the quick estimation of success probability and the corresponding complexity of key recovery. Next, a novel Multiple-Differential mechanism for CODCAs (MD-CODCA) is proposed. The first two differential mechanisms construct collision chains satisfying the given number of collisions from several sub-keys with the fewest candidates under a fixed probability provided by guessing theory, then exploit them to vote for the remaining sub-keys. This guarantees that the number of remaining chains is minimal, and makes MD-CODCA suitable for very high thresholds. Our third differential mechanism simply divides the key into several large non-overlapping “blocks” to further exploit intra-block collisions from the remaining candidates and properly ignore the inter-block collisions, thus facilitating the later key enumeration. The experimental results show that MD-CODCA significantly reduces the candidate space and lowers the complexity of collision detection, without considerably reducing the success probability of attacks. Changhai Ou, Chengju Zhou, Siew-Kei Lam, Guiyuan Jiang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | A Lightweight Detection Algorithm For Collision-Optimized Divide-and-Conquer AttacksabstractBy introducing collision information into divide-and-conquer attacks, several existing works transform the original candidate space, which may be too large to enumerate, into a significantly smaller collision space, thus making key recovery possible. However, the inefficient collision detection algorithms and fault tolerance mechanisms make them time-consuming and their success rate low. Moreover, they may still leave very huge chain spaces that makes it difficult for key recovery. In this article, we exploit collision attack to optimize Template Attack (TA), and propose a Lightweight Collision Detection (LCD) algorithm. The proposed method exploits a jump detection mechanism to efficiently reduce the repetitive collision detections on chains with the same prefix sub-chains. We then introduce guessing theory to reorder the collision detection of the sub-keys according to their guessing lengths, and provide us with an evaluation tool. Finally, we design a highly efficient fault tolerance mechanism for our LCD to allow flexible thresholds adjustment, and further optimize sieving mechanism to efficiently extract the best chains with the largest number of collisions. Experimental results fully demonstrate LCD's superiority. Changhai Ou, Siew-Kei Lam, Chengju Zhou, Guiyuan Jiang, Fan Zhang 0010 |
IEEE Trans. Computers | 3 |
| 2020 | A First Study of Compressive Sensing for Side-Channel Leakage SamplingabstractAn important prerequisite for side-channel attacks (SCAs) is leakage sampling where the side-channel measurements (i.e., power traces) of the cryptographic device are collected for further analysis. However, as the operating frequency of cryptographic devices continues to increase due to advancing technology, leakage sampling will impose higher requirements on the sampling rate and storage capacity of the sampling equipment. This article undertakes the first study to show that effective leakage sampling can be achieved without relying on sophisticated equipments through compressive sensing (CS). As long as the information is leaked in the low-frequency component, CS can obtain low-dimensional samples by simply projecting the high-dimensional signals onto the observation matrix. The power traces can then be reconstructed in a workstation for further analysis and storage. With this approach, the sampling rate to obtain power traces is no longer limited by the operating frequency of the cryptographic device and the Nyquist sampling theorem. Instead, it depends on the sparsity of the leakage signal. As such, CS can employ a much lower sampling rate and yet obtain equivalent leakage sampling performance, which significantly lowers the requirement of sampling equipments. The feasibility of our approach is verified theoretically and through experiments. Changhai Ou, Chengju Zhou, Siew-Kei Lam |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Group Cost-Sensitive BoostLR With Vector Form Decorrelated Filters for Pedestrian DetectionabstractPedestrian detection has achieved notable progress in the field of computer vision over the past decade. However, existing top-performing approaches suffer from high computational complexity which prohibits their realization on embedded platforms with low computational capabilities. In this paper, we propose a robust and fast pedestrian detection framework which is based on the Filtered Channel Feature (FCF) approach. The proposed framework exploits vector-form decorrelated filters to extract more discriminative channel features while benefiting from low computational complexity. A novel group cost-sensitive BoostLR (Boosting with Loss Regularization) algorithm is proposed to train the classifier. The proposed training strategy provides more emphasis to the harder samples by exploring the variations of negatives selected from different rounds in hard negative mining processing, and hence is able to boost the overall detection performance. In addition, the proposed method also benefits from the BoostLR framework to achieve better generalization. Experiments on the well-known Caltech, INRIA and CityPersons pedestrian detection datasets show that our proposed approach achieves the best detection performance among all of the state-of-the-art non-deep learning methods and can run one order of magnitude faster than classical FCF methods (e.g. Checkerboards). Chengju Zhou, Meiqing Wu, Siew-Kei Lam |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2017 | Group Cost-sensitive Boosting with Multi-scale Decorrelated Filters for Pedestrian Detection
Chengju Zhou, Meiqing Wu, Siew-Kei Lam |
BMVC | 1 |
| 2017 | Fast and Accurate Pedestrian Detection using Dual-Stage Group Cost-Sensitive RealBoost with Vector Form FiltersabstractDespite significant research efforts in pedestrian detection over the past decade, there is still a ten-fold performance gap between the state-of-the-art methods and human perception. Deep learning methods can provide good performance but suffers from high computational complexity which prohibits their deployment on affordable systems with limited computational resources. In this paper, we propose a pedestrian detection framework that provides a major fillip to the robustness and run-time efficiency of the recent top performing non-deep learning Filtered Channel Feature (FCF) approach. The proposed framework overcomes the computational bottleneck of existing FCF methods by exploiting vector form filters to efficiently extract more discriminative channel features for pedestrian detection. A novel dual-stage group cost-sensitive RealBoost algorithm is used to explore different costs among different types of misclassification in the boosting process in order to improve detection performance. In addition, we propose two strategies, selective classification and selective scale processing, to further accelerate the detection process at the channel feature level and image pyramid level respectively. Experiments on the Caltech and INRIA datasets show that the proposed method achieves the highest detection performance among all the state-of-the-art non-CNN methods and is about 148X faster than the existing best performing FCF method on the Caltech dataset. Chengju Zhou, Meiqing Wu, Siew-Kei Lam |
ACM Multimedia | 1 |
| 2015 | Multi-cue Augmented Face ClusteringabstractFace clustering is an important but challenging task since facial images always have huge variation due to change in facial expressions, head poses and partial occlusions, etc. Moreover, face clustering is actually an unsupervised problem which makes it more difficult to reach an accurate result. Fortunately, there are some cues that can be used to improve clustering performance. In this paper, two types of cues are employed. The first one is pairwise constraints: must-link and cannot-link constraints, which can be extracted from the temporal and spatial knowledge of data. The other is that each face is associated with a series of attributes (i.e, gender) which can contribute discrimination among faces. To take advantage of the above cues, we propose a new algorithm, Multi-cue Augmented Face Clustering (McAFC), which effectively incorporates the cues via graph-guided sparse subspace clustering technique. Specially, facial images from the same individual are encouraged to be connected while faces from different persons are restrained to be connected. Experiments on three face datasets from real-world videos show the improvements of our algorithm over the state-of-the-art methods. Chengju Zhou, Changqing Zhang 0002, Huazhu Fu, Rui Wang 0032, Xiaochun Cao |
ACM Multimedia | 1 |
| 2015 | Constrained Multi-View Video Face ClusteringabstractIn this paper, we focus on face clustering in videos. To promote the performance of video clustering by multiple intrinsic cues, i.e., pairwise constraints and multiple views, we propose a constrained multi-view video face clustering method under a unified graph-based model. First, unlike most existing video face clustering methods which only employ these constraints in the clustering step, we strengthen the pairwise constraints through the whole video face clustering framework, both in sparse subspace representation and spectral clustering. In the constrained sparse subspace representation, the sparse representation is forced to explore unknown relationships. In the constrained spectral clustering, the constraints are used to guide for learning more reasonable new representations. Second, our method considers both the video face pairwise constraints as well as the multi-view consistence simultaneously. In particular, the graph regularization enforces the pairwise constraints to be respected and the co-regularization penalizes the disagreement among different graphs of multiple views. Experiments on three real-world video benchmark data sets demonstrate the significant improvements of our method over the state-of-the-art methods. Xiaochun Cao, Changqing Zhang 0002, Chengju Zhou, Huazhu Fu, Hassan Foroosh |
IEEE Trans. Image Process. | 3 |
| 2014 | Video Face Clustering via Constrained Sparse RepresentationabstractIn this paper, we focus on the problem of clustering faces in videos. Different from traditional clustering on a collection of facial images, a video provides some inherent benefits: faces from a face track must belong to the same person and faces from a video frame can not be the same person. These benefits can be used to enhance the clustering performance. More precisely, we convert the above benefits into must-link and cannot-link constraints. These constraints are further effectively incorporated into our novel algorithm, Video Face Clustering via Constrained Sparse Representation (CS-VFC). The CS-VFC utilizes the constraints in two stages, including sparse representation and spectral clustering. Experiments on real-world videos show the improvements of our algorithm over the state-of-the-art methods. Chengju Zhou, Changqing Zhang 0002, Gaotao Shi, Xiaochun Cao |
ICME | 1 |