EDBT 2026 Demo / reviewers in the wild / expert
Xiuzhuang Zhou
dblp:85/7674
· DBLP profile ↗
50ranked-venue papers
14as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 8 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 8 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Security and privacy · 4 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | E2-Former: An Edge-Enhanced Transformer for UAV-Based Small Object DetectionabstractThe widespread deployment of low-altitude drones in Internet of Things (IoT) applications, such as smart transportation and urban security, demands efficient, real-time visual perception on resource-constrained edge devices. A critical challenge in this domain is the precise detection and localization of small objects against complex backgrounds. Low-altitude UAV imagery is particularly challenging due to blurred small-target edges, weak anti-interference for medium and large targets, and fragmented multi-scale features. Existing convolutional and general-purpose Transformer-based detectors exhibit significant limitations in localization accuracy and robustness under these conditions. To address these issues, we proposeE2-Former, a novel object detection model based on the Detection Transformer (DETR) framework, designed for real-time performance on computational-edge devices. Our model integrates three core components: an Edge-Enhanced Backbone that augments sensitivity to contours while preserving deep semantics; a Polarized Dynamic Multi-Feature Fusion (PDMF) Transformer that leverages dual-path polarized attention and frequency-domain modulation to enhance local-global feature modeling and suppress background noise; and an Edge-aware Path Aggregation Network (E-PAN) that uses bidirectional gating and multi-level context interaction to resolve feature fragmentation and promote fine-grained, cross-scale integration. On three challenging low-altitude datasets—VisDrone2019, UAVDT, and CODrone—E2- Former achieves leadingAP50scores of 48.9%, 41.2%, and 33.0%, respectively, outperforming existing mainstream methods in both detection accuracy and robustness. Practical deployment and scene detection experiments further validate its impressive performance in real-world scenarios, establishing a new paradigm for building efficient and reliable low-altitude drone perception systems. Yao Zhang 0026, Said M. Easa, Boxiang Xie, Lingfeng Lin, Xiuzhuang Zhou, Nianyin Zeng |
IEEE Internet Things J. | 6 |
| 2025 | A Semantic Knowledge Complementarity based Decoupling Framework for Semi-supervised Class-imbalanced Medical Image SegmentationabstractThe limited data annotations have made semi-supervised learning (SSL) increasingly popular in medical image analysis. However, the use of pseudo labels in SSL degrades the performance of decoders that heavily rely on high-accuracy annotations. This issue is particularly pronounced in class-imbalanced multi-organ segmentation tasks, where small organs may be under-segmented or even ignored. In this paper, we propose SKCDF, a semantic knowledge complementarity based decoupling framework for multi-organ segmentation in class-imbalanced medical images. SKCDF decouples the data flow based on the responsibilities of the encoder and decoder during model training to make the model effectively learn semantic features, while mitigating the negative impact of unlabeled data on the semantic segmentation task. We also design a semantic knowledge complementarity module that adopts labeled data to guide the generation of pseudo labels and enriches the semantic features of labeled data with unlabeled data, which improves the quality of generated pseudo labels and the robustness of the overall model. Furthermore, we design an auxiliary balanced segmentation head based training strategy to further enhance the segmentation performance of small organs. Experimental results on the Synapse and AMOS datasets show that our method significantly outperforms existing methods. Zheng Zhang 0038, Guanchun Yin, Bo Zhang 0032, Wu Liu 0005, Xiuzhuang Zhou, Wendong Wang 0003 |
CVPR | 5 |
| 2025 | PlaneRAS: Learning Planar Primitives for 3D Plane Recovery
Wenzhao Zheng, Linqing Zhao, Zelan Zhu, Jiwen Lu, Xiuzhuang Zhou |
ICCV | 6 |
| 2025 | WL-GAN: Learning to sample in generative latent space
Zeyi Hou, Ning Lang, Xiuzhuang Zhou |
Inf. Sci. | 3 |
| 2025 | Zero-X21: Scale-agnostic image feature conditioned INR for multi-modal and multi-planar anisotropic MRI inter-slice interpolation
Zibo Ma, Jianfei Huo, Guanchun Yin, Bo Zhang 0032, Xiuzhuang Zhou, Wendong Wang 0003 |
Pattern Recognit. Lett. | 5 |
| 2025 | Factorized neural radiance field for autonomous driving
Xiuzhuang Zhou |
Pattern Recognit. Lett. | 3 |
| 2025 | Eliminating Non-Overlapping Semantic Misalignment for Cross-Modal Medical RetrievalabstractIn recent years, increasing research has shown that fine-grained local alignment is crucial for the cross-modal medical image-report retrieval task. However, existing local alignment learning methods suffer from the misalignment of semantically non-overlapping features between different modalities, which in turn negatively affects the retrieval performance. To address this challenge, we propose a Global-Feature Guided Cross-modal Local Alignment (GFG-CMLA) method. Unlike prior methods that rely on explicit local attention or learned weighting mechanisms, our approach leverages global semantic features extracted from the cross-modal common semantic space to implicitly guide local alignment, adaptively focusing on semantically overlapping content while filtering out irrelevant local regions, thus mitigating misalignment interference without additional annotations or architectural complexity. We validated the effectiveness of the proposed method through ablation experiments on the MIMICCXR and CheXpert Plus dataset. Furthermore, comparisons with state-of-the-art local alignment methods indicate that our approach achieves superior cross-modal retrieval performance. Zeqiang Wei, Zeyi Hou, Xiuzhuang Zhou |
IEEE Signal Process. Lett. | 3 |
| 2025 | Conformal Depression PredictionabstractWhile existing depression prediction methods based on deep learning show promise, their practical application is hindered by the lack of trustworthiness, as these deep models are often deployed asblack boxmodels, leaving us uncertain on the confidence of their predictions. For high-risk clinical applications like depression prediction, uncertainty quantification is essential in decision-making. In this paper, we introduce conformal depression prediction (CDP), a depression prediction method with uncertainty quantification based on conformal prediction (CP), giving valid confidence intervals with theoretical coverage guarantees for the model predictions. CDP is a plug-and-play module that requires neither model retraining nor an assumption about the depression data distribution. As CDP provides only an average coverage guarantee across all inputs rather than per-input performance guarantee, we further propose CDP-ACC, an improved conformal prediction with approximate conditional coverage. CDP-ACC firstly estimates the prediction distribution through neighborhood relaxation, and then introduces a conformal score function by constructing nested sequences, so as to provide a tighter prediction interval adaptive to specific input. We empirically demonstrate the application of CDP in uncertainty-aware facial depression prediction, as well as the effectiveness and superiority of CDP-ACC on the AVEC 2013 and AVEC 2014 datasets. Shan Qu, Xiuzhuang Zhou |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | MemRank: Memory-Augmented Similarity Ranking for Video-Based Depression Severity EstimationabstractDeep learning-based methods have shown substantial promise in visual depression severity estimation. Nonetheless, their effectiveness is limited by the scarce availability of labeled depression data, potentially leading to overfitting during representation learning. One feasible approach to address this issue is to incorporate, in the training objective, regularization that considers the unique characteristics of depression data. Typical regularization includes the similarity ranking through ordered consistency between visual features and their target scores. However, previous ranking methods are limited to using only samples within a mini-batch, resulting in a decreased regularization effect in depression representation learning. To address this limitation, we propose MemRank, a global similarity ranking method that operates not only on mini-batch samples but also on a well-designed feature memory, which stores smoothed and dynamically updated feature prototypes at diverse levels of depression during training. Furthermore, we show that incorporating the feature memory in the regression loss enhances the stability of training a deep regressor, leading to improved depression predictions. Empirically and analytically, we show that our MemRank outperforms alternative ranking methods and achieves state-of-the-art results on two benchmark datasets. Zeqiang Wei, Guodong Guo, Xiuzhuang Zhou |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | CEUS-SAM: Cross-Modal Prompt-Based SAM Network for Breast CEUS Image SegmentationabstractThe precise segmentation of lesions in contrast-enhanced ultrasound (CEUS) videos, especially during the peak enhancement phase, is crucial for early breast cancer diagnosis. However, the dynamic contrast patterns and subtle differences in CEUS images challenge traditional methods. To overcome this, we propose the CEUS-SAM network, a deep learning framework leveraging the Segment Anything Model (SAM) for enhanced lesion segmentation. Our approach first trains on conventional ultrasound (US) data, generating segmentation masks as prompts for CEUS images. A key innovation, the Image Fusion Module (IFM), integrates cross-modal and multi-scale features from US and CEUS, improving tissue differentiation and lesion detection. The CEUS-SAM network significantly reduces manual effort with single-point prompts and minimizes inter-observer variability. Using a breast CEUS dataset with 135 video sequences, our method achieves a Dice score of 78.6% and an IoU score of 66.6%. The code and dataset are available at https://github.com/2284650586/CEUS-SAM. Min Xu 0003, Ximiao Zhang, Sihua Niu, Jiaan Zhu, Xiuzhuang Zhou |
BIBM | 6 |
| 2024 | RealNet: A Feature Selection Network with Realistic Synthetic Anomaly for Anomaly DetectionabstractSelf-supervised feature reconstruction methods have shown promising advances in industrial image anomaly de-tection and localization. Despite this progress, these meth-ods still face challenges in synthesizing realistic and di-verse anomaly samples, as well as addressing the feature redundancy and pre-training bias of pre-trained feature. In this work, we introduce RealNet, a feature reconstruction network with realistic synthetic anomaly and adaptive feature selection. It is incorporated with three key inno-vations: First, we propose Strength-controllable Diffusion Anomaly Synthesis (SDAS), a diffusion process-based syn-thesis strategy capable of generating samples with varying anomaly strengths that mimic the distribution of real anomalous samples. Second, we develop Anomaly-aware Features Selection (A FS), a method for selecting repre-sentative and discriminative pre-trained feature subsets to improve anomaly detection performance while controlling computational costs. Third, we introduce Reconstruction Residuals Selection (RRS), a strategy that adaptively selects discriminative residuals for comprehensive identification of anomalous regions across multiple levels of granularity. We assess RealNet onfour benchmark datasets, and our results demonstrate significant improvements in both Image AU-Rae and Pixel AUROC compared to the current state-of-the-art methods. The code, data, and models are available at https://github.com/cnulab/RealNet. Ximiao Zhang, Min Xu 0003, Xiuzhuang Zhou |
CVPR | 3 |
| 2024 | Norma: A Noise Robust Memory-Augmented Framework for Whole Slide Image Classification
Yu Bai 0020, Bo Zhang 0032, Zheng Zhang 0038, Zibo Ma, Wu Liu 0005, Xiuzhuang Zhou, Xiangyang Gong, Wendong Wang 0003 |
ECCV (51) | 7 |
| 2024 | Energy-Based Controllable Radiology Report Generation with Medical Knowledge
Zeyi Hou, Ruixin Yan, Ziye Yan, Ning Lang, Xiuzhuang Zhou |
MICCAI (5) | 5 |
| 2024 | MediCLIP: Adapting CLIP for Few-Shot Medical Image Anomaly Detection
Ximiao Zhang, Min Xu 0003, Dehui Qiu, Ruixin Yan, Ning Lang, Xiuzhuang Zhou |
MICCAI (11) | 6 |
| 2024 | Multi-label contrastive hashing
Zeqiang Wei, Zheng Zhang 0038, Xiuzhuang Zhou |
Pattern Recognit. | 4 |
| 2024 | DepressionMLP: A Multi-Layer Perceptron Architecture for Automatic Depression Level Prediction via Facial Keypoints and Action UnitsabstractPhysiological studies have confirmed that there are differences in facial activities between depressed and healthy individuals. Therefore, while protecting the privacy of subjects, substantial efforts are made to predict the depression severity of individuals by analyzing Facial Keypoints Representation Sequences (FKRS) and Action Units Representation Sequences (AURS). However, those works has struggled to examine the spatial distribution and temporal changes of Facial Keypoints (FKs) and Action Units (AUs) simultaneously, which is limited in extracting the facial dynamics characterizing depressive cues. Besides, those works don’t realize the complementarity of effective information extracted from FKRS and AURS, which reduces the prediction accuracy. To this end, we intend to use the recently proposed Multi-Layer Perceptrons with gating (gMLP) architecture to process FKRS and AURS for predicting depression levels. However, the channel projection in the gMLP disrupts the spatial distribution of FKs and AUs, leading to input and output sequences not having the same spatiotemporal attributes. This discrepancy hinders the additivity of residual connections in a physical sense. Therefore, we construct a novel MLP architecture named DepressionMLP. In this model, we propose the Dual Gating (DG) and Mutual Guidance (MG) modules. The DG module embeds cross-location and cross-frame gating results into the input sequence to maintain the physical properties of data to make up for the shortcomings of gMLP. The MG module takes the global information of FKRS (AURS) as a guidance mask to filter the AURS (FKRS) to achieve the interaction between FKRS and AURS. Experimental results on several benchmark datasets show the effectiveness of our method. Mingyue Niu, Ya Li 0001, Jianhua Tao 0001, Xiuzhuang Zhou, Björn W. Schuller |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Confidence-Calibrated Face and Kinship VerificationabstractIn this paper, we investigate the problem of prediction confidence in face and kinship verification. Most existing face and kinship verification methods focus on accuracy performance while ignoring confidence estimation for their prediction results. However, confidence estimation is essential for modeling reliability and trustworthiness in such high-risk tasks. To address this, we introduce an effective confidence measure that allows verification models to convert a similarity score into a confidence score for any given face pair. We further propose a confidence-calibrated approach, termed Angular Scaling Calibration (ASC). ASC is easy to implement and can be readily applied to existing verification models without model modifications, yielding accuracy-preserving and confidence-calibrated probabilistic verification models. In addition, we introduce the uncertainty in the calibrated confidence to boost the reliability and trustworthiness of the verification models in the presence of noisy data. To the best of our knowledge, our work presents the first comprehensive confidence-calibrated solution for modern face and kinship verification tasks. We conduct extensive experiments on four widely used face and kinship verification datasets, and the results demonstrate the effectiveness of our proposed approach. Code and models are available athttps://github.com/cnulab/ASC. Min Xu 0003, Ximiao Zhang, Xiuzhuang Zhou |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Diversity-Preserving Chest Radiographs Generation from Reports in One Stage
Zeyi Hou, Ruixin Yan, Qizheng Wang, Ning Lang, Xiuzhuang Zhou |
MICCAI (5) | 5 |
| 2022 | LSRML: A latent space regularization based meta-learning framework for MR image segmentation
Bo Zhang 0032, Yunpeng Tan, Zheng Zhang 0038, Xiuzhuang Zhou, Jingyun Wu, Yue Mi, Haiwen Huang, Wendong Wang 0003 |
Pattern Recognit. | 5 |
| 2022 | Joint Adaptation of ICP Proposal and Target Distribution for Probabilistic Surface RegistrationabstractNon-rigid surface registration plays a crucial role in various vision applications. Recent advances in probabilistic surface registration implied the effectiveness of Markov-chain Monte Carlo (MCMC) sampling with ICP-based proposals. Despite the progress, the ICP proposal is still less informed for sampling from highly multi-modal parameter space of deformation field, and hence inferred registration solutions can be of limited accuracy. In this letter, we propose to jointly learn the ICP proposal and target distribution in a unified framework, where the step-size of the proposal and the invariant distribution are both adaptively adjusted during sampling, by making full use of geometric and statistical characteristics underlying history samples. Under the conditions of adaptation diminishing, the two adaptation modules can benefit from each other, leading to boosted registration accuracy. Experimental results on surfaces extracted from the SICAS Medical Image Repository demonstrate the improvements in terms of registration accuracy and convergence speed compared to alternative algorithms. Zeyi Hou, Xiuzhuang Zhou |
IEEE Signal Process. Lett. | 2 |
| 2022 | Facial Depression Recognition by Deep Joint Label Distribution and Metric LearningabstractWhile existing prediction models built on popular deep architectures have shown promising results in facial depression recognition, they still lack sufficient discriminative power due to the issues of 1) limited amount of labeled depression data for deep representation learning and, 2) large variation in facial expression across different persons of the same depression score and the subtle difference in facial expression across different depression levels. In this article, we formulate the facial depression recognition as a label distribution learning (LDL) problem, and propose a deep joint label distribution and metric learning (DJ-LDML) method to address these issues. In DJ-LDML, LDL exploits label relevance inherent in depression data to implicitly increase the amount of training data associated with each depression level without actually enlarging the dataset, while deep metric learning (DML) aims at learning a deep ordinal embedding with a specifically designed label-aware histogram loss, allowing semantics similarity between video sequences (described by ordinal labels) to be preserved for discriminative feature learning. The two learning modules in our DJ-LDML work collaboratively to enhance the representation ability and discriminative power of the deeply learned spatiotemporal feature, leading to improved depression prediction. We empirically evaluate our method on two benchmark datasets and the results demonstrate the effectiveness of our formulation. Xiuzhuang Zhou, Zeqiang Wei, Min Xu 0003, Shan Qu, Guodong Guo |
IEEE Trans. Affect. Comput. | 1 |
| 2021 | Supervised Contrastive Learning for Facial Kinship RecognitionabstractVision-based kinship recognition aims to determine whether the face images have a kin relation. Compared to traditional solutions, the vision-based kinship recognition methods have the advantages of lower cost and being easy to implement. Therefore, such technique can be widely employed in lots of scenarios including missing children search and automatic management of family album. The Recognizing Families in the Wild (RFIW) Data Challenge provides a platform for evaluation of different kinship recognition approaches with ranked results. We propose a supervised contrastive learning approach to address three different kinship recognition tracks (i.e., kinship verification, tri-subject verification, and large-scale search-and-retrieval) announced in the RFIW 2021 with the 2021 FG. Our results on three tracks of 2021 RFIW challenge achieve the highest ranking, which demonstrate the superiority of the proposed solution. Ximiao Zhang, Min Xu 0003, Xiuzhuang Zhou, Guodong Guo |
FG | 3 |
| 2021 | Energy-Based Supervised Hashing for Multimorbidity Image Retrieval
Xiuzhuang Zhou, Zeqiang Wei, Guodong Guo |
MICCAI (5) | 2 |
| 2020 | Visually Interpretable Representation Learning for Depression Recognition from Facial ImagesabstractRecent evidence in mental health assessment have demonstrated that facial appearance could be highly indicative of depressive disorder. While previous methods based on the facial analysis promise to advance clinical diagnosis of depressive disorder in a more efficient and objective manner, challenges in visual representation of complex depression pattern prevent widespread practice of automated depression diagnosis. In this paper, we present a deep regression network termed DepressNet to learn a depression representation with visual explanation. Specifically, a deep convolutional neural network equipped with a global average pooling layer is first trained with facial depression data, which allows for identifying salient regions of input image in terms of its severity score based on the generated depression activation map (DAM). We then propose a multi-region DepressNet, with which multiple local deep regression models for different face regions are jointly leaned and their responses are fused to improve the overall recognition performance. We evaluate our method on two benchmark datasets, and the results show that our method significantly boosts state-of-the-art performance of the visual-based depression recognition. Most importantly, the DAM induced by our learned deep model may help reveal the visual depression pattern on faces and understand the insights of automated depression diagnosis. Xiuzhuang Zhou, Guodong Guo |
IEEE Trans. Affect. Comput. | 1 |
| 2018 | Consistency-Exclusivity Regularized Deep Metric Learning for General Kinship VerificationabstractWhile encouraging results have been made so far to advance kinship verification by using facial images, learning a robust genetic similarity measure remains challenging, especially in the setting of general kinship verification, wherein the gender labels of the test samples are unknown in advance. In this paper we present a deep metric learning method with a carefully designed two-stream neural network to jointly learn a pair of deep embeddings for parent-child images. In particular, the deep embeddings are first modeled to explicitly consist of the common and individual components, and then two additional constraints are introduced in deep metric learning: 1) value-aware consistency on the common components, and 2) position-aware exclusivity on the individual components. The proposed hierarchical consistency-exclusivity regularization enables our deep metric learning to harness the sharable and complementary patterns inherent in parent-child images. Empirically, we show improved performance over state of the art metric learning solutions to general kinship verification on two benchmarks. Xiuzhuang Zhou, Zheng Zhang 0038, Zeqiang Wei, Min Xu 0003 |
ICME | 1 |
| 2018 | Multiple face tracking and recognition with identity-specific localized metric learning
Xiuzhuang Zhou, Qian Chen 0033, Min Xu 0003 |
Pattern Recognit. | 1 |
| 2017 | Fast single image dehazing based on a regression model
Zhong Luan, Xiuzhuang Zhou, Zhuhong Shao, Guodong Guo, Xiaoming Liu 0002 |
Neurocomputing | 3 |
| 2017 | Learning spatially regularized similarity for robust visual tracking
Xiuzhuang Zhou, Qirun Huo, Min Xu 0003 |
Image Vis. Comput. | 1 |
| 2016 | ZigzagNet: Efficient Deep Learning for Real Object Recognition Based on 3D Models
Yida Wang 0001, Xiuzhuang Zhou, Weihong Deng |
ACCV (4) | 3 |
| 2016 | Sky detection by effective context inference
Zhong Luan, Xiuzhuang Zhou, Guodong Guo |
Neurocomputing | 4 |
| 2016 | Hybrid generative-discriminative learning for online tracking of sperm cell
Xiuzhuang Zhou, Min Xu 0003, Xiaoyan Fu |
Neurocomputing | 1 |
| 2016 | Kinship verification from facial images by scalable similarity fusion
Xiuzhuang Zhou, Haibin Yan |
Neurocomputing | 1 |
| 2015 | Neighborhood repulsed correlation metric learning for kinship verificationabstractIn this paper, we propose a new neighborhood repulsed correlation metric learning (NRCML) method for kinship verification. While several metric learning algorithms have been proposed in recent years and some of them have successfully applied to kinship verification, most existing metric learning methods are developed based on the Euclidian similarity metric, which is not powerful enough to measure the similarity of face samples. To address this, we propose a NRCML method by using the correlation similarity measure to learn a discriminative distance metric, under which positive pairs are pulled as close as possible and negative pairs lying in a neighborhood are repulsed as far as possible, simultaneously. Experimental results are presented to show the effectiveness of the proposed method. Haibin Yan, Xiuzhuang Zhou, Yongxin Ge |
VCIP | 2 |
| 2015 | Learning Compact Binary Face Descriptor for Face RecognitionabstractBinary feature descriptors such as local binary patterns (LBP) and its variations have been widely used in many face recognition systems due to their excellent robustness and strong discriminative power. However, most existing binary face descriptors are hand-crafted, which require strong prior knowledge to engineer them by hand. In this paper, we propose a compact binary face descriptor (CBFD) feature learning method for face representation and recognition. Given each face image, we first extract pixel difference vectors (PDVs) in local patches by computing the difference between each pixel and its neighboring pixels. Then, we learn a feature mapping to project these pixel difference vectors into low-dimensional binary vectors in an unsupervised manner, where 1) the variance of all binary codes in the training set is maximized, 2) the loss between the original real-valued codes and the learned binary codes is minimized, and 3) binary codes evenly distribute at each learned bin, so that the redundancy information in PDVs is removed and compact binary codes are obtained. Lastly, we cluster and pool these binary codes into a histogram feature as the final representation for each face image. Moreover, we propose a coupled CBFD (C-CBFD) method by reducing the modality gap of heterogeneous faces at the feature level to make our method applicable to heterogeneous face recognition. Extensive experimental results on five widely used face datasets show that our methods outperform state-of-the-art face descriptors. Jiwen Lu, Venice Erin Liong, Xiuzhuang Zhou, Jie Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Prototype-Based Discriminative Feature Learning for Kinship VerificationabstractIn this paper, we propose a new prototype-based discriminative feature learning (PDFL) method for kinship verification. Unlike most previous kinship verification methods which employ low-level hand-crafted descriptors such as local binary pattern and Gabor features for face representation, this paper aims to learn discriminative mid-level features to better characterize the kin relation of face images for kinship verification. To achieve this, we construct a set of face samples with unlabeled kin relation from the labeled face in the wild dataset as the reference set. Then, each sample in the training face kinship dataset is represented as a mid-level feature vector, where each entry is the corresponding decision value from one support vector machine hyperplane. Subsequently, we formulate an optimization function by minimizing the intraclass samples (with a kin relation) and maximizing the neighboring interclass samples (without a kin relation) with the mid-level features. To better use multiple low-level features for mid-level feature learning, we further propose a multiview PDFL method to learn multiple mid-level features to improve the verification performance. Experimental results on four publicly available kinship datasets show the superior performance of the proposed methods over both the state-of-the-art kinship verification methods and human ability in our kinship verification task. Haibin Yan, Jiwen Lu, Xiuzhuang Zhou |
IEEE Trans. Cybern. | 3 |
| 2014 | Kinship verification in the wild: The first kinship verification competitionabstractKinship verification from facial images in wild conditions is a relatively new and challenging problem in face analysis. Several datasets and algorithms have been proposed in recent years. However, most existing datasets are of small sizes and one standard evaluation protocol is still lack so that it is difficult to compare the performance of different kinship verification methods. In this paper, we present the Kinship Verification in the Wild Competition: the first kinship verification competition which is held in conjunction with the International Joint Conference on Biometrics 2014, Clearwater, Florida, USA. The key goal of this competition is to compare the performance of different methods on a new-collected dataset with the same evaluation protocol and develop the first standardized benchmark for kinship verification in the wild. Jiwen Lu, Junlin Hu 0001, Xiuzhuang Zhou, Jie Zhou 0001, Modesto Castrillón-Santana, Javier Lorenzo-Navarro, Lu Kou, Andrea Bottino, Tiago F. Vieira |
IJCB | 3 |
| 2014 | Multi-feature multi-manifold learning for single-sample face recognition
Haibin Yan, Jiwen Lu, Xiuzhuang Zhou |
Neurocomputing | 3 |
| 2014 | Neighborhood Repulsed Metric Learning for Kinship VerificationabstractKinship verification from facial images is an interesting and challenging problem in computer vision, and there are very limited attempts on tackle this problem in the literature. In this paper, we propose a new neighborhood repulsed metric learning (NRML) method for kinship verification. Motivated by the fact that interclass samples (without a kinship relation) with higher similarity usually lie in a neighborhood and are more easily misclassified than those with lower similarity, we aim to learn a distance metric under which the intraclass samples (with a kinship relation) are pulled as close as possible and interclass samples lying in a neighborhood are repulsed and pushed away as far as possible, simultaneously, such that more discriminative information can be exploited for verification. To make better use of multiple feature descriptors to extract complementary information, we further propose a multiview NRML (MNRML) method to seek a common distance metric to perform multiple feature fusion to improve the kinship verification performance. Experimental results are presented to demonstrate the efficacy of our proposed methods. Finally, we also test human ability in kinship verification from facial images and our experimental results show that our methods are comparable to that of human observers. Jiwen Lu, Xiuzhuang Zhou, Yap-Peng Tan, Jie Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Equidistant prototypes embedding for single sample based face recognition with generic learning and incremental learning
Weihong Deng, Jiani Hu, Xiuzhuang Zhou, Jun Guo 0002 |
Pattern Recognit. | 3 |
| 2014 | Discriminative Multimetric Learning for Kinship VerificationabstractIn this paper, we propose a new discriminative multimetric learning method for kinship verification via facial image analysis. Given each face image, we first extract multiple features using different face descriptors to characterize face images from different aspects because different feature descriptors can provide complementary information. Then, we jointly learn multiple distance metrics with these extracted multiple features under which the probability of a pair of face image with a kinship relation having a smaller distance than that of the pair without a kinship relation is maximized, and the correlation of different features of the same face sample is maximized, simultaneously, so that complementary and discriminative information is exploited for verification. Experimental results on four face kinship data sets show the effectiveness of our proposed method over the existing single-metric and multimetric learning methods. Haibin Yan, Jiwen Lu, Weihong Deng, Xiuzhuang Zhou |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2013 | Polygonal Approximation of Digital Planar Curves via Hybrid Monte Carlo OptimizationabstractThis letter presents a novel computing paradigm for polygonal approximation of digital planar curves. While the existing heuristic algorithms, such as genetic algorithm (GA) and particle swarm optimization (PSO), have achieved considerable success in solving the two types of polygonal approximation problems, more efficient optimization schemes are still desirable for practical applications. We propose to embed the split-and-merge local search in the Monte Carlo sampling framework, to combine strength of the local optimization and the global sampling. The proposed algorithm is essentially a well-designed basin hopping scheme that performs stochastic exploration in the reduced potential energy space. Experimental results on several benchmarks indicate that the proposed algorithm can achieve high approximation accuracy and is highly competitive to the state-of-the-art alternative algorithms with less computational cost. Xiuzhuang Zhou, Jiwen Lu |
IEEE Signal Process. Lett. | 1 |
| 2012 | Neighborhood repulsed metric learning for kinship verificationabstractKinship verification from facial images is a challenging problem in computer vision, and there is a very few attempts on tackling this problem in the literature. In this paper, we propose a new neighborhood repulsed metric learning (NRML) method for kinship verification. Motivated by the fact that interclass samples (without kinship relations) with higher similarity usually lie in a neighborhood and are more easily misclassified than those with lower similarity, we aim to learn a distance metric under which the intraclass samples (with kinship relations) are pushed as close as possible and interclass samples lying in a neighborhood are repulsed and pulled as far as possible, simultaneously, such that more discriminative information can be exploited for verification. Moreover, we propose a multiview NRM-L (MNRML) method to seek a common distance metric to make better use of multiple feature descriptors to further improve the verification performance. Experimental results are presented to demonstrate the efficacy of the proposed methods. Jiwen Lu, Junlin Hu 0001, Xiuzhuang Zhou, Yap-Peng Tan, Gang Wang 0012 |
CVPR | 3 |
| 2012 | Activity-based person identification using sparse coding and discriminative metric learningabstractThis paper presents a new activity-based person identification method using sparse coding and discriminative metric learning. Different from gait recognition where human walking activity is only utilized for person identification, we aim to recognize people from different activities such as running, jumping, skipping, and so on. For each activity video clip, we extract the binary human body mask using background substraction. Then, we cluster these body masks into a number of clusters by sparse coding with mean pooling to extract features for each video clip. Subsequently, we learn a discriminative distance metric under which intraclass (activities performed by the same person) variations are minimized and the interclass (activities performed by different persons) are maximized, simultaneously, such that more discriminative information can be exploited for recognition. Experimental results on a publicly available database are presented to show the efficacy of our proposed method. Jiwen Lu, Junlin Hu 0001, Xiuzhuang Zhou |
ACM Multimedia | 3 |
| 2012 | Gabor-based gradient orientation pyramid for kinship verification under uncontrolled environmentsabstractThis paper presents a Gabor-based Gradient Orientation Pyramid (GGOP) feature representation method for kinship verification from facial images. First, we perform Gabor wavelet on each face image to obtain a set of Gabor magnitude (GM) feature images from different scales and orientations. Then, we extract the Gradient Orientation Pyramid (GOP) feature of each GM feature image and perform multiple feature fusion for kinship verification. When combined with the discriminative support vector machine (SVM) classifier, GGOP demonstrates the best performance in our experiments, in comparison with several state-of-the-art face feature descriptors. Experimental results are presented to show the efficacy of our proposed approach. Moreover, the performance of our proposed method is also comparable to that of human observers. Xiuzhuang Zhou, Jiwen Lu, Junlin Hu 0001 |
ACM Multimedia | 1 |
| 2012 | Cost-Sensitive Semi-Supervised Discriminant Analysis for Face RecognitionabstractThis paper presents a cost-sensitive semi-supervised discriminant analysis method for face recognition. While a number of semi-supervised dimensionality reduction algorithms have been proposed in the literature and successfully applied to face recognition in recent years, most of them aim to seek low-dimensional feature representations to achieve low classification errors and assume the same loss from all misclassifications in the feature representation/extraction phase. In many real-world face recognition applications, however, this assumption may not hold as different misclassifications could lead to different losses. For example, it may cause inconvenience to a gallery person who is misrecognized as an impostor and not allowed to enter the room by a face recognition-based door locker, but it could result in a serious loss or damage if an impostor is misrecognized as a gallery person and allowed to enter the room. Motivated by this concern, we propose in this paper a new method to learn a discriminative feature subspace by making use of both labeled and unlabeled samples and exploring different cost information of all the training samples simultaneously. Experimental results are presented to demonstrate the efficacy of the proposed method. Jiwen Lu, Xiuzhuang Zhou, Yap-Peng Tan, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2012 | Abrupt Motion Tracking Via Intensively Adaptive Markov-Chain Monte Carlo SamplingabstractThe robust tracking of abrupt motion is a challenging task in computer vision due to its large motion uncertainty. While various particle filters and conventional Markov-chain Monte Carlo (MCMC) methods have been proposed for visual tracking, these methods often suffer from the well-known local-trap problem or from poor convergence rate. In this paper, we propose a novel sampling-based tracking scheme for the abrupt motion problem in the Bayesian filtering framework. To effectively handle the local-trap problem, we first introduce the stochastic approximation Monte Carlo (SAMC) sampling method into the Bayesian filter tracking framework, in which the filtering distribution is adaptively estimated as the sampling proceeds, and thus, a good approximation to the target distribution is achieved. In addition, we propose a new MCMC sampler with intensive adaptation to further improve the sampling efficiency, which combines a density-grid-based predictive model with the SAMC sampling, to give a proposal adaptation scheme. The proposed method is effective and computationally efficient in addressing the abrupt motion problem. We compare our approach with several alternative tracking algorithms, and extensive experimental results are presented to demonstrate the effectiveness and the efficiency of the proposed method in dealing with various types of abrupt motions. Xiuzhuang Zhou, Jiwen Lu, Jie Zhou 0001 |
IEEE Trans. Image Process. | 1 |
| 2011 | Kinship verification from facial images under uncontrolled conditionsabstractIn this paper, we present an automatic kinship verification system based on facial image analysis under uncontrolled conditions. While a large number of studies on human face analysis have been performed in the literature, there are a few attempts on automatic face analysis for kinship verification, possibly due to lacking of such publicly available databases and great challenges of this problem. To this end, we collect a kinship face database by searching 400+ pairs of public figures and celebrities from the internet, and automatically detect them with the Viola-Jones face detector. Then, we propose a new spatial pyramid learning-based (SPLE) feature descriptor for face representation and apply support vector machine (SVM) for kinship verification. The proposed system has the following three characteristics: 1) no manual human annotation of face landmarks is required and the kinship information is automatically obtained from the original pair of images; 2) both local appearance information and global spatial information have been effectively utilized in the proposed SPLE feature descriptor, and better performance can be obtained than state-of-the-art feature descriptors in our application; 3) the performance of our proposed system is comparable to that of human observers. Xiuzhuang Zhou, Junlin Hu 0001, Jiwen Lu |
ACM Multimedia | 1 |
| 2010 | Abrupt motion tracking via adaptive stochastic approximation Monte Carlo samplingabstractRobust tracking of abrupt motion is a challenging task in computer vision due to the large motion uncertainty. In this paper, we propose a stochastic approximation Monte Carlo (SAMC) based tracking scheme for abrupt motion problem in Bayesian filtering framework. In our tracking scheme, the particle weight is dynamically estimated by learning the density of states in simulations, and thus the local-trap problem suffered by the conventional MCMC sampling-based methods could be essentially avoided. In addition, we design an adaptive SAMC sampling method to further speed up the sampling process for tracking of abrupt motion. It combines the SAMC sampling and a density grid based statistical predictive model, to give a data-mining mode embedded global sampling scheme. It is computationally efficient and effective in dealing with abrupt motion difficulties. We compare it with alternative tracking methods. Extensive experimental results showed the effectiveness and efficiency of the proposed algorithm in dealing with various types of abrupt motions. Xiuzhuang Zhou |
CVPR | 1 |
| 2010 | Polygonal approximation of digital curves using adaptive MCMC samplingabstractPolygonal approximation (PA) of the digital planar curves is an important topic in computer vision community. In this paper, we address this problem in the energy-minimization framework. We present a novel stochastic search scheme, which combines a split-and-merge process and a stochastic approximation Monte Carlo (SAMC) sampling procedure for global optimization. The SAMC sampling method can effectively handle the local-trap problem suffered by many local search methods, while the split-and-merge process is used to construct a more informative proposal distribution, and thus further improves the overall sampling efficiency. Experimental results on various benchmarks show that the proposed algorithm can achieve high-quality solutions and comparable results to those of state-of-the-art methods. Xiuzhuang Zhou |
ICIP | 1 |
| 2010 | Efficient Polygonal Approximation of Digital Curves via Monte Carlo OptimizationabstractA novel stochastic searching scheme based on the Monte Carlo optimization is presented for polygonal approximation (PA) problem. We propose to combine the split-and-merge based local optimization and the Monte Carlo sampling, to give an efficient stochastic optimization scheme. Our approach, in essence, is a well-designed Basin-Hopping scheme, which performs stochastic hopping among the reduced energy peaks. Experiment results on various benchmarks show that our method achieves high-quality solutions with lower computational costs, and outperforms most of state-of-the-art algorithms for PA problem. Xiuzhuang Zhou |
ICPR | 1 |