Peng Zhang 0057

dblp:21/1048-57 · DBLP profile ↗
← Back
33ranked-venue papers
9as first author
22since 2021 · last 2026
0000-0001-6794-7352ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 9 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse Attention
abstract
Medical vision-language pre-training (VLP) offers significant potential for advancing medical image understanding by leveraging paired image-report data. However, existing methods are limited by False Negatives (FaNe) induced by semantically similar texts and insufficient fine-grained cross-modal alignment. To address these limitations, we propose FaNe, a semantic-enhanced VLP framework. To mitigate false negatives, we introduce a semantic-aware positive pair mining strategy based on text-text similarity with adaptive normalization. Furthermore, we design a text-conditioned sparse attention pooling module to enable fine-grained image-text alignment through localized visual representations guided by textual cues. To strengthen intra-modal discrimination, we develop a hard-negative aware contrastive loss that adaptively reweights semantically similar negatives. Extensive experiments on five downstream medical imaging benchmarks demonstrate that FaNe achieves state-of-the-art performance across image classification, object detection, and semantic segmentation, validating the effectiveness of our framework.
Peng Zhang 0057, Zhihui Lai 0001, Wenting Chen, Xu Wu 0001, Heng Kong
AAAI1
2026 A Lightweight Passive Depression Detection System Based on Interpretable Facial Feature Analysis
abstract
Automatic depression detection from facial videos is promising for IoT deployment but has long suffered from the black-box nature of deep models, which overlook how depression manifests on the face. This lack of interpretability leads to redundant feature learning, overly complex architectures, and consequently a trade-off between accuracy and deployability. We introduce a lightweight, interpretable system that explicitly models static–dynamic facial patterns, color channels, and regional textures. At its core is a Hybrid Static–Dynamic Model (HSDM) with detachable modules and identity decoupling, supported by two task-driven preprocessing steps (blue-channel suppression and 90° Gabor filtering). On AVEC2013/2014, our approach reduces mean absolute error by 10.1% on AVEC2014 while maintaining competitive results on AVEC2013. From a systems perspective, it achieves real-time inference on Jetson Nano (0.071M parameters, 0.21 GFLOPs, 17.69 ms latency, 56.54 samples/s throughput), outperforming larger SOTA models under identical conditions. The design facilitates on-device deployment and offers transparent decision paths through interpretable static–dynamic fusion, color synergy, and region-level analysis.
Peng Zhang 0057, Tianhuan Huang, Jian Zhao 0002, Wei Xiang 0001, Xianye Ben
IEEE Internet Things J.2
2026 Cross-scale gated embedding graph model for skeleton-based anomaly detection
Peng Zhang 0057, Tianhuan Huang, Xianye Ben, Lei Chen 0095
Multim. Syst.2
2026 ConsistencyTrack: A robust multi-object tracker with a generation strategy of consistency model
Lifan Jiang, Zhihui Wang 0003, Siqi Yin, Guangxiao Ma, Peng Zhang 0057, Boxi Wu 0001
Pattern Recognit.5
2026 PSO-HEAD: Pseudo-Supervision Guided Spatial Optimization for View-Consistent 3D Full-Head Reconstruction
abstract
Generative 3D human head reconstruction in \(360^{\circ}\) is attracting increasing attention because of its flexibility in downstream animation applications. Existing generative 3D head synthesis approaches are primarily limited to near-frontal face priors, which cause distorted artifacts at large view angles. In this article, we introduce a novel Pseudo-Supervision Guided Spatial Optimization (PSO-HEAD) framework that reconstructs 3D view-consistent full-head through explicitly introducing pseudo-label of back-head supervision for spatial texture and geometric optimization. Particularly, our PSO-HEAD introduces two key improvements, i.e., Pseudo-Supervision Augmented Inversion (PSA-Inversion) and Full-Head Aware Generative Enhancement (FAGE). PSA-Inversion augments plausible invisible back-head as pseudo-supervision to optimize the view-hallucinated latent code conditioned on the augmented camera poses via GAN inversion, enforcing 3D spatial consistency across both visible and invisible regions. Furthermore, FAGE fine-tunes the 3D GAN on a proposed auxiliary FK-Enhance dataset deriving from either generated or real-world high-quality back-head images, which therefore improves the generalization of our PSO-HEAD to diverse hairstyles or underrepresented regions. Benefiting from the improvements, our PSO-HEAD enables efficient \(360^{\circ}\) view-consistent full-head generation from single input images, particularly improving reconstruction fidelity of unobserved regions, which quantitatively and qualitatively outperforms the state-of-the-art methods.
Peng Zhang 0057, Chongxin Liang, Caifeng Shan
ACM Trans. Multim. Comput. Commun. Appl.1
2025 Semi-Supervised Dual-Threshold Contrastive Learning for Ultrasound Image Classification and Segmentation
abstract
Confidence-based pseudo-label selection usually generates overly confident yet incorrect predictions, due to the early misleadingness of model and overfitting inaccurate pseudo-labels in the learning process, which heavily degrades the performance of semi-supervised contrastive learning. Moreover, segmentation and classification tasks are treated independently and the affinity fails to be fully explored. To address these issues, we propose a novel semi-supervised dual-threshold contrastive learning strategy for ultrasound image classification and segmentation, named Hermes. This strategy combines the strengths of contrastive learning with semi-supervised learning, where the pseudo-labels assist contrastive learning by providing additional guidance. Specifically, an inter-task attention and saliency module is also developed to facilitate information sharing between the segmentation and classification tasks. Furthermore, an inter-task consistency learning strategy is designed to align tumor features across both tasks, avoiding negative transfer for reducing features discrepancy. To solve the lack of publicly available ultrasound datasets, we have collected the SZ-TUS dataset, a thyroid ultrasound image dataset. Extensive experiments on two public ultrasound datasets and one private dataset demonstrate that Hermes consistently outperforms several state-of-the-art methods across various semi-supervised settings. The code is available at https://github.com/Aventador8/Hermes.
Peng Zhang 0057, Zhihui Lai 0001, Heng Kong
ECAI1
2025 MSFIQA: A multi-scale feature fusion network based on human visual perception for no-reference image quality assessment
Tianfeng Xia, Yongcan Zhao, Peng Zhang 0057, Xianye Ben, Lei Chen 0095
J. Vis. Commun. Image Represent.3
2025 HIAN: A hybrid interactive attention network for multimodal sarcasm detection
Yongtang Bao, Peng Zhang 0057
Pattern Recognit.3
2025 Corruption-Invariant Person Re-Identification via Coarse-to-Fine Feature Alignment
abstract
Corruption-invariant Person Re-identification (CI-ReID) aims to build robust identity correspondence across non-overlapped cameras even when severe image corruptions occur. It is challenging as those corruptions contaminate intrinsic pedestrian characteristics and cause semantic misalignment in feature space. To address this issue, this paper proposes a coarse-to-fine semantic alignment framework that learns corruption-invariant pedestrian features for re-identification from the perspective of multi-modal feature alignment. In this framework, a Coarse-to-Fine Feature Alignment Transformer (CFAT) is introduced to extract and align features of pedestrian images with different corruptions. Specifically, the CFAT aligns features of corrupted samples to that of the corresponding clean samples in a knowledge distillation manner in the coarse alignment stage, i.e., a teacher network distils identity-related semantics from clean samples and supervises the student network learning semantic-consistent features from corrupted samples. To avoid information loss of the strict alignment, we propose to integrate a Bridge Feature Generation (BFG) module into CFAT to construct meaningful latent structures among modalities in the fine alignment stage. This enables seamless alignment of the same identity between corrupted and clean modalities, leading to better re-identification performance. To evaluate the effectiveness of the proposed method, extensive experiments are conducted on three public benchmark datasets, i.e., Market-1501, CUHK-03, and MSMT-17. The experimental results demonstrate our CFAT outputs state-of-the-arts with a large margin in various corrupted scenes.
Peng Zhang 0057, Caifeng Shan
IEEE Trans. Circuits Syst. Video Technol.2
2024 Identity Semantic Correspondence for Cloth-Changing Person Re-Identification
abstract
Cloth-changing Person Re-Identification (CC-ReID) aims at retrieving the same person who might change clothes across different locations. Although remarkable progress has been achieved in recent studies, most of the recent methods still lack sufficient emphasis on identity-related regions. To address these issues, we propose a novel Identity Semantic Correspondence framework (ISC) to fully utilize human semantic information, which includes dual-stream identity semantic correspondance networks, i.e., a Clothing-invariant Identity (CI) stream and a Fine-grained Identity Semantic Highlighting (FISH) stream. The CI mitigates the interference of human dressing and enhances clothing-invariant identity information by erasing clothing information. And, the FISH exploits fine-grained identity information by highlighting contribution of different body parts to identity with a part-aware weighting module. Additionally, an identity semantic consistency module is further proposed to extract the most representative and discriminative semantic features for each identity. Besides, we employ a mutual loss to transfer identity-related knowledge between components, which enables the original appearance module to be deployed independently during the inference stage. Extensive experiments on two CC-ReID benchmarks, including PRCC and VC-Clothes, are conducted to demonstrate the effectiveness of the proposed ISC method.
Yongtang Bao, Kun Zhan, Peng Zhang 0057
IJCNN5
2024 Efficient abnormal behavior detection with adaptive weight distribution
Yefeng Qin, Lei Chen 0095, Peng Zhang 0057, Xianye Ben
Neurocomputing4
2024 Boosting Micro-Expression Recognition via Self-Expression Reconstruction and Memory Contrastive Learning
abstract
Micro-expression (ME) is an instinctive reaction that is not controlled by thoughts. It reveals one's inner feelings, which is significant in sentiment analysis and lie detection. Since micro-expression is expressed as subtle facial changes within particular facial action units, learning discriminative and generalized features for Micro-expression Recognition (MER) is challenging. To achieve the purpose, this paper proposes a novel MER framework that simultaneously integrates supervised Prototype-based Memory Contrastive Learning (PMCL) for discriminative feature mining and adds Self-expression Reconstruction (SER) as an auxiliary task and regularization for better generalization. In particular, the proposed SER module is forced as a regularization by reconstructing input ME from the randomly dropped patch- wise features in the bottleneck. And, the PMCL module globally compares historical and current cluster agents learned from training instances to enhance intra-class compactness and inter-class separability. Extensive experiments are conducted on three benchmarks, e.g., SMIC, CASME II, and SAMM, under evaluation criteria of both Composite Database Evaluation (CDE) and Single Database Evaluation (SDE) protocols. The results show our method surpasses other state-of-the-art approaches under various evaluation metrics, achieving overall 86.30% unweighed F1-score and 88.30% unweighed average recall on the composite dataset. Furthermore, the ablation studies verify the effectiveness of our SER for better generalization and PMCL for better discrimination in learning feature representation from limited micro-expression samples.
Yongtang Bao, Peng Zhang 0057, Caifeng Shan, Xianye Ben
IEEE Trans. Affect. Comput.3
2024 Pedestrian Attribute Recognition via Spatio-temporal Relationship Learning for Visual Surveillance
abstract
Pedestrian attribute recognition (PAR) aims at predicting the visual attributes of a pedestrian image. PAR has been used as soft biometrics for visual surveillance and IoT security. Most of the current PAR methods are developed based on discrete images. However, it is challenging for the image-based method to handle the occlusion and action-related attributes in real-world applications. Recently, video-based PAR has attracted much attention in order to exploit the temporal cues in the video sequences for better PAR. Unfortunately, existing methods usually ignore the correlations among different attributes and the relations between attributes and spatio regions. To address this problem, we propose a novel method for video-based PAR by exploring the relationships among different attributes in both the spatio and temporal domains. More specifically, a spatio-temporal saliency module (STSM) is introduced to capture the key visual patterns from the video sequences, and a module for spatio-temporal attribute relationship learning (STARL) is proposed to mine the correlations among these patterns. Meanwhile, a large-scale benchmark for video-based PAR, RAP-Video, is built by extending the image-based dataset RAP-2, which contains 83,216 tracklets with 25 scenes. To the best of our knowledge, this is the largest dataset for video-based PAR. Extensive experiments are performed on the proposed benchmark as well as on MARS Attribute and DukeMTMC-Video Attribute. The superior performance demonstrates the effectiveness of the proposed method.
Da Li 0003, Zhang Zhang 0001, Peng Zhang 0057, Caifeng Shan, Jungong Han
ACM Trans. Multim. Comput. Commun. Appl.5
2023 Entropy Neural Estimation for Graph Contrastive Learning
abstract
Contrastive learning on graphs aims at extracting distinguishable high-level representations of nodes. We theoretically illustrate that the entropy of a dataset is approximated by maximizing the lower bound of the mutual information across different views of a graph, i.e., entropy is estimated by a neural network. Based on this finding, we propose a simple yet effective subset sampling strategy to contrast pairwise representations between views of a dataset. In particular, we randomly sample nodes and edges from a given graph to build the input subset for a view. Two views are fed into a parameter-shared Siamese network to extract the high-dimensional embeddings and estimate the information entropy of the entire graph. For the learning process, we propose to optimize the network using two objectives, simultaneously. Concretely, the input of the contrastive loss consists of positive and negative pairs. Our selection strategy of pairs is different from previous works and we present a novel strategy to enhance the representation ability by selecting nodes based on cross-view similarities. We enrich the diversity of the positive and negative pairs by selecting highly similar samples and totally different data with the guidance of cross-view similarity scores, respectively. We also introduce a cross-view consistency constraint on the representations generated from the different views. We conduct experiments on seven graph benchmarks, and the proposed approach achieves competitive performance compared to the current state-of-the-art methods. The source code is available at https://github.com/kunzhan/M-ILBO.
Yixuan Ma, Peng Zhang 0057, Kun Zhan
ACM Multimedia3
2023 Improving Semi-Supervised Semantic Segmentation with Dual-Level Siamese Structure Network
abstract
Semi-supervised semantic segmentation (SSS) is an important task that utilizes both labeled and unlabeled data to reduce expenses on labeling training examples. However, the effectiveness of SSS algorithms is limited by the difficulty of fully exploiting the potential of unlabeled data. To address this, we propose a dual-level Siamese structure network (DSSN) for pixel-wise contrastive learning. By aligning positive pairs with a pixel-wise contrastive loss using strong augmented views in both low-level image space and high-level feature space, the proposed DSSN is designed to maximize the utilization of available unlabeled data. Additionally, we introduce a novel class-aware pseudo-label selection strategy for weak-to-strong supervision, which addresses the limitations of most existing methods that do not perform selection or apply a predefined threshold for all classes. Specifically, our strategy selects the top high-confidence prediction of the weak view for each class to generate pseudo labels that supervise the strong augmented views. This strategy is capable of taking into account the class imbalance and improving the performance of long-tailed classes. Our proposed method achieves state-of-the-art results on two datasets, PASCAL VOC 2012 and Cityscapes, outperforming other SSS algorithms by a significant margin. The source code is available at https://github.com/kunzhan/DSSN.
Zhibo Tian, Peng Zhang 0057, Kun Zhan
ACM Multimedia3
2023 Deep Learning for Approximate Nearest Neighbour Search: A Survey and Future Directions
abstract
Approximate nearest neighbour search (ANNS) in high-dimensional space is an essential and fundamental operation in many applications from many domains such as multimedia database, information retrieval and computer vision. With the rapidly growing volume of data and the dramatically increasing demands of users, traditional heuristic-based ANNS solutions have been facing great challenges in terms of both efficiency and accuracy. Inspired by the recent successes of deep learning in many fields, substantial efforts have been devoted to applying deep learning techniques to ANNS for learning to index and learning to search, resulting in numerous algorithms that achieve state-of-the-art performance compared with conventional methods. In this survey paper, we comprehensively review the different types of deep learning-based ANNS methods according to two learning paradigms:learning to indexandlearning to search. We provide a comprehensive overview and analysis of these methods in a systematic manner. Based on the overview, we point out thatend-to-end learningwill be a new and promising research direction for deep learning-based ANNS, i.e., applying deep learning techniques to jointly learn the indexing and searching together, such that the underlying knowledge learned from data can directly contribute to the final searching performance. Finally, we conduct experiments and provide general performance analyses for the representative deep learning-based ANNS algorithms.
Mingjie Li 0004, Yuan-Gen Wang, Peng Zhang 0057, Hanpin Wang, Lisheng Fan, Enxia Li, Wei Wang 0011
IEEE Trans. Knowl. Data Eng.3
2023 Improving Disentangled Representation Learning for Gait Recognition Using Group Supervision
abstract
In decades, gait has been gathering extensive interest for the advantage that it can be measured from a distance without physical contact. However, for image/video-based gait recognition, its performance can be remarkably influenced by exterior factors, such as viewing angles and clothing changes. Thus, in this paper, a group-supervised disentangled representation learning network is proposed for gait recognition to extract features invariant to these factors. First, sequences are explicitly disentangled into pose, gait, appearance, and view features through a generic encoder-decoder framework. To ensure the feature adaptability and independency, a disentanglement swap module is specifically adopted during our encode-decoder process through a series of swap operations based on the feature attributes. Following the feature disentanglement, a disentanglement aggregation module is also specially proposed for pose, gait, and appearance features to enhance their effectiveness. Finally, the enhanced three features are concatenated together for gait recognition. Relevant experiments certify that compared with other disentangled representation learning-based gait recognition methods, our proposed method enables to obtain a more excellent recognition result, despite fewer gait frames being utilized.
Lingxiang Yao, Worapan Kusakunniran, Peng Zhang 0057, Qiang Wu 0001, Jian Zhang 0002
IEEE Trans. Multim.3
2022 Dual-branch self-attention network for pedestrian attribute recognition
Zhang Zhang 0001, Da Li 0003, Peng Zhang 0057, Caifeng Shan
Pattern Recognit. Lett.4
2022 Alleviating Modality Bias Training for Infrared-Visible Person Re-Identification
abstract
The task of infrared-visible person re-identification (IV-reID) is to recognize people across two modalities (i.e., RGB and IR). Existing cutting-edge approaches normally use a pair of images that have the same IDs (i.e., ID-tied cross-modality image pairs) and input them into an ImageNet-trained ResNet50. The ResNet50 backbone model can learn shared features across modalities to tolerate modality discrepancies between RGB and IR. This work will unveil a Modality Bias Training (MBT) problem that is less discussed in IV-reID, which will demonstrate that MBT significantly compromises the performance of IV-reID. Due to MBT, IR information can be overwhelmed by RGB information during training when the ResNet50 model is pretrained based on a large amount of RGB images from ImageNet. Thus, the trained models are more inclined to RGB information. Accordingly, the cross-modality generalization ability of the model is also compromised. To tackle this issue, we present a Dual-level Learning Strategy (DLS) that 1) enforces the focus of the network on ID-exclusive (rather than ID-tied) labels of cross-modality image pairs to mitigate the problem of MBT and 2) introduces third modality data that contain both RGB and IR information to further prevent the information from the IR modality from being overwhelmed during training. Our third modality images are generated by a generative adversarial network. A dynamic ID-exclusive Smooth (dIDeS) label is proposed for the generated third modality data. In experiments, comprehensive experiments are carried out to demonstrate the success of DLS in tackling the MBT issue exposed in IV-reID.
Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002, Peng Zhang 0057, Zhaoxiang Zhang 0001
IEEE Trans. Multim.5
2021 Beyond modality alignment: Learning part-level representation for visible-infrared person re-identification
Peng Zhang 0057, Qiang Wu 0001, Xunxiang Yao, Jingsong Xu
Image Vis. Comput.1
2021 Weighted Adaptive Image Super-Resolution Scheme Based on Local Fractal Feature and Image Roughness
abstract
Image super-resolution aims to reconstruct a high-resolution image from the known low-resolution version. During this process, it should keep the degree of image roughness non-decreasing, which reflects various texture features and appearance. However, this point is not well addressed in the current work. This work argues that reducing roughness during image super-resolution is the key reason causing various problems such as artificial texture and/or edge blur. In this work, keeping the image roughness non-decreasing during super-resolution is being well investigated for the first time to our best knowledge. Image super-resolution is cast as an optimization problem to keep image roughness non-decreasing. In order to tackle this problem, the image super-resolution is approached based on the theory of fractal, where adaptive fractal interpolation function is proposed. In this way, the rational fractal interpolation model is adaptive to every local region. Thus, the roughness of every image region can be best maintained while super-resolution is carried out through fractal interpolation. In this work, the image roughness is reflected by the fractal dimension, which is a key element affecting the construction of fractal interpolation model. That is, the image roughness is measurable using fractal dimension. Mathematically, the overall image super-resolution process can be converted into a fractal interpolation optimization problem where the local fractal dimension is maintained. Although adaptive super-resolution on image segments may best maintain image roughness using the proposed method, it still generates unnecessary block artifacts. To tackle this problem, this work proposes a fine-grained pixel-wise fractal function. Our extensive experimental results demonstrate that the proposed method achieves encouraging performance with the state-of-the-art super-resolution algorithms.
Xunxiang Yao, Qiang Wu 0001, Peng Zhang 0057, Fangxun Bao
IEEE Trans. Multim.3
2021 Learning Spatial-Temporal Representations Over Walking Tracklet for Long-Term Person Re-Identification in the Wild
abstract
Long-term person re-identification (re-ID) aims to build identity correspondence of the Target Subject of Interest (TSI) exposed under surveillance cameras over a long time interval. Compared to the conventional short-term re-ID studied by most existing works, it suffers an additional problem: significant dressing change observed with time lapsing. Unfortunately, this variation in long-term person re-ID case contradicts the assumption of prior short-term re-ID approaches, and thus causes significant difficulties if conventional short-term re-ID methods are applied. To address the problem, this paper proposes to learn hybrid feature representation via a two-stream network named SpTSkM, including a spatial-temporal stream and a skeleton motion stream. The former performs directly on image sequences, which tends to learn identity-related spatial-temporal patterns such as body geometric structure and body movement. The latter operates on normalized 3D skeletons by adapting graph convolutional network, which tends to learn pure motion patterns from skeleton sequences. Both streams extract fine-grained level time-gap stable information that is robust to appearance changes in long-term re-ID and meanwhile maintains sufficient discriminability to differentiate different people. The final matching metric is obtained by mixing information of the two streams in a score-level fusion strategy. In addition, we collect a Cloth-Varying vIDeo re-ID (CVID-reID) dataset particularly for long-term re-ID. It contains video tracklets of celebrities posted on the Internet. These videos are snapshots under extremely different scenarios that include highly dynamic background, diverse camera views and abundant cloth variations on each TSI. These factors cause CVID-reID more complicated and closer to practice. Our experiments demonstrate the difficulty of long-term person re-ID and also validate the effectiveness of the proposed SpTSkM, showing the best performance.
Peng Zhang 0057, Jingsong Xu, Qiang Wu 0001, Yan Huang 0023, Xianye Ben
IEEE Trans. Multim.1
2020 Coupled Bilinear Discriminant Projection for Cross-View Gait Recognition
abstract
A problem that hinders good performance of general gait recognition systems is that the appearance features of gaits are more affected-prone by views than identities, especially when the walking direction of the probe gait is different from the register gait. This problem cannot be solved by traditional projection learning methods because these methods can learn only one projection matrix, and thus for the same subject, it cannot transfer cross-view gait features into similar ones. This paper presents an innovative method to overcome this problem by aligning gait energy images (GEIs) across views with the coupled bilinear discriminant projection (CBDP). Specifically, the CBDP generates the aligned gait matrix features for two views with two sets of bilinear transformation matrices, so that the original GEIs' spatial structure information can be preserved. By iteratively maximizing the ratio of inter-class distance metric to intra-class distance metric, the CBDP can learn the optimal matrix subspace where the GEIs across views are aligned in both horizontal and vertical coordinates. Therefore, the CBDP is also able to avoid the under-sample problem. We also theoretically prove that the upper and lower bounds of the objective function sequence of the CBDP are both monotonically increasing, so the convergence of the CBDP is demonstrated. In the terms of accuracy, the comparative experiments on the CASIA (B) and OU-ISIR gait databases show that our method is superior to the state-of-the-art cross-view gait recognition methods. More impressively, encouraging performance is obtained by our method even in matching a lateral-view gait with a frontal-view gait.
Xianye Ben, Chen Gong 0002, Peng Zhang 0057, Qiang Wu 0001, Weixiao Meng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 Beyond Scalar Neuron: Adopting Vector-Neuron Capsules for Long-Term Person Re-Identification
abstract
Current person re-identification (re-ID) works mainly focus on the short-term scenario where a person is less likely to change clothes. However, in the long-term re-ID scenario, a person has a great chance to change clothes. A sophisticated re-ID system should take such changes into account. To facilitate the study of long-term re-ID, this paper introduces a large-scale re-ID dataset called “Celeb-reID” to the community. Unlike previous datasets, the same person can change clothes in the proposed Celeb-reID dataset. Images of Celeb-reID are acquired from the Internet using street snap-shots of celebrities. There is a total of 1,052 IDs with 34,186 images making Celeb-reID being the largest long-term re-ID dataset so far. To tackle the challenge of cloth changes, we propose to use vector-neuron (VN) capsules instead of the traditional scalar neurons (SN) to design our network. Compared with SN, one extra-dimensional information in VN can perceive cloth changes of the same person. We introduce a well-designed ReIDCaps network and integrate capsules to deal with the person re-ID task. Soft Embedding Attention (SEA) and Feature Sparse Representation (FSR) mechanisms are adopted in our network for performance boosting. Experiments are conducted on the proposed long-term re-ID dataset and two common short-term re-ID datasets. Comprehensive analyses are given to demonstrate the challenge exposed in our datasets. Experimental results show that our ReIDCaps can outperform existing state-of-the-art methods by a large margin in the long-term scenario.The new dataset and code will be released to facilitate future researches.
Yan Huang 0023, Jingsong Xu, Qiang Wu 0001, Yi Zhong 0002, Peng Zhang 0057, Zhaoxiang Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2020 Top-Push Constrained Modality-Adaptive Dictionary Learning for Cross-Modality Person Re-Identification
abstract
Person re-identification aims to match person captured by multiple non-overlapping cameras that mainly mean standard RGB cameras. In contemporary surveillance, cameras of different modalities such as infrared cameras and depth cameras are introduced because of their unique advantages in poor illumination scenarios. However, re-identifying the persons across such cameras of different modalities is extremely difficult and, unfortunately, seldom discussed. It is mainly caused by extremely different appearances of the person shown under such different camera modalities. In this paper, we tackle this challenging cross-modality people re-identification through a top-push constrained modality-adaptive dictionary learning. The proposed model asymmetrically projects the heterogeneous features from dissimilar modalities onto a common space. In this way, the modality-specific bias is mitigated. Thus, the heterogeneous data can be simultaneously enforced by a shared dictionary in a canonical space. Moreover, a top-push ranking graph regularization is embedded in the proposed model to improve the discriminability, which efficiently further boosts the matching accuracy. In order to implement the proposed model, an iterative process is developed in this paper to optimize these two processes jointly. Extensive experiments on the benchmark SYSU-MM01 and BIWI RGBD-ID person re-identification datasets show promising results which outperform state-of-the-art methods.
Peng Zhang 0057, Jingsong Xu, Qiang Wu 0001, Yan Huang 0023, Jian Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2019 VT-GAN: View Transformation GAN for Gait Recognition Across Views
abstract
Recognizing gaits without human cooperation is of importance in surveillance and forensics because of the benefits that gait is unique and collected remotely. However, change of camera view angle severely degrades the performance of gait recognition. To address the problem, previous methods usually learn mappings for each pair of views which incurs abundant independently built models. In this paper, we proposed a View Transformation Generative Adversarial Networks (VT-GAN) to achieve view transformation of gaits across two arbitrary views using only one uniform model. In specific, we generated gaits in target view conditioned on input images from any views and the corresponding target view indicator. In addition to the classical discriminator in GAN which makes the generated images look realistic, a view classifier is imposed. This controls the consistency of generated images and conditioned target view indicator and ensures to generate gaits in the specified target view. On the other hand, retaining identity information while performing view transformation is another challenge. To solve the issue, an identity distilling module with triplet loss is integrated, which constrains the generated images inheriting identity information from inputs and yields discriminative feature embeddings. The proposed VT-GAN generates visually promising gaits and achieves promising performances for cross-view gait recognition, which exhibits great effectiveness of the proposed VT-GAN.
Peng Zhang 0057, Qiang Wu 0001, Jingsong Xu
IJCNN1
2019 VN-GAN: Identity-preserved Variation Normalizing GAN for Gait Recognition
abstract
Gait is recognized as a unique biometric characteristic to identify a walking person remotely across surveillance networks. However, the performance of gait recognition severely suffers challenges from view angle diversity. To address the problem, an identity-preserved Variation Normalizing Generative Adversarial Network (VN-GAN) is proposed for learning purely identity-related representations. It adopts a coarse-to-fine manner which firstly generates initial coarse images by normalizing view to an identical one and then refines the coarse images by injecting identity-related information. In specific, Siamese structure with discriminators for both camera view angles and human identities is utilized to achieve variation normalization and identity preservation of two stages, respectively. In addition to discriminators, reconstruction loss and identity-preserving loss are integrated, which forces the generated images to be the same in view and to be discriminative in identity. This ensures to generate identity-related images in an identical view of good visual effect for gait recognition. Extensive experiments on benchmark datasets demonstrate that the proposed VN-GAN can generate visually interpretable results and achieve promising performance for gait recognition.
Peng Zhang 0057, Qiang Wu 0001, Jingsong Xu
IJCNN1
2019 Adaptive rational fractal interpolation function for image super-resolution via local fractal analysis
Xunxiang Yao, Qiang Wu 0001, Peng Zhang 0057, Fangxun Bao
Image Vis. Comput.3
2019 A general tensor representation framework for cross-view gait recognition
Xianye Ben, Peng Zhang 0057, Zhihui Lai 0001, Xinliang Zhai, Weixiao Meng 0001
Pattern Recognit.2
2019 Coupled Patch Alignment for Matching Cross-View Gaits
abstract
Gait recognition has attracted growing attention in recent years as the gait of humans has a strong discriminative ability even under low resolution at a distance. Unfortunately, the performance of gait recognition can be largely affected by view change. To address this problem, we propose a Coupled Patch Alignment (CPA) algorithm that effectively matches a pair of gaits across different views. To realize CPA, we first build a certain amount of patches, and each of them is made up of a sample as well as its intra-class and inter-class nearest-neighbors. Then we design an objective function for each patch to balance the cross-view intra-class compactness and the cross-view inter-class separability. Finally, all the local independent patches are combined to render a unified objective function. Theoretically, we show that the proposed CPA has a close relationship with Canonical Correlation Analysis (CCA). Algorithmically, we extend CPA to "Multi-dimensional Patch Alignment" (MPA) that can handle an arbitrary number of views. Comprehensive experiments on CASIA(B), USF and OU-ISIR gait databases firmly demonstrate the effectiveness of our methods over other existing popular methods in terms of cross-view gait recognition.
Xianye Ben, Chen Gong 0002, Peng Zhang 0057, Xitong Jia, Qiang Wu 0001, Weixiao Meng 0001
IEEE Trans. Image Process.3
2018 Long-Term Person Re-identification Using True Motion from Videos
abstract
Most person re-identification approaches and benchmarks assume that pedestrians go across the surveillance network without significant appearance changes in a brief period, which explicitly restricts person re-identification to a short-term event and incurs inter-sample similarity measurement by appearance matching. However, pedestrians are likely to reappear in the surveillance network after a long-time interval (long-term) and change their wearing in many real-world scenarios. These scenarios inevitably cause appearances between subjects more ambiguous and indistinguishable. In this paper we consider these scenarios and propose a unified feature representation based on true motion cues from videos named FIne moTion encoDing (FITD). Our hypothesis is that people keep constant motion patterns under non-distraction walking condition. Therefore, the motion characteristics are more reliable than static appearance feature to describe a walking person. Particularly, we extract motion patterns hierarchically by encoding trajectory-aligned descriptors with Fisher vectors in a spatial-aligned pyramid. To verify benefits of the proposed FITD, we collect a new dataset typically for the long-term situations. Extensive experiments demonstrate the merits of our FITD especially for the long-term scenarios.
Peng Zhang 0057, Qiang Wu 0001, Jingsong Xu, Jian Zhang 0002
WACV1
2016 On the distance metric learning between cross-domain gaits
Xianye Ben, Peng Zhang 0057, Weixiao Meng 0001, Wenhe Liu
Neurocomputing2
2016 Gait recognition and micro-expression recognition based on maximum margin projection with tensor representation
Xianye Ben, Peng Zhang 0057, Guodong Ge
Neural Comput. Appl.2