Shanshan Zhang 0001

dblp:34/3535-1 · also Shan-Shan Zhang 0001 · DBLP profile ↗
← Back
74ranked-venue papers
11as first author
45since 2021 · last 2026
0000-0003-4013-6300ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 52 · 8 first-author · 29 since 2021Artificial intelligence and machine learning · 43 · 8 first-author · 26 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 One-shot novel view and pose human image synthesis via 3D prior guided diffusion model
Shenjian Gong, Kangkan Wang, Shanshan Zhang 0001, Jian Yang 0003
Pattern Recognit.3
2026 Unleashing the potential of traditional person re-ID methods to clothes changed scenario via curriculum learning
Shanshan Zhang 0001, Jian Yang 0003
Pattern Recognit.2
2026 Shape-prior-based point cloud completion for single-stage fully sparse 3D object detection
Mingqian Ji, Jian Yang 0003, Shanshan Zhang 0001
Pattern Recognit. Lett.4
2026 Siamese Feature Decoupling and Adaptive Prototype Alignment for Clothes Changed Person Re-Identification
abstract
The core to tackle clothes changed person re-identification is to extract clothes-irrelevant features. Previous methods typically decoupled the entire image features into clothes-relevant and ID-relevant ones via an auxiliary task of clothes classification, yet in this way, it is not guaranteed that the ID-relevant features fully get rid of clothes-relevant information, as the entire image feature also includes other information, such as posture. To enhance the purity of ID-relevant features, this letter proposes a simple yet effective siamese feature decoupling and adaptive prototype alignment framework (Decoupling and Alignment,D&A), which involves a reference branch and an alignment branch. Specifically, the reference branch first uses human parsing to remove clothes information at the image level to ensure that the ID prototype subsequently extracted by the image encoder is purely ID-relevant without clothes; the alignment branch excludes clothes information from the ID-relevant features by dual-token decoupling, and then aligns the ID-relevant features with the ID prototypes in the reference branch to enhance the purity of the identity information obtained after decoupling. Furthermore, we introduce an adaptive update strategy to update the ID prototypes to prevent the ID prototypes and the decoupled features from shifting in feature space. Extensive experiments on widely used benchmarks (PRCC, LTCC, and VC-Clothes) validate the effectiveness of our method, which achieves new state-of-the-art results across most metrics.
Jian Yang 0003, Shanshan Zhang 0001
IEEE Signal Process. Lett.3
2026 Toward Robust Proactive Deepfake Detection via Orthogonal Moment Watermarking
Chunpeng Wang 0001, Xianqiu Xu, Shanshan Zhang 0001, Bin Ma 0003, Qi Li 0029, Yunan Liu 0001
IEEE Trans. Dependable Secur. Comput.3
2026 Can Watermarks Be Removed Like Noise? A Watermarking Attack Network Using Residual Diffusion Model
abstract
Digital image watermarking is a critical technology for image copyright protection. The concurrent evolution of watermarking attacks and defenses has spurred rapid advancements in the field. However, watermarking attack methods have lagged behind, often facing two primary challenges: limited watermark removal ability and quality degradation of the attacked image. In this paper, we introduce a Watermarking Attack method based on the Residual Diffusion Model, termed WARDM. Our WARDM treats watermark information as noise and leverages the powerful image reconstruction capabilities of the diffusion model to effectively remove the watermark. Specifically, we construct a Markov chain based on the residuals between the host and watermarked images, and employ reverse propagation to reconstruct the original host image. To optimally balance watermark removal ability and image quality, we incorporate a noise schedule into WARDM that controls both the velocity and intensity of noise at each stage of the Markov chain. Extensive experiments demonstrate the superior performance of WARDM in both watermark removal capability and visual quality preservation, achieving an improvement of 5.39% in PSNR over state-of-the-art methods. Moreover, our method demonstrates strong generalization, effectively executing attacks across a variety of watermarking techniques.
Chunpeng Wang 0001, Shanshan Zhang 0001, Yunan Liu 0001, Yuli Wang, Qi Li 0029
IEEE Trans. Dependable Secur. Comput.3
2026 Language-Driven Visual Data Generation for Zero-Shot HOI Detection
abstract
Zero-shot human-object interaction (HOI) detection aims to recognize both seen and unseen interaction categories while detecting humans and objects in an image. However, due to the absence of training samples for unseen categories, existing methods often overfit on seen HOIs and struggle to generalize to unseen ones. To address this issue, we introduce a novel Language-Driven Visual Data Generation (LD-VDG) approach that generates pseudo visual features from textual semantics of unseen HOIs. This provides an innovative solution enabling generalization to unseen HOIs without relying on visual samples. Specifically, we first design a text-to-vision (T-V) adapter to align HOI text and visual features, trained on seen HOIs with paired image-text data. For unseen HOIs, we guide the large language model to produce multiple fine-grained textual descriptions based on HOI labels, which are then encoded by the vision-language model and transformed into pseudo visual features via the T-V adapter. After that, these pseudo features together with real features from seen HOIs are jointly used to train a transformer-based HOI detector. In this way, our method enables effective recognition of unseen HOIs by leveraging language-driven visual representations. Experimental results on standard datasets demonstrate that the proposed LD-VDG outperforms previous methods. In particular, it achieves superior performance on unseen categories under various zero-shot settings.
Pei Geng, Shanshan Zhang 0001, Jian Yang 0003
IEEE Trans. Image Process.2
2026 Focus on Finding Deepfakes: A Robust Proactive Detection Method Based on Orthogonal Moment Watermarking
abstract
Deepfake detection remains a challenging research topic, especially when the quality of forged images degrades, leading to unreliable detection results. In this paper, we propose a watermarking-based proactive method for robust proactive deepfake detection. First, we embed a watermark into the Fractional-order Quaternion Exponent Moments (FrQEMs) space of the host face image, achieving a balance between imperceptibility and robustness of the watermarking algorithm. Then, we introduce the Frequency Mamba (FreMamba) block to enhance feature extraction by leveraging correlations between frequency-domain subbands, thereby enabling the extraction of more discriminative feature representations. Finally, at the detection stage, we construct a dual-branch framework comprising a watermark extractor and a forgery discriminator. Through knowledge distillation, the watermark extractor guides the forgery discriminator to perceive forgery traces. Specifically, the integrity of the extracted watermark is compromised only when the host image is subjected to a deepfake attack, while conventional attacks do not affect the integrity. Experimental results on benchmark datasets demonstrate that the proposed method achieves superior deepfake detection accuracy. In particular, when images are subjected to conventional attacks, our method surpasses state-of-the-art approaches by more than 5.3% in terms of ACC.
Chunpeng Wang 0001, Shanshan Zhang 0001, Jie Gui, Qi Li 0029, Yunan Liu 0001
IEEE Trans. Image Process.3
2026 DepthFusion: Depth-Aware Hybrid Feature Fusion for LiDAR-Camera 3D Object Detection
abstract
State-of-the-art LiDAR-camera 3D object detectors usually focus on feature fusion, but often neglect the role of depth in modulating the contribution of different modalities. In this work, we observe that the importance of each modality varies with depth through statistical analysis and visualization. Motivated by this finding, we propose a Depth-Aware Hybrid Feature Fusion (DepthFusion) strategy that explicitly encodes metric depth into the BEV feature space, guiding the fusion weights of point cloud and RGB image features at both global and local levels. Specifically, the Depth-GFusion module adaptively adjusts the weights of image BEV features in multi-modal global features via explicit depth encoding. Furthermore, to recover information lost when projecting raw features to BEV space, the Depth-LFusion module dynamically adjusts the weights of voxel and multi-view image features in local features based on depth. Extensive experiments on the nuScenes and KITTI datasets demonstrate that DepthFusion outperforms previous state-of-the-art methods. Moreover, DepthFusion shows superior robustness to various corruptions on the nuScenes-C dataset.
Mingqian Ji, Jian Yang 0003, Shanshan Zhang 0001
IEEE Trans. Multim.3
2025 HORP: Human-Object Relation Priors Guided HOI Detection
abstract
Human-Object Interaction (HOI) detection aims to predict thetriplets, where the core challenge lies in recognizing the interaction of each human-object pair. Despite recent progress thanks to more advanced model architectures, HOI performance remains unsatisfactory. In this work, we first perform some failure analysis and find that the accuracy of the no-interaction category is extremely low, largely hindering the improvement of overall performance. We further look into the error types and find the mis-classification between no-interaction and with-interaction ones can be handled by human-object relation priors. Specifically, to better distinguish no-interaction from direct interactions, we propose 3D location prior, which indicates the distance between human and object; as of no-interaction vs. indirect interactions, we propose gaze area prior, which denotes whether human can see the object or not. The above two types of human-object relation priors are represented by text and are combined with the original visual features, generating multi-modal cues for interaction recognition. Experimental results on the HICO-DET and V-COCO datasets demonstrate that our proposed human-object relation priors are effective and our method HORP surpasses previous methods under various settings and scenarios. In particular, the usage of our priors significantly enhances the model’s recognition ability for the no-interaction category. Code is available at https://github.com/namegp/HORP.
Pei Geng, Jian Yang 0003, Shanshan Zhang 0001
CVPR3
2025 OcRFDet: Object-Centric Radiance Fields for Multi-View 3D Object Detection in Autonomous Driving
abstract
Current multi-view 3D object detection methods typically transfer 2D features into 3D space using depth estimation or 3D position encoder, but in a fully data-driven and implicit manner, which limits the detection performance. Inspired by the success of radiance fields on 3D reconstruction, we assume they can be used to enhance the detector's ability of 3D geometry estimation. However, we observe a decline in detection performance, when we directly use them for 3D rendering as an auxiliary task. From our analysis, we find the performance drop is caused by the strong responses on the background when rendering the whole scene. To address this problem, we propose object-centric radiance fields, focusing on modeling foreground objects while discarding background noises. Specifically, we employ Object-centric Radiance Fields (OcRF) to enhance 3D voxel features via an auxiliary task of rendering foreground objects. We further use opacity - the side-product of rendering- to enhance the 2D foreground BEV features via Height-aware Opacity-based Attention (HOA), where attention maps at different height levels are generated separately via multiple networks in parallel. Extensive experiments on the nuScenes validation and test datasets demonstrate that our OcRFDet achieves superior performance, outperforming previous state-of-the-art methods with 57.2$\%$ mAP and 64.8$\%$ NDS on the nuScenes test benchmark. Code will be available at https://github.com/Mingqj/OcRFDet.
Mingqian Ji, Shanshan Zhang 0001, Jian Yang 0003
ICCV2
2025 Enhancing Pseudo-Boxes via Data-Level LiDAR-Camera Fusion for Unsupervised 3D Object Detection
abstract
Existing LiDAR-based 3D object detectors typically rely on manually annotated labels for training to achieve good performance. However, obtaining high-quality 3D labels is time-consuming and labor-intensive. To address this issue, recent works explore unsupervised 3D object detection by introducing RGB images as an auxiliary modal to assist pseudo-box generation. However, these methods simply integrate pseudo-boxes generated by LiDAR point clouds and RGB images. Yet, such a label-level fusion strategy brings limited improvements to the quality of pseudo-boxes, as it overlooks the complementary nature in terms of LiDAR and RGB image data. To overcome the above limitations, we propose a novel data-level fusion framework that integrates RGB images and LiDAR data at an early stage. Specifically, we utilize vision foundation models for instance segmentation and depth estimation on images and introduce a bi-directional fusion method, where real points acquire category labels from the 2D space, while 2D pixels are projected onto 3D to enhance real point density. To mitigate noise from depth and segmentation estimations, we propose a local and global filtering method, which applies local radius filtering to suppress depth estimation errors and global statistical filtering to remove segmentation-induced outliers. Furthermore, we propose a data-level fusion based dynamic self-evolution strategy, which iteratively refines pseudo-boxes under a dense representation, significantly improving localization accuracy. Extensive experiments on the nuScenes dataset demonstrate that the detector trained by our method significantly outperforms that trained by previous state-of-the-art methods with 28.4% mAP on the nuScenes validation benchmark.
Mingqian Ji, Jian Yang 0003, Shanshan Zhang 0001
ACM Multimedia3
2025 Hierarchical Meta Alignment for cross-domain object detection
Yang Li 0190, Shanshan Zhang 0001, Yunan Liu 0001, Jian Yang 0003
Eng. Appl. Artif. Intell.2
2025 Spatially adaptive pyramid feature fusion for scale-aware crowd counting
Shenjian Gong, Zhaoliang Yao, Wangmeng Zuo, Jian Yang 0003, Pongchi Yuen, Shanshan Zhang 0001
Pattern Recognit.6
2025 Adaptive and Background-Aware Match for Class-Agnostic Counting
abstract
Class-Agnostic Counting (CAC) aims to count object instances in an image by simply specifying a few exemplar boxes of interest. The key challenge for CAC is how to tailor a desirable interaction between exemplar and query features. Previous CAC methods implement such interaction by solely leveraging standard global feature convolution. We find this interaction leads to under-match caused by intra-class diversity and over-match on background, which harms counting performance severely. In this work, we propose a novel feature interaction method called Adaptive and Background-aware Match (ABM) against high intra-class diversity and noisy background. Concretely, given exemplar and query features, we improve the original high-dimensional coupled spaces match to Adaptive Orthogonal subspaces Match (AOM), avoiding under-match caused by intra-class diversity. Moreover, Background-Specific Match (BSM) employs interaction between the learnable background prototype and query features to provide global background priors, making the match be aware of background regions. Additionally, we find the high scale variance among different query images leads to bad counting performance for extremely small scale objects. Object-Scale Unify (OSU) is proposed to take the size of the exemplars as the scale prior and resize query images so that all objects are at a uniform average scale. Extensive experiments on FSC-147 show that our method performs better. We also conduct extensive ablation studies to demonstrate the effectiveness of each component of our proposed method.
Shenjian Gong, Jian Yang 0003, Shanshan Zhang 0001
IEEE Signal Process. Lett.3
2024 Divide and Conquer: Hybrid Pre-training for Person Search
abstract
Large-scale pre-training has proven to be an effective method for improving performance across different tasks. Current person search methods use ImageNet pre-trained models for feature extraction, yet it is not an optimal solution due to the gap between the pre-training task and person search task (as a downstream task). Therefore, in this paper, we focus on pre-training for person search, which involves detecting and re-identifying individuals simultaneously. Although labeled data for person search is scarce, datasets for two sub-tasks person detection and re-identification are relatively abundant. To this end, we propose a hybrid pre-training framework specifically designed for person search using sub-task data only. It consists of a hybrid learning paradigm that handles data with different kinds of supervisions, and an intra-task alignment module that alleviates domain discrepancy under limited resources. To the best of our knowledge, this is the first work that investigates how to support full-task pre-training using sub-task data. Extensive experiments demonstrate that our pre-trained model can achieve significant improvements across diverse protocols, such as person search method, fine-tuning data, pre-training data and model backbone. For example, our model improves ResNet50 based NAE by 10.3% relative improvement w.r.t. mAP. Our code and pre-trained models are released for plug-and-play usage to the person search community (https://github.com/personsearch/PretrainPS).
Yanling Tian, Yunan Liu 0001, Jian Yang 0003, Shanshan Zhang 0001
AAAI5
2024 Distilling Knowledge from Large-Scale Image Models for Object Detection
Wenhai Wang, Xiang Li 0041, Jian Yang 0003, Jifeng Dai, Yu Qiao 0001, Shanshan Zhang 0001
ECCV (84)8
2024 Adaptive Pedestrian Trajectory Prediction via Target-Directed Angle Augmentation
abstract
Pedestrian trajectory prediction is an important task for many applications such as autonomous driving and surveillance systems. Yet the prediction performance drops dramatically when applying a model trained on the source domain to a new target domain. Therefore, it is of great importance to adapt a predictor to a new domain. Previous works mainly focus on feature-level alignment to solve this problem. In contrast, we solve it from a new perspective of instance-level alignment. Specifically, we first point out one key factor of the domain gaps, i.e., trajectory angles, and then augment the source training data by target-directed orientation augmentation so that its distribution matches with that of the target data. In this way, the trajectory predictor trained on the aligned source data performs better on the target domain. Experiments on standard baselines show that our method improves the state of the art by a large margin. The source code is available at https://github.com/NeoKH/PTP-DA.
Jie Xu 0021, Shenjian Gong, Jian Yang 0003, Shanshan Zhang 0001
ICASSP5
2024 Reliable hybrid knowledge distillation for multi-source domain adaptive object detection
Yang Li 0190, Shanshan Zhang 0001, Yunan Liu 0001, Jian Yang 0003
Knowl. Based Syst.2
2024 From Simple to Complex Scenes: Learning Robust Feature Representations for Accurate Human Parsing
abstract
Human parsing has attracted considerable research interest due to its broad potential applications in the computer vision community. In this paper, we explore several useful properties, including high-resolution representation, auxiliary guidance, and model robustness, which collectively contribute to a novel method for accurate human parsing in both simple and complex scenes. Starting from simple scenes: we propose the boundary-aware hybrid resolution network (BHRN), an advanced human parsing network. BHRN utilizes deconvolutional layers and multi-scale supervision to generate rich high-resolution representations. Additionally, it includes an edge perceiving branch designed to enhance the fineness of part boundaries. Building on BHRN, we construct a dual-task mutual learning (DTML) framework. It not only provides implicit guidance to assist the parser by incorporating boundary features, but also explicitly maintains the high-order consistency between the parsing prediction and the ground truth. Toward complex scenes: we develop a domain transform method to enhance the model robustness. By transforming the input space from the spatial domain to the polar harmonic Fourier moment domain, the mapping relationship to the output semantic space is highly stable. This transformation yields robust representations for both clean and corrupted data. When evaluated on standard benchmark datasets, our method achieves superior performance compared to state-of-the-art human parsing methods. Furthermore, our domain transform strategy significantly improves the robustness of DTML dramatically in most complex scenes.
Yunan Liu 0001, Chunpeng Wang 0001, Mingyu Lu, Jian Yang 0003, Jie Gui, Shanshan Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Loose to compact feature alignment for domain adaptive object detection
Yang Li 0190, Shanshan Zhang 0001, Yunan Liu 0001, Jian Yang 0003
Pattern Recognit. Lett.2
2024 Adaptive Teaching for Cross-Domain Crowd Counting
abstract
The main challenge of Unsupervised Domain Adaptation (UDA) crowd counting is the large domain gap between a synthetic domain with annotations (source) and a real-world domain of interest without annotations (target). Previous mainstream UDA crowd counting methods either employ feature alignment or a semi-supervised learning paradigm via pseudo-labels. We for the first time combine both of their advantages and propose an Adversarial Mean Teacher (AMT) framework. On the one hand, we optimize the student model with domain adversarial learning. On the other hand, we feed perturbed target images to the teacher model to generate pseudo-labels. Furthermore, to improve the quality of the pseudo-labels, we propose an Adaptive Teaching (AT) module, consisting of pseudo-label refinement and credible pseudo-label selection. Concretely, we first generate two candidate pseudo-labels from the prediction of the teacher model and obtain a refined pseudo-label by mixing them at the pixel-level. Moreover, we introduce an auxiliary task of foreground-background classification to assist credible region selection and only activate supervision signals on those regions. Extensive experiments on four real-world crowd counting benchmarks demonstrate the effectiveness of our method namely Cross-Domain Adaptive Teacher (CDAT).
Shenjian Gong, Jian Yang 0003, Shanshan Zhang 0001
IEEE Trans. Multim.3
2023 Adaptive Decoupled Pose Knowledge Distillation
abstract
Existing state-of-the-art human pose estimation approaches require heavy computational resources for accurate prediction. One promising technique to obtain an accurate yet lightweight pose estimator is Knowledge Distillation (KD), which distills the pose knowledge from a powerful teacher model to a lightweight student model. However, existing human pose KD methods focus more on designing paired student and teacher network architectures, yet ignore the mechanism of pose knowledge distillation. In this work, we reformulate the human pose KD to a coarse to fine process and decouple the classical KD loss into three terms: Binary Keypoint vs. Non-Keypoint Distillation (BiKD), Keypoint Area Distillation (KAD) and Non-keypoint Area Distillation (NAD). Observing the decoupled formulation, we point out an important limitation of the classical pose KD, i.e. the bias between different loss terms limits the performance gain of the student network. To address the biased knowledge distillation problem, we present a novel KD method named Adaptive Decoupled Pose knowledge Distillation (ADPD), enabling BiKD, KAD and NAD to play their roles more effectively and flexibly. Extensive experiments on two standard human pose datasets, MPII and MS COCO, demonstrate that our proposed method outperforms previous KD methods and is generalizable to different teacher-student pairs. The code will be available at https://github.com/SuperJay1996/ADPD.
Jie Xu 0021, Shanshan Zhang 0001, Jian Yang 0003
ACM Multimedia2
2023 OKGR: Occluded Keypoint Generation and Refinement for 3D Object Detection
Mingqian Ji, Jian Yang 0003, Shanshan Zhang 0001
PRCV (12)3
2023 Human Co-Parsing Guided Alignment for Occluded Person Re-Identification
abstract
Occluded person re-identification (ReID) is a challenging task due to more background noises and incomplete foreground information. Although existing human parsing-based ReID methods can tackle this problem with semantic alignment at the finest pixel level, their performance is heavily affected by the human parsing model. Most supervised methods propose to train an extra human parsing model aside from the ReID model with cross-domain human parts annotation, suffering from expensive annotation cost and domain gap; Unsupervised methods integrate a feature clustering-based human parsing process into the ReID model, but lacking supervision signals brings less satisfactory segmentation results. In this paper, we argue that the pre-existing information in the ReID training dataset can be directly used as supervision signals to train the human parsing model without any extra annotation. By integrating a weakly supervised human co-parsing network into the ReID network, we propose a novel framework that exploits shared information across different images of the same pedestrian, called the Human Co-parsing Guided Alignment (HCGA) framework. Specifically, the human co-parsing network is weakly supervised by three consistency criteria, namely global semantics, local space, and background. By feeding the semantic information and deep features from the person ReID network into the guided alignment module, features of the foreground and human parts can then be obtained for effective occluded person ReID. Experiment results on two occluded and two holistic datasets demonstrate the superiority of our method. Especially on Occluded-DukeMTMC, it achieves 70.2% Rank-1 accuracy and 57.5% mAP.
Shuguang Dou, Cairong Zhao, Xinyang Jiang, Shanshan Zhang 0001, Wei-Shi Zheng 0001, Wangmeng Zuo
IEEE Trans. Image Process.4
2022 Keypoint Message Passing for Video-Based Person Re-identification
abstract
Video-based person re-identification~(re-ID) is an important technique in visual surveillance systems which aims to match video snippets of people captured by different cameras. Existing methods are mostly based on convolutional neural networks~(CNNs), whose building blocks either process local neighbor pixels at a time, or, when 3D convolutions are used to model temporal information, suffer from the misalignment problem caused by person movement. In this paper, we propose to overcome the limitations of normal convolutions with a human-oriented graph method. Specifically, features located at person joint keypoints are extracted and connected as a spatial-temporal graph. These keypoint features are then updated by message passing from their connected nodes with a graph convolutional network~(GCN). During training, the GCN can be attached to any CNN-based person re-ID model to assist representation learning on feature maps, whilst it can be dropped after training for better inference speed. Our method brings significant improvements over the CNN-based baseline model on the MARS dataset with generated person keypoints and a newly annotated dataset: PoseTrackReID. It also defines a new state-of-the-art method in terms of top-1 accuracy and mean average precision in comparison to prior works.
Andreas Doering, Shanshan Zhang 0001, Jian Yang 0003, Juergen Gall, Bernt Schiele
AAAI3
2022 Knowledge Distillation for Object Detection via Rank Mimicking and Prediction-Guided Feature Imitation
abstract
Knowledge Distillation (KD) is a widely-used technology to inherit information from cumbersome teacher models to compact student models, consequently realizing model compression and acceleration. Compared with image classification, object detection is a more complex task, and designing specific KD methods for object detection is non-trivial. In this work, we elaborately study the behaviour difference between the teacher and student detection models, and obtain two intriguing observations: First, the teacher and student rank their detected candidate boxes quite differently, which results in their precision discrepancy. Second, there is a considerable gap between the feature response differences and prediction differences between teacher and student, indicating that equally imitating all the feature maps of the teacher is the sub-optimal choice for improving the student's accuracy. Based on the two observations, we propose Rank Mimicking (RM) and Prediction-guided Feature Imitation (PFI) for distilling one-stage detectors, respectively. RM takes the rank of candidate boxes from teachers as a new form of knowledge to distill, which consistently outperforms the traditional soft label distillation. PFI attempts to correlate feature differences with prediction differences, making feature imitation directly help to improve the student's accuracy. On MS COCO and PASCAL VOC benchmarks, extensive experiments are conducted on various detectors with different backbones to validate the effectiveness of our method. Specifically, RetinaNet with ResNet50 achieves 40.4% mAP on MS COCO, which is 3.5% higher than its baseline, and also outperforms previous KD methods.
Xiang Li 0041, Shanshan Zhang 0001, Yichao Wu, Ding Liang
AAAI4
2022 PoseTrack21: A Dataset for Person Search, Multi-Object Tracking and Multi-Person Pose Tracking
abstract
Current research evaluates person search, multi-object tracking and multi-person pose estimation as separate tasks and on different datasets although these tasks are very akin to each other and comprise similar sub-tasks, e.g. person detection or appearance-based association of detected persons. Consequently, approaches on these respective tasks are eligible to complement each other. Therefore, we introduce PoseTrack21, a large-scale dataset for person search, multi-object tracking and multi-person pose tracking in real-world scenarios with a high diversity of poses. The dataset provides rich annotations like human pose annotations including annotations of joint occlusions, bounding box annotations even for small persons, and person-ids within and across video sequences. The dataset allows to evaluate multi-object tracking and multi-person pose tracking jointly with person re-identification or exploit structural knowledge of human poses to improve person search and tracking, particularly in the context of severe occlusions. With PoseTrack21, we want to encourage researchers to work on joint approaches that perform reasonably well on all three tasks.
Andreas Doering, Shanshan Zhang 0001, Bernt Schiele, Juergen Gall
CVPR3
2022 Bi-level Alignment for Cross-Domain Crowd Counting
abstract
Recently, crowd density estimation has received increasing attention. The main challenge for this task is to achieve high-quality manual annotations on a large amount of training data. To avoid reliance on such annotations, previous works apply unsupervised domain adaptation (UDA) techniques by transferring knowledge learned from easily accessible synthetic data to real-world datasets. However, current state-of-the-art methods either rely on external data for training an auxiliary task or apply an expensive coarse-to-fine estimation. In this work, we aim to develop a new adversarial learning based method, which is simple and efficient to apply. To reduce the domain gap between the synthetic and real data, we design a bi-level alignment framework (BLA) consisting of (1) task-driven data alignment and (2) fine-grained feature alignment. In contrast to previous domain augmentation methods, we introduce AutoML to search for an optimal transform on source, which well serves for the downstream task. On the other hand, we do fine-grained alignment for foreground and background separately to alleviate the alignment difficulty. We evaluate our approach on five real-world crowd counting benchmarks, where we outperform existing approaches by a large margin. Also, our approach is simple, easy to implement and efficient to apply. The code is publicly available at https://github.com/Yankeegsj/BLA.
Shenjian Gong, Shanshan Zhang 0001, Jian Yang 0003, Dengxin Dai, Bernt Schiele
CVPR2
2022 Class-Agnostic Object Counting Robust to Intraclass Diversity
Shenjian Gong, Shanshan Zhang 0001, Jian Yang 0003, Dengxin Dai, Bernt Schiele
ECCV (33)2
2022 PseCo: Pseudo Labeling and Consistency Training for Semi-Supervised Object Detection
Xiang Li 0041, Yichao Wu, Ding Liang, Shanshan Zhang 0001
ECCV (9)6
2022 Grouped Adaptive Loss Weighting for Person Search
abstract
Person search is an integrated task of multiple sub-tasks such as foreground/background classification, bounding box regression and person re-identification. Therefore, person search is a typical multi-task learning problem, especially when solved in an end-to-end manner. Recently, some works enhance person search features by exploiting various auxiliary information, e.g. person joint keypoints, body part position, attributes, etc., which brings in more tasks and further complexifies a person search model. The inconsistent convergence rate of each task could potentially harm the model optimization. A straightforward solution is to manually assign different weights to different tasks, compensating for the diverse convergence rates. However, given the special case of person search, i.e. with a large number of tasks, it is impractical to weight the tasks manually. To this end, we propose a Grouped Adaptive Loss Weighting (GALW) method which adjusts the weight of each task automatically and dynamically. Specifically, we group tasks according to their convergence rates. Tasks within the same group share the same learnable weight, which is dynamically assigned by considering the loss uncertainty. Experimental results on two typical benchmarks, CUHK-SYSU and PRW, demonstrate the effectiveness of our method.
Yanling Tian, Yunan Liu 0001, Shanshan Zhang 0001, Jian Yang 0003
ACM Multimedia4
2022 DTG-SSOD: Dense Teacher Guidance for Semi-Supervised Object Detection
abstract
The Mean-Teacher (MT) scheme is widely adopted in semi-supervised object detection (SSOD). In MT, sparse pseudo labels, offered by the final predictions of the teacher (e.g., after Non Maximum Suppression (NMS) post-processing), are adopted for the dense supervision for the student via hand-crafted label assignment. However, the "sparse-to-dense'' paradigm complicates the pipeline of SSOD, and simultaneously neglects the powerful direct, dense teacher supervision. In this paper, we attempt to directly leverage the dense guidance of teacher to supervise student training, i.e., the "dense-to-dense'' paradigm. Specifically, we propose the Inverse NMS Clustering (INC) and Rank Matching (RM) to instantiate the dense supervision, without the widely used, conventional sparse pseudo labels. INC leads the student to group candidate boxes into clusters in NMS as the teacher does, which is implemented by learning grouping information revealed in NMS procedure of the teacher. After obtaining the same grouping scheme as the teacher via INC, the student further imitates the rank distribution of the teacher over clustered candidates through Rank Matching. With the proposed INC and RM, we integrate Dense Teacher Guidance into Semi-Supervised Object Detection (termed "DTG-SSOD''), successfully abandoning sparse pseudo labels and enabling more informative learning on unlabeled data. On COCO benchmark, our DTG-SSOD achieves state-of-the-art performance under various labelling ratios. For example, under 10% labelling ratio, DTG-SSOD improves the supervised baseline from 26.9 to 35.9 mAP, outperforming the previous best method Soft Teacher by 1.9 points.
Xiang Li 0041, Yichao Wu, Ding Liang, Shanshan Zhang 0001
NeurIPS6
2021 Hierarchical Information Passing Based Noise-Tolerant Hybrid Learning for Semi-Supervised Human Parsing
abstract
Deep learning based human parsing methods usually require a large amount of training data to reach high performance. However, it is costly and time-consuming to obtain manually annotated high quality labels for a large scale dataset. To alleviate annotation efforts, we propose a new semi-supervised human parsing method for which we only need a small number of labels for training. First, we generate high quality pseudo labels on unlabeled images using a hierarchical information passing network (HIPN), which reasons human part segmentation in a coarse to fine manner. Furthermore, we develop a noise-tolerant hybrid learning method, which takes advantage of positive and negative learning to better handle noisy pseudo labels. When evaluated on standard human parsing benchmarks, our HIPN achieves a new state-of-the-art performance. Moreover, our noise-tolerant hybrid learning method further improves the performance and outperforms the state-of-the-art semi-supervised method (i.e. GRN) by 4.47 points w.r.t mIoU on the LIP dataset.
Yunan Liu 0001, Shanshan Zhang 0001, Jian Yang 0003, Pong C. Yuen
AAAI2
2021 Adaptive Multi-Domain Learning for Outdoor 3d Human Pose and Shape Estimation
abstract
It is an extremely challenging task to estimate 3D human pose and shape in outdoor scenes for which we can hardly obtain precise ground truth data for training. Previous methods usually use multiple datasets collected at different scenes to train their models, including those collected in laboratories with precise ground truth and those collected at outdoor scenes with estimated or even no ground truth. Since data from different scenes are included in training, it is necessary to handle the domain difference problem, which unfortunately has never been considered by previous works. In this paper, we first point out this problem and then address it via a novel cascade multi-domain learning module (CMDL), where multiple adapters are employed to extract more discriminative features for different domains. We show that our method with CMDL outperforms previous methods in outdoor scenes. In principle, the proposed CMDL module can be easily applied on top of any arbitrary 3D human pose and shape approach.
Zhaoyang Gui, Shanshan Zhang 0001, Kangkan Wang, Jian Yang 0003, Pong C. Yuen
ICASSP2
2021 Seeking Similarities over Differences: Similarity-based Domain Alignment for Adaptive Object Detection
abstract
In order to robustly deploy object detectors across a wide range of scenarios, they should be adaptable to shifts in the input distribution without the need to constantly an-notate new data. This has motivated research in Unsupervised Domain Adaptation (UDA) algorithms for detection. UDA methods learn to adapt from labeled source domains to unlabeled target domains, by inducing alignment between detector features from source and target domains. Yet, there is no consensus on what features to align and how to do the alignment. In our work, we propose a framework that generalizes the different components commonly used by UDA methods laying the ground for an in-depth analysis of the UDA design space. Specifically, we propose a novel UDA algorithm, ViSGA, a direct implementation of our framework, that leverages the best design choices and introduces a simple but effective method to aggregate features at instance-level based on visual similarity before inducing group alignment via adversarial training. We show that both similarity-based grouping and adversarial training allows our model to focus on coarsely aligning feature groups, without being forced to match all instances across loosely aligned domains. Finally, we examine the applicability of ViSGA to the setting where labeled data are gathered from different sources. Experiments show that not only our method outperforms previous single-source approaches on Sim2Real and Adverse Weather, but also generalizes well to the multi-source setting.
Farzaneh Rezaeianaran, Rakshith Shetty, Rahaf Aljundi, Daniel Olmeda Reino, Shanshan Zhang 0001, Bernt Schiele
ICCV5
2021 Tiny Person Pose Estimation via Image and Feature Super Resolution
Jie Xu 0021, Yunan Liu 0001, Lin Zhao 0003, Shanshan Zhang 0001, Jian Yang 0003
ICIG (3)4
2021 Improving Pedestrian Detection from a Long-tailed Domain Perspective
abstract
Although pedestrian detection has developed a lot recently, there still exists some challenging scenarios, such as small-scale, occlusion and low-light. Current works usually focus on one of these scenarios independently and propose specific methods. However, different challenges may happen at a time simultaneously and change across time, making a specific method infeasible in practice. Therefore we are motivated to design a method which is able to handle various challenges and to obtain reasonable performance across different scenarios. In this paper, we first propose Instance Domain Compactness (IDC) to measure the difference of each instance in the feature space and handle hard cases from a novel long-tailed domain perspective. Specifically, we first propose a Feature Augmentation Module (FAM) to augment the tail instances in the feature space, thereby increasing the number and diversity of tail samples. Besides, a IDC-guided loss weighting module (IDCW) is formulated to adaptively re-weight the loss of each sample so as to balance the optimization procedure. Extensive analysis and experiments illustrate that our method improves the generalization of the model without any extra parameters and achieves comparable results across different challenging scenarios on both CityPersons and Caltech datasets.
Mengyuan Ding, Shanshan Zhang 0001, Jian Yang 0003
ACM Multimedia2
2021 Learning to Adapt via Latent Domains for Adaptive Semantic Segmentation
abstract
Domain adaptive semantic segmentation aims to transfer knowledge learned from labeled source domain to unlabeled target domain. To narrow down the domain gap and ease adaptation difficulty, some recent methods translate source images to target-like images (latent domains), which are used as supplement or substitute to the original source data. Nevertheless, these methods neglect to explicitly model the relationship of knowledge transferring across different domains. Alternatively, in this work we break through the standard “source-target” one pair adaptation framework and construct multiple adaptation pairs (e.g. “source-latent” and “latent-target”). The purpose is to use the meta-knowledge (how to adapt) learned from one pair as guidance to assist the adaptation of another pair under a meta-learning framework. Furthermore, we extend our method to a more practical setting of open compound domain adaptation (a.k.a multiple-target domain adaptation), where the target is a compound of multiple domains without domain labels. In this setting, we embed an additional pair of “latent-latent” to reduce the domain gap between the source and different latent domains, allowing the model to adapt well on multiple target domains simultaneously. When evaluated on standard benchmarks, our method is superior to the state-of-the-art methods in both the single target and multiple-target domain adaptation settings.
Yunan Liu 0001, Shanshan Zhang 0001, Yang Li 0190, Jian Yang 0003
NeurIPS2
2021 Single image super-resolution via hybrid resolution NSST prediction
Yunan Liu 0001, Shanshan Zhang 0001, Chunpeng Wang 0001, Jie Xu 0021
Comput. Vis. Image Underst.2
2021 Norm-Aware Embedding for Efficient Person Search and Tracking
Shanshan Zhang 0001, Jian Yang 0003, Bernt Schiele
Int. J. Comput. Vis.2
2021 Guided Attention in CNNs for Occluded Pedestrian Detection and Re-identification
Shanshan Zhang 0001, Jian Yang 0003, Bernt Schiele
Int. J. Comput. Vis.1
2021 Self-Fusion Convolutional Neural Networks
Shenjian Gong, Shanshan Zhang 0001, Jian Yang 0003, Pong C. Yuen
Pattern Recognit. Lett.2
2021 An Accurate and Lightweight Method for Human Body Image Super-Resolution
abstract
In this paper, we propose a new method to super-resolve low resolution human body images by learning efficient multi-scale features and exploiting useful human body prior. Specifically, we propose a lightweight multi-scale block (LMSB) as basic module of a coherent framework, which contains an image reconstruction branch and a prior estimation branch. In the image reconstruction branch, the LMSB aggregates features of multiple receptive fields so as to gather rich context information for low-to-high resolution mapping. In the prior estimation branch, we adopt the human parsing maps and nonsubsampled shearlet transform (NSST) sub-bands to represent the human body prior, which is expected to enhance the details of reconstructed human body images. When evaluated on the newly collected HumanSR dataset, our method outperforms state-of-the-art image super-resolution methods with ∼ 8× fewer parameters; moreover, our method significantly improves the performance of human image analysis tasks (e.g. human parsing and pose estimation) for low-resolution inputs.
Yunan Liu 0001, Shanshan Zhang 0001, Jie Xu 0021, Jian Yang 0003, Yu-Wing Tai
IEEE Trans. Image Process.2
2021 Incremental Generative Occlusion Adversarial Suppression Network for Person ReID
abstract
Person re-identification (re-id) suffers from the significant challenge of occlusion, where an image contains occlusions and less discriminative pedestrian information. However, certain work consistently attempts to design complex modules to capture implicit information (including human pose landmarks, mask maps, and spatial information). The network, consequently, focuses on discriminative features learning on human non-occluded body regions and realizes effective matching under spatial misalignment. Few studies have focused on data augmentation, given that existing single-based data augmentation methods bring limited performance improvement. To address the occlusion problem, we propose a novel Incremental Generative Occlusion Adversarial Suppression (IGOAS) network. It consists of 1) an incremental generative occlusion block, generating easy-to-hard occlusion data, that makes the network more robust to occlusion by gradually learning harder occlusion instead of hardest occlusion directly. And 2) a global-adversarial suppression (G&A) framework with a global branch and an adversarial suppression branch. The global branch extracts steady global features of the images. The adversarial suppression branch, embedded with two occlusion suppression module, minimizes the generated occlusion's response and strengthens attentive feature representation on human non-occluded body regions. Finally, we get a more discriminative pedestrian feature descriptor by concatenating two branches' features, which is robust to the occlusion problem. The experiments on the occluded dataset show the competitive performance of IGOAS. On Occluded-DukeMTMC, it achieves 60.1% Rank-1 accuracy and 49.4% mAP.
Cairong Zhao, Xinbi Lv, Shuguang Dou, Shanshan Zhang 0001, Jun Wu 0006, Liang Wang 0001
IEEE Trans. Image Process.4
2020 Hierarchical Online Instance Matching for Person Search
abstract
Person Search is a challenging task which requires to retrieve a person's image and the corresponding position from an image dataset. It consists of two sub-tasks: pedestrian detection and person re-identification (re-ID). One of the key challenges is to properly combine the two sub-tasks into a unified framework. Existing works usually adopt a straightforward strategy by concatenating a detector and a re-ID model directly, either into an integrated model or into separated models. We argue that simply concatenating detection and re-ID is a sub-optimal solution, and we propose a Hierarchical Online Instance Matching (HOIM) loss which exploits the hierarchical relationship between detection and re-ID to guide the learning of our network. Our novel HOIM loss function harmonizes the objectives of the two sub-tasks and encourages better feature learning. In addition, we improve the loss update policy by introducing Selective Memory Refreshment (SMR) for unlabeled persons, which takes advantage of the potential discrimination power of unlabeled data. From the experiments on two standard person search benchmarks, i.e. CUHK-SYSU and PRW, we achieve state-of-the-art performance, which justifies the effectiveness of our proposed HOIM loss on learning robust features.
Shanshan Zhang 0001, Wanli Ouyang, Jian Yang 0003, Bernt Schiele
AAAI2
2020 Unified Density-Aware Image Dehazing and Object Detection in Real-World Hazy Scenes
Zhengxi Zhang, Yunan Liu 0001, Shanshan Zhang 0001, Jian Yang 0003
ACCV (4)4
2020 Norm-Aware Embedding for Efficient Person Search
abstract
Person Search is a practically relevant task that aims to jointly solve Person Detection and Person Re-identification (re-ID). Specifically, it requires to find and locate all instances with the same identity as the query person in a set of panoramic gallery images. One major challenge comes from the contradictory goals of the two sub-tasks, i.e., person detection focuses on finding the commonness of all persons while person re-ID handles the differences among multiple identities. Therefore, it is crucial to reconcile the relationship between the two sub-tasks in a joint person search model. To this end, We present a novel approach called Norm-Aware Embedding to disentangle the person embedding into norm and angle for detection and re-ID respectively, allowing for both effective and efficient multi-task training. We further extend the proposal-level person embedding to pixel-level, whose discrimination ability is less affected by mis-alignment. We outperform other one-step methods by a large margin and achieve comparable performance to two-step methods on both CUHK-SYSU and PRW. Also, Our method is easy to train and resource-friendly, running at 12 fps on a single GPU.
Shanshan Zhang 0001, Jian Yang 0003, Bernt Schiele
CVPR2
2020 Learning a Dynamic High-Resolution Network for Multi-Scale Pedestrian Detection
abstract
Pedestrian detection is a canonical instance of object detection in computer vision. In practice, scale variation is one of the key challenges, resulting in unbalanced performance across different scales. Recently, the High-Resolution Network (HRNet) has become popular because high-resolution feature representations are more friendly to small objects. However, when we apply HRNet to pedestrian detection, we observe that it improves for small pedestrians on one hand, but hurts the performance for larger ones on the other hand. To overcome this problem, we propose a learnable Dynamic HRNet (DHRNet) aiming to generate different network paths adaptive to different scales. Specifically, we construct a parallel multi-branch architecture and add a soft conditional gate module allowing for dynamic feature fusion. Both branches share all the same parameters except the soft gate module. Experimental results on CityPersons and Caltech benchmarks indicate that our proposed dynamic HRNet is more capable of dealing with pedestrians of various scales, and thus improves the performance across different scales consistently.
Mengyuan Ding, Shanshan Zhang 0001, Jian Yang 0003
ICPR2
2020 Nighttime Pedestrian Detection Based on Feature Attention and Transformation
abstract
Pedestrian detection at nighttime is an important yet challenging task, which is fundamental for many practical applications, e.g. autonomous driving, video surveillance. To address this problem, in this work we start with some analysis, from which we find that the nighttime features have much more noise than that of daytime, resulting in low discrimination ability. Besides, we also observe some pedestrian examples are under adverse illumination conditions, and they can hardly provide sufficient information for accurate detection. Based on these findings, we propose the Feature Attention Module (FAM) and Feature Transformation Module (FTM) to enhance nighttime features. In FAM, guided by progressive segmentation supervision, hierarchical feature attention is produced to enhance multilevel features. On the other hand, FTM is introduced to enforce features from adverse illumination to approach that from better illumination. Based on feature attention and transformation (FAT) mechanism, a two-stage detector called FATNet is constructed for nighttime pedestrian detection. We conduct extensive experiments on nighttime datasets of EuroCity Persons (Night) and NightOwls to demonstrate the effectiveness of our method. On both datasets, our method achieves significant improvements to the baseline and also outperforms state-of-the-art detectors.
Shanshan Zhang 0001, Jian Yang 0003
ICPR2
2020 Learning Hierarchical Graph for Occluded Pedestrian Detection
abstract
Although pedestrian detection has made significant progress with the help of deep convolution neural networks, it is still a challenging problem to detect occluded pedestrians since the occluded ones can not provide sufficient information for classification and regression. In this paper, we propose a novel Hierarchical Graph Pedestrian Detector (HGPD), which integrates semantic and spatial relation information to construct two graphs named intra-proposal graph and inter-proposal graph, without relying on extra cues w.r.t visible regions. In order to capture the occlusion patterns and enhance features from visible regions, the intra-proposal graph considers body parts as nodes and assigns corresponding edge weights based on semantic relations between body parts. On the other hand, the inter-proposal graph adopts spatial relations between neighbouring proposals to provide additional proposal-wise context information for each proposal, which alleviates the lack of information caused by occlusion. We conduct extensive experiments on standard benchmarks of CityPersons and Caltech to demonstrate the effectiveness of our method. On CityPersons, our approach outperforms the baseline method by a large margin of 5.24pp on the heavy occlusion set, and surpasses all previous methods; on Caltech, we establish a new state of the art of 3.78% MR. Code is available at https://github.com/ligang-cs/PedestrianDetection-HGPD.
Jian Li 0062, Shanshan Zhang 0001, Jian Yang 0003
ACM Multimedia3
2020 Hybrid Resolution Network Using Edge Guided Region Mutual Information Loss for Human Parsing
abstract
In this paper, we propose a new method for human parsing, which effectively maintains high-resolution representations and leverages body edge details to improve the performance. First, we propose a hybrid resolution network (HyRN) for human parsing and body edge detection. In our HyRN, we adopt deconvolution operation and auxiliary supervision to increase the discrimination ability of features from each scale. Second, considering the close relationship between human parsing and body edge detection, we propose a dual-task cascaded framework (DTCF), which implicitly integrates parsing and edge features to progressively refine the parsing results. Third, we develop an edge guided region mutual information loss, which uses the edge detection results to explicitly maintain the high order consistency between parsing prediction and ground truth around body edge pixels. When evaluated on standard benchmarks, our proposed HyRN achieves competitive accuracy compared with state-of-the-art human parsing methods. Moreover, our DTCF further improves the performance and outperforms the established baseline approach by 3.42 points w.t.r mIoU on the LIP dataset.
Yunan Liu 0001, Shanshan Zhang 0001, Jian Yang 0003
ACM Multimedia3
2020 Perceiving heavily occluded human poses by assigning unbiased score
Lin Zhao 0003, Jie Xu 0021, Shanshan Zhang 0001, Chen Gong 0002, Jian Yang 0003, Xinbo Gao 0001
Inf. Sci.3
2020 Accurate quaternion radial harmonic Fourier moments for color image reconstruction and object recognition
Yunan Liu 0001, Shanshan Zhang 0001, Houjun Wang, Jian Yang 0003
Pattern Anal. Appl.2
2020 Integrating prediction and reconstruction for anomaly detection
Lin Zhao 0003, Shanshan Zhang 0001, Chen Gong 0002, Jian Yang 0003
Pattern Recognit. Lett.3
2020 Multi-task learning for object keypoints detection and classification
Jie Xu 0021, Lin Zhao 0003, Shanshan Zhang 0001, Chen Gong 0002, Jian Yang 0003
Pattern Recognit. Lett.3
2020 Color image watermark decoder by modeling quaternion polar harmonic transform with BKF distribution
Yunan Liu 0001, Shanshan Zhang 0001, Jian Yang 0003
Signal Process. Image Commun.2
2020 Person Search by Separated Modeling and A Mask-Guided Two-Stream CNN Model
abstract
In this work, we tackle the problem of person search, which is a challenging task consisted of pedestrian detection and person re-identification (re-ID). Instead of sharing representations in a single joint model, we find that separating detector and re-ID feature extraction yields better performance. In order to extract more representative features for each identity, we segment out the foreground person from the original image patch. We propose a simple yet effective re-ID method, which models foreground person and original image patches individually, and obtains enriched representations from two separate CNN streams. We also propose a Confidence Weighted Stream Attention method which further re-adjusts the relative importance of the two streams by incorporating the detection confidence. Furthermore, we simplify the whole pipeline by incorporating semantic segmentation into the re-ID network, which is trained by bounding boxes as weakly-annotated masks and identification labels simultaneously. From the experiments on two standard person search benchmarks i.e. CUHK-SYSU and PRW, we achieve mAP of 83.3% and 32.8% respectively, surpassing the state of the art by a large margin. The extensive ablation study and model inspection further justifies our motivation.
Shanshan Zhang 0001, Wanli Ouyang, Jian Yang 0003, Ying Tai
IEEE Trans. Image Process.2
2019 Coarse-to-Fine 3D Human Pose Estimation
Yu Guo 0006, Lin Zhao 0003, Shanshan Zhang 0001, Jian Yang 0003
ICIG (3)3
2019 Online Multi-object Tracking Using Single Object Tracker and Markov Clustering
Jiao Zhu, Shanshan Zhang 0001, Jian Yang 0003
ICIG (3)2
2019 Improving image retrieval by integrating shape and texture features
Yunan Liu 0001, Shanshan Zhang 0001, Si-Miao Wang
Multim. Tools Appl.2
2019 Feature Affinity-Based Pseudo Labeling for Semi-Supervised Person Re-Identification
abstract
Vision-based person re-identification aims to match a person's identity across multiple images, which is a fundamental task in multimedia content analysis and retrieval. Deep neural networks have recently manifested great potential in this task. However, a major bottleneck of existing supervised deep networks is their reliance on a large amount of annotated training data. Manual labeling for person identities in large-scale surveillance camera systems is quite challenging and incurs significant costs. Some recent studies adopt generative model outputs as training data augmentation. To more effectively use these synthetic data for an improved feature learning and re-identification performance, this paper proposes a novel feature affinity-based pseudo labeling method with two possible label encodings. To the best of our knowledge, this is the first study that employs pseudo-labeling by measuring the affinity of unlabeled samples with the underlying clusters of labeled data samples using the intermediate feature representations from deep networks. We propose training the network with the joint supervision of cross-entropy loss together with a center regularization term, which not only ensures discriminative feature representation learning but also simultaneously predicts pseudo-labels for unlabeled data. We show that both label encodings can be learned in a unified manner and help improve the overall performance. Our extensive experiments on three person re-identification datasets: Market-1501, DukeMTMC-reID, and CUHK03, demonstrate significant performance boost over the state-of-the-art person re-identification approaches.
Guodong Ding, Shanshan Zhang 0001, Salman Khan 0001, Zhenmin Tang, Jian Zhang 0002, Fatih Porikli
IEEE Trans. Multim.2
2018 NightOwls: A Pedestrians at Night Dataset
Lukás Neumann, Michelle Karg, Shanshan Zhang 0001, Christian Scharfenberger, Eric Piegert, Sarah Mistr, Olga Prokofyeva, Robert Thiel, Andrea Vedaldi, Andrew Zisserman, Bernt Schiele
ACCV (1)3
2018 Occluded Pedestrian Detection Through Guided Attention in CNNs
abstract
Pedestrian detection has progressed significantly in the last years. However, occluded people are notoriously hard to detect, as their appearance varies substantially depending on a wide range of occlusion patterns. In this paper, we aim to propose a simple and compact method based on the FasterRCNN architecture for occluded pedestrian detection. We start with interpreting CNN channel features of a pedestrian detector, and we find that different channels activate responses for different body parts respectively. These findings motivate us to employ an attention mechanism across channels to represent various occlusion patterns in one single model, as each occlusion pattern can be formulated as some specific combination of body parts. Therefore, an attention network with self or external guidances is proposed as an add-on to the baseline FasterRCNN detector. When evaluating on the heavy occlusion subset, we achieve a significant improvement of 8pp to the baseline FasterRCNN detector on CityPersons and on Caltech we outperform the state-of-the-art method by 4pp.
Shanshan Zhang 0001, Jian Yang 0003, Bernt Schiele
CVPR1
2018 Person Search via a Mask-Guided Two-Stream CNN Model
Shanshan Zhang 0001, Wanli Ouyang, Jian Yang 0003, Ying Tai
ECCV (7)2
2018 Towards Reaching Human Performance in Pedestrian Detection
abstract
Encouraged by the recent progress in pedestrian detection, we investigate the gap between current state-of-the-art methods and the "perfect single frame detector". We enable our analysis by creating a human baseline for pedestrian detection (over the Caltech pedestrian dataset). After manually clustering the frequent errors of a top detector, we characterise both localisation and background-versus-foreground errors. To address localisation errors we study the impact of training annotation noise on the detector performance, and show that we can improve results even with a small portion of sanitised training data. To address background/foreground discrimination, we study convnets for pedestrian detection, and discuss which factors affect their performance. Other than our in-depth analysis, we report top performance on the Caltech pedestrian dataset, and provide a new sanitised set of training and test annotations.
Shanshan Zhang 0001, Rodrigo Benenson, Mohamed Omran, Jan Hosang, Bernt Schiele
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 CityPersons: A Diverse Dataset for Pedestrian Detection
abstract
Convnets have enabled significant progress in pedestrian detection recently, but there are still open questions regarding suitable architectures and training data. We revisit CNN design and point out key adaptations, enabling plain FasterRCNN to obtain state-of-the-art results on the Caltech dataset. To achieve further improvement from more and better data, we introduce CityPersons, a new set of person annotations on top of the Cityscapes dataset. The diversity of CityPersons allows us for the first time to train one single CNN model that generalizes well over multiple benchmarks. Moreover, with additional training with CityPersons, we obtain top results using FasterRCNN on Caltech, improving especially for more difficult cases (heavy occlusion and small scale) and providing higher localization quality.
Shanshan Zhang 0001, Rodrigo Benenson, Bernt Schiele
CVPR1
2016 How Far are We from Solving Pedestrian Detection?
abstract
Encouraged by the recent progress in pedestrian detection, we investigate the gap between current state-of-the-art methods and the "perfect single frame detector". We enable our analysis by creating a human baseline for pedestrian detection (over the Caltech dataset), and by manually clustering the recurrent errors of a top detector. Our results characterise both localisation and background-versusforeground errors. To address localisation errors we study the impact of training annotation noise on the detector performance, and show that we can improve even with a small portion of sanitised training data. To address background/foreground discrimination, we study convnets for pedestrian detection, and discuss which factors affect their performance. Other than our in-depth analysis, we report top performance on the Caltech dataset, and provide a new sanitised set of training and test annotations.
Shanshan Zhang 0001, Rodrigo Benenson, Mohamed Omran, Jan Hosang, Bernt Schiele
CVPR1
2016 Fast moving pedestrian detection based on motion segmentation and new motion features
Shanshan Zhang 0001, Dominik A. Klein, Christian Bauckhage, Armin B. Cremers
Multim. Tools Appl.1
2015 Filtered channel features for pedestrian detection
abstract
This paper starts from the observation that multiple top performing pedestrian detectors can be modelled by using an intermediate layer filtering low-level features in combination with a boosted decision forest. Based on this observation we propose a unifying framework and experimentally explore different filter families. We report extensive results enabling a systematic analysis. Using filtered channel features we obtain top performance on the challenging Caltech and KITTI datasets, while using only HOG+LUV as low-level features. When adding optical flow features we further improve detection quality and report the best known results on the Caltech dataset, reaching 93% recall at 1 FPPI.
Shanshan Zhang 0001, Rodrigo Benenson, Bernt Schiele
CVPR1
2015 Exploring Human Vision Driven Features for Pedestrian Detection
abstract
Motivated by the center-surround mechanism in the human visual attention system, we propose to use average contrast maps for the challenge of pedestrian detection in street scenes due to the observation that pedestrians indeed exhibit discriminative contrast texture. Our main contributions are the first to design a local statistical multichannel descriptor to incorporate both color and gradient information. Second, we introduce a multidirection and multiscale contrast scheme based on grid cells to integrate expressive local variations. Contributing to the issue of selecting most discriminative features for assessing and classification, we perform extensive comparisons with respect to statistical descriptors, contrast measurements, and scale structures. By this way, we obtain reasonable results under various configurations. Empirical findings from applying our optimized detector on the INRIA and Caltech pedestrian datasets show that our features yield state-of-the-art performance in pedestrian detection.
Shanshan Zhang 0001, Christian Bauckhage, Dominik A. Klein, Armin B. Cremers
IEEE Trans. Circuits Syst. Video Technol.1
2015 Efficient Pedestrian Detection via Rectangular Features Based on a Statistical Shape Model
abstract
Automatic pedestrian detection for advanced driver assistance systems (ADASs) is still a challenging task. Major reasons are dynamic and complex backgrounds in street scenes and variations in clothing or postures of pedestrians. We propose a simple yet effective detector for robust pedestrian detection. Observing that pedestrians usually appear upright in video data, we employ a statistical model of the upright human body in which the head, upper body, and lower body are treated as three distinct components. Our main contribution is to systematically design a pool of rectangular features that are tailored to this shape model. As we incorporate different kinds of low-level measurements, the resulting multimodal and multichannel Haar-like features represent characteristic differences between parts of the human body but are robust against variations in clothing or environmental settings. Our approach avoids exhaustive searches over all possible configurations of rectangular features nor does it rely on random sampling. It thus marks a middle ground among recently published techniques and yields efficient low-dimensional yet highly discriminative features. Experimental results on the well-established INRIA, Caltech, and KITTI pedestrian data sets show that our detector reaches state-of-the-art performance at low computational costs and that our features are robust against occlusions.
Shanshan Zhang 0001, Christian Bauckhage, Armin B. Cremers
IEEE Trans. Intell. Transp. Syst.1
2014 Informed Haar-Like Features Improve Pedestrian Detection
abstract
We propose a simple yet effective detector for pedestrian detection. The basic idea is to incorporate common sense and everyday knowledge into the design of simple and computationally efficient features. As pedestrians usually appear up-right in image or video data, the problem of pedestrian detection is considerably simpler than general purpose people detection. We therefore employ a statistical model of the up-right human body where the head, the upper body, and the lower body are treated as three distinct components. Our main contribution is to systematically design a pool of rectangular templates that are tailored to this shape model. As we incorporate different kinds of low-level measurements, the resulting multi-modal & multi-channel Haar-like features represent characteristic differences between parts of the human body yet are robust against variations in clothing or environmental settings. Our approach avoids exhaustive searches over all possible configurations of rectangle features and neither relies on random sampling. It thus marks a middle ground among recently published techniques and yields efficient low-dimensional yet highly discriminative features. Experimental results on the INRIA and Caltech pedestrian datasets show that our detector reaches state-of-the-art performance at low computational costs and that our features are robust against occlusions.
Shanshan Zhang 0001, Christian Bauckhage, Armin B. Cremers
CVPR1
2014 Center-Surround Contrast Features for Pedestrian Detection
abstract
Inspired by the human vision system, in this paper we propose a specifically organized kind of center-surround contrast features and show their suitability for pedestrian detection. These contrasts are computed from a novel combination of both local color and gradient statistics aggregated quickly for arbitrary sized square cells. We exploit our contrast features in a rich multi-scale and -direction fashion between each central cell and its neighbors and boost the significant ones for pedestrian detection. Experimental results on the INRIA and Caltech pedestrian datasets show that our method achieves state-of-the-art performance.
Shanshan Zhang 0001, Dominik A. Klein, Christian Bauckhage, Armin B. Cremers
ICPR1