Zhenjun Han

dblp:11/2938 · DBLP profile ↗
← Back
53ranked-venue papers
6as first author
29since 2021 · last 2026
0000-0002-9970-5152ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 4 first-author · 20 since 2021Artificial intelligence and machine learning · 27 · 3 first-author · 18 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 SeqCount: A sequence modeling framework for class-agnostic counting
abstract
Class-agnostic counting aims to count the number of objects in any category with only a few exemplars. It is crucial for solving the challenge of counting any visual class without re-finetuning, which in turn lowers deployment costs across diverse scenarios. Existing methods count the number of exemplar objects by integrating them over the density map smoothed by Gaussian kernels. However, designing generic kernels to generate density maps is challenging due to the different sizes and shapes of objects. To solve this problem, we propose SeqCount, which eliminates the need for density maps and treats object counting as a sequence generation problem. Specifically, we consider an input image as $N\times N$ patches and propose a serialization scheme. Then, we use an encoding-decoding structure to exploit the correlation among the patches. Experimental results on five challenging datasets demonstrate that our method performs favorably against the state-of-the-art models.
Guorong Li, Xinyan Liu 0008, Zhenjun Han, Yuankai Qi
J. Vis. Commun. Image Represent.5
2026 Consistency-Aware Anchor Pyramid Network for Crowd Localization
abstract
Crowd localization aims to predict the positions of humans in images of crowded scenes. While existing methods have made significant progress, two primary challenges remain: (i) a fixed number of evenly distributed anchors can cause excessive or insufficient predictions across regions in an image with varying crowd densities, and (ii) ranking inconsistency of predictions between the testing and training phases leads to the model being sub-optimal in inference. To address these issues, we propose a Consistency-Aware Anchor Pyramid Network (CAAPN) comprising two key components: an Adaptive Anchor Generator (AAG) and a Localizer with Augmented Matching (LAM). The AAG module adaptively generates anchors based on estimated crowd density in local regions to alleviate the anchor deficiency or excess problem. It also considers the spatial distribution prior to heads for better performance. The LAM module is designed to augment the predictions which are used to optimize the neural network during training by introducing an extra set of target candidates and correctly matching them to the ground truth. The proposed method achieves favorable performance against state-of-the-art approaches on five challenging datasets: ShanghaiTech A and B, UCF-QNRF, JHU-CROWD++, and NWPU-Crowd.
Xinyan Liu 0008, Guorong Li, Yuankai Qi, Zhenjun Han, Anton van den Hengel, Nicu Sebe, Ming-Hsuan Yang 0001, Qingming Huang
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 SAPNet++: Evolving Point-Prompted Instance Segmentation With Semantic and Spatial Awareness
abstract
Single-point annotation is increasingly prominent in visual tasks for labeling cost reduction. However, it challenges tasks requiring high precision, such as the point-prompted instance segmentation (PPIS) task, which aims to estimate precise masks using single-point prompts to train a segmentation network. Due to the constraints of point annotations, granularity ambiguity and boundary uncertainty arise i.e., the difficulty distinguishing between different levels of detail (e.g., whole object vs. parts) and the challenge of precisely delineating object boundaries. Previous works have usually inherited the paradigm of mask generation along with proposal selection to achieve PPIS. However, proposal selection relies solely on category information, failing to resolve the ambiguity of different granularity. Furthermore, mask generators offer only finite discrete solutions that often deviate from actual masks, particularly at boundaries. To address these issues, we propose the Semantic-Aware Point-Prompted Instance Segmentation Network (SAPNet). It integrates Point Distance Guidance and Box Mining Strategy to tackle group and local issues caused by the point's granularity ambiguity. Additionally, we incorporate completeness scores within proposals to add spatial granularity awareness, enhancing multiple instance learning (MIL) in proposal selection termed S-MIL. The Multi-level Affinity Refinement conveys pixel and semantic clues, narrowing boundary uncertainty during mask refinement. These modules culminate in SAPNet++, mitigating point prompt's granularity ambiguity and boundary uncertainty and significantly improving segmentation performance. Extensive experiments on four challenging datasets validate the effectiveness of our methods, highlighting the potential to advance PPIS.
Zhaoyang Wei, Xumeng Han, Xuehui Yu, Xue Yang 0005, Guorong Li, Zhenjun Han, Jianbin Jiao
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Boosting Segment Anything Model Towards Open-Vocabulary Learning
abstract
The recent Segment Anything Model (SAM) has emerged as a new paradigmatic vision foundation model, showcasing potent zero-shot generalization and flexible prompting. Despite SAM finding applications and adaptations in various domains, its primary limitation lies in the inability to grasp object semantics. In this paper, we present Sambor to seamlessly integrate SAM with the open-vocabulary object detector in an end-to-end framework. While retaining all the remarkable capabilities inherent to SAM, we boost it to detect arbitrary objects from human inputs like category names or reference expressions. Building upon the SAM image encoder, we introduce a novel SideFormer module designed to acquire SAM features adept at perceiving objects and inject comprehensive semantic information for recognition. In addition, we devise an Open-set RPN that leverages SAM proposals to assist in finding potential objects. Consequently, Sambor enables the open-vocabulary detector to equally focus on generalizing both localization and classification sub-tasks. Our approach demonstrates superior zero-shot performance across benchmarks, including COCO and LVIS, proving highly competitive against previous state-of-the-art methods. We aspire for this work to serve as a meaningful endeavor in endowing SAM to recognize diverse object categories and advancing open-vocabulary learning with the support of vision foundation models.
Xumeng Han, Longhui Wei, Xuehui Yu, Zhiyang Dou, Kuiran Wang, Yingfei Sun, Zhenjun Han, Qi Tian 0001
AAAI8
2025 SAM-CP: Marrying SAM with Composable Prompts for Versatile Segmentation
abstract
The Segment Anything model (SAM) has shown a generalized ability to group image pixels into patches, but applying it to semantic-aware segmentation still faces major challenges. This paper presents SAM-CP, a simple approach that establishes two types of composable prompts beyond SAM and composes them for versatile segmentation. Specifically, given a set of classes (in texts) and a set of SAM patches, the Type-I prompt judges whether a SAM patch aligns with a text label, and the Type-II prompt judges whether two SAM patches with the same text label also belong to the same instance. To decrease the complexity in dealing with a large number of semantic classes and patches, we establish a unified framework that calculates the affinity between (semantic and instance) queries and SAM patches, and then merges patches with high affinity to the query. Experiments show that SAM-CP achieves semantic, instance, and panoptic segmentation in both open and closed domains. In particular, it achieves state-of-the-art performance in open-vocabulary segmentation. Our research offers a novel and generalized methodology for equipping vision foundation models like SAM with multi-grained semantic perception abilities. Codes are released on https://github.com/ucas-vg/SAM-CP.
Pengfei Chen 0004, Lingxi Xie, Xinyue Huo, Xuehui Yu, Xiaopeng Zhang 0008, Yingfei Sun, Zhenjun Han, Qi Tian 0001
ICLR7
2025 VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
abstract
With the rapid development of MLLMs, evaluating their visual capabilities has become increasingly crucial. Current benchmarks primarily fall into two main types: basic perception benchmarks,which focus on local details but lack deep reasoning (e.g., ''what is in the image?''), and mainstream reasoning benchmarks, which concentrate on prominent image elements but may fail to assess subtle clues requiring intricate analysis. However, profound visual understanding and complex reasoning depend more on interpreting subtle, inconspicuous local details than on perceiving salient, macro-level objects. These details, though occupying minimal image area, often contain richer, more critical information for robust analysis. To bridge this gap, we introduce the VER-Bench, a novel framework to evaluate MLLMs' ability to: 1) identify fine-grained visual clues, often occupying, on average, just 0.25% of the image area; 2) integrate these clues with world knowledge for complex reasoning. Comprising 374 carefully designed questions across Geospatial, Temporal, Situational, Intent, System State, and Symbolic reasoning, each question in VER-Bench is accompanied by structured evidence: visual clues and question-related reasoning derived from them. VER-Bench reveals current models' limitations in extracting subtle visual evidence and constructing evidence-based reasoning chains, highlighting the need to enhance models' capabilities in fine-grained visual evidence extraction, integration, and reasoning for genuine visual understanding and human-like analysis. The dataset is available at https://github.com/verbta/ACMMM-25-Materials.
Chenhui Qiang, Zhaoyang Wei, Xumeng Han, Siyao Li, Xiangyuan Lan, Jianbin Jiao, Zhenjun Han
ACM Multimedia8
2025 P2Object: Single Point Supervised Object Detection and Instance Segmentation
Pengfei Chen 0004, Xuehui Yu, Xumeng Han, Kuiran Wang, Guorong Li, Lingxi Xie, Zhenjun Han, Jianbin Jiao
Int. J. Comput. Vis.7
2025 Wholly-WOOD: Wholly Leveraging Diversified-Quality Labels for Weakly-Supervised Oriented Object Detection
abstract
Accurately estimating the orientation of visual objects with compact rotated bounding boxes (RBoxes) has become a prominent demand, which challenges existing object detection paradigms that only use horizontal bounding boxes (HBoxes). To equip the detectors with orientation awareness, supervised regression/classification modules have been introduced at the high cost of rotation annotation. Meanwhile, some existing datasets with oriented objects are already annotated with horizontal boxes or even single points. It becomes attractive yet remains open for effectively utilizing weaker single point and horizontal annotations to train an oriented object detector (OOD). We develop Wholly-WOOD, a weakly-supervised OOD framework, capable of wholly leveraging various labeling forms (Points, HBoxes, RBoxes, and their combination) in a unified fashion. By only using HBox for training, our Wholly-WOOD achieves performance very close to that of the RBox-trained counterpart on remote sensing and other areas, significantly reducing the tedious efforts on labor-intensive annotation for oriented objects.
Yi Yu 0010, Xue Yang 0005, Yansheng Li 0001, Zhenjun Han, Feipeng Da, Junchi Yan
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 ClickTrack: Towards real-time interactive single object tracking
Kuiran Wang, Xuehui Yu, Wenwen Yu, Guorong Li, Xiangyuan Lan, Qixiang Ye, Jianbin Jiao, Zhenjun Han
Pattern Recognit.8
2025 ViMoE: An Empirical Study of Designing Vision Mixture-of-Experts
abstract
Mixture-of-Experts (MoE) models embody the divide-and-conquer concept and are a promising approach for increasing model capacity, demonstrating excellent scalability across multiple domains. In this paper, we integrate the MoE structure into the classic Vision Transformer (ViT), naming it ViMoE, and explore the potential of applying MoE to vision through a comprehensive study on image classification and semantic segmentation. However, we observe that the performance is sensitive to the configuration of MoE layers, making it challenging to obtain optimal results without careful design. The underlying cause is that inappropriate MoE layers lead to unreliable routing and hinder experts from effectively acquiring helpful information. To address this, we introduce a shared expert to learn and capture common knowledge, serving as an effective way to construct a stable ViMoE. Furthermore, we demonstrate how to analyze expert routing behavior, revealing which MoE layers are capable of specializing in handling specific information and which are not. This provides guidance for retaining the critical layers while removing redundancies, thereby advancing ViMoE to be more efficient without sacrificing accuracy. We aspire for this work to offer new insights into the design of vision MoE models and provide valuable empirical guidance for future research.
Xumeng Han, Longhui Wei, Zhiyang Dou, Yingfei Sun, Zhenjun Han, Qi Tian 0001
IEEE Trans. Image Process.5
2024 Weakly Supervised Video Individual Counting
abstract
Video Individual Counting (VIC) aims to predict the number of unique individuals in a single video. Existing methods learn representations based on trajectory labels for individuals, which are annotation-expensive. To provide a more realistic reflection of the underlying practical challenge, we introduce a weakly supervised VIC task, wherein trajectory labels are not provided. Instead, two types of labels are provided to indicate traffic entering the field of view (inflow) and leaving the field view (outflow). We also propose the first solution as a baseline that formulates the task as a weakly supervised contrastive learning problem under group-level matching. In doing so, we devise an end-to-end trainable soft contrastive loss to drive the network to distin-guish inflow, outflow, and the remaining. To facilitate future study in this direction, we generate annotations from the existing VIC datasets Sense Crowd and CroHD and also build a new dataset, UAVVIC. Extensive results show that our baseline weakly supervised method outperforms supervised methods, and thus, little information is lost in the transition to the more practically relevant weakly supervised task. The code and trained model can be found at CGNet.
Xinyan Liu 0008, Guorong Li, Yuankai Qi, Ziheng Yan, Zhenjun Han, Anton van den Hengel, Ming-Hsuan Yang 0001, Qingming Huang
CVPR5
2024 Semantic-aware SAM for Point-Prompted Instance Segmentation
abstract
Single-point annotation in visual tasks, with the goal of minimizing labelling costs, is becoming increasingly prominent in research. Recently, visual foundation models, such as Segment Anything (SAM), have gained widespread usage due to their robust zero-shot capabilities and exceptional annotation performance. However, SAM's class-agnostic output and high confidence in local segmentation introduce semantic ambiguity, posing a challenge for precise category-specific segmentation. In this paper, we introduce a cost-effective category-specific segmenter using SAM. To tackle this challenge, we have devised a Semantic-Aware Instance Segmentation Network (SAPNet) that integrates Multiple Instance Learning (MIL) with matching capability and SAM with point prompts. SAPNet strategically selects the most representative mask proposals generated by SAM to supervise segmentation, with a specific focus on object category information. Moreover, we introduce the Point Distance Guidance and Box Mining Strategy to mitigate inherent challenges: group and local issues in weakly supervised segmentation. These strategies serve to further enhance the overall segmentation performance. The experimental results on Pascal VOC and COCO demonstrate the promising performance of our proposed SAPNet, emphasizing its semantic matching capabilities and its potential to advance point-prompted instance segmentation. The code is available at https://github.com/zhaoyangwei123/SAPNet.
Zhaoyang Wei, Pengfei Chen 0004, Xuehui Yu, Guorong Li, Jianbin Jiao, Zhenjun Han
CVPR6
2024 P2Seg: Pointly-supervised Segmentation via Mutual Distillation
abstract
Point-level Supervised Instance Segmentation (PSIS) aims to enhance the applicability and scalability of instance segmentation by utilizing low-cost yet instance-informative annotations. Existing PSIS methods usually rely on positional information to distinguish objects, but predicting precise boundaries remains challenging due to the lack of contour annotations. Nevertheless, weakly supervised semantic segmentation methods are proficient in utilizing intra-class feature consistency to capture the boundary contours of the same semantic regions. In this paper, we design a Mutual Distillation Module (MDM) to leverage the complementary strengths of both instance position and semantic information and achieve accurate instance-level object perception. The MDM consists of Semantic to Instance (S2I) and Istance to Semantic (I2S). S2I is guided by the precise boundaries of semantic regions to learn the association between annotated points and instance contours. I2S leverages discriminative relationships between instances to facilitate the differentiation of various objects within the semantic map. Extensive experiments substantiate the efficacy of MDM in fostering the synergy between instance and semantic information, consequently improving the quality of instance-level object representations. Our method achieves 55.7 mAP50 and 17.6 mAP on the PASCAL VOC and MS COCO datasets, significantly outperforming recent PSIS methods and several box-supervised instance segmentation competitors.
Xuehui Yu, Xumeng Han, Wenwen Yu, Zhixun Huang, Jianbin Jiao, Zhenjun Han
ICLR7
2024 CPR++: Object Localization via Single Coarse Point Supervision
abstract
Point-based object localization (POL), which pursues high-performance object sensing under low-cost data annotation, has attracted increased attention. However, the point annotation mode inevitably introduces semantic variance due to the inconsistency of annotated points. Existing POL heavily rely on strict annotation rules, which are difficult to define and apply, to handle the problem. In this study, we propose coarse point refinement (CPR), which to our best knowledge is the first attempt to alleviate semantic variance from an algorithmic perspective. CPR reduces the semantic variance by selecting a semantic centre point in a neighbourhood region to replace the initial annotated point. Furthermore, We design a sampling region estimation module to dynamically compute a sampling region for each object and use a cascaded structure to achieve end-to-end optimization. We further integrate a variance regularization into the structure to concentrate the predicted scores, yielding CPR++. We observe that CPR++ can obtain scale information and further reduce the semantic variance in a global region, thus guaranteeing high-performance object localization. Extensive experiments on four challenging datasets validate the effectiveness of both CPR and CPR++. We hope our work can inspire more research on designing algorithms rather than annotation rules to address the semantic variance problem in POL.
Xuehui Yu, Pengfei Chen 0004, Kuiran Wang, Xumeng Han, Guorong Li, Zhenjun Han, Qixiang Ye, Jianbin Jiao
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Save the Tiny, Save the All: Hierarchical Activation Network for Tiny Object Detection
abstract
Tiny object detection (TOD) remains a challenging problem due to the extremely small size and weak feature presentations of tiny objects. Many effective methods have improved the detection of small objects below$32\times 32$pixels to some extent, but the performance is still poor for the tiny objects below$16\times 16$pixels. In this paper, we find that the aliasing between the features and object scales, namely feature-scale-aliasing, leads to the misalignment between feature subspaces and detection subspaces, and thus results in the interference of features, especially for tiny objects. To alleviate this, we propose a Hierarchical Activation (HA) method to obtain scale-specific feature subspaces by activating object features at different scales hierarchically. To this end, we design a Scale-Guided Feature Activation (SGFA) to decompose the original object-aliasing feature spaces into a group of scale-specific feature subspaces by scale-guided activation maps. Then, Scale-Specific Feature re-Coupling (SSFC) is used to enhance the feature subspaces by adaptively aggregating the feature subspaces from different groups. In addition, we propose to complement the scale-specific detailed information by a designed Detailed Information Compensation (DIC) method. Implementing HA, a multi-scale keypoint-based detector is constructed to improve the tiny object detection, referred to as Hierarchical Activation Network (HANet). Extensive experiments are carried out on three tiny object detection datasets, e.g., TinyPerson, AI-TOD, and TinyCOCO. Our HANet achieves 58.45%$AP_{50}^{all}$, 22.1%$AP$, and 15.76%$AP$on TinyPerson, AI-TOD, and TinyCOCO, respectively, showing a significant performance gain over the competitors.
Guangqian Guo, Pengfei Chen 0004, Xuehui Yu, Zhenjun Han, Qixiang Ye, Shan Gao 0003
IEEE Trans. Circuits Syst. Video Technol.4
2024 Self Supervised Progressive Network for High Performance Video Object Segmentation
abstract
Recently, self-supervised video object segmentation (VOS) has attracted much interest. However, most proxy tasks are proposed to train only a single backbone, which relies on a point-to-point correspondence strategy to propagate masks through a video sequence. Due to its simple pipeline, the performance of the single backbone paradigm is still unsatisfactory. Instead of following the previous literature, we propose our self-supervised progressive network (SSPNet) which consists of a memory retrieval module (MRM) and collaborative refinement module (CRM). The MRM can perform point-to-point correspondence and produce a propagated coarse mask for a query frame through self-supervised pixel-level and frame-level similarity learning. The CRM, which is trained via cycle consistency region tracking, aggregates the reference & query information and learns the collaborative relationship among them implicitly to refine the coarse mask. Furthermore, to learn semantic knowledge from unlabeled data, we also design two novel mask-generation strategies to provide the training data with meaningful semantic information for the CRM. Extensive experiments conducted on DAVIS-17, YouTube- VOS and SegTrack v2 demonstrate that our method surpasses the state-of-the-art self-supervised methods and narrows the gap with the fully supervised methods.
Guorong Li, Dexiang Hong, Kai Xu 0013, Bineng Zhong 0001, Li Su 0003, Zhenjun Han, Qingming Huang
IEEE Trans. Neural Networks Learn. Syst.6
2023 Spatial Self-Distillation for Object Detection with Inaccurate Bounding Boxes
abstract
Object detection via inaccurate bounding boxes supervision has boosted a broad interest due to the expensive high-quality annotation data or the occasional inevitability of low annotation quality (e.g. tiny objects). The previous works usually utilize multiple instance learning (MIL), which highly depends on category information, to select and refine a low-quality box. Those methods suffer from object drift, group prediction and part domination problems without exploring spatial information. In this paper, we heuristically propose a Spatial Self-Distillation based Object Detector (SSD-Det) to mine spatial information to refine the inaccurate box in a self-distillation fashion. SSD-Det utilizes a Spatial Position Self-Distillation (SPSD) module to exploit spatial information and an interactive structure to combine spatial information and category information, thus constructing a high-quality proposal bag. To further improve the selection procedure, a Spatial Identity Self-Distillation (SISD) module is introduced in SSD-Det to obtain spatial confidence to help select the best proposals. Experiments on MS-COCO and VOC datasets with noisy box annotation verify our method’s effectiveness and achieve state-of-the-art performance. The code is available at https://github.com/ucas-vg/PointTinyBenchmark/tree/SSD-Det.
Pengfei Chen 0004, Xuehui Yu, Guorong Li, Zhenjun Han, Jianbin Jiao
ICCV5
2023 Rethinking Sampling Strategies for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (re-ID) remains a challenging task. While extensive research has focused on the framework design and loss function, this paper shows that sampling strategy plays an equally important role. We analyze the reasons for the performance differences between various sampling strategies under the same framework and loss function. We suggest that deteriorated over-fitting is an important factor causing poor performance, and enhancing statistical stability can rectify this problem. Inspired by that, a simple yet effective approach is proposed, termed group sampling, which gathers samples from the same class into groups. The model is thereby trained using normalized group samples, which helps alleviate the negative impact of individual samples. Group sampling updates the pipeline of pseudo-label generation by guaranteeing that samples are more efficiently classified into the correct classes. It regulates the representation learning process, enhancing statistical stability for feature representation in a progressive fashion. Extensive experiments on Market-1501, DukeMTMC-reID and MSMT17 show that group sampling achieves performance comparable to state-of-the-art methods and outperforms the current techniques under purely camera-agnostic settings. Code has been available at https://github.com/ucas-vg/GroupSampling.
Xumeng Han, Xuehui Yu, Guorong Li, Jian Zhao 0006, Gang Pan 0002, Qixiang Ye, Jianbin Jiao, Zhenjun Han
IEEE Trans. Image Process.8
2023 Anti-UAV: A Large-Scale Benchmark for Vision-Based UAV Tracking
abstract
Unmanned Aerial Vehicles (UAV) have many applications in both commerce and recreation. However, irresponsibly operated UAVs will pose a threat to public safety. Therefore, developing our understanding of UAVs and their uses is of particular interest. This paper considers tracking UAVs, which provide multifaceted information around location, paths and trajectories. To facilitate research on this topic, we introduce a new benchmark, herein referred to as Anti-UAV, which provides a novel direction for UAV tracking with more than 300 video pairs containing over 580 k manually annotated bounding boxes. Addressing anti-UAV research challenges could help to design anti-UAV systems, which in turn may improve surveillance. Accordingly, we have proposed a simple yet effective approach, called dual-flow semantic consistency (DFSC) is proposed for UAV tracking. Modulated by the semantic flow across video sequences, tracker learns more robust class-level semantic information and obtains more discriminative instance-level features. Experiments highlight significant performance gain with the proposed approach over state-of-the-art trackers and the challenging aspects of Anti-UAV. The Anti-UAV benchmark and the code for the proposed approach have been made publicly available athttps://github.com/ucas-vg/Anti-UAVandhttps://github.com/ZhaoJ9014/Anti-UAV.
Kuiran Wang, Xiaoke Peng, Xuehui Yu, Qiang Wang 0051, Junliang Xing, Guorong Li, Guodong Guo, Qixiang Ye, Jianbin Jiao, Jian Zhao 0006, Zhenjun Han
IEEE Trans. Multim.12
2022 Object Localization under Single Coarse Point Supervision
abstract
Point-based object localization (POL), which pursues high-performance object sensing under low-cost data annotation, has attracted increased attention. However, the point annotation mode inevitably introduces semantic variance for the inconsistency of annotated points. Existing POL methods heavily reply on accurate keypoint annotations which are difficult to define. In this study, we propose a POL method using coarse point annotations, relaxing the supervision signals from accurate key points to freely spotted points. To this end, we propose a coarse point refinement (CPR) approach, which to our best knowledge is the first attempt to alleviate semantic variance from the perspective of algorithm. CPR constructs point bags, selects semantic-correlated points, and produces semantic center points through multiple instance learning (MIL). In this way, CPR defines a weakly supervised evolution procedure, which ensures training high-performance object localizer under coarse point supervision. Experimental results on COCO, DOTA and our proposed SeaPerson dataset validate the effectiveness of the CPR approach. The dataset and code will be available at https://github.com/ucas-vg/PointTinyBenchmark/
Xuehui Yu, Pengfei Chen 0004, Najmul Hassan, Guorong Li, Junchi Yan, Humphrey Shi, Qixiang Ye, Zhenjun Han
CVPR9
2022 Point-to-Box Network for Accurate Object Detection via Single Point Supervision
Pengfei Chen 0004, Xuehui Yu, Xumeng Han, Najmul Hassan, Kai Wang 0058, Jiachen Li 0003, Jian Zhao 0006, Humphrey Shi, Zhenjun Han, Qixiang Ye
ECCV (9)9
2022 End-to-End Weakly Supervised Object Detection with Sparse Proposal Evolution
Mingxiang Liao, Fang Wan 0001, Zhenjun Han, Jialing Zou, Yuze Wang 0004, Bailan Feng, Qixiang Ye
ECCV (9)4
2022 Multi-Attention Network for Compressed Video Referring Object Segmentation
abstract
Referring video object segmentation aims to segment the object referred by a given language expression. Existing works typically require compressed video bitstream to be decoded to RGB frames before being segmented, which increases computation and storage requirements and ultimately slows the inference down. This may hamper its application in real-world computing resource limited scenarios, such as autonomous cars and drones. To alleviate this problem, in this paper, we explore the referring object segmenta- tion task on compressed videos, namely on the original video data flow. Besides the inherent difficulty of the video referring object segmentation task itself, obtaining discriminative representation from compressed video is also rather challenging. To address this problem, we propose a multi-attention network which consists of dual-path dual-attention module and a query-based cross-modal Transformer module. Specifically, the dual-path dual-attention module is designed to extract effective representation from compressed data in three modalities, i.e., I-frame, Motion Vector and Residual. The query-based cross-modal Transformer firstly models the corre- lation between linguistic and visual modalities, and then the fused multi-modality features are used to guide object queries to generate a content-aware dynamic kernel and to predict final segmentation masks. Different from previous works, we propose to learn just one kernel, which thus removes the complicated post mask-matching procedure of existing methods. Extensive promising experimental results on three challenging datasets show the effectiveness of our method compared against several state-of-the-art methods which are proposed for processing RGB data. Source code is available at: https://github.com/DexiangHong/MANet.
Weidong Chen 0013, Dexiang Hong, Yuankai Qi, Zhenjun Han, Shuhui Wang, Laiyun Qing, Qingming Huang, Guorong Li
ACM Multimedia4
2022 Dynamic Perception Framework for Fine-Grained Recognition
abstract
Fine-grained recognition poses the challenge of discriminating categories with only small subtle visual differences, which can be easily overwhelmed by diverse appearance within categories. Conventional approaches generally locate discriminative parts and then recognize the part-based features. However, we find that tuning the effective receptive field (ERF) of the network to the task plays the key role, which enables significant regions to contribute more to the output. Inspired by the receptive field stimulation mechanism of the visual cortex, we propose a Dynamic Perception framework as a solution. Our framework adapts the ERF by considering the image space and the kernel space simultaneously. In the image space, the Spatial Selective Sampling module is adopted to enlarge informative regions locally. In the kernel space, Spatial Selective Kernel convolution is introduced to adapt different kernel sizes for regions of interest and backgrounds by embedding spatial attention in the multi-path convolution. Extensive experiments on challenging benchmarks, including CUB-200-2011, FGVC-Aircraft, and Stanford Cars, demonstrate that our method yields a performance boost over the state-of-the-art methods.
Yao Ding 0006, Zhenjun Han, Yanzhao Zhou, Yi Zhu 0004, Jie Chen 0001, Qixiang Ye, Jianbin Jiao
IEEE Trans. Circuits Syst. Video Technol.2
2022 From Coarse to Fine: Hierarchical Structure-aware Video Summarization
abstract
Hierarchical structure is a common characteristic for some kinds of videos (e.g., sports videos, game videos): The videos are composed of several actions hierarchically and there exist temporal dependencies among segments with different scales, where action labels can be enumerated. Our ideas are based on two observations: First, the actions are the fundamental units for people to understand these videos. Second, the humans summarize a video by iteratively observing and refining, i.e., observing segments in video and hierarchically refining the boundaries of important actions. Based on the above insights, we generate action proposals to construct the structure of the video and formulate the summarization process as a hierarchical refining process. We also train a hierarchical summarization network with deep Q-learning (HQSN) to achieve the refining process and explore temporal dependency. Besides, we collect a new dataset that consists of structured game videos with fine-grain actions and importance annotations. The experimental results demonstrate the effectiveness of the proposed method.
Wenxu Li, Gang Pan 0002, Zhenjun Han
ACM Trans. Multim. Comput. Commun. Appl.5
2021 SM+: Refined Scale Match for Tiny Person Detection
abstract
Detecting tiny objects (e.g., less than 20 × 20 pixels) in large-scale images is an important yet open problem. Modern CNN-based detectors are challenged by the scale mismatch between the dataset for network pre-training and the target dataset for detector training. In this paper, we investigate the scale alignment between pre-training and target datasets, and propose a new refined Scale Match method (termed SM+) for tiny person detection. SM+ improves the scale match from image level to instance level, and effectively promotes the similarity between pre-training and target dataset. Moreover, considering SM+ possibly destroys the image structure, a new probabilistic structure inpainting (PSI) method is proposed for the background processing. Experiments conducted across various detectors show that SM+ noticeably improves the performance on TinyPerson, and outperforms the state-of-the-art detectors with a significant margin.
Xuehui Yu, Xiaoke Peng, Yuqi Gong, Zhenjun Han
ICASSP5
2021 TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object Localization
abstract
Weakly supervised object localization (WSOL) is a challenging problem when given image category labels but requires to learn object localization models. Optimizing a convolutional neural network (CNN) for classification tends to activate local discriminative regions while ignoring complete object extent, causing the partial activation issue. In this paper, we argue that partial activation is caused by the intrinsic characteristics of CNN, where the convolution operations produce local receptive fields and experience difficulty to capture long-range feature dependency among pixels. We introduce the token semantic coupled attention map (TS-CAM) to take full advantage of the self-attention mechanism in visual transformer for long-range dependency extraction. TS-CAM first splits an image into a sequence of patch tokens for spatial embedding, which produce attention maps of long-range visual dependency to avoid partial activation. TS-CAM then re-allocates category-related semantics for patch tokens, enabling each of them to be aware of object categories. TS-CAM finally couples the patch tokens with the semantic-agnostic attention map to achieve semantic-aware localization. Experiments on the ILSVRC/CUB-200-2011 datasets show that TS-CAM outperforms its CNN-CAM counterparts by 7.1%/27.1% for WSOL, achieving state-of-the-art performance. Code is available at https://github.com/vasgaowei/TS-CAM
Wei Gao 0050, Fang Wan 0001, Xingjia Pan, Zhiliang Peng, Qi Tian 0001, Zhenjun Han, Bolei Zhou, Qixiang Ye
ICCV6
2021 Exploiting sample correlation for crowd counting with multi-expert network
abstract
Crowd counting is a difficult task because of the diversity of scenes. Most of the existing crowd counting methods adopt complex structures with massive backbones to enhance the generalization ability. Unfortunately, the performance of existing methods on large-scale data sets is not satisfactory. In order to handle various scenarios with less complex network, we explored how to efficiently use the multi-expert model for crowd counting tasks. We mainly focus on how to train more efficient expert networks and how to choose the most suitable expert. Specifically, we propose a task-driven similarity metric based on sample’s mutual enhancement, referred as co-fine-tune similarity, which can find a more efficient subset of data for training the expert network. Similar samples are considered as a cluster which is used to obtain parameters of an expert. Besides, to make better use of the proposed method, we design a simple network called FPN with Deconvolution Counting Network, which is a more suitable base model for the multi-expert counting network. Experimental results show that multiple experts FDC (MFDC) achieves the best performance on four public data sets, including the large scale NWPU-Crowd data set. Furthermore, the MFDC trained on an extensive dense crowd data set can generalize well on the other data sets without extra training or fine-tuning.1
Xinyan Liu 0008, Guorong Li, Zhenjun Han, Weigang Zhang, Qingming Huang, Nicu Sebe
ICCV3
2021 Effective Fusion Factor in FPN for Tiny Object Detection
abstract
FPN-based detectors have made significant progress in general object detection, e.g., MS COCO and PASCAL VOC. However, these detectors fail in certain application scenarios, e.g., tiny object detection. In this paper, we argue that the top-down connections between adjacent layers in FPN bring two-side influences for tiny object detection, not only positive. We propose a novel concept, fusion factor, to control information that deep layers deliver to shallow layers, for adapting FPN to tiny object detection. After series of experiments and analysis, we explore how to estimate an effective value of fusion factor for a particular dataset by a statistical method. The estimation is dependent on the number of objects distributed in each layer. Comprehensive experiments are conducted on tiny object detection datasets, e.g., TinyPerson and Tiny CityPersons. Our results show that when configuring FPN with a proper fusion factor, the network is able to achieve significant performance gains over the baseline on tiny object detection datasets. Codes and models will be released.
Yuqi Gong, Xuehui Yu, Yao Ding 0006, Xiaoke Peng, Jian Zhao 0006, Zhenjun Han
WACV6
2020 Scale Match for Tiny Person Detection
abstract
Visual object detection has achieved unprecedented advance with the rise of deep convolutional neural networks. However, detecting tiny objects (for example tiny persons less than 20 pixels) in large-scale images remains not well investigated. The extremely small objects raise a grand challenge about feature representation while the massive and complex backgrounds aggregate the risk of false alarms. In this paper, we introduce a new benchmark, referred to as TinyPerson, opening up a promising direction for tiny object detection in a long distance and with massive backgrounds. We experimentally find that the scale mismatch between the dataset for network pre-training and the dataset for detector learning could deteriorate the feature representation and the detectors. Accordingly, we propose a simple yet effective Scale Match approach to align the object scales between the two datasets for favorable tiny-object representation. Experiments show the significant performance gain of our proposed approach over state-of-the-art detectors, and the challenging aspects of TinyPerson related to real-world scenarios. The TinyPerson benchmark and the code for our approach will be publicly available1.
Xuehui Yu, Yuqi Gong, Qixiang Ye, Zhenjun Han
WACV5
2020 Adaptive Discriminative Deep Correlation Filter for Visual Object Tracking
abstract
Correlation filter trackers building on deep convolution neural networks (CNNs) contribute efficient visual object trackers but remain challenged with severe target appearance variations. The reason for this is that CNNs trained for image classification tasks are less discriminative to the dynamic variations of targets and backgrounds. In this paper, we propose an adaptive discriminative deep correlation filter (adaDDCF), which, by incorporating discriminative feature fine-tuning with adaptive appearance modeling, pursues stable object tracking in complex backgrounds. In adaDDCF, a convolutional Fisher discriminative analysis (FDA) layer is implemented for positive and negative instance mining and scene-specific feature learning. A correlation layer is then embedded to learn the correlation response of consecutive frames for target appearance modeling. With an online learning procedure using forward-backward propagation, the FDA layer and the correlation layer are effectively coupled, leading to effective and discriminative fine-tuning for the proposed tracker, which consequently alleviates the target drifting problem. Extensive experiments on the challenging benchmarks OTB2013, OTB2015, and OTB50 demonstrate that the proposed adaDDCF tracker outperforms many state-of-the-art trackers.
Zhenjun Han, Qixiang Ye
IEEE Trans. Circuits Syst. Video Technol.1
2020 CircleNet: Reciprocating Feature Adaptation for Robust Pedestrian Detection
abstract
Pedestrian detection in the wild remains a challenging problem especially when the scene contains significant occlusion and/or low resolution of the pedestrians to be detected. Existing methods are unable to adapt to these difficult cases while maintaining acceptable performance. In this paper we propose a novel feature learning model, referred to as CircleNet, to achieve feature adaptation by mimicking the process humans looking at low resolution and occluded objects: focusing on it again, at a finer scale, if the object can not be identified clearly for the first time. CircleNet is implemented as a set of feature pyramids and uses weight sharing path augmentation for better feature fusion. It targets at reciprocating feature adaptation and iterative object detection using multiple top-down and bottom-up pathways. To take full advantage of the feature adaptation capability in CircleNet, we design an instance decomposition training strategy to focus on detecting pedestrian instances of various resolutions and different occlusion levels in each cycle. Specifically, CircleNet implements feature ensemble with the idea of hard negative boosting in an end-to-end manner. Experiments on two pedestrian detection datasets, Caltech and CityPersons, show that CircleNet improves the performance of occluded and low-resolution pedestrians with significant margins while maintaining good performance on normal instances.
Tianliang Zhang 0003, Zhenjun Han, Huijuan Xu 0001, Baochang Zhang 0001, Qixiang Ye
IEEE Trans. Intell. Transp. Syst.2
2020 Spatial Preserved Graph Convolution Networks for Person Re-identification
abstract
Person Re-identification is a very challenging task due to inter-class ambiguity caused by similar appearances, and large intra-class diversity caused by viewpoints, illuminations, and poses. To address these challenges, in this article, a graph convolution network based model for person re-identification is proposed to learn more discriminative feature embeddings, where a graph-structured relationship between person images and person parts are together integrated. Graph convolution networks extract common characteristics of the same person, while pyramid feature embedding exploits parts relations and learns stable representation with each person image. We achieve a very competitive performance respectively on three widely used datasets, indicating that the proposed approach significantly outperforms the baseline methods and achieves the state-of-the-art performance.
Zhaoju Li, Zongwei Zhou, Zhenjun Han, Junliang Xing, Jianbin Jiao
ACM Trans. Multim. Comput. Commun. Appl.4
2019 Min-Entropy Latent Model for Weakly Supervised Object Detection
abstract
Weakly supervised object detection is a challenging task when provided with image category supervision but required to learn, at the same time, object locations and object detectors. The inconsistency between the weak supervision and learning objectives introduces significant randomness to object locations and ambiguity to detectors. In this paper, a min-entropy latent model (MELM) is proposed for weakly supervised object detection. Min-entropy serves as a model to learn object locations and a metric to measure the randomness of object localization during learning. It aims to principally reduce the variance of learned instances and alleviate the ambiguity of detectors. MELM is decomposed into three components including proposal clique partition, object clique discovery, and object localization. MELM is optimized with a recurrent learning algorithm, which leverages continuation optimization to solve the challenging non-convexity problem. Experiments demonstrate that MELM significantly improves the performance of weakly supervised object detection, weakly supervised object localization, and image classification, against the state-of-the-art approaches.
Fang Wan 0001, Pengxu Wei, Zhenjun Han, Jianbin Jiao, Qixiang Ye
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 High performance person re-identification via a boosting ranking ensemble
Zhaoju Li, Zhenjun Han, Junliang Xing, Qixiang Ye, Xuehui Yu, Jianbin Jiao
Pattern Recognit.2
2018 Min-Entropy Latent Model for Weakly Supervised Object Detection
abstract
Weakly supervised object detection is a challenging task when provided with image category supervision but required to learn, at the same time, object locations and object detectors. The inconsistency between the weak supervision and learning objectives introduces randomness to object locations and ambiguity to detectors. In this paper, a min-entropy latent model (MELM) is proposed for weakly supervised object detection. Min-entropy is used as a metric to measure the randomness of object localization during learning, as well as serving as a model to learn object locations. It aims to principally reduce the variance of positive instances and alleviate the ambiguity of detectors. MELM is deployed as two sub-models, which respectively discovers and localizes objects by minimizing the global and local entropy. MELM is unified with feature learning and optimized with a recurrent learning algorithm, which progressively transfers the weak supervision to object locations. Experiments demonstrate that MELM significantly improves the performance of weakly supervised detection, weakly supervised localization, and image classification, against the state-of-the-art approaches.
Fang Wan 0001, Pengxu Wei, Jianbin Jiao, Zhenjun Han, Qixiang Ye
CVPR4
2017 A scalable convolutional neural network for task-specified scenarios via knowledge distillation
abstract
In this paper, we explore the redundancy in convolutional neural network, which scales with the complexity of vision tasks. Considering that many front-end visual systems are interested in only a limited range of visual targets, the removing of task-specified network redundancy can promote a wide range of potential applications. We propose a task-specified knowledge distillation algorithm to derive a simplified model with pre-set computation cost and minimized accuracy loss, which suits the resource constraint front-end systems well. Experiments on the MNIST and CIFAR10 datasets demonstrate the feasibility of the proposed approach as well as the existence of task-specified redundancy.
Qixiang Ye, Zhenjun Han, Jianbin Jiao
ICASSP4
2017 Unsupervised person re-identification via re-ranking enhanced sample-specific metric learning
abstract
Despite of the great progress of image-based person re-identification, most existing methods use supervised metric learning to build re-identification models and thus require repeated human effort to annotate sample pairs from non-overlapping cameras. In this paper, we propose an unsupervised sample-specific metric learning approach (SSML) to alleviate this problem. Specifically, using samples those are negatives (with a high probability) to the query samples, we train a local metric for each query sample following the max-margin learning theory. Moreover, a KNN intersection re-ranking (KIRR) method is used to further decrease the ambiguity of samples and aggregate the re-identification performance. With experiments on three widely used person re-identification datasets: VIPeR, CUHK01, and PRID, we demonstrate that the proposed approach is simple but effective.
Zhenjun Han, Zhaoju Li
ICIP2
2017 Beyond Group: Multiple Person Tracking via Minimal Topology-Energy-Variation
abstract
Tracking multiple persons is a challenging task when persons move in groups and occlude each other. Existing group-based methods have extensively investigated how to make group division more accurately in a tracking-by-detection framework; however, few of them quantify the group dynamics from the perspective of targets' spatial topology or consider the group in a dynamic view. Inspired by the sociological properties of pedestrians, we propose a novel socio-topology model with a topology-energy function to factor the group dynamics of moving persons and groups. In this model, minimizing the topology-energy-variance in a two-level energy form is expected to produce smooth topology transitions, stable group tracking, and accurate target association. To search for the strong minimum in energy variation, we design the discrete group-tracklet jump moves embedded in the gradient descent method, which ensures that the moves reduce the energy variation of group and trajectory alternately in the varying topology dimension. Experimental results on both RGB and RGB-D data sets show the superiority of our proposed model for multiple person tracking in crowd scenes.
Shan Gao 0003, Qixiang Ye, Junliang Xing, Arjan Kuijper, Zhenjun Han, Jianbin Jiao, Xiangyang Ji
IEEE Trans. Image Process.5
2016 Person re-identification via adaboost ranking ensemble
abstract
Matching specific persons across scenes, known as person re-identification, is an important yet unsolved computer vision problem. Feature representation and metric learning are two fundamental factors in person re-identification. However, current person re-identification methods, which use single handcrafted feature with corresponding metric, could be not powerful enough when facing illumination, viewpoint and pose variations. Thus it inevitably produces suboptimal ranking lists. In this paper, we propose incorporating multiple features with metrics to build weak learners, and aggregate the base ranking lists by AdaBoost Ranking. Experiments on two commonly used datasets, VIPeR and CUHK01, show that our proposed approach greatly improves recognition rates over the state-of-the-art methods.
Zhaoju Li, Zhenjun Han, Qixiang Ye
ICIP2
2016 Weakly supervised object detection with correlation and part suppression
abstract
In weakly supervised object detection, conventional methods treat object location in each image as a latent variable and use non-convex optimization to solve the latent variable. However, as the optimization objective is image-level instead of sample-level, the learning procedure tends to choose object parts as false positive samples. Furthermore, when multiple classes of objects appear in the same images, the models could invite class-correlations and lose discriminative capability. In this paper, we propose a simple but effective suppression strategy that mines hard negative samples in the learning procedure to ease the above problems. We propose using a spatial-voting strategy to help finding negative samples to suppress the impact of object parts. We also use regions from class-correlated images as negative samples to suppress the impact of class-correlations. Experiments show that our approach significantly improves the baseline by 6% and achieves state-of-the-art performance.
Fang Wan 0001, Pengxu Wei, Zhenjun Han, Kun Fu 0001, Qixiang Ye
ICIP3
2015 Real-Time Multipedestrian Tracking in Traffic Scenes via an RGB-D-Based Layered Graph Model
abstract
Multipedestrian tracking in traffic scenes is challenging due to cluttered backgrounds and serious occlusions. In this paper, we propose a layered graph model in image (RGB) and depth (D) domains for real-time robust multipedestrian tracking. The motivation is to investigate high-level constraints in RGB-D data association and to improve the optimization from the trajectory level to the layer level. To construct a layered graph, we define constraints in the depth domain so that pedestrian objects in the image domain are assigned to proper layers. We use pedestrian detection responses in the RGB domain as graph nodes, and we integrate 3-D motion, appearance, and depth features as graph edges. An online updating depth factor is defined to describe the depth relationships among the observations in and out of the layers, and the occlusion issue is processed with an analytical layer-level strategy. With a heuristic label switching algorithm, multiple pedestrian objects are optimally associated and tracked. Experiments and comparison on five public data sets show that our proposed approach significantly reduces pedestrian's ID switch and improves tracking accuracy in the cases of serious occlusions.
Shan Gao 0003, Zhenjun Han, Ce Li 0005, Qixiang Ye, Jianbin Jiao
IEEE Trans. Intell. Transp. Syst.2
2014 Depth Structure Association for RGB-D Multi-target Tracking
abstract
Multi-target tracking in outdoor scenes plays an important role in many computer vision applications. Most previous work on visual information based multi-target tracking does not incorporate depth information and the absence of depth information often leads to mismatching or tracking failures. In this paper, we propose a Depth Structure Association (DSA) approach for RGB-D data based multi-target tracking. DSA encodes depth information in a chain structure, the structure is used by DSA together with appearance and motion information to address object occlusion issues in outdoor scenes. Additionally, the use of DSA has the advantages of regulating a much smaller solution space, greatly reducing the computational complexity. Experimental results on three datasets demonstrate that our DSA approach can significantly reduce object mismatch and tracking failure for long term occlusions.
Shan Gao 0003, Zhenjun Han, David S. Doermann, Jianbin Jiao
ICPR2
2014 Locality-Constrained Sparse Reconstruction for Trajectory Classification
abstract
Trajectory classification has been extensively investigated in recent years, however, problems remain when processing incomplete trajectories of noises and local variations. In this paper, we propose a Locality-constrained Sparse Reconstruction (LSR) approach that explores both sparsity and local adaptability for robust trajectory classification. A trajectory dictionary with locality constrains is constructed with track lets partitioned from collected trajectories by control points of cubic B-spline curves. On the dictionary, the proposed LSR is used to calculate a discriminate code matrix. Then, a loss weighted decoding strategy is employed to perform multi-class trajectory classification. In addition, the approach can be used for anomalous trajectory detection with a thresholding strategy. Experiments on two datasets show that the results of the LSR approach improve the state of the art.
Ce Li 0005, Zhenjun Han, Qixiang Ye, Shan Gao 0003, Lijin Pang, Jianbin Jiao
ICPR2
2013 Robust Visual Object Tracking via Sparse Representation and Reconstruction
Zhenjun Han, Qixiang Ye, Jianbin Jiao
CAIP (2)1
2013 Visual abnormal behavior detection based on trajectory sparse reconstruction analysis
Ce Li 0005, Zhenjun Han, Qixiang Ye, Jianbin Jiao
Neurocomputing2
2013 Human Detection in Images via Piecewise Linear Support Vector Machines
abstract
Human detection in images is challenged by the view and posture variation problem. In this paper, we propose a piecewise linear support vector machine (PL-SVM) method to tackle this problem. The motivation is to exploit the piecewise discriminative function to construct a nonlinear classification boundary that can discriminate multiview and multiposture human bodies from the backgrounds in a high-dimensional feature space. A PL-SVM training is designed as an iterative procedure of feature space division and linear SVM training, aiming at the margin maximization of local linear SVMs. Each piecewise SVM model is responsible for a subspace, corresponding to a human cluster of a special view or posture. In the PL-SVM, a cascaded detector is proposed with block orientation features and a histogram of oriented gradient features. Extensive experiments show that compared with several recent SVM methods, our method reaches the state of the art in both detection accuracy and computational efficiency, and it performs best when dealing with low-resolution human regions in clutter backgrounds.
Qixiang Ye, Zhenjun Han, Jianbin Jiao, Jianzhuang Liu
IEEE Trans. Image Process.2
2011 Abnormal Behavior Detection via Sparse Reconstruction Analysis of Trajectory
abstract
This paper proposes a new method for abnormal behavior detection in surveillance videos via sparse reconstruction analysis. The motion trajectories of objects are firstly defined as fixed-length parametric vectors based on approximating cubic B-spline curves. Then the vectors are classified as behavior patterns and finally distinguished between normal and abnormal behaviors based on sparse reconstruction analysis, in which a classifier is constructed with sparse linear reconstruction coefficients by computing L1-norm minimization and sparse reconstruction residuals learning from labeled training samples. Experimental results on public dataset show the effectiveness of the proposed approach.
Ce Li 0005, Zhenjun Han, Qixiang Ye, Jianbin Jiao
ICIG2
2011 Fast Pedestrian Detection with Laser and Image Data Fusion
abstract
In this paper, we proposed a pedestrian detection system based on laser and image data fusion. The high speed of laser data based location and precise of image based classification are fully explored. First, laser scanner point data is clustered into segments, each of which implies a pedestrian candidate. Then, the segments are projected to the image domain to form regions of interest (ROI) on the image, given camera calibration parameters. Finally two SVM classifiers on Histogram of Oriented Gradient (HOG) features are used to precisely locate pedestrians on the ROI. Experiments report over 30 times higher speed than the state-of-the-art method and a comparable detection rate.
Jixiang Liang, Qixiang Ye, Zhenjun Han, Jianbin Jiao
ICIG4
2011 A fast object tracking approach based on sparse representation
abstract
This paper proposes a new approach based on object sparse representation (OSR) for object tracking. The OSR method implemented by L1-norm minimization is robust to the partial occlusion and deterioration in object images. Firstly, we dynamically construct a set of samples in a predicted searching window in a new video frame, on which the sparse representation of the tracked object can be calculated by the OSR method. This procedure can automatically select the subset of the samples as a basis which most compactly expresses the object with small residuals and rejects all other possible but less compact representations. In terms of this sparse and compact representation, the instantaneous tracking result is achieved in the new video frame. Extensive comparative experiments demonstrate the effectiveness of the proposed approach especially in occlusion context.
Zhenjun Han, Jianbin Jiao, Qixiang Ye
ICIP1
2011 Combined feature evaluation for adaptive visual object tracking
Zhenjun Han, Qixiang Ye, Jianbin Jiao
Comput. Vis. Image Underst.1
2011 Visual object tracking via sample-based Adaptive Sparse Representation (AdaSR)
Zhenjun Han, Jianbin Jiao, Baochang Zhang 0001, Qixiang Ye, Jianzhuang Liu
Pattern Recognit.1
2008 Online feature evaluation for object tracking using Kalman Filter
abstract
An online feature evaluation method for visual object tracking is put forward in this paper. Firstly, a combined feature set is built using color histogram (HC) bins and gradient orientation histogram (HOG) bins considering the color and contour representation of an object respectively. Then a novel method is proposed to evaluate the features’ weights in a tracking process using Kalman Filter, which is used to comprise the inter-frame predication and single-frame measurement of features’ discriminative power. In this way, we extend the traditional filter framework from modeling motion states to modeling feature evaluation. Experiments show this method can greatly improve the tracking stabilization when objects go across complex backgrounds.
Zhenjun Han, Qixiang Ye, Jianbin Jiao
ICPR1