Chunluan Zhou

dblp:127/5662 · DBLP profile ↗
← Back
28ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0003-0284-6256ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 19 · 8 first-author · 7 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EBPersons: A Dataset for Person Detection at the Edges of Buildings
abstract
With the increasing prevalence of buildings, incidents of falling from heights have become more and more frequent. Accurately detecting individuals at the edges of buildings through surveillance videos is crucial for timely intervention and accident prevention. However, this task, termed Person Detection at the Edges of Buildings (PDEB), presents significant challenges including variations in lighting conditions, occlusions, and small size of person instances. Existing person detection datasets are inadequate for PDEB due to domain gaps. To address this issue, we construct EBPersons, a completely new dataset specifically designed for PDEB. Comprising 1,314 videos captured across over 300 diverse building scenes with diverse lighting conditions, EBPersons provides a rich and challenging benchmark for PDEB research. Furthermore, we propose a baseline method specifically designed for PDEB, named STASH, which includes three key components: a Scale Match strategy to improve small object detection, a Temporal ROI Align Operator to leverage temporal context, and a Sequential-level Semantics Aggregation head to enhance feature representation. Extensive experiments are conducted on EBPersons to compare our method with other detectors, including generic object detectors, pedestrian detectors, and video object detectors. The results demonstrate the superior performance of the proposed STASH, providing a strong baseline for future research on PDEB. Our EBPer sons dataset and the baseline code are publicly available at https://ebpersons.github.io/.
Zitao Gao, Bing Qu, Chunluan Zhou, Junsong Yuan 0001, Zhigang Tu 0001
IEEE Trans. Multim.3
2025 SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories
abstract
While MLLMs have demonstrated impressive image understanding capabilities, they still struggle with pixel-level comprehension, limiting their practical applications. Current evaluation tasks such as VQA and visual grounding remain too coarse to assess fine-grained pixel comprehension accurately. Although segmentation is foundational for pixel-level understanding, existing methods often require MLLMs to generate implicit tokens, decoded through external pixel decoders. This approach disrupts the MLLM’s text output space, potentially compromising language capabilities and reducing flexibility and extensibility while failing to reflect the model’s intrinsic pixel-level understanding. Thus, we introduce the Human-Like Mask Annotation Task (HLMAT), a new paradigm where MLLMs mimic human annotators using interactive segmentation tools. Modelling segmentation as a multi-step Markov Decision Process, HLMAT enables MLLMs to iteratively generate text-based click points, achieving high-quality masks without architectural changes or implicit tokens. Through this setup, we develop SegAgent, a model fine-tuned on human-like annotation trajectories, which achieves performance comparable to SoTA methods and supports additional tasks like mask refinement and annotation filtering. HLMAT provides a protocol for assessing fine-grained pixel understanding in MLLMs and introduces a vision-centric, multi-step decision-making task that facilitates the exploration of MLLMs’ visual reasoning abilities. Our adaptations of policy improvement method StaR and PRM guided tree search further enhance model robustness in complex segmentation tasks, laying a foundation for future advancements in fine-grained visual perception and multi-step decision-making for MLLMs. Code can be found at https://github.com/aim-uofa/SegAgent.
Muzhi Zhu, Yuzhuo Tian, Hao Chen 0041, Chunluan Zhou, Qingpei Guo, Yang Liu 0357, Ming Yang 0007, Chunhua Shen
CVPR4
2025 FADE: A Dataset for Detecting Falling Objects Around Buildings in Video
abstract
Objects falling from buildings, a frequently occurring event in daily life, can cause severe injuries to pedestrians due to the high impact force they exert. Surveillance cameras are often installed around buildings to detect falling objects, but such detection remains challenging due to the small size and fast motion of the objects. Moreover, the field of falling object detection around buildings (FODB) lacks a large-scale dataset for training learning-based detection methods and for standardized evaluation. To address these challenges, we propose a large and diverse video benchmark dataset named FADE. Specifically, FADE contains 2,611 videos from 25 scenes, featuring 8 falling object categories, 4 weather conditions, and 4 video resolutions. Additionally, we develop a novel detection method for FODB that effectively leverages motion information and generates small-sized yet high-quality detection proposals. The efficacy of our method is evaluated on the proposed FADE dataset by comparing it with state-of-the-art approaches in generic object detection, video object detection, and moving object detection. The dataset and code are publicly available at https://fadedataset.github.io/FADE.github.io/.
Zhigang Tu 0001, Zhengbo Zhang, Zitao Gao, Chunluan Zhou, Junsong Yuan 0001, Bo Du 0001
IEEE Trans. Inf. Forensics Secur.4
2024 EVE: Efficient Zero-Shot Text-Based Video Editing With Depth Map Guidance and Temporal Consistency Constraints
Xingning Dong, Tian Gan 0002, Chunluan Zhou, Ming Yang 0007, Qingpei Guo
IJCAI4
2024 M2-RAAP: A Multi-Modal Recipe for Advancing Adaptation-based Pre-training towards Effective and Efficient Zero-shot Video-text Retrieval
abstract
We present a Recipe for Effective and Efficient zero-shot video-text Retrieval, dubbed M2-RAAP. Upon popular image-text models like CLIP, most current adaptation-based video-text pre-training methods are confronted by three major issues, i.e., noisy data corpus, time-consuming pre-training, and limited performance gain. Towards this end, we conduct a comprehensive study including four critical steps in video-text pre-training. Specifically, we investigate 1) data filtering and refinement, 2) video input type selection, 3) temporal modeling, and 4) video feature enhancement. We then summarize this empirical study into the M2-RAAP recipe, where our technical contributions lie in 1) the data filtering and text re-writing pipeline resulting in 1M high-quality bilingual video-text pairs, 2) the promotion of video inputs with key-frames to accelerate pre-training, and 3) the Auxiliary-Caption-Guided (ACG) strategy to enhance video features. We conduct extensive experiments by adapting three image-text foundation models on two refined video-text datasets from different languages, validating the robustness and reproducibility of M2-RAAP for adaptation-based pre-training. Results demonstrate that M2-RAAP yields superior performance with significantly less data (-90%) and time consumption (-95%), establishing a new SOTA on four English zero-shot retrieval datasets and two Chinese ones. Codebase and refined bilingual data annotations are available at https://github.com/alipay/Ant-Multi-Modal-Framework/tree/main/prj/M2_RAAP.
Xingning Dong, Zipeng Feng, Chunluan Zhou, Xuzheng Yu, Ming Yang 0007, Qingpei Guo
SIGIR3
2023 Generalized Relation Modeling for Transformer Tracking
abstract
Compared with previous two-stream trackers, the recent one-stream tracking pipeline, which allows earlier interaction between the template and search region, has achieved a remarkable performance gain. However, existing one-stream trackers always let the template interact with all parts inside the search region throughout all the encoder layers. This could potentially lead to target-background confusion when the extracted feature representations are not sufficiently discriminative. To alleviate this issue, we propose a generalized relation modeling method based on adaptive token division. The proposed method is a generalized formulation of attention-based relation modeling for Transformer tracking, which inherits the merits of both previous two-stream and one-stream pipelines whilst enabling more flexible relation modeling by selecting appropriate search tokens to interact with template tokens. An attention masking strategy and the Gumbel-Softmax technique are introduced to facilitate the parallel computation and end-to-end learning of the token division module. Extensive experiments show that our method is superior to the two-stream and one-stream pipelines and achieves state-of-the-art performance on six challenging benchmarks with a real-time running speed. Code and models are publicly available at https://github.com/Little-Podi/GRM.
Shenyuan Gao, Chunluan Zhou, Jun Zhang 0004
CVPR2
2023 SOAR: Scene-debiasing Open-set Action Recognition
abstract
Deep learning models have a risk of utilizing spurious clues to make predictions, such as recognizing actions based on the background scene. This issue can severely degrade the open-set action recognition performance when the testing samples have different scene distributions from the training samples. To mitigate this problem, we propose a novel method, called Scene-debiasing Open-set Action Recognition (SOAR), which features an adversarial scene reconstruction module and an adaptive adversarial scene classification module. The former prevents the decoder from reconstructing the video background given video features, and thus helps reduce the background information in feature learning. The latter aims to confuse scene type classification given video features, with a specific emphasis on the action foreground, and helps to learn scene-invariant information. In addition, we design an experiment to quantify the scene bias. The results indicate that the current open-set action recognizers are biased toward the scene, and our proposed SOAR method better mitigates such bias. Furthermore, our extensive experiments demonstrate that our method outperforms state-of-the-art methods, and the ablation studies confirm the effectiveness of our proposed modules.
Yuanhao Zhai 0001, Ziyi Liu 0001, Zhenyu Wu 0002, Chunluan Zhou, David S. Doermann, Junsong Yuan 0001, Gang Hua 0001
ICCV5
2023 Cyclic Self-Training With Proposal Weight Modulation for Cross-Supervised Object Detection
abstract
Weakly-supervised object detection (WSOD), which requires only image-level annotations for training detectors, has gained enormous attention. Despite recent rapid advance in WSOD, there remains a large performance gap compared with fully-supervised object detection. To narrow the performance gap, we study cross-supervised object detection (CSOD), where existing classes (base classes) have instance-level annotations while newly added classes (novel classes) only need image-level annotations. For improving localization accuracy, we propose a Cyclic Self-Training (CST) method to introduce instance-level supervision into a commonly used WSOD method, online instance classifier refinement (OICR). Our proposed CST consists of forward pseudo labeling and backward pseudo labeling. Specifically, OICR exploits the forward pseudo labeling to generate pseudo ground-truth bounding-boxes for all classes, thus enabling instance classifier training. Then, the backward pseudo labeling is designed to generate pseudo ground-truth bounding-boxes of higher quality for novel classes by fusing the predictions of the instance classifiers. As a result, both novel and base classes will have bounding-box annotations for training, alleviating the supervision inconsistency between base and novel classes. In the forward pseudo labeling, the generated pseudo ground-truths may be misaligned with objects and thus introduce poor-quality examples for training the ICs. To reduce the impacts of these poor-quality training examples, we propose a Proposal Weight Modulation (PWM) module learned in a class-agnostic and contrastive manner by exploiting bounding-box annotations of base classes. Experiments on PASCAL VOC and MS COCO datasets demonstrate the superiority of our proposed method.
Yunqiu Xu, Chunluan Zhou, Xin Yu 0002, Yi Yang 0001
IEEE Trans. Image Process.2
2022 AiATrack: Attention in Attention for Transformer Visual Tracking
Shenyuan Gao, Chunluan Zhou, Xinggang Wang, Junsong Yuan 0001
ECCV (22)2
2022 Distilling Inter-Class Distance for Semantic Segmentation
abstract
Knowledge distillation is widely adopted in semantic segmentation to reduce the computation cost. The previous knowledge distillation methods for semantic segmentation focus on pixel-wise feature alignment and intra-class feature variation distillation, neglecting to transfer the knowledge of the inter-class distance in the feature space, which is important for semantic segmentation such a pixel-wise classification task. To address this issue, we propose an Inter-class Distance Distillation (IDD) method to transfer the inter-class distance in the feature space from the teacher network to the student network. Furthermore, semantic segmentation is a position-dependent task, thus we exploit a position information distillation module to help the student network encode more position information. Extensive experiments on three popular datasets: Cityscapes, Pascal VOC and ADE20K show that our method is helpful to improve the accuracy of semantic segmentation models and achieves the state-of-the-art performance. E.g. it boosts the benchmark model (``PSPNet+ResNet18") by 7.50% in accuracy on the Cityscapes dataset.
Zhengbo Zhang, Chunluan Zhou, Zhigang Tu 0001
IJCAI2
2021 Learning Dynamics via Graph Neural Networks for Human Pose Estimation and Tracking
abstract
Multi-person pose estimation and tracking serve as crucial steps for video understanding. Most state-of-the-art approaches rely on first estimating poses in each frame and only then implementing data association and refinement. Despite the promising results achieved, such a strategy is inevitably prone to missed detections especially in heavily-cluttered scenes, since this tracking-by-detection paradigm is, by nature, largely dependent on visual evidences that are absent in the case of occlusion. In this paper, we propose a novel online approach to learning the pose dynamics, which are independent of pose detections in current fame, and hence may serve as a robust estimation even in challenging scenarios including occlusion. Specifically, we derive this prediction of dynamics through a graph neural network (GNN) that explicitly accounts for both spatial-temporal and visual information. It takes as input the historical pose tracklets and directly predicts the corresponding poses in the following frame for each tracklet. The predicted poses will then be aggregated with the detected poses, if any, at the same frame so as to produce the final pose, potentially recovering the occluded joints missed by the estimator. Experiments on PoseTrack 2017 and Pose-Track 2018 datasets demonstrate that the proposed method achieves results superior to the state of the art on both human pose estimation and tracking tasks.
Yiding Yang, Zhou Ren, Chunluan Zhou, Xinchao Wang, Gang Hua 0001
CVPR4
2021 Pyramidal Multiple Instance Detection Network With Mask Guided Self-Correction for Weakly Supervised Object Detection
abstract
Weakly supervised object detection has attracted more and more attention as it only needs image-level annotations for training object detectors. A popular solution to this task is to train a multiple instance detection network (MIDN) which integrates multiple instance learning into a deep convolutional neural network. One major issue of the MIDN is that it is prone to be stuck at local discriminative regions. To address this local optimum issue, we propose a pyramidal MIDN (P-MIDN) comprised of a sequence of multiple MIDNs. In particular, one MIDN performs proposal removal for its subsequent MIDN to reduce the exposure of local discriminative proposal regions to the latter during training. In this manner, it allows our MIDNs to focus on proposals which cover objects more completely. Furthermore, we integrate the P-MIDN into an online instance classifier refinement (OICR) framework. Combined with the P-MIDN, a mask guided self-correction (MGSC) method is proposed to generate high-quality pseudo ground-truths for training the OICR. Experimental results on PASCAL VOC 2007, PASCAL VOC 2010, PASCAL VOC 2012, ILSVRC 2013 DET and MS-COCO benchmarks demonstrate that our approach achieves state-of-the-art performance.
Yunqiu Xu, Chunluan Zhou, Xin Yu 0002, Bin Xiao 0002, Yi Yang 0001
IEEE Trans. Image Process.2
2020 Temporal-Context Enhanced Detection of Heavily Occluded Pedestrians
abstract
State-of-the-art pedestrian detectors have performed promisingly on non-occluded pedestrians, yet they are still confronted by heavy occlusions. Although many previous works have attempted to alleviate the pedestrian occlusion issue, most of them rest on still images. In this paper, we exploit the local temporal context of pedestrians in videos and propose a tube feature aggregation network (TFAN) aiming at enhancing pedestrian detectors against severe occlusions. Specifically, for an occluded pedestrian in the current frame, we iteratively search for its relevant counterparts along temporal axis to form a tube. Then, features from the tube are aggregated according to an adaptive weight to enhance the feature representations of the occluded pedestrian. Furthermore, we devise a temporally discriminative embedding module (TDEM) and a part-based relation module (PRM), respectively, which adapts our approach to better handle tube drifting and heavy occlusions. Extensive experiments are conducted on three datasets, Caltech, NightOwls and KAIST, showing that our proposed method is significantly effective for heavily occluded pedestrian detection. Moreover, we achieve the state-of-the-art performance on the Caltech and NightOwls datasets.
Jialian Wu, Chunluan Zhou, Ming Yang 0007, Qian Zhang 0009, Junsong Yuan 0001
CVPR2
2020 Temporal Keypoint Matching and Refinement Network for Pose Estimation and Tracking
Chunluan Zhou, Zhou Ren, Gang Hua 0001
ECCV (22)1
2020 Self-Mimic Learning for Small-scale Pedestrian Detection
abstract
Detecting small-scale pedestrians is one of the most challenging problems in pedestrian detection. Due to the lack of visual details, the representations of small-scale pedestrians tend to be weak to be distinguished from background clutters. In this paper, we conduct an in-depth analysis of the small-scale pedestrian detection problem, which reveals that weak representations of small-scale pedestrians are the main cause for a classifier to miss them. To address this issue, we propose a novel Self-Mimic Learning (SML) method to improve the detection performance on small-scale pedestrians. We enhance the representations of small-scale pedestrians by mimicking the rich representations from large-scale pedestrians. Specifically, we design a mimic loss to force the feature representations of small-scale pedestrians to approach those of large-scale pedestrians. The proposed SML is a general component that can be readily incorporated into both one-stage and two-stage detectors, with no additional network layers and incurring no extra computational cost during inference. Extensive experiments on both the CityPersons and Caltech datasets show that the detector trained with the mimic loss is significantly effective for small-scale pedestrian detection and achieves state-of-the-art results on CityPersons and Caltech, respectively.
Jialian Wu, Chunluan Zhou, Qian Zhang 0009, Ming Yang 0007, Junsong Yuan 0001
ACM Multimedia2
2020 Detecting spatiotemporal irregularities in videos via a 3D convolutional autoencoder
Mengjia Yan 0003, Jingjing Meng, Chunluan Zhou, Zhigang Tu 0001, Yap-Peng Tan, Junsong Yuan 0001
J. Vis. Commun. Image Represent.3
2020 Occlusion Pattern Discovery for Object Detection and Occlusion Reasoning
abstract
Despite recent progress of object category detection in real scenes, detecting objects that are partially or heavily occluded remains a challenging problem due to the uncertainty and diversity of occlusion situations which could cause large intra-category appearance variance. To learn these occlusion situations, we propose a novel approach to discover occlusion patterns that cannot only boost occluded object detection but also provide occlusion reasoning. Our approach is based on a classic deformable part model (DPM) trained on fully observed object examples. Each occlusion pattern contains only a subset of visible parts, thus the total number of occlusion patterns are exponential to the number of parts, i.e., m parts will generate 2mocclusion patterns to compose an occlusion pattern pool. From this occlusion pattern pool, we look for a small group of occlusion patterns that are: (1) representative patterns that can well explain training examples and (2) discriminative patterns that have high detection performance individually. To select such occlusion patterns, we formulate occlusion pattern discovery as a facility location problem, which can be solved effectively by greedy search. The discovered occlusion patterns are themselves DPMs and can be used as object detectors when properly tuned. They can also be combined with the state-of-the-art detectors (e.g. Faster R-CNN) for improving detection performance and achieving part-level occlusion reasoning. The effectiveness of the proposed approach is validated on Pascal VOC2007 and VOC2010 datasets.
Chunluan Zhou, Junsong Yuan 0001
IEEE Trans. Circuits Syst. Video Technol.1
2019 Discriminative Feature Transformation for Occluded Pedestrian Detection
abstract
Despite promising performance achieved by deep con- volutional neural networks for non-occluded pedestrian de- tection, it remains a great challenge to detect partially oc- cluded pedestrians. Compared with non-occluded pedes- trian examples, it is generally more difficult to distinguish occluded pedestrian examples from background in featue space due to the missing of occluded parts. In this paper, we propose a discriminative feature transformation which en- forces feature separability of pedestrian and non-pedestrian examples to handle occlusions for pedestrian detection. Specifically, in feature space it makes pedestrian exam- ples approach the centroid of easily classified non-occluded pedestrian examples and pushes non-pedestrian examples close to the centroid of easily classified non-pedestrian ex- amples. Such a feature transformation partially compen- sates the missing contribution of occluded parts in feature space, therefore improving the performance for occluded pedestrian detection. We implement our approach in the Fast R-CNN framework by adding one transformation net- work branch. We validate the proposed approach on two widely used pedestrian detection datasets: Caltech and CityPersons. Experimental results show that our approach achieves promising performance for both non-occluded and occluded pedestrian detection.
Chunluan Zhou, Ming Yang 0007, Junsong Yuan 0001
ICCV1
2019 Attention to Head Locations for Crowd Counting
Youmei Zhang, Chunluan Zhou, Faliang Chang, Alex Chichung Kot, Wei Zhang 0021
ICIG (2)2
2019 Multi-resolution attention convolutional neural network for crowd counting
abstract
Estimating crowd counts remains a challenging task due to the problems of scale variations, non-uniform distribution and complex backgrounds. In this paper, we propose a multi-resolution attention convolutional neural network (MRA-CNN) to address this challenging task. Except for the counting task, we exploit an additional density-level classification task during training and combine features learned for the two tasks, thus forming multi-scale, multi-contextual features to cope with the scale variation and non-uniform distribution. Besides, we utilize a multi-resolution attention (MRA) model to generate score maps, where head locations are with higher scores to guide the network to focus on head regions and suppress non-head regions regardless of the complex backgrounds. During the generation of score maps, atrous convolution layers are used to expand the receptive field with fewer parameters, thus getting higher-level features and providing the MRA model more comprehensive information. Experiments on ShanghaiTech, WorldExpo’10 and UCF datasets demonstrate the effectiveness of our method.
Youmei Zhang, Chunluan Zhou, Faliang Chang, Alex Chichung Kot
Neurocomputing2
2019 A scale adaptive network for crowd counting
Youmei Zhang, Chunluan Zhou, Faliang Chang, Alex Chichung Kot
Neurocomputing2
2019 Multi-label learning of part detectors for occluded pedestrian detection
Chunluan Zhou, Junsong Yuan 0001
Pattern Recognit.1
2018 Actor-Action Semantic Segmentation with Region Masks
Kang Dang, Chunluan Zhou, Zhigang Tu 0001, Michael Hoy, Justin Dauwels, Junsong Yuan 0001
BMVC2
2018 Bi-box Regression for Pedestrian Detection and Occlusion Estimation
Chunluan Zhou, Junsong Yuan 0001
ECCV (1)1
2017 Multi-label Learning of Part Detectors for Heavily Occluded Pedestrian Detection
abstract
Detecting pedestrians that are partially occluded remains a challenging problem due to variations and uncertainties of partial occlusion patterns. Following a commonly used framework of handling partial occlusions by part detection, we propose a multi-label learning approach to jointly learn part detectors to capture partial occlusion patterns. The part detectors share a set of decision trees via boosting to exploit part correlations and also reduce the computational cost of applying these part detectors. The learned decision trees capture the overall distribution of all the parts. When used as a pedestrian detector individually, our part detectors learned jointly show better performance than their counterparts learned separately in different occlusion situations. The learned part detectors can be further integrated to better detect partially occluded pedestrians. Experiments on the Caltech dataset show state-of-the-art performance of our approach for detecting heavily occluded pedestrians.
Chunluan Zhou, Junsong Yuan 0001
ICCV1
2016 Learning to Integrate Occlusion-Specific Detectors for Heavily Occluded Pedestrian Detection
Chunluan Zhou, Junsong Yuan 0001
ACCV (2)1
2014 Non-rectangular Part Discovery for Object Detection
Chunluan Zhou, Junsong Yuan 0001
BMVC1
2012 Arbitrary-Shape Object Localization Using Adaptive Image Grids
Chunluan Zhou, Junsong Yuan 0001
ACCV (1)1