Wenhui Jiang 0001

dblp:179/1020-1 · DBLP profile ↗
← Back
28ranked-venue papers
10as first author
24since 2021 · last 2026
0000-0002-4144-6725ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 16 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021
YearPublicationVenuePosition
2026 DSRAS: Dual-Stage Reasoning and Answer Selection for Video-Text Visual Question Answering
abstract
Video text-based visual question answering (Video TextVQA) aims to answer questions by spatio-temporal joint reasoning over textual and visual information in a video. Existing methods have achieved remarkable progress using uniform sampling strategies. However, uniform frame sampling may introduce noisy frames and miss keyframes. Meanwhile, current methods treat all video frames equally, which is suboptimal, especially since video text question answering tasks primarily target questions that include text within the video. To address the aforementioned issues, we introduce a novel Dual-Stage Reasoning and Answer Selection (DSRAS) model, which not only adaptively focuses on and extracts keyframes, but also significantly enhances attention to video frames containing text through an answer selection mechanism. Specifically, we propose a Dual-Stage Reasoning (DSR) module to achieve adaptive frame selection. Then, we introduce an Answer Selection Module (ASM) to guide our model to focus on keyframes containing textual information. Extensive experiments demonstrate that our model outperforms existing approaches on the RoadTextVQA and M4-ViteVQA datasets.
Chengyang Fang, Xiankun Wan, Wenhui Jiang 0001, Yuming Fang 0001
IEEE Signal Process. Lett.3
2026 Text-Conditional Visual-Language Alignment for Video Captioning
abstract
Video captioning remains a challenging task due to the diverse video content and the complex relationships between visual and textual elements. Recent efforts predominantly focus on multimodal architecture designs trained with paired video-caption data. Nonetheless, the learning paradigm suffers from the “one-to-many” corresponding problem, since one source video is mapped to multiple caption annotations. The difficulty of video captioning is further exacerbated by the poor-written captions, which mislead the captioner with irrelevant information. Essentially, the problem stems from the inadequate alignment between video and caption. In this work, we propose a Text-Conditional Alignment Transformer, which fully exploits the rich information provided by diverse labeled captions, and avoids the impacts of label ambiguity and noise. To alleviate the challenge of the “one-to-many” correspondence, we introduce Text-conditioned Video Encoding, which diversifies the video representation by emphasizing the spatial-temporal visual areas relevant to the given descriptions while filtering out redundant visual information. The refined video representation is well-aligned to match the corresponding text description, and naturally converts the “one-to-many” mapping to “one-to-one” mapping. To deal with the noisy annotations, we propose Quality-aware Caption Decoding. We first dynamically measure the qualities of different captions corresponding to the same video in a reference-free manner. Then the estimated qualities are further utilized as auxiliary signals, guiding the model to perform quality-aligned learning from noisy captions. We conduct extensive experiments on MSR-VTT, MSVD, VATEX and ActivityNet-Entities datasets, and demonstrate their consistent performance improvements compared to state-of-the-arts.
Wenhui Jiang 0001, Wenbin Guan, Zhizhen Li, Yuming Fang 0001, Yuxin Peng 0001, Yang Liu 0293
IEEE Trans. Circuits Syst. Video Technol.1
2025 What Happens in the Surroundings: A Benchmark for 360° image Captioning
abstract
Image captioning has been widely studied by the computer vision and natural language processing communities. However, conventional image captioning models are mainly built upon 2D images with narrow field-of-views. To comprehensively analyze what happens in real scenes, using 360° cameras to capture 360° images has been a popular research trend. However, 360° image captioning is rarely studied due to the lack of related datasets. To bridge the research gap, we introduce a novel dataset for 360° image captioning, namely 360IC (360° Image Captioning). It contains 1250 360° images from rich scenes, and each image is manually labeled with at least three detailed descriptions, which will greatly promote the research of 360° image captioning. We also propose a Multi-View Transformer Network (MVTransNet) for 360° image captioning based on multi-view analysis and fusion, which deals with the characteristics of large resolution, wide field-of-view and complex visual content of 360° images. Specifically, it builds a hierarchical architecture to model the spatial dependency of images with larger content, thus forming rich visual features of 360° images and making the generated descriptions more accurate. Extensive experiments on 360IC show that the proposed network outperforms other competing methods considerably. Our dataset will be released soon.
Wenhui Jiang 0001, Tiancong Xu, Zichen Li, Yuming Fang 0001
IJCNN1
2025 Weak-shot Keypoint Estimation via Keyness and Correspondence Transfer
abstract
Keypoint estimation is a fundamental task in computer vision, but generally requires large-scale annotated data for training. Few-shot and unsupervised keypoint estimation are prevalent economical paradigms, but the former still requires annotations for extensive novel classes while the latter only supports for single class. In this paper, we focus on the task of weak-shot keypoint estimation, where multiple novel classes are learned from unlabeled images with the help of labeled base classes. The key problem is what to transfer from base classes to novel classes, and we propose to transfer keyness and correspondence, which essentially belong to comparing entities and thus are class-agnostic and class-wise transferable. The keyness compares which pixel in the local region is more key, which can guide the keypoints of novel classes to move towards the local maximum (i.e., obtaining keypoints). The correspondence compares whether the two pixels belongs to the same semantic part, which can activate the keypoints of novel classes by reinforcing the consistency between corresponding points on two paired images. By transferring keyness and correspondence, our framework achieves favourable performance for weak-shot keypoint estimation. Extensive experiments and analyses on large-scale benchmark MP-100 demonstrate our effectiveness.
Junjie Chen 0008, Zeyu Luo, Zezheng Liu, Wenhui Jiang 0001, Li Niu 0002, Yuming Fang 0001
NeurIPS4
2025 Opinion-unaware blind stereoscopic image quality assessment: A comprehensive study
Jiebin Yan, Yuming Fang 0001, Xuelin Liu, Wenhui Jiang 0001, Yang Liu 0293
Pattern Recognit.4
2025 Learning Stage-wise Fusion Transformer for light field saliency detection
Wenhui Jiang 0001, Qi Shu, Hongwei Cheng, Yuming Fang 0001, Yifan Zuo 0001
Pattern Recognit. Lett.1
2025 Separate, Locate, and Align: Determine Context Relation of Scene Text From Multiple Perspectives in TextVQA
abstract
Text-based Visual Question Answering (TextVQA) focuses on answering questions about the scene text in images. Most works in this field uses transformer based models to modeling the interaction of question and scene texts which means the scene texts will be treated as a natural language sentence and concatenated in reading order as a part of input. However, they ignore the fact that different from words in natural language sentence which have inherent context relation, the context relation of scene texts in images need to be determined. To tackle this problem, we propose a novel method named Separate, Locate and Align (SLA) that discriminate the context relation of scene texts from semantic, visual and spatial aspects. Specifically, based on scene texts with similar visual information (e.g. background color, font color, font style, etc.) having semantic contextual relations, we propose a Text Semantic Separate (TSS) module to discriminate the semantic relation between different scene texts according to their visual contextual information. Then, we introduce a Spatial Circle Position (SCP) module that helps the model discriminate the spatial relation between different scene texts. Last, we design a Visual Alignment (VA) module to help the model distinguish the visual relationships between different scene texts according to the color distribution differences. Extensive experiments show that our method outperforms existing alternatives on TextVQA and ST-VQA datasets without pre-training tasks.
Chengyang Fang, Wenhui Jiang 0001, Yuming Fang 0001, Yuxin Peng 0001, Yang Liu 0293
IEEE Trans. Circuits Syst. Video Technol.2
2025 Learning Comprehensive Visual Grounding for Video Captioning
abstract
The grounding accuracy of existing video captioners is still behind the expectation. The majority of existing methods perform grounded video captioning on sparse entity annotations. However, grounded captioning models rely on deliberate grounding annotations as supervision, which are relatively hard to obtain. Moreover, the captioning accuracy often suffers from degenerated object appearances on the annotated area such as motion blur and video defocus, and these models seldom consider the complex interactions among entities. In this paper, we propose a comprehensive visual grounding network to improve video captioning, by using inexpensive pseudo annotation while avoiding the need to collect large amounts of manual annotations. Specifically, the network consists of spatial-temporal entity grounding and action grounding. The proposed entity grounding encourages the attention mechanism to focus on informative spatial areas across video frames. The action grounding dynamically associates the verbs to related subjects and the corresponding context, which keeps fine-grained spatial and temporal details for action prediction. Both entity grounding and action grounding are formulated as a unified task guided by a soft grounding supervision. More importantly, the grounding objective is supervised by pseudo annotations automatically produced by a grounding annotation generation module, thus our model can be easily applied to the challenging dataset without any grounding annotation provided. We conduct extensive experiments on three benchmark datasets and demonstrate significant performance improvements of +2.4 CIDEr on MSR-VTT, +4.7 CIDEr on MSVD, and +5.1 CIDEr on ActivityNet-Entities compared to state-of-the-arts.
Wenhui Jiang 0001, Linxin Liu, Yuming Fang 0001, Yibo Cheng, Yuxin Peng 0001, Yang Liu 0293
IEEE Trans. Circuits Syst. Video Technol.1
2025 Omnidirectional Image Quality Captioning: A Large-Scale Database and a New Model
abstract
The fast growing application of omnidirectional images calls for effective approaches for omnidirectional image quality assessment (OIQA). Existing OIQA methods have been developed and tested on homogeneously distorted omnidirectional images, but it is hard to transfer their success directly to the heterogeneously distorted omnidirectional images. In this paper, we conduct the largest study so far on OIQA, where we establish a large-scale database called OIQ-10K containing 10,000 omnidirectional images with both homogeneous and heterogeneous distortions. A comprehensive psychophysical study is elaborated to collect human opinions for each omnidirectional image, together with the spatial distributions (within local regions or globally) of distortions, and the head and eye movements of the subjects. Furthermore, we propose a novel multitask-derived adaptive feature-tailoring OIQA model named IQCaption360, which is capable of generating a quality caption for an omnidirectional image in a manner of textual template. Extensive experiments demonstrate the effectiveness of IQCaption360, which outperforms state-of-the-art methods by a significant margin on the proposed OIQ-10K database. The OIQ-10K database and the related source codes are available at https://github.com/WenJuing/IQCaption360.
Jiebin Yan, Ziwen Tan, Yuming Fang 0001, Junjie Chen 0008, Wenhui Jiang 0001, Zhou Wang 0001
IEEE Trans. Image Process.5
2025 Learning Guided Implicit Depth Function With Scale-Aware Feature Fusion
abstract
Recently, the single image super-resolution based on implicit image function is a hot topic, which learns a universal model for arbitrary upsampling scales. By contrast, color-guided depth map super-resolution is less explored based on implicit function learning. The related research faces three questions. First, is it also necessary and applicable to fuse the depth feature and the color feature in the encoder with continuous upsampling scales? Second, is the scale information in the encoder as important as that in the decoder? Third, how to efficiently and effectively model the affinity of location distance and content similarity within cross domains in the decoder? This paper proposes a transformer-based network to answer the above questions, which includes a depth super-resolution branch and a guidance extraction branch. Specifically, in the encoder, the effective implicit cross transformer is designed to fuse the guidance from the color feature with continuous coordinate mapping. In addition, the unrelated guidance is filtered out by correlation evaluation in the high-dimension feature space. Unlike the scale only introduced in the decoder, this paper additionally embeds the scale into the position encoding and the feed-forward network in the encoder to learn the scale-aware feature representation. In the decoder, the high-resolution depth feature is reconstructed by using the internal prior and the external guidance. The internal prior is implemented by implicit self-attention in the depth super-resolution branch, and the external guidance is exploited via implicit cross-attention between both branches. Finally, the above decoded features are complementary to generate the high-resolution depth map. The sufficient experiments on the synthetic and real datasets for in-distribution and out-of-distribution upsampling scales validate the improved performance. The code and the models are public via https://github.com/NaNRan13/GIDF.
Yifan Zuo 0001, Yuming Fang 0001, Jiebin Yan, Wenhui Jiang 0001, Yuxin Peng 0001, Yan Huang 0023
IEEE Trans. Image Process.7
2024 Comprehensive Visual Grounding for Video Description
abstract
The grounding accuracy of existing video captioners is still behind the expectation. The majority of existing methods perform grounded video captioning on sparse entity annotations, whereas the captioning accuracy often suffers from degenerated object appearances on the annotated area such as motion blur and video defocus. Moreover, these methods seldom consider the complex interactions among entities. In this paper, we propose a comprehensive visual grounding network to improve video captioning, by explicitly linking the entities and actions to the visual clues across the video frames. Specifically, the network consists of spatial-temporal entity grounding and action grounding. The proposed entity grounding encourages the attention mechanism to focus on informative spatial areas across video frames, albeit the entity is annotated in only one frame of a video. The action grounding dynamically associates the verbs to related subjects and the corresponding context, which keeps fine-grained spatial and temporal details for action prediction. Both entity grounding and action grounding are formulated as a unified task guided by a soft grounding supervision, which brings architecture simplification and improves training efficiency as well. We conduct extensive experiments on two challenging datasets, and demonstrate significant performance improvements of +2.3 CIDEr on ActivityNet-Entities and +2.2 CIDEr on MSR-VTT compared to state-of-the-arts.
Wenhui Jiang 0001, Yibo Cheng, Linxin Liu, Yuming Fang 0001, Yuxin Peng 0001, Yang Liu 0293
AAAI1
2024 PosCap: Boosting Video Captioning with Part-of-Speech Guidance
Jingfu Xiao, Wenhui Jiang 0001, Yuming Fang 0001
PRCV (10)3
2024 CFNet: Conditional filter learning with dynamic noise estimation for real image denoising
Yifan Zuo 0001, Wenhao Yao, Yifeng Zeng, Yuming Fang 0001, Yan Huang 0023, Wenhui Jiang 0001
Knowl. Based Syst.7
2023 Feature Adaptive YOLO for Remote Sensing Detection in Adverse Weather Conditions
abstract
Target detection in remote sensing has been one of the most challenging tasks in the past few decades. However, the detection performance in adverse weather conditions still needs to be satisfactory, mainly caused by the low-quality image features and the fuzzy boundary information. This work proposes a novel framework called Feature Adaptive YOLO (FA-YOLO). Specifically, we present a Hierarchical Feature Enhancement Module (HFEM), which adaptively performs feature-level enhancement to tackle the adverse impacts of different weather conditions. Then, we propose an Adaptive receptive Field enhancement Module (AFM) that dynamically adjusts the receptive field of the features and thus can enrich the context information for feature augmentation. In addition, we introduce Deformable Gated Head (DG-Head) which reduces the clutter caused by adverse weather. Experimental results on RTTS and two synthetic datasets demonstrate that our proposed FA-YOLO significantly outperforms other state-of-the-art target detection models.
Chaojun Ni, Wenhui Jiang 0001, Qishou Zhu, Yuming Fang 0001
VCIP2
2023 UDNet: Uncertainty-aware deep network for salient object detection
Yuming Fang 0001, Jiebin Yan, Wenhui Jiang 0001, Yang Liu 0293
Pattern Recognit.4
2022 Informative Attention Supervision for Grounded Video Description
abstract
Attention supervision encourages grounded video description models (GVDMs) to focus on the related visual content when generating words. Thus, it improves the description performance of GVDMs. However, existing GVDMs often fail to focus on small but informative regions because these regions are considered as negative by using the intersection-over-union (IoU) based attention groundtruth sampling method. Moreover, the prevailing attention loss functions enforce the GVDMs to focus equally on all sampled regions when the GVDMs generate words, which may make it difficult for the model to attend to informative regions and thus degrade the quality of the generated sentences. To alleviate the above problems, we propose an informative attention supervision method including a novel attention groundtruth sampling method and a group-based weak grounding supervision. Specifically, our attention groundtruth sample method captures small proposal regions that overlap with the entity boxes. The proposed grounding supervision allows the GVDMs to dynamically focus on some of the most informative attention regions instead of all of them. Our approach yields competitive results on the ActivityNet Entities dataset without bells and whistles, surpassing previous methods without increasing inference costs.
Boyang Wan, Wenhui Jiang 0001, Yuming Fang 0001
ICASSP2
2022 Bilinear CNNs for Blind Quality Assessment of Fine-Grained Images
abstract
Most of the existing image quality assessment (IQA) studies focus on discriminable images, whose relative visual quality could be easily determined by human beings (we also call this issue coarse-grained (CG)-IQA). The effective models designed for CG-IQA struggle for quality assessment of the images with subtle differences (often exist in many real applications), which is also called fine-grained (FG) IQA problem. Thus, we make the first, to the best of our knowledge, attempt to build a novel blind IQA (BIQA) model for the images with FG distortion, aiming to fill the gap between objective IQA model and real applications. Specifically, the proposed model mainly consists of a feature extraction module (a sequence of convolution layers), a squeeze-and-excitation module, and a bilinear pooling module, whose objectives are extracting quality-aware features, enhancing features' representation ability, and discriminability. We conduct extensive experiments on a public FG-IQA database, and demonstrate the superiority of the proposed method and the effectiveness of each module.
Jiebin Yan, Yuming Fang 0001, Wenhui Jiang 0001
MMSP4
2022 Dual-stream Self-attention Network for Image Captioning
abstract
Self-attention based encoder-decoder models achieve dominant performance in image captioning. However, most existing image captioning models (ICMs) only focus on modeling the relation between spatial tokens, while channel-wise attention is neglected for getting visual representation. Considering that different channels of visual representation usually denote different visual objects, it may lead to poor performance in terms of object and attribute words in the captioning sentences generated by the ICMs. In this paper, we propose a novel dual-stream self-attention module (DSM) to alleviate the above issue. Specifically, we propose a parallel self-attention based module that simultaneously encodes visual information from the spatial and channel dimensions. Besides, to obtain channel-wise visual features effectively and efficiently, we introduce a group self-attention block with linear computational complexity. To validate the effectiveness of our model, we conduct extensive experiments on the standard IC benchmarks including MSCOCO and Flickr30k. Without bells and whistles, the proposed model performs new SOTAs containing 135.4 CIDEr score on MSCOCO and 70.8 CIDEr score on Flickr30k.
Boyang Wan, Wenhui Jiang 0001, Yuming Fang 0001, Wenying Wen, Hantao Liu
VCIP2
2022 Revisiting image captioning via maximum discrepancy competition
Boyang Wan, Wenhui Jiang 0001, Yuming Fang 0001, Minwei Zhu, Yang Liu 0293
Pattern Recognit.2
2022 Visual Cluster Grounding for Image Captioning
abstract
Attention mechanisms have been extensively adopted in vision and language tasks such as image captioning. It encourages a captioning model to dynamically ground appropriate image regions when generating words or phrases, and it is critical to alleviate the problems of object hallucinations and language bias. However, current studies show that the grounding accuracy of existing captioners is still far from satisfactory. Recently, much effort is devoted to improving the grounding accuracy by linking the words to the full content of objects in images. However, due to the noisy grounding annotations and large variations of object appearance, such strict word-object alignment regularization may not be optimal for improving captioning performance. In this paper, to improve the performance of both grounding and captioning, we propose a novel grounding model which implicitly links the words to the evidence in the image. The proposed model encourages the captioner to dynamically focus on informative regions of the objects, which could be either discriminative parts or full object content. With slacked constraints, the proposed captioning model can capture correct linguistic characteristics and visual relevance, and then generate more grounded image captions. In addition, we propose a novel quantitative metric for evaluating the correctness of the soft attention mechanism by considering the overall contribution of all object proposals when generating certain words. The proposed grounding model can be seamlessly plugged into most attention-based architectures without introducing inference complexity. We conduct extensive experiments on Flickr30k (Young et al., 2014) and MS COCO datasets (Lin et al., 2014), demonstrating that the proposed method consistently improves image captioning in both grounding and captioning. Besides, the proposed attention evaluation metric shows better consistency with the captioning performance.
Wenhui Jiang 0001, Minwei Zhu, Yuming Fang 0001, Guangming Shi, Yang Liu 0293
IEEE Trans. Image Process.1
2021 Anomaly detection in video sequences: A benchmark and computational model
abstract
Abstract Anomaly detection has attracted considerable search attention. However, existing anomaly detection databases encounter two major problems. Firstly, they are limited in scale. Secondly, training sets contain only video‐level labels indicating the existence of an abnormal event during the full video while lacking annotations of precise time durations. To tackle these problems, we contribute a new L arge‐scale A nomaly D etection ( LAD ) database as the benchmark for anomaly detection in video sequences, which is featured in two aspects. 1) It contains 2000 video sequences including normal and abnormal video clips with 14 anomaly categories including crash, fire, violence etc . with large scene varieties, making it the largest anomaly analysis database to date. 2) It provides the annotation data, including video‐level labels (abnormal/normal video, anomaly type) and frame‐level labels (abnormal/normal video frame) to facilitate anomaly detection. Leveraging the above benefits from the LAD database, we further formulate anomaly detection as a fully supervised learning problem and propose a multi‐task deep neural network to solve it. We firstly obtain the local spatiotemporal contextual feature by using an Inflated 3D convolutional (I3D) network. Then we construct a recurrent convolutional neural network fed the local spatiotemporal contextual feature to extract the spatiotemporal contextual feature. With the global spatiotemporal contextual feature, the anomaly type and score can be computed simultaneously by a multi‐task neural network. Experimental results show that the proposed method outperforms the state‐of‐the‐art anomaly detection methods on our database and other public databases of anomaly detection. Supplementary materials are available at http://sim.jxufe.cn/JDMKL/ymfang/anomaly‐detection.html .
Boyang Wan, Wenhui Jiang 0001, Yuming Fang 0001, Zhiyuan Luo 0002, Guanqun Ding
IET Image Process.2
2021 Dynamic proposal sampling for weakly supervised object detection
Wenhui Jiang 0001, Zhicheng Zhao 0001, Yuming Fang 0001
Neurocomputing1
2021 Visual attention prediction for Autism Spectrum Disorder with hierarchical semantic fusion
Yuming Fang 0001, Yifan Zuo 0001, Wenhui Jiang 0001, Hanqin Huang, Jiebin Yan
Signal Process. Image Commun.4
2021 Superpixel-Based Quality Assessment of Multi-Exposure Image Fusion for Both Static and Dynamic Scenes
abstract
Multi-exposure image fusion (MEF) algorithms have been used to merge a stack of low dynamic range images with various exposure levels into a well-perceived image. However, little work has been dedicated to predicting the visual quality of fused images. In this work, we propose a novel and efficient objective image quality assessment (IQA) model for MEF images of both static and dynamic scenes based on superpixels and an information theory adaptive pooling strategy. First, with the help of superpixels, we divide fused images into large- and small-changed regions using the structural inconsistency map between each exposure and fused images. Then, we compute the quality maps based on the Laplacian pyramid for large- and small-changed regions separately. Finally, an information theory induced adaptive pooling strategy is proposed to compute the perceptual quality of the fused image. Experimental results on three public databases of MEF images demonstrate the proposed model achieves promising performance and yields a relatively low computational complexity. Additionally, we also demonstrate the potential application for parameter tuning of MEF algorithms.
Yuming Fang 0001, Yan Zeng 0001, Wenhui Jiang 0001, Hanwei Zhu, Jiebin Yan
IEEE Trans. Image Process.3
2018 Weakly supervised detection with decoupled attention-based deep representation
Wenhui Jiang 0001, Zhicheng Zhao 0001
Multim. Tools Appl.1
2016 ALADDIN: A locality aligned deep model for instance search
abstract
Most instance search systems are based on modeling local features. It remains a challenge to apply deep learning techniques into this task because of the asymmetrical similarity between the query region and dataset images. In this paper, we propose ALADDIN, A Locality Aligned Deep moDel for INstance search. This model deals with the asymmetrical similarity by searching query instances at the scale of aligned target regions instead of the whole image. Towards discriminative region representations, we utilize a deep convolutional network which captures both intra-class and inter-class distinctions of the regions. In addition, we propose a semi-supervised method to collect appropriate data to train the network. Extensive experiments confirm that our method is more suitable for generic instance search than most conventional methods, and outperforms the best CNNs-based method in both accuracy and efficiency.
Wenhui Jiang 0001, Zhicheng Zhao 0001, Anni Cai
ICASSP1
2016 Bayes pooling of visual phrases for object retrieval
Wenhui Jiang 0001, Zhicheng Zhao 0001
Multim. Tools Appl.1
2015 Part-based deep network for pedestrian detection in surveillance videos
abstract
Accurate pedestrian detection in highly crowded surveillance videos is a challenging task, since the regions of pedestrians in the videos may be largely occluded by other pedestrians. In this paper, we propose an effective part-based deep network cascade (HsNet) to solve this problem. In this model, the part-based scheme effectively restrains the appearance variations of pedestrians caused by heavy occlusion. The deep network captures discriminative information of visible body parts. In addition, the cascade architecture enables very fast detection. We make experiments on one of the largest surveillance video dataset, namely TRECVid SED Pedestrian Dataset (SED-PD). It is shown that in highly crowded surveillance videos, our proposed method achieves very competitive performance compared with state-of-the-art methods. More importantly, our method is significantly faster.
Qi Chen 0014, Wenhui Jiang 0001, Yanyun Zhao, Zhicheng Zhao 0001
VCIP2