EDBT 2026 Demo / reviewers in the wild / expert
Jiale Cao
dblp:167/3851
· DBLP profile ↗
77ranked-venue papers
15as first author
63since 2021 · last 2026
0000-0002-5160-6841ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 11 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 41 · 10 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multispectral remote sensing object detection via selective cross-modal interaction and aggregation
Minghao Cui, Jing Nie 0001, Hanqing Sun 0001, Jin Xie 0005, Jiale Cao, Yanwei Pang, Xuelong Li 0001 |
Neural Networks | 5 |
| 2026 | Parameter-Efficient Fine-Tuning for Continual Learning: A Neural Tangent Kernel PerspectiveabstractParameter-efficient fine-tuning for continual learning (PEFT-CL) has shown promise in adapting pre-trained models to sequential tasks while mitigating catastrophic forgetting problem. However, understanding the mechanisms that dictate continual performance in this paradigm remains elusive. To unravel this mystery, we undertake a rigorous analysis of PEFT-CL dynamics to derive relevant metrics for continual scenarios using Neural Tangent Kernel (NTK) theory. With the aid of NTK as a mathematical analysis tool, we recast the challenge of test-time forgetting into the quantifiable generalization gaps during training, identifying three key factors that influence these gaps and the performance of PEFT-CL: training sample size, task-level feature orthogonality, and regularization. To address these challenges, we introduce NTK-CL, a novel framework that eliminates task-specific parameter storage while adaptively generating task-relevant features. Aligning with theoretical guidance, NTK-CL triples the feature representation of each sample, theoretically and empirically reducing the magnitude of both task-interplay and task-specific generalization gaps. Grounded in NTK analysis, our framework imposes an adaptive exponential moving average mechanism and constraints on task-level feature orthogonality, maintaining intra-task NTK forms while attenuating inter-task NTK forms. Ultimately, by fine-tuning optimizable parameters with appropriate regularization, NTK-CL achieves state-of-the-art performance on established PEFT-CL benchmarks. This work provides a theoretical foundation for understanding and improving PEFT-CL models, offering insights into the interplay between feature representation, task orthogonality, and generalization, contributing to the development of more efficient continual learning systems. Jingren Liu, Zhong Ji, Yunlong Yu 0001, Jiale Cao, Yanwei Pang, Jungong Han, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | iSeg: An Iterative Refinement-Based Framework for Training-Free SegmentationabstractStable Diffusion has demonstrated strong image synthesis ability to given text descriptions, suggesting it to contain strong semantic clue for grouping objects. The researchers have explored employing Stable Diffusion for training-free segmentation. Most existing approaches refine cross-attention map by self-attention map once, demonstrating that self-attention map contains useful semantic information to improve segmentation. To fully utilize self-attention map, we present a deep experimental analysis on iteratively refining cross-attention map with self-attention map, and propose an effective iterative refinement framework for training-free segmentation, named iSeg. Our iSeg introduces an entropy-reduced self-attention module that utilizes a gradient descent scheme to reduce the entropy of self-attention map, thereby suppressing the weak responses corresponding to irrelevant global information. Leveraging the entropy-reduced self-attention module, our iSeg stably improves cross-attention map with iterative refinement. Further, we design a category-enhanced cross-attention module to generate accurate cross-attention map, providing a better initial input for iterative refinement. Extensive experiments across different datasets and diverse segmentation tasks (weakly-supervised semantic segmentation, open-vocabulary semantic segmentation, unsupervised segmentation, and mask generation on synthetic dataset) reveal the merits of proposed contributions, leading to promising performance. For unsupervised semantic segmentation on Cityscapes, our iSeg achieves an absolute gain of 3.8% in terms of mIoU compared to the best existing training-free approach in literature. Moreover, our proposed iSeg can support segmentation with different kinds of images and interactions, and also be used as a post-processing, or in different frameworks, to improve training-free segmentation. Jiale Cao, Jin Xie 0005, Fahad Shahbaz Khan, Yanwei Pang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | SED++: A Simple Encoder-Decoder for Improved Open-Vocabulary Semantic SegmentationabstractOpen-vocabulary semantic segmentation aims to partition an image into distinct semantic regions based on an open set of categories. Existing approaches primarily rely on image-level pre-trained vision-language models to perform this pixel-level task. In this paper, we propose SED, a simple yet effective encoder-decoder architecture for open-vocabulary semantic segmentation leveraging pre-trained vision-language models. SED consists of a hierarchical image encoder, a text encoder, and a gradual fusion decoder. The hierarchical image encoder and text encoder collaboratively generate a cost volume, which is progressively decoded by the gradual fusion decoder to produce segmentation results. In contrast to a plain encoder, the hierarchical encoder better captures image detail information while maintaining linear computational complexity with respect to input size. The gradual fusion decoder adopts a top-down structure to progressively integrate high-resolution features with the cost volume. Furthermore, a category early rejection strategy is introduced in gradual fusion decoder to filter out non-existent categories at different layers, significantly improving inference efficiency. Based on SED, we further introduce two modules, including non-label text embedding and additional category early rejection in the encoder. Moreover, we extend our method with minimal decoder modification for open-vocabulary video semantic segmentation. Extensive experiments on multiple datasets validate the effectiveness and efficiency of our proposed method. With ConvNeXt-B, our method achieves an mIoU of 34.9% on the ADE20 K with 150 classes (i.e., A-150) at an inference speed of 69 ms per image on a single A6000 GPU, and has an mIoU score of 40.2% on video segmentation dataset VSPW. Jiale Cao, Jin Xie 0005, Fahad Shahbaz Khan, Yanwei Pang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | SNNSIR: A fully Spiking Neural Network for Stereo Image Restoration
Jin Xie 0005, Jing Nie 0001, Jiale Cao, Yanwei Pang |
Pattern Recognit. | 4 |
| 2026 | Global context guided refinement and aggregation network for lightweight surface defect detection
Xiaoheng Jiang, Yang Lu 0016, Lisha Cui, Jiale Cao, Mingliang Xu 0001 |
Pattern Recognit. | 5 |
| 2026 | HG-LMM: Unleashing High-Quality Pixel Grounding Capabilities in Frozen Large Multimodal ModelsabstractLarge Multimodal Models (LMMs) have demonstrated remarkable capabilities in multimodal understanding and conversation. Recently, some researchers have explored fine-tuning LMM for pixel grounding, leading to catastrophic loss of their inherent conversational capabilities. To preserve the conversational capability, some researchers explore to freeze LMM, but employ heavy segmenter SAM for high-quality grounding. In this paper, we propose a novel approach, named HG-LMM, to fully exploit the inherent features of frozen LMM for high-quality pixel grounding. Our HG-LMM introduces two main modules: an LLM-guided instance-aware feature generation (LIFG) and a layer-wise detail and semantic injection (LDSI). The LIFG module employs the output text embeddings of LMM belonging to grounded instances to generate multi-level instance-aware feature maps from the image encoder. Afterwards, we employ the LDSI module to inject more detail and semantic information into these instance-aware feature maps. With these instance-aware feature maps, we employ a simple top-down fusion to predict the segmentation masks of different instances. We perform experiments on various tasks, including referring expression segmentation, panoramic narrative grounding, reasoning segmentation, grounded conversation generation, and visual chain-of-thought reasoning. When using DeepSeekVL-1.3B, our HG-LMM is 12.6% better than F-LMM without SAM in terms of segmentation accuracy on the all set of PNG dataset. Compared to F-LMM with SAM, our HG-LMM achieves comparable segmentation accuracy while being 2.7 times faster. We release our source code and models at https://github.com/WenjieLi2008/HG-LMM. Jiale Cao, Jin Xie 0005, Ai-Ping Yang, Yanwei Pang |
IEEE Trans. Image Process. | 2 |
| 2026 | Frequency-Decomposed Interaction Network for Stereo Image RestorationabstractStereo image restoration in adverse environments, such as low-light conditions, rain, and low resolution, requires effective exploitation of cross-view complementary information to recover degraded visual content. In monocular image restoration, frequency decomposition has proven effective, where high-frequency components aid in recovering fine textures and reducing blur, while low-frequency components facilitate noise suppression and illumination correction. However, existing stereo restoration methods have yet to explore cross-view interactions by frequency decomposition, which is a promising direction for enhancing restoration quality. To address this, we propose a frequency-aware framework comprising a Frequency Decomposition Module (FDM), Detail Interaction Module (DIM), Structural Interaction Module (SIM), and Adaptive Fusion Module (AFM). FDM employs learnable filters to decompose the image into high- and low-frequency components. DIM enhances the high-frequency branch by capturing local detail cues through deformable convolution. SIM processes the low-frequency branch by modeling global structural correlations via a cross-view row-wise attention mechanism. Finally, AFM adaptively fuses the complementary frequency-specific information to generate high-quality restored images. Extensive experiments demonstrate the efficacy and generalizability of our framework across three diverse stereo restoration tasks, where it achieves state-of-the-art performance in low-light enhancement, rain removal, alongside highly competitive results in super-resolution. Our code is available at https://github.com/C2022J/FDIN. Xianmin Tian, Jin Xie 0005, Jing Nie 0001, Jiale Cao, Yanwei Pang, Xuelong Li 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | SSLFusion: Scale and Space Aligned Latent Fusion Model for Multimodal 3D Object DetectionabstractMultimodal 3D object detection based on deep neural networks has indeed made significant progress. However, it still faces challenges due to the misalignment of scale and spatial information between features extracted from 2D images and those derived from 3D point clouds. Existing methods usually aggregate multimodal features at a single stage. However, leveraging multi-stage cross-modal features is crucial for detecting objects of various scales. Therefore, these methods often struggle to integrate features across different scales and modalities effectively, thereby restricting the accuracy of detection. Additionally, the time-consuming Query-Key-Value-based (QKV-based) cross-attention operations often utilized in existing methods aid in reasoning the location and existence of objects by capturing non-local contexts. However, this approach tends to increase computational complexity. To address these challenges, we present SSLFusion, a novel Scale & Space Aligned Latent Fusion Model, consisting of a scale-aligned fusion strategy (SAF), a 3D-to-2D space alignment module (SAM), and a latent cross-modal fusion module (LFM). SAF mitigates scale misalignment between modalities by aggregating features from both images and point clouds across multiple levels. SAM is designed to reduce the inter-modal gap between features from images and point clouds by incorporating 3D coordinate information into 2D image features. Additionally, LFM captures cross-modal non-local contexts in the latent space without utilizing the QKV-based attention operations, thus mitigating computational complexity. Experiments on the KITTI and DENSE datasets demonstrate that our SSLFusion outperforms state-of-the-art methods. Our approach obtains an absolute gain of 2.15% in 3D AP, compared with the state-of-art method GraphAlign on the moderate level of the KITTI test set. Bonan Ding, Jin Xie 0005, Jing Nie 0001, Jiale Cao |
AAAI | 4 |
| 2025 | VideoGLaMM : A Large Multimodal Model for Pixel-Level Visual Grounding in VideosabstractFine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing video-based Large Multimodal Models (LMMs) handle basic conversations but struggle with precise pixel-level grounding in videos. To address this, we introduce VideoGLaMM, a LMM designed for fine-grained pixel-level grounding in videos based on user-provided textual inputs. Our design seamlessly connects three key components: a Large Language Model, a dual vision encoder that emphasizes both spatial and temporal details, and a spatio-temporal decoder for accurate mask generation. This connection is facilitated via tunable V→L and L→V adapters that enable close Vision-Language (VL) alignment. The architecture is trained to synchronize both spatial and temporal elements of video content with textual instructions. To enable fine-grained grounding, we curate a multimodal dataset featuring detailed visually-grounded conversations using a semiautomatic annotation pipeline, resulting in a diverse set of 38k video-QA triplets along with 83k objects and 671k masks. We evaluate VideoGLaMM on three challenging tasks: Grounded Conversation Generation, Visual Grounding, and Referring Video Segmentation. Experimental results show that our model consistently outperforms existing approaches across all three tasks. Shehan Munasinghe, Hanan Gani, Jiale Cao, Eric P. Xing, Fahad Shahbaz Khan, Salman Khan 0001 |
CVPR | 4 |
| 2025 | Wavelet and Prototype Augmented Query-based Transformer for Pixel-level Surface Defect DetectionabstractAs an important part of intelligent manufacturing, pixel-level surface defect detection (SDD) aims to locate defect areas through mask prediction. Previous methods adopt the image-independent static convolution to indiscriminately classify per-pixel features for mask prediction, which leads to suboptimal results for some challenging scenes such as weak defects and cluttered backgrounds. In this paper, inspired by query-based methods, we propose a Wavelet and Prototype Augmented Query-based Transformer (WP-Former) for surface defect detection. Specifically, a set of dynamic queries for mask prediction is updated through the dual-domain transformer decoder. Firstly, a Wavelet-enhanced Cross-Attention (WCA) is proposed, which aggregates meaningful high-and low-frequency information of image features in the wavelet domain to refine queries. WCA enhances the representation of high-frequency components by capturing multi-scale relationships between different frequency components, enabling queries to focus more on defect details. Secondly, a Prototype-guided Cross-Attention (PCA) is proposed to refine queries through meta-prototypes in the spatial domain. The prototypes aggregate semantically meaningful tokens from image features, facilitating queries to aggregate crucial defect information under the cluttered backgrounds. Extensive experiments on three defect detection datasets (i.e., ESDIs-SOD, CrackSeg9k, and ZJU-Leaper) demonstrate that the proposed method achieves state-of-the-art performance in defect detection. The code will be available at https://github.com/yfhdm/WPFormer. Xiaoheng Jiang, Yang Lu 0016, Jiale Cao, Dong Chen 0017, Mingliang Xu 0001 |
CVPR | 4 |
| 2025 | CLIPeR: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic SegmentationabstractContrastive Language-Image Pre-training (CLIP) exhibits strong zero-shot classification ability on various image-level tasks, leading to the research to adapt CLIP for pixel-level open-vocabulary semantic segmentation without additional training. The key is to improve spatial representation of image-level CLIP, such as replacing self-attention map at last layer with self-self attention map or vision foundation model based attention map. In this paper, we present a novel hierarchical framework, named CLIPer, that hierarchically improves spatial representation of CLIP. The proposed CLIPer includes an early-layer fusion module and a fine-grained compensation module. We observe that, the embeddings and attention maps at early layers can preserve spatial structural information. Inspired by this, we design the early-layer fusion module to generate segmentation map with better spatial coherence. Afterwards, we employ a fine-grained compensation module to compensate the local details using the self-attention maps of diffusion model. We conduct the experiments on seven segmentation datasets. Our proposed CLIPer achieves the state-of-the-art performance on these datasets. For instance, using ViT-L, CLIPer has the mIoU of 69.8% and 43.3% on VOC and COCO Object, outperforming ProxyCLIP by 9.2% and 4.1% respectively. Jiale Cao, Jin Xie 0005, Xiaoheng Jiang, Yanwei Pang |
ICCV | 2 |
| 2025 | Glad: A Streaming Scene Generator for Autonomous DrivingabstractThe generation and simulation of diverse real-world scenes have significant application value in the field of autonomous driving, especially for the corner cases. Recently, researchers have explored employing neural radiance fields or diffusion models to generate novel views or synthetic data under driving scenes. However, these approaches suffer from unseen scenes or restricted video length, thus lacking sufficient adaptability for data generation and simulation. To address these issues, we propose a simple yet effective framework, named Glad, to generate video data in a frame-by-frame style. To ensure the temporal consistency of synthetic video, we introduce a latent variable propagation module, which views the latent features of previous frame as noise prior and injects it into the latent features of current frame. In addition, we design a streaming data sampler to orderly sample the original image in a video clip at continuous iterations. Given the reference frame, our Glad can be viewed as a streaming simulator by generating the videos for specific scenes. Extensive experiments are performed on the widely-used nuScenes dataset. Experimental results demonstrate that our proposed Glad achieves promising performance, serving as a strong baseline for online video generation. We will release the source code and models publicly. Yingfei Liu, Tiancai Wang, Jiale Cao, Xiangyu Zhang 0005 |
ICLR | 4 |
| 2025 | Multi-scale Progressive Low-Light Image Enhancement Networks
Bing Liu 0016, Jiale Cao |
PRCV (9) | 2 |
| 2025 | Dual-Domain Low-Light Image Enhancement Network via Frequency Interaction and Structure-Guided Attention
Bing Liu 0016, Jiale Cao, Peng Liu 0013 |
PRCV (8) | 2 |
| 2025 | SGD: Street View Synthesis with Gaussian Splatting and Diffusion PriorabstractNovel View Synthesis (NVS) for street scenes plays a critical role in the autonomous driving simulation. Current mainstream methods, such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), struggle to maintain rendering quality at the viewpoint that deviates significantly from the training viewpoints. This issue stems from the sparse training views captured by a fixed camera on a moving vehicle. To tackle this problem, we propose a novel approach that enhances the capacity of 3DGS by leveraging prior from a Diffusion Model along with complementary multi-modal data. Specifically, we first fine-tune a Diffusion Model by adding images from adjacent frames as condition, meanwhile exploiting depth data from LiDAR point clouds to supply additional spatial information. Then we apply the fine-tuned Diffusion Model to regularize the 3DGS at unseen views during training. Experimental results validate the effectiveness of our method compared with current state-of-the-art models, and demonstrate its advance in rendering images from broader views. Zhongrui Yu, Haoran Wang 0004, Jinze Yang, Jiale Cao, Zhong Ji, Mingming Sun 0001 |
WACV | 5 |
| 2025 | DANet: spatial gene expression prediction from H&E histology images through dynamic alignmentabstractPredicting spatial gene expression from Hematoxylin and Eosin histology images offers a promising approach to significantly reduce the time and cost associated with gene expression sequencing, thereby facilitating a deeper understanding of tissue architecture and disease mechanisms. Achieving accurate gene expression prediction requires the extraction of highly refined features from pathological images; however, existing methods often struggle to effectively capture fine-grained local details and model gene-gene correlations. Moreover, in bimodal contrastive learning, dynamically and efficiently aligning heterogeneous modalities remains a critical challenge. To address these issues, we propose a novel method for predicting gene expression. First, we introduce a dense connective structure that enables efficient feature reuse, thereby enhancing the capturing and mining of local refinement features. Second, we leverage the state space models to uncover underlying patterns and capture dependencies within 1D gene expression data, enabling more accurate modeling of gene-gene correlations. Furthermore, we design the Residual Kolmogorov-Arnold Network (RKAN) that uses a learnable activation function to dynamically adjust bimodal mappings based on input characteristics. Through continuous parameter updates during contrastive training, RKAN progressively refines the alignment between modalities. Extensive experiments conducted on two publicly available datasets, GSE240429 and HER2+, demonstrate the effectiveness of our approach and its significant improvements over existing methods. Source codes are available at https://github.com/202324131016T/DANet. Jin Xie 0005, Jing Nie 0001, Jiale Cao, Yuansong Zeng |
Briefings Bioinform. | 4 |
| 2025 | Mitigating forgetting in the adaptation of CLIP for few-shot classification
Jiale Cao, Yuanheng Liu, Zhong Ji, Jingren Liu, Ai-Ping Yang, Yanwei Pang |
Comput. Vis. Image Underst. | 1 |
| 2025 | Bi-orientated rectification few-shot segmentation network based on fine-grained prototypes
Ai-Ping Yang, Zijia Sang, Yaran Zhou, Jiale Cao |
Neurocomputing | 4 |
| 2025 | Cross-Scale Atomic Feature Enhanced Network for high-fidelity Single Image Super-Resolution
Ai-Ping Yang, Chenhui Yu, Jinbin Wang, Zihao Wei, Jiale Cao |
Multim. Syst. | 5 |
| 2025 | Token-aware and step-aware acceleration for Stable Diffusion
Ting Zhen, Jiale Cao, Xuebin Sun, Zhong Ji, Yanwei Pang |
Pattern Recognit. | 2 |
| 2025 | Multi-Granularity Language-Guided Training for Multi-Object TrackingabstractMost existing multi-object tracking methods typically learn visual tracking features via maximizing dis-similarities of different instances and minimizing similarities of the same instance. While such a feature learning scheme achieves promising performance, learning discriminative features solely based on visual information is challenging especially in case of environmental interference such as occlusion, blur and domain variance. In this work, we argue that multi-modal language-driven features provide complementary information to classical visual features, thereby aiding in improving the robustness to such environmental interference. To this end, we propose a new multi-object tracking framework, named LG-MOT, that explicitly leverages language information at different levels of granularity (scene-and instance-level) and combines it with standard visual features to obtain discriminative representations. To develop LG-MOT, we annotate existing MOT datasets with scene-and instance-level language descriptions. We then encode both scene-and instance-level language information into high-dimensional embeddings, which are utilized to guide the visual features during training. At inference, our LG-MOT uses the standard visual features without relying on annotated language descriptions. Extensive experiments on three benchmarks, MOT17, DanceTrack and SportsMOT, reveal the merits of the proposed contributions leading to state-of-the-art performance. On the DanceTrack test set, our LG-MOT achieves an absolute gain of 2.2% in terms of target object association (IDF1 score), compared to the baseline using only visual features. Further, our LG-MOT exhibits strong cross-domain generalizability. Source code and pre-trained models are available at https://github.com/WesLee88524/LG-MOT. Jiale Cao, Muzammal Naseer, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001, Fahad Shahbaz Khan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance SegmentationabstractOpen-vocabulary video instance segmentation strives to segment and track instances belonging to an open set of categories in a videos. The vision-language model Contrastive Language-Image Pre-training (CLIP) has shown robust zero-shot classification ability in image-level open-vocabulary tasks. In this paper, we propose a simple encoder-decoder network, called CLIP-VIS, to adapt CLIP for open-vocabulary video instance segmentation. Our CLIP-VIS adopts frozen CLIP and introduces three modules, including class-agnostic mask generation, temporal topK-enhanced matching, and weighted open-vocabulary classification. Given a set of initial queries, class-agnostic mask generation introduces a pixel decoder and a transformer decoder on CLIP pre-trained image encoder to predict query masks and corresponding object scores and mask IoU scores. Then, temporal topK-enhanced matching performs query matching across frames using the K mostly matched frames. Finally, weighted open-vocabulary classification first employs mask pooling to generate query visual features from CLIP pre-trained image encoder, and second performs weighted classification using object scores and mask IoU scores. Our CLIP-VIS does not require the annotations of instance categories and identities. The experiments are performed on various video instance segmentation datasets, which demonstrate the effectiveness of our proposed method, especially for novel categories. When using ConvNeXt-B as backbone, our CLIP-VIS achieves the AP and APn scores of 32.2% and 40.2% on the validation set of LV-VIS dataset, which outperforms OV2Seg by 11.1% and 23.9% respectively. We will release the source code and models athttps://github.com/zwq456/CLIP-VIS.git. Jiale Cao, Jin Xie 0005, Shuangming Yang, Yanwei Pang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Video Instance Segmentation Without Using Mask and Identity SupervisionabstractVideo instance segmentation (VIS) is a challenging vision problem in which the task is to simultaneously detect, segment, and track all the object instances in a video. Most existing VIS approaches rely on pixel-level mask supervision within a frame as well as instance-level identity annotation across frames. However, obtaining these ‘mask and identity’ annotations is time-consuming and expensive. We propose the first mask-identity-free VIS framework that neither utilizes mask annotations nor requires identity supervision. Accordingly, we introduce a query contrast and exchange network (QCEN) comprising instance query contrast and query-exchanged mask learning. The instance query contrast first performs cross-frame instance matching and then conducts query feature contrastive learning. The query-exchanged mask learning exploits both intra-video and inter-video query exchange properties: exchanging queries of an identical instance from different frames within a video results in consistent instance masks, whereas exchanging queries across videos results in all-zero background masks. Extensive experiments on three benchmarks (YouTube-VIS 2019, YouTube-VIS 2021, and OVIS) reveal the merits of the proposed approach, which significantly reduces the performance gap between the identify-free baseline and our mask-identify-free VIS method. On the YouTube-VIS 2019 validation set, our mask-identity-free approach achieves 91.4% of the stronger-supervision-based baseline performance when utilizing the same ImageNet pre-trained model. Jiale Cao, Hanqing Sun 0001, Rao Muhammad Anwer, Jin Xie 0005, Fahad Shahbaz Khan, Yanwei Pang |
IEEE Trans. Multim. | 2 |
| 2025 | Implicit and Explicit Language Guidance for Diffusion-Based Visual PerceptionabstractText-to-image diffusion models have shown powerful ability on conditional image synthesis. With large-scale vision-language pre-training, diffusion models are able to generate high-quality images with rich textures and reasonable structures under different text prompts. However, adapting pre-trained diffusion models for visual perception is an open problem. In this paper, we propose an implicit and explicit language guidance framework for diffusion-based visual perception, named IEDP. Our IEDP comprises an implicit language guidance branch and an explicit language guidance branch. The implicit branch employs a frozen CLIP image encoder to directly generate implicit text embeddings that are fed to the diffusion model without explicit text prompts. The explicit branch uses the ground-truth labels of corresponding images as text prompts to condition feature extraction in diffusion model. During training, we jointly train the diffusion model by sharing the model weights of these two branches. As a result, the implicit and explicit branches can jointly guide feature learning. During inference, we employ only implicit branch for final prediction, which does not require any ground-truth labels. Experiments are performed on two typical perception tasks, including semantic segmentation and depth estimation. Our IEDP achieves promising performance on both tasks. For semantic segmentation, our IEDP has the mIoU$^\text{ss}$score of 55.9% on ADE20K validation set, which outperforms the baseline method VPD by 2.2%. For depth estimation, our IEDP outperforms the baseline method VPD with a relative gain of 11.0%. Hefeng Wang, Jiale Cao, Jin Xie 0005, Ai-Ping Yang, Yanwei Pang |
IEEE Trans. Multim. | 2 |
| 2025 | DefectSAM: Hierarchically Adapting SAM for Pixel-Wise Surface Defect DetectionabstractSegment anything model (SAM) has recently demonstrated powerful segmentation ability for natural scene images (NSIs). However, the SAM exhibits limited performance in defect detection owing to the weak appearance of defects and cluttered backgrounds in industrial images. In this article, we propose a hierarchically adapting SAM for pixel-wise surface defect detection, named DefectSAM, which effectively modulates and decodes multilevel features of the encoder to capture defect information. Specifically, we introduce a learnable feature adaptation component between the image encoder and the decoder to modulate each level of features via the dual-feature adaptation unit. The dual-feature adaptation unit mainly includes the correlation-gated feature adaptation (CGFA) module and the mask-guided feature adaptation (MGFA) module. The CGFA exploits cross correlation spatial gating maps to adaptively incorporate a convolutional feature pyramid and Transformer features during feature adaptation, which is beneficial for capturing defect details. Moreover, the MGFA utilizes the mask prediction of high-level features as semantic guidance to select top-confidence foreground and background tokens for feature adaptation, focusing more on defect details and suppressing background noise. Extensive experiments on three defect detection datasets (i.e., MVTec AD, CrackSeg9k, ZJU-Leaper, and Magnetic tile) demonstrate that the proposed method achieves state-of-the-art performance with few learnable parameters, which greatly improves the generalization of SAM in defect detection. Xiaoheng Jiang, Yang Lu 0016, Jiale Cao, Mingliang Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | SSNet: a joint learning network for semantic segmentation and disparity estimation
Dayu Jia, Yanwei Pang, Jiale Cao |
Vis. Comput. | 3 |
| 2024 | An Evolutionary Algorithm with Feasibility Tracking Strategy for Constrained Multi-Objective Optimization ProblemsabstractIn recent years, many constrained multi-objective evolutionary algorithms have been proposed to address com-plex constrained multi-objective optimization problems and have shown significant performance. However, when facing challenging constrained multi-objective optimization problem, some algorithms struggle to explore all feasible regions. To address this issue, this paper proposes a feasibility tracking strategy. By setting an expansion value and controlling constraint relaxation, adaptive adjustment of the reference vectors simulates an ex-panded narrow feasible region. After the preliminary exploration of the feasible region, a new set of reference vectors is generated to ensure the existence of feasible regions in the direction of the new reference vectors, thereby increasing the density of individuals within the feasible region. Experimental results on a series of benchmark problems demonstrate that the proposed algorithm is quite effective compared to other state-of-the-art constrained multi-objective evolutionary algorithms. Lei Yang 0040, Jinglin Tian, Jiale Cao, Kangshun Li, Chaoda Peng |
CEC | 3 |
| 2024 | SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic SegmentationabstractOpen-vocabulary semantic segmentation strives to distinguish pixels into different semantic groups from an open set of categories. Most existing methods explore utilizing pre-trained vision-language models, in which the key is to adapt the image-level model for pixel-level segmentation task. In this paper, we propose a simple encoder-decoder, named SED, for open-vocabulary semantic segmentation, which comprises a hierarchical encoder-based cost map generation and a gradual fusion decoder with category early rejection. The hierarchical encoder-based cost map generation employs hierarchical backbone, instead of plain transformer, to predict pixel-level image-text cost map. Compared to plain transformer, hierarchical backbone better captures local spatial information and has linear computational complexity with respect to input size. Our gradual fusion decoder employs a top-down structure to combine cost map and the feature maps of different backbone levels for segmentation. To accelerate inference speed, we introduce a category early rejection scheme in the decoder that rejects many no-existing categories at the early layer of decoder, resulting in at most 4.7 times acceleration without accuracy degradation. Experiments are performed on multiple open-vocabulary semantic segmentation datasets, which demonstrates the efficacy of our SED method. When using ConvNeXt-B, our SED method achieves mIoU score of 31.6% on ADE20K with 150 categories at 82 millisecond (ms) per image on a single A6000. Our source code is available at https://github.com/xb534/SED. Jiale Cao, Jin Xie 0005, Fahad Shahbaz Khan, Yanwei Pang |
CVPR | 2 |
| 2024 | DB-SAM: Delving into High Quality Universal Medical Image Segmentation
Jiale Cao, Huazhu Fu, Fahad Shahbaz Khan, Rao Muhammad Anwer |
MICCAI (12) | 2 |
| 2024 | Localization-aware logit mimicking for object detection in adverse weather conditions
Peiyun Luo, Jing Nie 0001, Jin Xie 0005, Jiale Cao, Xiaohong Zhang 0002 |
Image Vis. Comput. | 4 |
| 2024 | C2BG-Net: Cross-modality and cross-scale balance network with global semantics for multi-modal 3D object detection
Bonan Ding, Jin Xie 0005, Jing Nie 0001, Jiale Cao |
Neural Networks | 5 |
| 2024 | Deep intra-image contrastive learning for weakly supervised one-step person search
Jiabei Wang, Yanwei Pang, Jiale Cao, Hanqing Sun 0001, Xuelong Li 0001 |
Pattern Recognit. | 3 |
| 2024 | Multi-query and multi-level enhanced network for semantic segmentation
Jiale Cao, Rao Muhammad Anwer, Jin Xie 0005, Jing Nie 0001, Ai-Ping Yang, Yanwei Pang |
Pattern Recognit. | 2 |
| 2024 | ESGN: Efficient Stereo Geometry Network for Fast 3D Object DetectionabstractFast stereo based 3D object detectors have made great progress recently. However, they suffer from the inferior accuracy. We argue that the main reason is due to the poor geometry-aware feature representation in 3D space. To solve this problem, we propose an efficient stereo geometry network (ESGN). The key in our ESGN is an efficient geometry-aware feature generation (EGFG) module. Our EGFG module first uses a stereo correlation and reprojection module to construct multi-scale stereo volumes in camera frustum space, second employs a multi-scale bird’s eye view (BEV) projection and fusion module to generate multiple geometry-aware features. In these two steps, we adopt deep multi-scale information fusion for discriminative geometry-aware feature generation, without any complex aggregation networks. In addition, we introduce a deep geometry-aware feature distillation scheme to guide stereo feature learning with a LiDAR-based detector. The experiments are performed on the classical KITTI dataset. On KITTI test set, our ESGN outperforms the fast state-of-art-art detector YOLOStereo3D by 5.14% on mAP3d at$62ms$. To the best of our knowledge, our ESGN achieves a best trade-off between accuracy and speed. We hope that our efficient stereo geometry network can provide more possible directions for fast 3D object detection. Aqi Gao, Yanwei Pang, Jing Nie 0001, Jiale Cao, Yishun Guo, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Toward Generalizable Multispectral Pedestrian DetectionabstractMultispectral pedestrian detection has achieved great success in past years, which can be used in autonomous driving for intelligent transportation system. Most existing multispectral pedestrian detection approaches are developed on the assumption that training and test data belong to an identical distribution, which does not guarantee a good generalization to cross-domain (unseen) data. In this paper, we aim to develop a generalizable multispectral pedestrian detector, which achieves a favorable performance on both intra-dataset evaluation and cross-dataset evaluation. To achieve this goal, we conduct intra-dataset and cross-dataset experiments using single-modal and multi-modal data. By deep analysis, we find that, compared to visible or multi-modal data, thermal data not only has a best cross-dataset generalization, but also generates high-quality proposals on intra-dataset and cross-dataset evaluations. Inspired by this, we propose a novel thermal-first and fusion-second network (called TFNet) for multispectral pedestrian detection. In our TFNet, we first employ a thermal-based proposal network to extract candidate pedestrian proposals. After that, we design a transformer fusion based head network to further classify/regress these proposals. Experiments are performed on three public datasets. The comprehensive results demonstrate the effectiveness of our proposed TFNet on both intra-dataset and cross-dataset evaluations. We hope that our simple design can promote the future study on generalizable multispectral pedestrian detection. Fuchen Chu, Jiale Cao, Zhanjie Song, Yanwei Pang, Xuelong Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Transformer-Based Stereo-Aware 3D Object Detection From Binocular ImagesabstractTransformers have shown promising progress in various visual object detection tasks, including monocular 2D/3D detection and surround-view 3D detection. More importantly, the attention mechanism in the Transformer model and the 3D information extraction in binocular stereo are both similarity-based. However, directly applying existing Transformer-based detectors to binocular stereo 3D object detection leads to slow convergence and significant precision drops. We argue that a key cause of that defect is that existing Transformers ignore the binocular-stereo-specific image correspondence information. In this paper, we explore the model design of Transformers in binocular 3D object detection, focusing particularly on extracting and encoding task-specific image correspondence information. To achieve this goal, we present TS3D, a Transformer-based Stereo-aware 3D object detector. In the TS3D, a Disparity-Aware Positional Encoding (DAPE) module is proposed to embed the image correspondence information into stereo features. The correspondence is encoded as normalized sub-pixel-level disparity and is used in conjunction with sinusoidal 2D positional encoding to provide the 3D location information of the scene. To enrich multi-scale stereo features, we propose a Stereo Preserving Feature Pyramid Network (SPFPN). The SPFPN is designed to preserve the correspondence information while fusing intra-scale and aggregating cross-scale stereo features. Our proposed TS3D achieves a 41.29% Moderate Car detection average precision on the KITTI test set and takes 88 ms to detect objects from each binocular image pair. It is competitive with advanced counterparts in terms of both precision and inference speed. Hanqing Sun 0001, Yanwei Pang, Jiale Cao, Jin Xie 0005, Xuelong Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Multi-feature self-attention super-resolution network
Ai-Ping Yang, Zihao Wei, Jinbin Wang, Jiale Cao, Zhong Ji, Yanwei Pang |
Vis. Comput. | 4 |
| 2023 | A Spatial-Temporal Deformable Attention Based Framework for Breast Lesion Detection in Videos
Jiale Cao, Huazhu Fu, Rao Muhammad Anwer, Fahad Shahbaz Khan |
MICCAI (2) | 2 |
| 2023 | Attentive Alignment Network for Multispectral Pedestrian DetectionabstractMultispectral pedestrian detection is of great importance in various around-the-clock applications, i.e., self-driving and video surveillance. Fusing the features from RGB images and thermal infrared (TIR) images to explore the complementary information between different modalities is one of the most effective manners to improve multispectral pedestrian detection performance. However, the misalignment between different modalities in spatial dimension and modality reliability would introduce harmful information during feature fusion, limiting the performance of multispectral pedestrian detection. To address the above issues, we propose an attentive alignment network, consisting of an attentive position alignment (APA) module and an attentive modality alignment (AMA) module. Our APA module emphasizes pedestrian regions while aligning the pedestrian regions between different modalities. Our AMA module utilizes a channel-wise attention mechanism with illumination guidance to eliminate the imbalance between different modalities. The experiments are conducted on two widely used multispectral detection datasets, KASIT and CVC-14. Our approach surpasses the current state-of-the-art performance on both datasets. Nuo Chen 0003, Jin Xie 0005, Jing Nie 0001, Jiale Cao, Yanwei Pang |
ACM Multimedia | 4 |
| 2023 | Spatial attention-guided deformable fusion network for salient object detection
Ai-Ping Yang, Simeng Cheng, Jiale Cao, Zhong Ji, Yanwei Pang |
Multim. Syst. | 4 |
| 2023 | Context and detail interaction network for stereo rain streak and raindrop removal
Jing Nie 0001, Jin Xie 0005, Jiale Cao, Yanwei Pang |
Neural Networks | 3 |
| 2023 | Visual-quality-driven unsupervised image dehazing
Ai-Ping Yang, Jinbin Wang, Jiale Cao, Zhong Ji, Yanwei Pang |
Neural Networks | 5 |
| 2023 | SipMaskv2: Enhanced Fast Image and Video Instance SegmentationabstractWe propose a fast single-stage method for both image and video instance segmentation, called SipMask, that preserves the instance spatial information by performing multiple sub-region mask predictions. The main module in our method is a light-weight spatial preservation (SP) module that generates a separate set of spatial coefficients for the sub-regions within a bounding-box, enabling a better delineation of spatially adjacent instances. To better correlate mask prediction with object detection, we further propose a mask alignment weighting loss and a feature alignment scheme. In addition, we identify two issues that impede the performance of single-stage instance segmentation and introduce two modules, including a sample selection scheme and an instance refinement module, to address these two issues. Experiments are performed on both image instance segmentation dataset MS COCO and video instance segmentation dataset YouTube-VIS. On MS COCO test-dev set, our method achieves a state-of-the-art performance. In terms of real-time capabilities, it outperforms YOLACT by a gain of 3.0% (mask AP) under the similar settings, while operating at a comparable speed. On YouTube-VIS validation set, our method also achieves promising results. The source code is available at https://github.com/JialeCao001/SipMask. Jiale Cao, Yanwei Pang, Rao Muhammad Anwer, Hisham Cholakkal, Fahad Shahbaz Khan, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Semantic-aware self-supervised depth estimation for stereo 3D detection
Hanqing Sun 0001, Jiale Cao, Yanwei Pang |
Pattern Recognit. Lett. | 2 |
| 2023 | Real-Time Stereo 3D Car Detection With Shape-Aware Non-Uniform SamplingabstractPseudo-LiDAR based stereo 3D detectors have gained popularity due to their high accuracy. However, these methods need dense depth supervision and suffer from inferior speed. To solve these two issues, a recently introduced RTS3D builds an efficient 4D feature-consistency embedding (FCE) space for the object intermediate representation without depth supervision, where FCE space performs uniform sampling to generate feature sampling points, which ignores the importance of different object regions. In this paper, we observe that, compared with the inner region, the outer region of the object plays a more important role for accurate 3D detection. To fully exploit the useful information from the outer region, we propose a novel shape-aware non-uniform sampling strategy. Instead of uniform sampling, our proposed non-uniform sampling strategy performs dense sampling in outer region and sparse sampling in inner region. Therefore, more points are sampled from the outer region and more useful features are extracted for 3D detection. In addition, we design a high-level semantic enhanced FCE module to exploit more contextual information and suppress noise better. As a result, it further improves feature discrimination of each sampling point. Experimental results on the KITTI dataset show the effectiveness of the proposed method. Compared with the baseline RTS3D, our proposed method has 2.57% improvement on car AP$_{3d}$almost without extra network parameters. Moreover, our proposed method outperforms the state-of-the-art methods without extra supervision at a real-time speed. Aqi Gao, Jiale Cao, Yanwei Pang, Xuelong Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Latent Feature Pyramid Network for Object DetectionabstractObject detection methods based on Convolution Neural Networks (CNN) usually utilize feature pyramid networks to detect objects with various scales. The state-of-the-art feature pyramid networks improve detection accuracy by enhancing multi-level feature representations. Fusing multi-level features is the most effective manner to enhance the feature representations. However, the existing feature pyramid networks usually fuse multi-level features by element-wise operations. It leads to the lack of long-range dependencies in the feature fusion. To address the problem, we propose a simple yet efficient feature pyramid network named latent feature pyramid network (LFPN). LFPN can enhance the feature representations by modeling inner-scale and cross-scale long-range dependencies through conducting inner-scale and cross-scale feature fusion in the latent space. Comprehensive experiments are performed on two challenge object detection datasets: MS COCO and Pascal VOC. The experimental results show consistent improvements on various feature pyramid networks, backbones, and object detectors, which demonstrates the effectiveness and generality of our LFPN. Jin Xie 0005, Yanwei Pang, Jing Nie 0001, Jiale Cao, Jungong Han |
IEEE Trans. Multim. | 4 |
| 2023 | Hierarchical Regression and Classification for Accurate Object DetectionabstractAccurate object detection requires correct classification and high-quality localization. Currently, most of the single shot detectors (SSDs) conduct simultaneous classification and regression using a fully convolutional network. Despite high efficiency, this structure has some inappropriate designs for accurate object detection. The first one is the mismatch of bounding box classification, where the classification results of the default bounding boxes are improperly treated as the results of the regressed bounding boxes during the inference. The second one is that only one-time regression is not good enough for high-quality object localization. To solve the problem of classification mismatch, we propose a novel reg-offset-cls (ROC) module including three hierarchical steps: the regression of the default bounding box, the prediction of new feature sampling locations, and the classification of the regressed bounding box with more accurate features. For high-quality localization, we stack two ROC modules together. The input of the second ROC module is the output of the first ROC module. In addition, we inject a feature enhanced (FE) module between two stacked ROC modules to extract more contextual information. The experiments on three different datasets (i.e., MS COCO, PASCAL VOC, and UAVDT) are performed to demonstrate the effectiveness and superiority of our method. Without any bells or whistles, our proposed method outperforms state-of-the-art one-stage methods at a real-time speed. The source code is available at https://github.com/JialeCao001/HSD. Jiale Cao, Yanwei Pang, Jungong Han, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Complementary Feature Pyramid Network for Object DetectionabstractThe way of constructing a robust feature pyramid is crucial for object detection. However, existing feature pyramid methods, which aggregate multi-level features by using element-wise sum or concatenation, are inefficient to construct a robust feature pyramid. The reason is that these methods cannot be effective in discriminating the relevant semantics of objects. In this article, we propose a Complementary Feature Pyramid Network (CFPN) to aggregate multi-level features selectively and efficiently by exploring complementary information between multi-level features. Specifically, a Spatial Complementary Module (SCM) and a Channel Complementary Module (CCM) are designed and embedded in CFPN to enhance useful information and suppress irrelevant information during feature fusions along spatial and channel dimensions, respectively. CFPN is a generic feature extractor, as evidenced by its seamless integration into single-stage, two-stage, and end-to-end object detectors. Experiments conducted on the COCO and Pascal VOC datasets demonstrate that integrating our CFPN into RetinaNet, Faster RCNN, Cascade RCNN, and Sparse RCNN obtains consistent performance improvements with negligible overheads. Code and models are available at: https://github.com/VIPLab-CQU/CFPN . Jin Xie 0005, Yanwei Pang, Jing Nie 0001, Jiale Cao, Jungong Han |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | PSTR: End-to-End One-Step Person Search With TransformersabstractWe propose a novel one-step transformer-based person search framework, PSTR, that jointly performs person detection and re-identification (re-id) in a single architecture. PSTR comprises a person search-specialized (PSS) module that contains a detection encoder-decoder for person detection along with a discriminative re-id decoder for person re-id. The discriminative re-id decoder utilizes a multi-level supervision scheme with a shared decoder for discriminative re-id feature learning and also comprises a part attention block to encode relationship between different parts of a person. We further introduce a simple multi-scale scheme to support re-id across person instances at different scales. PSTR jointly achieves the diverse objectives of object-level recognition (detection) and instance-level matching (re-id). To the best of our knowledge, we are the first to propose an end-to-end one-step transformer-based person search framework. Experiments are performed on two popular benchmarks: CUHK-SYSU and PRW. Our extensive ablations reveal the merits of the proposed contributions. Further, the proposed PSTR sets a new state-of-the-art on both benchmarks. On the challenging PRW benchmark, PSTR achieves a mean average precision (mAP) score of 56.5%. The source code is available at https://github.com/JialeCao001/PSTR. Jiale Cao, Yanwei Pang, Rao Muhammad Anwer, Hisham Cholakkal, Jin Xie 0005, Mubarak Shah, Fahad Shahbaz Khan |
CVPR | 1 |
| 2022 | Video Instance Segmentation via Multi-Scale Spatio-Temporal Split Attention Transformer
Omkar Thawakar, Sanath Narayan, Jiale Cao, Hisham Cholakkal, Rao Muhammad Anwer, Muhammad Haris Khan, Salman Khan 0001, Michael Felsberg, Fahad Shahbaz Khan |
ECCV (29) | 3 |
| 2022 | Learning a Dynamic Cross-Modal Network for Multispectral Pedestrian DetectionabstractMultispectral pedestrian detection that enables continuous (day and night) localization of pedestrians has numerous applications. Existing approaches typically aggregate multispectral features by a simple element-wise operation. However, such a local feature aggregation scheme ignores the rich non-local contextual information. Further, we argue that a local tight correspondence across modalities is desired for multi-modal feature aggregation. To address these issues, we introduce a multispectral pedestrian detection framework that comprises a novel dynamic cross-modal network (DCMNet), which strives to adaptively utilize the local and non-local complementary information between multi-modal features. The proposed DCMNet consists of a local and a non-local feature aggregation module. The local module employs dynamically learned convolutions to capture local relevant information across modalities. On the other hand, the non-local module captures non-local cross-modal information by first projecting features from both modalities into the latent space and then obtaining dynamic latent feature nodes for feature aggregation. Comprehensive experiments are performed on two challenging benchmarks: KAIST and LLVIP. Experiments reveal the benefits of the proposed DCMNet, leading to consistently improved detection performance on diverse detection paradigms and backbones. When using the same backbone, our proposed detector achieves absolute gains of 1.74% and 1.90% over the baseline Cascade RCNN on the KAIST and LLVIP datasets. Jin Xie 0005, Rao Muhammad Anwer, Hisham Cholakkal, Jing Nie 0001, Jiale Cao, Jorma Laaksonen, Fahad Shahbaz Khan |
ACM Multimedia | 5 |
| 2022 | Multi-stream densely connected network for semantic segmentationabstractAbstract Semantic segmentation is a challenging task in computer vision which is widely used in autonomous driving and scene understanding. State‐of‐the‐art semantic segmentation networks, like DeepLab and PSPNet, make full use of multiple feature information to improve spatial resolution. However, the feature resolution in the scale‐axis is not dense enough for practical applications. To tackle this problem, a multi‐stream network is designed with atrous convolutional layers at multiple rates to capture objects and context at multiple scales. Furthermore, intra‐connections and inter‐connections are designed to fuse multi‐scale features densely which produce a feature pyramid with much larger scale diversity and larger receptive field by involving small quantity of computation. The proposed module can be easily used in other methods and it helps to increase the performance. Compared with existing methods, the proposed network, called Multi‐stream Densely Connected Network, reaches competitive results on ADE20K dataset, PASCAL VOC 2012 dataset, and Cityscapes dataset. Dayu Jia, Jiale Cao, Yanwei Pang |
IET Comput. Vis. | 2 |
| 2022 | Improving 2D object detection with binocular images for outdoor surveillance
Fuchen Chu, Yanwei Pang, Jiale Cao, Jing Nie 0001, Xuelong Li 0001 |
Neurocomputing | 3 |
| 2022 | PCNet: Paired channel feature volume network for accurate and efficient depth estimation
Dayu Jia, Yanwei Pang, Jiale Cao |
Neurocomputing | 3 |
| 2022 | Saliency detection network with two-stream encoder and interactive decoder
Ai-Ping Yang, Simeng Cheng, Shangyang Song, Jinbin Wang, Zhong Ji, Yanwei Pang, Jiale Cao |
Neurocomputing | 7 |
| 2022 | Non-linear perceptual multi-scale network for single image super-resolution
Ai-Ping Yang, Jinbin Wang, Zhong Ji, Yanwei Pang, Jiale Cao, Zihao Wei |
Neural Networks | 6 |
| 2022 | From Handcrafted to Deep Features for Pedestrian Detection: A SurveyabstractPedestrian detection is an important but challenging problem in computer vision, especially in human-centric tasks. Over the past decade, significant improvement has been witnessed with the help of handcrafted features and deep features. Here we present a comprehensive survey on recent advances in pedestrian detection. First, we provide a detailed review of single-spectral pedestrian detection that includes handcrafted features based methods and deep features based approaches. For handcrafted features based methods, we present an extensive review of approaches and find that handcrafted features with large freedom degrees in shape and space have better performance. In the case of deep features based approaches, we split them into pure CNN based methods and those employing both handcrafted and CNN based features. We give the statistical analysis and tendency of these methods, where feature enhanced, part-aware, and post-processing methods have attracted main attention. In addition to single-spectral pedestrian detection, we also review multi-spectral pedestrian detection, which provides more robust features for illumination variance. Furthermore, we introduce some related datasets and evaluation metrics, and a deep experimental analysis. We conclude this survey by emphasizing open problems that need to be addressed and highlighting various future directions. Researchers can track an up-to-date list at https://github.com/JialeCao001/PedSurvey. Jiale Cao, Yanwei Pang, Jin Xie 0005, Fahad Shahbaz Khan, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Continuous Seizure Detection Based on Transformer and Long-Term iEEGabstractAutomatic seizure detection algorithms are necessary for patients with refractory epilepsy. Many excellent algorithms have achieved good results in seizure detection. Still, most of them are based on discontinuous intracranial electroencephalogram (iEEG) and ignore the impact of different channels on detection. This study aimed to evaluate the proposed algorithm using continuous, long-term iEEG to show its applicability in clinical routine. In this study, we introduced the ability of the transformer network to calculate the attention between the channels of input signals into seizure detection. We proposed an end-to-end model that included convolution and transformer layers. The model did not need feature engineering or format transformation of the original multi-channel time series. Through evaluation on two datasets, we demonstrated experimentally that the transformer layer could improve the performance of the seizure detection algorithm. For the SWEC-ETHZ iEEG dataset, we achieved 97.5% event-based sensitivity, 0.06/h FDR, and 13.7 s latency. For the TJU-HH iEEG dataset, we achieved 98.1% event-based sensitivity, 0.22/h FDR, and 9.9 s latency. In addition, statistics showed that the model allocated more attention to the channels close to the seizure onset zone within 20 s after the seizure onset, which improved the explainability of the model. This paper provides a new method to improve the performance and explainability of automatic seizure detection. Weipeng Jin, Xiaopeng Si, Jiale Cao, Shaoya Yin, Dong Ming |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | Co-mining: Self-Supervised Learning for Sparsely Annotated Object DetectionabstractObject detectors usually achieve promising results with the supervision of complete instance annotations. However, their performance is far from satisfactory with sparse instance annotations. Most existing methods for sparsely annotated object detection either re-weight the loss of hard negative samples or convert the unlabeled instances into ignored regions to reduce the interference of false negatives. We argue that these strategies are insufficient since they can at most alleviate the negative effect caused by missing annotations. In this paper, we propose a simple but effective mechanism, called Co-mining, for sparsely annotated object detection. In our Co-mining, two branches of a siamese network predict the pseudo-label sets for each other. To enhance multi-view learning and better mine unlabeled instances, the original image and corresponding augmented image are used as the inputs of two branches of the siamese network, respectively. Co-mining can serve as a general training mechanism applied to most of modern object detectors. Experiments are performed on MS COCO dataset with three different sparsely annotated settings using two typical frameworks: anchor-based detector RetinaNet and anchor-free detector FCOS. Experimental results show that our Co-mining with RetinaNet achieves 1.4%∼2.1% improvements compared with different baselines and surpasses existing methods under the same sparsely annotated setting. Tiancai Wang, Tong Yang 0005, Jiale Cao, Xiangyu Zhang 0005 |
AAAI | 3 |
| 2021 | Track To Detect and Segment: An Online Multi-Object TrackerabstractMost online multi-object trackers perform object detection stand-alone in a neural net without any input from tracking. In this paper, we present a new online joint detection and tracking model, TraDeS (TRAck to DEtect and Segment), exploiting tracking clues to assist detection end-to-end. TraDeS infers object tracking offset by a cost volume, which is used to propagate previous object features for improving current object detection and segmentation. Effectiveness and superiority of TraDeS are shown on 4 datasets, including MOT (2D tracking), nuScenes (3D tracking), MOTS and Youtube-VIS (instance segmentation tracking). Project page: https://jialianwu.com/projects/TraDeS.html. Jialian Wu, Jiale Cao, Liangchen Song, Yu Wang 0032, Ming Yang 0007, Junsong Yuan 0001 |
CVPR | 2 |
| 2021 | Improving Single Shot Object Detection With Feature Scale UnmixingabstractDue to the advantages of real-time detection and improved performance, single-shot detectors have gained great attention recently. To solve the complex scale variations, single-shot detectors make scale-aware predictions based on multiple pyramid layers. Typically, small objects are detected on shallow layers while large objects are detected on deep layers. However, the features in the pyramid are not scale-aware enough, which limits the detection performance. Two common problems in single-shot detectors caused by object scale variations can be observed: (1) false negative problem, i.e., small objects are easily missed due to the weak features; (2) part-false positive problem, i.e., the salient part of a large object is sometimes detected as an object. With this observation, a new Neighbor Erasing and Transferring (NET) mechanism is proposed for feature scale-unmixing to explore scale-aware features in this paper. In NET, a Neighbor Erasing Module (NEM) is designed to erase the salient features of large objects and emphasize the features of small objects in shallow layers. A Neighbor Transferring Module (NTM) is introduced to transfer the erased features and highlight large objects in deep layers. With this mechanism, a single-shot network called NETNet is constructed for scale-aware object detection. In addition, we propose to aggregate nearest neighboring pyramid features to enhance our NET. Experiments on MS COCO dataset and UAVDT dataset demonstrate the effectiveness of our method. NETNet obtains 38.5% AP at a speed of 27 FPS and 32.0% AP at a speed of 55 FPS on MS COCO dataset. As a result, NETNet achieves a better trade-off for real-time and accurate object detection. Yazhao Li, Yanwei Pang, Jiale Cao, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | TJU-DHD: A Diverse High-Resolution Dataset for Object DetectionabstractVehicles, pedestrians, and riders are the most important and interesting objects for the perception modules of self-driving vehicles and video surveillance. However, the state-of-the-art performance of detecting such important objects (esp. small objects) is far from satisfying the demand of practical systems. Large-scale, rich-diversity, and high-resolution datasets play an important role in developing better object detection methods to satisfy the demand. Existing public large-scale datasets such as MS COCO collected from websites do not focus on the specific scenarios. Moreover, the popular datasets (e.g., KITTI and Citypersons) collected from the specific scenarios are limited in the number of images and instances, the resolution, and the diversity. To attempt to solve the problem, we build a diverse high-resolution dataset (called TJU-DHD). The dataset contains 115354 high-resolution images (52% images have a resolution of 1624×1200 pixels and 48% images have a resolution of at least 2, 560×1.440 pixels) and 709 330 labeled objects in total with a large variance in scale and appearance. Meanwhile, the dataset has a rich diversity in season variance, illumination variance, and weather variance. In addition, a new diverse pedestrian dataset is further built. With the four different detectors (i.e., the one-stage RetinaNet, anchor-free FCOS, two-stage FPN, and Cascade R-CNN), experiments about object detection and pedestrian detection are conducted. We hope that the newly built dataset can help promote the research on object detection and pedestrian detection in these two scenes. The dataset is available at https://github.com/tjubiit/TJU-DHD. Yanwei Pang, Jiale Cao, Yazhao Li, Jin Xie 0005, Hanqing Sun 0001, Jinfeng Gong |
IEEE Trans. Image Process. | 2 |
| 2020 | D2Det: Towards High Quality Object Detection and Instance SegmentationabstractWe propose a novel two-stage detection method, D2Det, that collectively addresses both precise localization and accurate classification. For precise localization, we introduce a dense local regression that predicts multiple dense box offsets for an object proposal. Different from traditional regression and keypoint-based localization employed in two-stage detectors, our dense local regression is not limited to a quantized set of keypoints within a fixed region and has the ability to regress position-sensitive real number dense offsets, leading to more precise localization. The dense local regression is further improved by a binary overlap prediction strategy that reduces the influence of background region on the final box regression. For accurate classification, we introduce a discriminative RoI pooling scheme that samples from various sub-regions of a proposal and performs adaptive weighting to obtain discriminative features. On MS COCO test-dev, our D2Det outperforms existing two-stage methods, with a single-model performance of 45.4 AP, using ResNet101 backbone. When using multi-scale training and inference, D2Det obtains AP of 50.1. In addition to detection, we adapt D2Det for instance segmentation, achieving a mask AP of 40.2 with a two-fold speedup, compared to the state-of-the-art. We also demonstrate the effectiveness of our D2Det on airborne sensors by performing experiments for object detection in UAV images (UAVDT dataset) and instance segmentation in satellite images (iSAID dataset). Source code is available at https://github.com/JialeCao001/D2Det. Jiale Cao, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan, Yanwei Pang, Ling Shao 0001 |
CVPR | 1 |
| 2020 | NETNet: Neighbor Erasing and Transferring Network for Better Single Shot Object DetectionabstractDue to the advantages of real-time detection and improved performance, single-shot detectors have gained great attention recently. To solve the complex scale variations, single-shot detectors make scale-aware predictions based on multiple pyramid layers. However, the features in the pyramid are not scale-aware enough, which limits the detection performance. Two common problems in single-shot detectors caused by object scale variations can be observed: (1) small objects are easily missed; (2) the salient part of a large object is sometimes detected as an object. With this observation, we propose a new Neighbor Erasing and Transferring (NET) mechanism to reconfigure the pyramid features and explore scale-aware features. In NET, a Neighbor Erasing Module (NEM) is designed to erase the salient features of large objects and emphasize the features of small objects in shallow layers. A Neighbor Transferring Module (NTM) is introduced to transfer the erased features and highlight large objects in deep layers. With this mechanism, a single-shot network called NETNet is constructed for scale-aware object detection. In addition, we propose to aggregate nearest neighboring pyramid features to enhance our NET. NETNet achieves 38.5% AP at a speed of 27 FPS and 32.0% AP at a speed of 55 FPS on MS COCO dataset. As a result, NETNet achieves a better trade-off for real-time and accurate object detection. Yazhao Li, Yanwei Pang, Jianbing Shen, Jiale Cao, Ling Shao 0001 |
CVPR | 4 |
| 2020 | SipMask: Spatial Information Preservation for Fast Image and Video Instance Segmentation
Jiale Cao, Rao Muhammad Anwer, Hisham Cholakkal, Fahad Shahbaz Khan, Yanwei Pang, Ling Shao 0001 |
ECCV (14) | 1 |
| 2020 | High-Level Semantic Networks for Multi-Scale Object DetectionabstractTo better solve scale variance problem, deep multi-scale methods usually detect objects of different scales by different in-network layers. However, the semantic levels of features from different layers are usually inconsistent. In this paper, we propose a multi-branch and high-level semantic network by gradually splitting a base network into multiple different branches. As a result, the different branches have same depth and the output features of different branches have similarly high-level semantics. Due to the difference of receptive fields, the different branches are suitable to detect objects of different scales. Meanwhile, the multi-branch network does not introduce additional parameters by sharing the convolutional weights of different branches. To further improve detection performance, skip-layer connections are used to add context to the branch of relatively small receptive field, and dilated convolution is incorporated to enlarge the resolutions of output feature maps. When they are embedded into Faster RCNN architecture, the weighted scores of proposal generation network and proposal classification network are further proposed. Experiments on three pedestrian datasets (i.e., the KITTI dataset, the Caltech dataset, and the Citypersons dataset), one face dataset (i.e., the WIDER FACE dataset), and two general object datasets (i.e., the COCO benchmark and the PASCAL VOC dataset) demonstrate the effectiveness and generality of proposed method. On these datasets, our method achieves state-of-the-art performance. Jiale Cao, Yanwei Pang, Shengjie Zhao 0001, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Taking a Look at Small-Scale Pedestrians and Occluded PedestriansabstractSmall-scale pedestrian detection and occluded pedestrian detection are two challenging tasks. However, most state-of-the-art methods merely handle one single task each time, thus giving rise to relatively poor performance when the two tasks, in practice, are required simultaneously. In this paper, it is found that small-scale pedestrian detection and occluded pedestrian detection actually have a common problem, i.e., an inaccurate location problem. Therefore, solving this problem enables to improve the performance of both tasks. To this end, we pay more attention to the predicted bounding box with worse location precision and extract more contextual information around objects, where two modules (i.e., location bootstrap and semantic transition) are proposed. The location bootstrap is used to reweight regression loss, where the loss of the predicted bounding box far from the corresponding ground-truth is upweighted and the loss of the predicted bounding box near the corresponding ground-truth is downweighted. Additionally, the semantic transition adds more contextual information and relieves semantic inconsistency of the skip-layer fusion. Since the location bootstrap is not used at the test stage and the semantic transition is lightweight, the proposed method does not add many extra computational costs during inference. Experiments on the challenging CityPersons and Caltech datasets show that the proposed method outperforms the state-of-the-art methods on the small-scale pedestrians and occluded pedestrians (e.g., 5.20% and 4.73% improvements on the Caltech). Jiale Cao, Yanwei Pang, Jungong Han, Bolin Gao, Xuelong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Triply Supervised Decoder Networks for Joint Detection and SegmentationabstractJoint object detection and semantic segmentation is essential in many fields such as self-driving cars. An initial attempt towards this goal is to simply share a single network for multi-task learning. We argue that it does not make full use of the fact that detection and segmentation are mutually beneficial. In this paper, we propose a framework called TripleNet to deeply boost these two tasks. On the one hand, to deeply join the two tasks at different scales, triple supervisions including detection-oriented supervision and class-aware/agnostic segmentation supervisions are imposed on each layer of the decoder. Class-agnostic segmentation provides an objectness prior to detection and segmentation. On the other hand, to further intercross the two tasks and refine the features in each scale, two light-weight modules (i.e., the inner-connected module and the attention skip-layer fusion) are incorporated. Because segmentation supervision on each decoder layer are not performed at the test stage and two added modules are light-weight, the proposed TripleNet can run at a real-time speed (16 fps). Experiments on the VOC 2007/2012 and COCO datasets show that TripleNet outperforms all the other one-stage methods on both two tasks (e.g., 81.9% mAP and 83.3% mIoU on VOC 2012, and 37.1% mAP and 59.6% mIoU on COCO) by a single network. Jiale Cao, Yanwei Pang, Xuelong Li 0001 |
CVPR | 1 |
| 2019 | Hierarchical Shot DetectorabstractSingle shot detector simultaneously predicts object categories and regression offsets of the default boxes. Despite of high efficiency, this structure has some inappropriate designs: (1) The classification result of the default box is improperly assigned to that of the regressed box during inference, (2) Only regression once is not good enough for accurate object detection. To solve the first problem, a novel reg-offset-cls (ROC) module is proposed. It contains three hierarchical steps: box regression, the feature sampling location predication, and the regressed box classification with the features of offset locations. To further solve the second problem, a hierarchical shot detector (HSD) is proposed, which stacks two ROC modules and one feature enhanced module. The second ROC treats the regressed boxes and the feature sampling locations of features in the first ROC as the inputs. Meanwhile, the feature enhanced module injected between two ROCs aims to extract the local and non-local context. Experiments on the MS COCO and PASCAL VOC datasets demonstrate the superiority of proposed HSD. Without the bells or whistles, HSD outperforms all one-stage methods at real-time speed. Jiale Cao, Yanwei Pang, Jungong Han, Xuelong Li 0001 |
ICCV | 1 |
| 2019 | JCS-Net: Joint Classification and Super-Resolution Network for Small-Scale Pedestrian Detection in Surveillance ImagesabstractWhile convolutional neural network (CNN)-based pedestrian detection methods have proven to be successful in various applications, detecting small-scale pedestrians from surveillance images is still challenging. The major reason is that the small-scale pedestrians lack much detailed information compared to the large-scale pedestrians. To solve this problem, we propose to utilize the relationship between the large-scale pedestrians and the corresponding small-scale pedestrians to help recover the detailed information of the small-scale pedestrians, thus improving the performance of detecting small-scale pedestrians. Specifically, a unified network (called JCS-Net) is proposed for small-scale pedestrian detection, which integrates the classification task and the super-resolution task in a unified framework. As a result, the super-resolution and classification are fully engaged, and the super-resolution sub-network can recover some useful detailed information for the subsequent classification. Based on HOG+LUV and JCS-Net, multi-layer channel features (MCF) are constructed to train the detector. The experimental results on the Caltech pedestrian dataset and the KITTI benchmark demonstrate the effectiveness of the proposed method. To further enhance the detection, multi-scale MCF based on JCS-Net for pedestrian detection is also proposed, which achieves the state-of-the-art performance. Yanwei Pang, Jiale Cao, Jian Wang 0087, Jungong Han |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2018 | Randomly translational activation inspired by the input distributions of ReLU
Jiale Cao, Yanwei Pang, Xuelong Li 0001, Jingkun Liang |
Neurocomputing | 1 |
| 2017 | Learning Sampling Distributions for Efficient Object DetectionabstractObject detection is an important task in computer vision and machine intelligence systems. Multistage particle windows (MPW), proposed by Gualdi et al., is an algorithm of fast and accurate object detection. By sampling particle windows (PWs) from a proposal distribution (PD), MPW avoids exhaustively scanning the image. Despite its success, it is unknown how to determine the number of stages and the number of PWs in each stage. Moreover, it has to generate too many PWs in the initialization step and it unnecessarily regenerates too many PWs around object-like regions. In this paper, we attempt to solve the problems of MPW. An important fact we used is that there is a large probability for a randomly generated PW not to contain the object because the object is a sparse event relative to the huge number of candidate windows. Therefore, we design a PD so as to efficiently reject the huge number of nonobject windows. Specifically, we propose the concepts of rejection, acceptance, and ambiguity windows and regions. Then, the concepts are used to form and update a dented uniform distribution and a dented Gaussian distribution. This contrasts to MPW which utilizes only on region of support. The PD of MPW is acceptance-oriented whereas the PD of our method (called iPW) is rejection-oriented. Experimental results on human and face detection demonstrate the efficiency and the effectiveness of the iPW algorithm. The source code is publicly accessible. Yanwei Pang, Jiale Cao, Xuelong Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2017 | Cascade Learning by Optimally PartitioningabstractCascaded AdaBoost classifier is a well-known efficient object detection algorithm. The cascade structure has many parameters to be determined. Most of existing cascade learning algorithms are designed by assigning detection rate and false positive rate to each stage either dynamically or statically. Their objective functions are not directly related to minimum computation cost. These algorithms are not guaranteed to have optimal solution in the sense of minimizing computation cost. On the assumption that a strong classifier is given, in this paper, we propose an optimal cascade learning algorithm (iCascade) which iteratively partitions the strong classifiers into two parts until predefined number of stages are generated. iCascade searches the optimal partition point of each stage by directly minimizing the computation cost of the cascade. Theorems are provided to guarantee the existence of the unique optimal solution. Theorems are also given for the proposed efficient algorithm of searching optimal parameters . Once a new stage is added, the parameter for each stage decreases gradually as iteration proceeds, which we call decreasing phenomenon. Moreover, with the goal of minimizing computation cost, we develop an effective algorithm for setting the optimal threshold of each stage. In addition, we prove in theory why more new weak classifiers in the current stage are required compared to that of the previous stage. Experimental results on face detection and pedestrian detection demonstrate the effectiveness and efficiency of the proposed algorithm. Yanwei Pang, Jiale Cao, Xuelong Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2017 | Learning Multilayer Channel Features for Pedestrian DetectionabstractPedestrian detection based on the combination of convolutional neural network (CNN) and traditional handcrafted features (i.e., HOG+LUV) has achieved great success. In general, HOG+LUV are used to generate the candidate proposals and then CNN classifies these proposals. Despite its success, there is still room for improvement. For example, CNN classifies these proposals by the fully connected layer features, while proposal scores and the features in the inner-layers of CNN are ignored. In this paper, we propose a unifying framework called multi-layer channel features (MCF) to overcome the drawback. It first integrates HOG+LUV with each layer of CNN into a multi-layer image channels. Based on the multi-layer image channels, a multi-stage cascade AdaBoost is then learned. The weak classifiers in each stage of the multi-stage cascade are learned from the image channels of corresponding layer. Experiments on Caltech data set, INRIA data set, ETH data set, TUD-Brussels data set, and KITTI data set are conducted. With more abundant features, an MCF achieves the state of the art on Caltech pedestrian data set (i.e., 10.40% miss rate). Using new and accurate annotations, an MCF achieves 7.98% miss rate. As many non-pedestrian detection windows can be quickly rejected by the first few stages, it accelerates detection speed by 1.43 times. By eliminating the highly overlapped detection windows with lower scores after the first stage, it is 4.07 times faster than negligible performance loss. Jiale Cao, Yanwei Pang, Xuelong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Pedestrian Detection Inspired by Appearance Constancy and Shape SymmetryabstractThe discrimination and simplicity of features are very important for effective and efficient pedestrian detection. However, most state-of-the-art methods are unable to achieve good tradeoff between accuracy and efficiency. Inspired by some simple inherent attributes of pedestrians (i.e., appearance constancy and shape symmetry), we propose two new types of non-neighboring features (NNF): side-inner difference features (SIDF) and symmetrical similarity features (SSF). SIDF can characterize the difference between the background and pedestrian and the difference between the pedestrian contour and its inner part. SSF can capture the symmetrical similarity of pedestrian shape. However, it's difficult for neighboring features to have such above characterization abilities. Finally, we propose to combine both non-neighboring and neighboring features for pedestrian detection. It's found that nonneighboring features can further decrease the average miss rate by 4.44%. Experimental results on INRIA and Caltech pedestrian datasets demonstrate the effectiveness and efficiency of the proposed method. Compared to the state-of the-art methods without using CNN, our method achieves the best detection performance on Caltech, outperforming the second best method (i.e., Checkerboards) by 1.63%. Jiale Cao, Yanwei Pang, Xuelong Li 0001 |
CVPR | 1 |
| 2016 | Pedestrian Detection Inspired by Appearance Constancy and Shape SymmetryabstractMost state-of-the-art methods in pedestrian detection are unable to achieve a good trade-off between accuracy and efficiency. For example, ACF has a fast speed but a relatively low detection rate, while checkerboards have a high detection rate but a slow speed. Inspired by some simple inherent attributes of pedestrians (i.e., appearance constancy and shape symmetry), we propose two new types of non-neighboring features: side-inner difference features (SIDF) and symmetrical similarity features (SSFs). SIDF can characterize the difference between the background and pedestrian and the difference between the pedestrian contour and its inner part. SSF can capture the symmetrical similarity of pedestrian shape. However, it is difficult for neighboring features to have such above characterization abilities. Finally, we propose to combine both non-neighboring features and neighboring features for pedestrian detection. It is found that non-neighboring features can further decrease the log-average miss rate by 4.44%. The relationship between our proposed method and some state-of-the-art methods is also given. Experimental results on INRIA, Caltech, and KITTI data sets demonstrate the effectiveness and efficiency of the proposed method. Compared with the state-of-the-art methods without using CNN, our method achieves the best detection performance on Caltech, outperforming the second best method (i.e., checkerboards) by 2.27%. Using the new annotations of Caltech, it can achieve 11.87% miss rate, which outperforms other methods. Jiale Cao, Yanwei Pang, Xuelong Li 0001 |
IEEE Trans. Image Process. | 1 |