Yanwei Pang

dblp:35/5889 · DBLP profile ↗
← Back
260ranked-venue papers
46as first author
121since 2021 · last 2026
0000-0001-6670-3727ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 135 · 26 first-author · 64 since 2021Graphics, computer vision, multimedia, augmented reality and games · 114 · 17 first-author · 46 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 2 first-author · 13 since 2021Databases, data management, data science and information retrieval · 7Human-computer interaction and ubiquitous computing · 7 · 3 first-author · 2 since 2021Security and privacy · 2 · 2 first-authorComputer networks · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 iEBAKER: Improved remote sensing image-text retrieval framework via eliminate before align and keyword explicit reasoning
Yan Zhang 0135, Zhong Ji, Changxu Meng, Yanwei Pang
Expert Syst. Appl.4
2026 Mamba-based multi-slice unrolled network for accelerated prostate MR imaging
Xuebin Sun, Jinxi Wang, Yanwei Pang
Image Vis. Comput.5
2026 Multispectral remote sensing object detection via selective cross-modal interaction and aggregation
Minghao Cui, Jing Nie 0001, Hanqing Sun 0001, Jin Xie 0005, Jiale Cao, Yanwei Pang, Xuelong Li 0001
Neural Networks6
2026 SD2-SNN: Self-distillation and structural decomposition framework for SNNs in continual learning
Zhenhao Xie, Xia Xiao 0001, Yanwei Pang, Zhong Ji
Neural Networks4
2026 Parameter-Efficient Fine-Tuning for Continual Learning: A Neural Tangent Kernel Perspective
abstract
Parameter-efficient fine-tuning for continual learning (PEFT-CL) has shown promise in adapting pre-trained models to sequential tasks while mitigating catastrophic forgetting problem. However, understanding the mechanisms that dictate continual performance in this paradigm remains elusive. To unravel this mystery, we undertake a rigorous analysis of PEFT-CL dynamics to derive relevant metrics for continual scenarios using Neural Tangent Kernel (NTK) theory. With the aid of NTK as a mathematical analysis tool, we recast the challenge of test-time forgetting into the quantifiable generalization gaps during training, identifying three key factors that influence these gaps and the performance of PEFT-CL: training sample size, task-level feature orthogonality, and regularization. To address these challenges, we introduce NTK-CL, a novel framework that eliminates task-specific parameter storage while adaptively generating task-relevant features. Aligning with theoretical guidance, NTK-CL triples the feature representation of each sample, theoretically and empirically reducing the magnitude of both task-interplay and task-specific generalization gaps. Grounded in NTK analysis, our framework imposes an adaptive exponential moving average mechanism and constraints on task-level feature orthogonality, maintaining intra-task NTK forms while attenuating inter-task NTK forms. Ultimately, by fine-tuning optimizable parameters with appropriate regularization, NTK-CL achieves state-of-the-art performance on established PEFT-CL benchmarks. This work provides a theoretical foundation for understanding and improving PEFT-CL models, offering insights into the interplay between feature representation, task orthogonality, and generalization, contributing to the development of more efficient continual learning systems.
Jingren Liu, Zhong Ji, Yunlong Yu 0001, Jiale Cao, Yanwei Pang, Jungong Han, Xuelong Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 iSeg: An Iterative Refinement-Based Framework for Training-Free Segmentation
abstract
Stable Diffusion has demonstrated strong image synthesis ability to given text descriptions, suggesting it to contain strong semantic clue for grouping objects. The researchers have explored employing Stable Diffusion for training-free segmentation. Most existing approaches refine cross-attention map by self-attention map once, demonstrating that self-attention map contains useful semantic information to improve segmentation. To fully utilize self-attention map, we present a deep experimental analysis on iteratively refining cross-attention map with self-attention map, and propose an effective iterative refinement framework for training-free segmentation, named iSeg. Our iSeg introduces an entropy-reduced self-attention module that utilizes a gradient descent scheme to reduce the entropy of self-attention map, thereby suppressing the weak responses corresponding to irrelevant global information. Leveraging the entropy-reduced self-attention module, our iSeg stably improves cross-attention map with iterative refinement. Further, we design a category-enhanced cross-attention module to generate accurate cross-attention map, providing a better initial input for iterative refinement. Extensive experiments across different datasets and diverse segmentation tasks (weakly-supervised semantic segmentation, open-vocabulary semantic segmentation, unsupervised segmentation, and mask generation on synthetic dataset) reveal the merits of proposed contributions, leading to promising performance. For unsupervised semantic segmentation on Cityscapes, our iSeg achieves an absolute gain of 3.8% in terms of mIoU compared to the best existing training-free approach in literature. Moreover, our proposed iSeg can support segmentation with different kinds of images and interactions, and also be used as a post-processing, or in different frameworks, to improve training-free segmentation.
Jiale Cao, Jin Xie 0005, Fahad Shahbaz Khan, Yanwei Pang
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 SED++: A Simple Encoder-Decoder for Improved Open-Vocabulary Semantic Segmentation
abstract
Open-vocabulary semantic segmentation aims to partition an image into distinct semantic regions based on an open set of categories. Existing approaches primarily rely on image-level pre-trained vision-language models to perform this pixel-level task. In this paper, we propose SED, a simple yet effective encoder-decoder architecture for open-vocabulary semantic segmentation leveraging pre-trained vision-language models. SED consists of a hierarchical image encoder, a text encoder, and a gradual fusion decoder. The hierarchical image encoder and text encoder collaboratively generate a cost volume, which is progressively decoded by the gradual fusion decoder to produce segmentation results. In contrast to a plain encoder, the hierarchical encoder better captures image detail information while maintaining linear computational complexity with respect to input size. The gradual fusion decoder adopts a top-down structure to progressively integrate high-resolution features with the cost volume. Furthermore, a category early rejection strategy is introduced in gradual fusion decoder to filter out non-existent categories at different layers, significantly improving inference efficiency. Based on SED, we further introduce two modules, including non-label text embedding and additional category early rejection in the encoder. Moreover, we extend our method with minimal decoder modification for open-vocabulary video semantic segmentation. Extensive experiments on multiple datasets validate the effectiveness and efficiency of our proposed method. With ConvNeXt-B, our method achieves an mIoU of 34.9% on the ADE20 K with 150 classes (i.e., A-150) at an inference speed of 69 ms per image on a single A6000 GPU, and has an mIoU score of 40.2% on video segmentation dataset VSPW.
Jiale Cao, Jin Xie 0005, Fahad Shahbaz Khan, Yanwei Pang
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 SNNSIR: A fully Spiking Neural Network for Stereo Image Restoration
Jin Xie 0005, Jing Nie 0001, Jiale Cao, Yanwei Pang
Pattern Recognit.5
2026 A Fresh Look at Generalized Category Discovery Through Non-Negative Matrix Factorization
abstract
Generalized Category Discovery (GCD) aims to classify both base and novel images using labeled base data. However, current approaches inadequately address the intrinsic optimization of the co-occurrence matrix A¯ based on cosine similarity, failing to achieve zero base-novel regions and adequate sparsity in base and novel domains. To address these deficiencies, we propose a Non-Negative Generalized Category Discovery (NN-GCD) framework. By establishing within the Symmetric Non-negative Matrix Factorization (SNMF) framework: (i) the equivalence between ideal k-means clustering and ideal SNMF, and (ii) the equivalence between SNMF solvers and Non-negative Contrastive Learning (NCL) optimization, we reformulate both the optimization of A¯ and k-means clustering as an NCL optimization problem. Moreover, to satisfy the non-negative constraints and make a GCD model converge to a near-ideal region, we propose a GELU activation function and an NMF NCE loss. To transition A¯ from a near-ideal state to the desired A¯∗, we introduce a hybrid sparse regularization approach to impose sparsity constraints. Experimental results show NN-GCD outperforms state-of-the-art methods on GCD benchmarks, achieving an average accuracy of 66.9% on the Semantic Shift Benchmark, surpassing prior counterparts by 2.5%.
Zhong Ji, Jingren Liu, Yanwei Pang
IEEE Trans. Circuits Syst. Video Technol.4
2026 Vision-Language Enhancement Network Based on Decoupling-Joint Adaptation for Few-Shot Action Recognition
abstract
Learning robust and generalizable feature extractors to generate discriminative prototypes is crucial for few-shot action recognition. However, most existing methods rely on fine-tuning large pre-trained image models, easily leading to transferability and overfitting issues. In this paper, we propose a novel vision-language enhancement network based on decoupling-joint adaptation (VEDA) for few-shot action recognition, which decouples visual features into temporal and spatial branches, followed by a joint operation that integrates these two branches using an adapter-tuning paradigm. VEDA can gradually equip the model with spatio-temporal reasoning capabilities. Since relying exclusively on local frame feature matching results in inaccurate performance, we design a video-level relation module (VLR) to enhance video context awareness through global feature matching. In addition, we design a vision-language fusion module (VLF) that introduces multimodal information to alleviate the data scarcity issue. Simultaneously, we apply adapter-tuning to both visual and textual branches to enhance the generalization ability. Based on the proposed components above, our network can extract both informative and discriminative prototypes, resulting in excellent recognition performance. Experimental results on five challenging benchmarks demonstrate the effectiveness of the proposed VEDA. The code will be released soon at https://github.com/ReverseSuzhou/VEDA.
Suzhou Que, Hanyu Guo, Kaiwen Du, Yan Yan 0001, Yanwei Pang, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.5
2026 Cyclic Pseudo-Label Generation and Refinement for Weakly Supervised Referring Expression Grounding
abstract
Weakly supervised Referring Expression Grounding (WREG) aims at grounding the target region based on a given expression, where the mapping between regions and expressions is unknown during training. Recent WREG methods leverage the strategy of generating pseudo-labels utilizing Vision-Language Pre-training (VLP) to avoid the cross-modal heterogeneous gaps arising from the two-stage reconstruction strategy. However, mainstream VLPs are trained with image-text alignment data, which makes the generated labels inapplicable to REG task. Furthermore, due to the constraints of WREG data, it is challenging to ensure the quality of the pseudo-labels. To this end, we propose a Cyclic Pseudo-label Generation and Refinement (CPGR) method to alleviate the above limitations. Specifically, we cycle through the process of Generation-Refinement-Grounding to alleviate the impact of missing region annotations. We perform REG task-adaptive fine-tuning on BLIP-2 to generate REG-style descriptions with Region-Centrality. Then, we design a Pseudo-label Refinement module by utilizing cross-modal token attention to enhance the reliability of pseudo-labels and ensure their Reference-Discrimination. Experiments on five benchmark datasets demonstrate that our proposed method outperforms the current state-of-the-art weakly supervised methods. Our code and models will be released at https://github.com/5jiahe/CPGR.
Jiahe Wu, Zhong Ji, Yanwei Pang, Jungong Han
IEEE Trans. Circuits Syst. Video Technol.4
2026 Underlying Semantic Diffusion for Effective and Efficient In-Context Learning
abstract
Diffusion models have emerged as a powerful framework for tasks like image controllable generation and dense prediction. However, existing models often struggle to capture underlying semantics (e.g., edges, textures, shapes) and effectively utilize in-context learning, limiting their contextual understanding and image generation quality. Furthermore, high computational costs and slow inference speeds hinder their real-time applications. To address these challenges, we propose Underlying Semantic Diffusion (US-Diffusion), an enhanced diffusion model that improves underlying semantics learning, computational efficiency, and in-context learning capabilities on multi-task scenarios. We introduce Separate & Gather Adapter (SGA), which decouples input conditions for different tasks while sharing the architecture, enabling better in-context learning and generalization across diverse visual domains. We also present a Feedback-Aided Learning (FAL) framework, which leverages feedback signals to guide the model in capturing semantic details and dynamically adapting to task-specific contextual cues. Furthermore, we propose a plug-and-play Efficient Sampling Strategy (ESS) for dense sampling at time steps with high-noise levels, which aims at optimizing training and inference efficiency while maintaining strong in-context learning performance. Experimental results demonstrate that US-Diffusion outperforms the state-of-the-art method, achieving an average reduction of 7.47 in FID on Map2Image tasks and an average reduction of 0.026 in RMSE on Image2Map tasks, while achieving approximately $9.45\times $ faster inference speed. Our method also demonstrates superior training efficiency and in-context learning capabilities, excelling in new datasets and tasks, highlighting its robustness and adaptability across diverse visual domains. The source code will be released at https://github.com/dragon-cao/US-Diffusion.
Zhong Ji, Weilong Cao, Yan Zhang 0135, Yanwei Pang, Jungong Han
IEEE Trans. Image Process.4
2026 Interpretable Few-Shot Image Classification via Prototypical Concept-Guided Mixture of LoRA Experts
abstract
Self-Explainable Models (SEMs) rely on Prototypical Concept Learning (PCL) to enable their visual recognition processes more interpretable, but they often struggle in data-scarce settings where insufficient training samples lead to suboptimal performance. To address this limitation, we propose a Few-Shot Prototypical Concept Classification (FSPCC) framework that systematically mitigates two key challenges under low-data regimes: parametric imbalance and representation misalignment. Specifically, our approach leverages a Mixture of LoRA Experts (MoLE) for parameter-efficient adaptation, ensuring a balanced allocation of trainable parameters between the backbone and the PCL module. Meanwhile, cross-module concept guidance enforces tight alignment between the backbone's feature representations and the prototypical concept activation patterns. In addition, we incorporate a multi-level feature preservation strategy that fuses spatial and semantic cues across various layers, thereby enriching the learned representations and mitigating the challenges posed by limited data availability. Finally, to enhance interpretability and minimize concept overlap, we introduce a geometry-aware concept discrimination loss that enforces orthogonality among concepts, encouraging more disentangled and transparent decision boundaries. Experimental results on six popular benchmarks (CUB-200-2011, mini-ImageNet, CIFAR-FS, Stanford Cars, FGVC-Aircraft, and DTD) demonstrate that our approach consistently outperforms existing SEMs by a notable margin, with 4.2%-8.7% relative gains in 5-way 5-shot classification. These findings highlight the efficacy of coupling concept learning with few-shot adaptation to achieve both higher accuracy and clearer model interpretability, paving the way for more transparent visual recognition systems.
Zhong Ji, Rongshuai Wei, Jingren Liu, Yanwei Pang, Jungong Han
IEEE Trans. Image Process.4
2026 HG-LMM: Unleashing High-Quality Pixel Grounding Capabilities in Frozen Large Multimodal Models
abstract
Large Multimodal Models (LMMs) have demonstrated remarkable capabilities in multimodal understanding and conversation. Recently, some researchers have explored fine-tuning LMM for pixel grounding, leading to catastrophic loss of their inherent conversational capabilities. To preserve the conversational capability, some researchers explore to freeze LMM, but employ heavy segmenter SAM for high-quality grounding. In this paper, we propose a novel approach, named HG-LMM, to fully exploit the inherent features of frozen LMM for high-quality pixel grounding. Our HG-LMM introduces two main modules: an LLM-guided instance-aware feature generation (LIFG) and a layer-wise detail and semantic injection (LDSI). The LIFG module employs the output text embeddings of LMM belonging to grounded instances to generate multi-level instance-aware feature maps from the image encoder. Afterwards, we employ the LDSI module to inject more detail and semantic information into these instance-aware feature maps. With these instance-aware feature maps, we employ a simple top-down fusion to predict the segmentation masks of different instances. We perform experiments on various tasks, including referring expression segmentation, panoramic narrative grounding, reasoning segmentation, grounded conversation generation, and visual chain-of-thought reasoning. When using DeepSeekVL-1.3B, our HG-LMM is 12.6% better than F-LMM without SAM in terms of segmentation accuracy on the all set of PNG dataset. Compared to F-LMM with SAM, our HG-LMM achieves comparable segmentation accuracy while being 2.7 times faster. We release our source code and models at https://github.com/WenjieLi2008/HG-LMM.
Jiale Cao, Jin Xie 0005, Ai-Ping Yang, Yanwei Pang
IEEE Trans. Image Process.5
2026 Frequency-Decomposed Interaction Network for Stereo Image Restoration
abstract
Stereo image restoration in adverse environments, such as low-light conditions, rain, and low resolution, requires effective exploitation of cross-view complementary information to recover degraded visual content. In monocular image restoration, frequency decomposition has proven effective, where high-frequency components aid in recovering fine textures and reducing blur, while low-frequency components facilitate noise suppression and illumination correction. However, existing stereo restoration methods have yet to explore cross-view interactions by frequency decomposition, which is a promising direction for enhancing restoration quality. To address this, we propose a frequency-aware framework comprising a Frequency Decomposition Module (FDM), Detail Interaction Module (DIM), Structural Interaction Module (SIM), and Adaptive Fusion Module (AFM). FDM employs learnable filters to decompose the image into high- and low-frequency components. DIM enhances the high-frequency branch by capturing local detail cues through deformable convolution. SIM processes the low-frequency branch by modeling global structural correlations via a cross-view row-wise attention mechanism. Finally, AFM adaptively fuses the complementary frequency-specific information to generate high-quality restored images. Extensive experiments demonstrate the efficacy and generalizability of our framework across three diverse stereo restoration tasks, where it achieves state-of-the-art performance in low-light enhancement, rain removal, alongside highly competitive results in super-resolution. Our code is available at https://github.com/C2022J/FDIN.
Xianmin Tian, Jin Xie 0005, Jing Nie 0001, Jiale Cao, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Image Process.6
2026 Multi-Stage Knowledge Integration of Vision-Language Models for Continual Learning
abstract
Vision Language Models (VLMs), pre-trained on large-scale image-text datasets, enable zero-shot predictions for unseen data but may underperform on specific unseen tasks. Continual learning (CL) can help VLMs effectively adapt to new data distributions without joint training, but faces challenges of catastrophic forgetting and generalization forgetting. Although significant progress has been achieved by distillation-based methods, they exhibit two severe limitations. One is the popularly adopted single-teacher paradigm fails to impart comprehensive knowledge, The other is the existing methods inadequately leverage the multimodal information in the original training dataset, instead they rely on additional data for distillation, which increases computational and storage overhead. To mitigate both limitations, by drawing on Knowledge Integration Theory (KIT), we propose a Multi-Stage Knowledge Integration network (MulKI) to emulate the human learning process in distillation methods. MulKI achieves this through four stages, including Eliciting Ideas, Adding New Ideas, Distinguishing Ideas, and Making Connections. During the four stages, we first leverage prototypes to align across modalities, eliciting cross-modal knowledge, then adding new knowledge by constructing fine-grained intra- and inter-modality relationships with prototypes. After that, knowledge from two teacher models is adaptively distinguished and re-weighted. Finally, we connect between models from intra- and inter-task, integrating preceding and new knowledge. Our method demonstrates significant improvements in maintaining zero-shot capabilities while supporting continual learning across diverse downstream tasks, showcasing its potential in adapting VLMs to evolving data distributions.
Zhong Ji, Jingren Liu, Yanwei Pang, Jungong Han
IEEE Trans. Image Process.4
2025 CLIPeR: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation
abstract
Contrastive Language-Image Pre-training (CLIP) exhibits strong zero-shot classification ability on various image-level tasks, leading to the research to adapt CLIP for pixel-level open-vocabulary semantic segmentation without additional training. The key is to improve spatial representation of image-level CLIP, such as replacing self-attention map at last layer with self-self attention map or vision foundation model based attention map. In this paper, we present a novel hierarchical framework, named CLIPer, that hierarchically improves spatial representation of CLIP. The proposed CLIPer includes an early-layer fusion module and a fine-grained compensation module. We observe that, the embeddings and attention maps at early layers can preserve spatial structural information. Inspired by this, we design the early-layer fusion module to generate segmentation map with better spatial coherence. Afterwards, we employ a fine-grained compensation module to compensate the local details using the self-attention maps of diffusion model. We conduct the experiments on seven segmentation datasets. Our proposed CLIPer achieves the state-of-the-art performance on these datasets. For instance, using ViT-L, CLIPer has the mIoU of 69.8% and 43.3% on VOC and COCO Object, outperforming ProxyCLIP by 9.2% and 4.1% respectively.
Jiale Cao, Jin Xie 0005, Xiaoheng Jiang, Yanwei Pang
ICCV5
2025 HGOT: Self-supervised Heterogeneous Graph Neural Network with Optimal Transport
abstract
Heterogeneous Graph Neural Networks (HGNNs), have demonstrated excellent capabilities in processing heterogeneous information networks. Self-supervised learning on heterogeneous graphs, especially contrastive self-supervised strategy, shows great potential when there are no labels. However, this approach requires the use of carefully designed graph augmentation strategies and the selection of positive and negative samples. Determining the exact level of similarity between sample pairs is non-trivial.To solve this problem, we propose a novel self-supervised Heterogeneous graph neural network with Optimal Transport (HGOT) method which is designed to facilitate self-supervised learning for heterogeneous graphs without graph augmentation strategies. Different from traditional contrastive self-supervised learning, HGOT employs the optimal transport mechanism to relieve the laborious sampling process of positive and negative samples. Specifically, we design an aggregating view (central view) to integrate the semantic information contained in the views represented by different meta-paths (branch views). Then, we introduce an optimal transport plan to identify the transport relationship between the semantics contained in the branch view and the central view. This allows the optimal transport plan between graphs to align with the representations, forcing the encoder to learn node representations that are more similar to the graph space and of higher quality. Extensive experiments on four real-world datasets demonstrate that our proposed HGOT model can achieve state-of-the-art performance on various downstream tasks. In particular, in the node classification task, HGOT achieves an average of more than 6\% improvement in accuracy compared with state-of-the-art methods.
Yanbei Liu, Chongxu Wang, Zhitao Xiao, Lei Geng, Yanwei Pang, Xiao Wang 0017
ICML5
2025 Concept agent network for zero-base generalized few-shot learning
Xuan Wang 0016, Zhong Ji, Xiyao Liu 0002, Yanwei Pang, Xuelong Li 0001
Appl. Intell.4
2025 Mitigating forgetting in the adaptation of CLIP for few-shot classification
Jiale Cao, Yuanheng Liu, Zhong Ji, Jingren Liu, Ai-Ping Yang, Yanwei Pang
Comput. Vis. Image Underst.6
2025 fRAKI: k-space deep learning with offline data-universal and online scan-specific priors
Xuebin Sun, Yanwei Pang
Neurocomputing5
2025 Video Wire Inpainting via Hierarchical Feature Mixture
Zhong Ji, Yimu Su, Yan Zhang 0135, Shuangming Yang, Yanwei Pang
Image Vis. Comput.5
2025 Hierarchical and complementary experts transformer with momentum invariance for image-text retrieval
Yan Zhang 0135, Zhong Ji, Yanwei Pang, Jungong Han
Knowl. Based Syst.3
2025 The state-of-the-art in cardiac MRI reconstruction: Results of the CMRxRecon challenge in MICCAI 2023
Chen Qin, Shuo Wang 0011, Fanwen Wang, Yan Li 0064, Zi Wang 0005, Kunyuan Guo, Ouyang Cheng, Michael Tänzer, Longyu Sun, Mengting Sun, Zhang Shi, Sha Hua, Hao Li 0082, Zhensen Chen, Bingyu Xin, Dimitris N. Metaxas, George Yiasemis, Jonas Teuwen, Weitian Chen, Yidong Zhao, Yanwei Pang, Artem Razumov, Dmitry V. Dylov, Quan Dou, Yuyang Xue, Yuning Du, Julia Dietlmeier, Carles García-Cabrera, Ziad Al-Haj Hemidi, Nora Vogt, Ying-Hua Chu, Weibo Chen, Wenjia Bai, Xiahai Zhuang, Harry Qin, Lianming Wu, Guang Yang 0006, Xiaobo Qu 0001, He Wang 0016, Chengyan Wang
Medical Image Anal.27
2025 Token-aware and step-aware acceleration for Stable Diffusion
Ting Zhen, Jiale Cao, Xuebin Sun, Zhong Ji, Yanwei Pang
Pattern Recognit.6
2025 CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation
abstract
Open-vocabulary video instance segmentation strives to segment and track instances belonging to an open set of categories in a videos. The vision-language model Contrastive Language-Image Pre-training (CLIP) has shown robust zero-shot classification ability in image-level open-vocabulary tasks. In this paper, we propose a simple encoder-decoder network, called CLIP-VIS, to adapt CLIP for open-vocabulary video instance segmentation. Our CLIP-VIS adopts frozen CLIP and introduces three modules, including class-agnostic mask generation, temporal topK-enhanced matching, and weighted open-vocabulary classification. Given a set of initial queries, class-agnostic mask generation introduces a pixel decoder and a transformer decoder on CLIP pre-trained image encoder to predict query masks and corresponding object scores and mask IoU scores. Then, temporal topK-enhanced matching performs query matching across frames using the K mostly matched frames. Finally, weighted open-vocabulary classification first employs mask pooling to generate query visual features from CLIP pre-trained image encoder, and second performs weighted classification using object scores and mask IoU scores. Our CLIP-VIS does not require the annotations of instance categories and identities. The experiments are performed on various video instance segmentation datasets, which demonstrate the effectiveness of our proposed method, especially for novel categories. When using ConvNeXt-B as backbone, our CLIP-VIS achieves the AP and APn scores of 32.2% and 40.2% on the validation set of LV-VIS dataset, which outperforms OV2Seg by 11.1% and 23.9% respectively. We will release the source code and models athttps://github.com/zwq456/CLIP-VIS.git.
Jiale Cao, Jin Xie 0005, Shuangming Yang, Yanwei Pang
IEEE Trans. Circuits Syst. Video Technol.5
2025 Raformer: Redundancy-Aware Transformer for Video Wire Inpainting
abstract
Video Wire Inpainting (VWI) is a prominent application in video inpainting, aimed at flawlessly removing wires in films or TV series, offering significant time and labor savings compared to manual frame-by-frame removal. However, wire removal poses greater challenges due to the wires being longer and slimmer than objects typically targeted in general video inpainting tasks, and often intersecting with people and background objects irregularly, which adds complexity to the inpainting process. Recognizing the limitations posed by existing video wire datasets, which are characterized by their small size, poor quality, and limited variety of scenes, we introduce a new VWI dataset with a novel mask generation strategy, namely Wire Removal Video Dataset 2 (WRV2) and Pseudo Wire-Shaped (PWS) Masks. WRV2 dataset comprises over 4,000 videos with an average length of 80 frames, designed to facilitate the development and efficacy of inpainting models. Building upon this, our research proposes the Redundancy-Aware Transformer (Raformer) method that addresses the unique challenges of wire removal in video inpainting. Unlike conventional approaches that indiscriminately process all frame patches, Raformer employs a novel strategy to selectively bypass redundant parts, such as static background segments devoid of valuable information for inpainting. At the core of Raformer is the Redundancy-Aware Attention (RAA) module, which isolates and accentuates essential content through a coarse-grained, window-based attention mechanism. This is complemented by a Soft Feature Alignment (SFA) module, which refines these features and achieves end-to-end feature alignment. Extensive experiments on both the traditional video inpainting datasets and our proposed WRV2 dataset demonstrate that Raformer outperforms other state-of-the-art methods. Our codes and the WRV2 dataset will be made available at: https://github.com/Suyimu/WRV2.
Zhong Ji, Yimu Su, Yan Zhang 0135, Yanwei Pang, Jungong Han
IEEE Trans. Image Process.5
2025 Frequency-Spatial Complementation: Unified Channel-Specific Style Attack for Cross-Domain Few-Shot Learning
abstract
Cross-Domain Few-Shot Learning (CD-FSL) addresses the challenges of recognizing targets with out-of-domain data when only a few instances are available. Many current CD-FSL approaches primarily focus on enhancing the generalization capabilities of models in spatial domain, which neglects the role of the frequency domain in domain generalization. To take advantage of frequency domain in processing global information, we propose a Frequency-Spatial Complementation (FSC) model, which combines frequency domain information with spatial domain information to learn domain-invariant information from attacked data style. Specifically, we design a Frequency and Spatial Fusion (FusionFS) module to enhance the ability of the model to capture style-related information. Besides, we propose two attack strategies, i.e., the Gradient-guided Unified Style Attack (GUSA) strategy and the Channel-specific Attack Intensity Calculation (CAIC) strategy, which conduct targeted attacks on different channels to provide more diversified style data during the training phase, especially in single-source domain scenarios where the source domain data style is homogeneous. Extensive experiments across eight target domains demonstrate that our method significantly improves the model's performance under various styles.
Zhong Ji, Zhilong Wang 0001, Xiyao Liu 0002, Yunlong Yu 0001, Yanwei Pang, Jungong Han
IEEE Trans. Image Process.5
2025 Visual Semantic Contextualization Network for Multi-Query Image Retrieval
abstract
Multi-Query Image Retrieval (MQIR) aims to establish connections between vision and language by exploring fine-grained region-query alignments. It is still a challenging task owing to its intrinsical ambiguity, where a query matches with multiple semantically similar regions and introduces misleading noises. Although researchers have made great efforts to alleviate the ambiguity in many retrieval-related tasks, there are few attempts considering this bottleneck in MQIR, which greatly limits present performance. To this end, we propose a novel Visual Semantic Contextualization Network (VSCN) to mitigate ambiguity by capturing the contextual knowledge within each image-text pair. Specifically, we first develop a Context Semantic Perception (CSP) module to capture the dual-level context, where a visual context transformer explores the intra-context within regions, and a cross-modal context transformer mines the inter-context among concatenated visual-linguistic embeddings. Then, to yield superior contextual understanding, we strengthen the connotations in context via a Context Semantic Interaction (CSI) module. Particularly, knowledge distillation is first employed to transfer the CLIP-guided semantic into the regional intra-context to complement the potential background information. Then, the intra-context & inter-context interaction is conducted via the self-attention mechanism to link the dual-level context and obtain the interacted contextual knowledge. Our method is evaluated on the Visual Genome dataset and substantially outperforms the state-of-the-art methods (30.3% improvements on Recall@1 in the first round). Our source codes will be released athttps://github.com/zhli-cs/VSCN.
Zhong Ji, Zhihao Li 0006, Yan Zhang 0135, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Multim.4
2025 Video Instance Segmentation Without Using Mask and Identity Supervision
abstract
Video instance segmentation (VIS) is a challenging vision problem in which the task is to simultaneously detect, segment, and track all the object instances in a video. Most existing VIS approaches rely on pixel-level mask supervision within a frame as well as instance-level identity annotation across frames. However, obtaining these ‘mask and identity’ annotations is time-consuming and expensive. We propose the first mask-identity-free VIS framework that neither utilizes mask annotations nor requires identity supervision. Accordingly, we introduce a query contrast and exchange network (QCEN) comprising instance query contrast and query-exchanged mask learning. The instance query contrast first performs cross-frame instance matching and then conducts query feature contrastive learning. The query-exchanged mask learning exploits both intra-video and inter-video query exchange properties: exchanging queries of an identical instance from different frames within a video results in consistent instance masks, whereas exchanging queries across videos results in all-zero background masks. Extensive experiments on three benchmarks (YouTube-VIS 2019, YouTube-VIS 2021, and OVIS) reveal the merits of the proposed approach, which significantly reduces the performance gap between the identify-free baseline and our mask-identify-free VIS method. On the YouTube-VIS 2019 validation set, our mask-identity-free approach achieves 91.4% of the stronger-supervision-based baseline performance when utilizing the same ImageNet pre-trained model.
Jiale Cao, Hanqing Sun 0001, Rao Muhammad Anwer, Jin Xie 0005, Fahad Shahbaz Khan, Yanwei Pang
IEEE Trans. Multim.7
2025 Implicit and Explicit Language Guidance for Diffusion-Based Visual Perception
abstract
Text-to-image diffusion models have shown powerful ability on conditional image synthesis. With large-scale vision-language pre-training, diffusion models are able to generate high-quality images with rich textures and reasonable structures under different text prompts. However, adapting pre-trained diffusion models for visual perception is an open problem. In this paper, we propose an implicit and explicit language guidance framework for diffusion-based visual perception, named IEDP. Our IEDP comprises an implicit language guidance branch and an explicit language guidance branch. The implicit branch employs a frozen CLIP image encoder to directly generate implicit text embeddings that are fed to the diffusion model without explicit text prompts. The explicit branch uses the ground-truth labels of corresponding images as text prompts to condition feature extraction in diffusion model. During training, we jointly train the diffusion model by sharing the model weights of these two branches. As a result, the implicit and explicit branches can jointly guide feature learning. During inference, we employ only implicit branch for final prediction, which does not require any ground-truth labels. Experiments are performed on two typical perception tasks, including semantic segmentation and depth estimation. Our IEDP achieves promising performance on both tasks. For semantic segmentation, our IEDP has the mIoU$^\text{ss}$score of 55.9% on ADE20K validation set, which outperforms the baseline method VPD by 2.2%. For depth estimation, our IEDP outperforms the baseline method VPD with a relative gain of 11.0%.
Hefeng Wang, Jiale Cao, Jin Xie 0005, Ai-Ping Yang, Yanwei Pang
IEEE Trans. Multim.5
2025 Imbalance Mitigation for Continual Learning via Knowledge Decoupling and Dual Enhanced Contrastive Learning
abstract
Continual learning (CL) aims at studying how to learn new knowledge continuously from data streams without catastrophically forgetting the previous knowledge. One of the key problems is catastrophic forgetting, that is, the performance of the model on previous tasks declines significantly after learning the subsequent task. Several studies addressed it by replaying samples stored in the buffer when training new tasks. However, the data imbalance between old and new task samples results in two serious problems: information suppression and weak feature discriminability. The former refers to the information in the sufficient new task samples suppressing that in the old task samples, which is harmful to maintaining the knowledge since the biased output worsens the consistency of the same sample's output at different moments. The latter refers to the feature representation being biased to the new task, which lacks discrimination to distinguish both old and new tasks. To this end, we build an imbalance mitigation for CL (IMCL) framework that incorporates a decoupled knowledge distillation (DKD) approach and a dual enhanced contrastive learning (DECL) approach to tackle both problems. Specifically, the DKD approach alleviates the suppression of the new task on the old tasks by decoupling the model output probability during the replay stage, which better maintains the knowledge of old tasks. The DECL approach enhances both low- and high-level features and fuses the enhanced features to construct contrastive loss to effectively distinguish different tasks. Extensive experiments on three popular datasets show that our method achieves promising performance under task incremental learning (Task-IL), class incremental learning (Class-IL), and domain incremental learning (Domain-IL) settings.
Zhong Ji, Zhanyu Jiao, Qiang Wang 0056, Yanwei Pang, Jungong Han
IEEE Trans. Neural Networks Learn. Syst.4
2025 Graph Neural Networks With Adaptive Confidence Discrimination
abstract
Graph neural networks (GNNs) have demonstrated remarkable success for semisupervised node classification. However, these GNNs are still limited to the conventionally semisupervised framework and cannot fully leverage the potential value of large numbers of unlabeled samples. The pseudolabeling method in semisupervised learning (SSL) is widely recognized because it can clearly leverage unlabeled samples. Nevertheless, the existing pseudolabeling methods usually utilize a fixed threshold for all classes and only use a portion of unlabeled samples (ones with high prediction confidence), which leads to class imbalance and low data utilization. To solve these problems, we propose GNNs with adaptive confidence discrimination (ACDGNN) to fully utilize unlabeled samples for facilitating semisupervised node classification. Specifically, an adaptive confidence discrimination module is designed to divide all unlabeled nodes into two subsets by comparing their confidence scores with the adaptive confidence threshold at each training epoch. Then, different constraint strategies for two subset nodes are employed. Unlabeled nodes with high confidence are used to iteratively expand the label set, while ones with low confidence learn discriminative features by applying contrastive learning. Validated by extensive experiments, the proposed ACDGNN delivers significant accuracy gains over the previous SOTAs: an average improvement of 2.0% on all datasets and 5.7% on the Flickr dataset in particular.
Yanbei Liu, Shichuan Zhao, Xiao Wang 0017, Lei Geng, Zhitao Xiao, Shuai Ma 0001, Yanwei Pang
IEEE Trans. Neural Networks Learn. Syst.8
2025 Multiscale Subgraph Adversarial Contrastive Learning
abstract
Graph contrastive learning (GCL), as a typical self-supervised learning paradigm, has been able to achieve promising performance without labels and gradually attracts much attention. Graph-level method aims to learn representations of each graph by contrasting two augmented graphs. Previous studies usually simply apply contrastive learning to keep the embeddings of augmented views from the same anchor graph (positive pairs) close to each other, as well as separate the embeddings of augmented views from different anchor graphs (negative pairs). However, it is well-known that the structure of graph is always complex and multiscale, which gives rise to a fundamental question: after graph augmentation, will the previous assumption still hold in reality? Through experimental analytics, we find that the semantic information of two augmented graphs from the same anchor graph may be not consistent, and whether two augmented graphs are positive or negative sample pairs is highly correlated with the multiscale structure of the graph. Based on this observation, we then propose a multiscale subgraph contrastive learning method, named MSSGCL, which can characterize the fine-grained semantic information. Specifically, we generate global and local views at different scales based on subgraph sampling and construct multiple contrastive relationships according to their semantic associations to provide richer self-supervised information. Furthermore, to further improve the generalization performance of the model, we propose an extended model called MSSGCL++. It adopts an asymmetric structure to avoid pushing semantically similar negative samples far away. We further introduce adversarial training to perturb the augmented view and thus construct a more difficult self-supervised training task. Finally, a min-max saddle point problem is optimized and the "free" strategy is used to speed up the training process. Extensive experiments and parametric analysis on 16 real-world graph classification datasets confirm the effectiveness of our proposed approach. Compared with state of the art (SOTA) method, our method achieves improvements of 2% and 1.6% in unsupervised and transfer learning settings, respectively.
Yanbei Liu, Zhitao Xiao, Lei Geng, Xiao Wang 0017, Yanwei Pang, Jerry Chun-Wei Lin
IEEE Trans. Neural Networks Learn. Syst.6
2025 Guest Editorial Special Issue on Effective Feature Fusion in Deep Neural Networks
Yanwei Pang, Fahad Shahbaz Khan, Fabio Cuzzolin
IEEE Trans. Neural Networks Learn. Syst.1
2025 SSNet: a joint learning network for semantic segmentation and disparity estimation
Dayu Jia, Yanwei Pang, Jiale Cao
Vis. Comput.2
2024 SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic Segmentation
abstract
Open-vocabulary semantic segmentation strives to distinguish pixels into different semantic groups from an open set of categories. Most existing methods explore utilizing pre-trained vision-language models, in which the key is to adapt the image-level model for pixel-level segmentation task. In this paper, we propose a simple encoder-decoder, named SED, for open-vocabulary semantic segmentation, which comprises a hierarchical encoder-based cost map generation and a gradual fusion decoder with category early rejection. The hierarchical encoder-based cost map generation employs hierarchical backbone, instead of plain transformer, to predict pixel-level image-text cost map. Compared to plain transformer, hierarchical backbone better captures local spatial information and has linear computational complexity with respect to input size. Our gradual fusion decoder employs a top-down structure to combine cost map and the feature maps of different backbone levels for segmentation. To accelerate inference speed, we introduce a category early rejection scheme in the decoder that rejects many no-existing categories at the early layer of decoder, resulting in at most 4.7 times acceleration without accuracy degradation. Experiments are performed on multiple open-vocabulary semantic segmentation datasets, which demonstrates the efficacy of our SED method. When using ConvNeXt-B, our SED method achieves mIoU score of 31.6% on ADE20K with 150 categories at 82 millisecond (ms) per image on a single A6000. Our source code is available at https://github.com/xb534/SED.
Jiale Cao, Jin Xie 0005, Fahad Shahbaz Khan, Yanwei Pang
CVPR5
2024 On the Approximation Risk of Few-Shot Class-Incremental Learning
Xuan Wang 0016, Zhong Ji, Xiyao Liu 0002, Yanwei Pang, Jungong Han
ECCV (51)4
2024 Eliminate Before Align: A Remote Sensing Image-Text Retrieval Framework with Keyword Explicit Reasoning
abstract
Mountains of researches center around the Remote Sensing Image-Text Retrieval (RSITR), aiming at retrieving the corresponding targets based on the given query. Among them, the transfer of Foundation Models (FMs), such as CLIP, to remote sensing domain shows promising results. However, existing FM-based approaches neglect the negative impact of weakly correlated sample pairs and the key distinctions among remote sensing texts, leading to biased and superficial exploration of sample pairs. To address these challenges, we propose a novel Eliminate Before Align strategy with Keyword Explicit Reasoning framework (EBAKER) for RSITR. Specifically, we devise an innovative Eliminate Before Align (EBA) strategy to filter out the weakly correlated sample pairs to mitigate their deviations from optimal embedding space during alignment. Moreover, we introduce a Keyword Explicit Reasoning (KER) module to facilitate the positive role of subtle key concept differences. Without bells and whistles, our method achieves a one-step transformation from FM to RSITR task, obviating the necessity for extra pretraining on remote sensing data. Extensive experiments on three popular benchmark datasets validate that our proposed EBAKER method outperform the state-of-the-art methods with fewer training data. Our source code will be released soon.
Zhong Ji, Changxu Meng, Yan Zhang 0135, Haoran Wang 0004, Yanwei Pang, Jungong Han
ACM Multimedia5
2024 Uncertainty-aware enhanced dark experience replay for continual learning
Qiang Wang 0056, Zhong Ji, Yanwei Pang, Zhongfei Zhang
Appl. Intell.3
2024 Modality-experts coordinated adaptation for large multimodal models
Yan Zhang 0135, Zhong Ji, Yanwei Pang, Jungong Han, Xuelong Li 0001
Sci. China Inf. Sci.3
2024 Dual states based reinforcement learning for fast MR scan and image reconstruction
Yanwei Pang, Xuebin Sun, Yonghong Hou, Zhenghan Yang, Zhenchang Wang
Neurocomputing2
2024 A cognition-driven framework for few-shot class-incremental learning
Xuan Wang 0016, Zhong Ji, Yanwei Pang, Yunlong Yu 0001
Neurocomputing3
2024 Towards Unsupervised Referring Expression Comprehension with Visual Semantic Parsing
Zhong Ji, Di Wang 0026, Yanwei Pang, Xuelong Li 0001
Knowl. Based Syst.4
2024 Diversity-Infused Network for Unsupervised Few-Shot Remote Sensing Scene Classification
abstract
Few-shot Remote Sensing Scene Classification (RSSC) confronts challenges due to its dependence on extensive labeled datasets. Addressing this, we propose the Diversity-Infused Network (DIN), an unsupervised paradigm for few-shot RSSC, utilizing unlabeled data in training and adapting to novel classes with limited labeled samples. Within an augmentation-based framework, DIN includes a Random Augmentation Sampling (RAS) strategy for task diversity in the meta-training stage, and a Channel-Driven Metric Learning (CDML) module to decode complex channel interactions, enhancing information diversity. Additionally, DIN presents a Multi-Mutual Information (Multi-MI) objective function to balance the architecture and reduce the unreliability and potential biases from pseudo-training. Experimental results demonstrate that DIN surpasses other unsupervised approaches by over 11% in 1-shot and nearly 9% in 5-shot settings on WHU-RS19, and closely approaches supervised methods, with less than 1% difference in both settings.
Liyuan Hou 0002, Zhong Ji, Xuan Wang 0016, Yunlong Yu 0001, Yanwei Pang
IEEE Geosci. Remote. Sens. Lett.5
2024 Supervised biadjacency networks for stereo matching
Hanqing Sun 0001, Jungong Han, Yanwei Pang, Xuelong Li 0001
Multim. Tools Appl.3
2024 Hierarchical matching and reasoning for multi-query image retrieval
Zhong Ji, Zhihao Li 0006, Yan Zhang 0135, Haoran Wang 0004, Yanwei Pang, Xuelong Li 0001
Neural Networks5
2024 Multi-task hierarchical convolutional network for visual-semantic cross-modal retrieval
Zhong Ji, Zhigang Lin, Haoran Wang 0004, Yanwei Pang, Xuelong Li 0001
Pattern Recognit.4
2024 Cross-scale contrastive triplet networks for graph representation learning
Yanbei Liu, Wanjin Shan, Xiao Wang 0017, Zhitao Xiao, Lei Geng, Fang Zhang 0001, Dongdong Du, Yanwei Pang
Pattern Recognit.8
2024 Deep intra-image contrastive learning for weakly supervised one-step person search
Jiabei Wang, Yanwei Pang, Jiale Cao, Hanqing Sun 0001, Xuelong Li 0001
Pattern Recognit.2
2024 Multi-query and multi-level enhanced network for semantic segmentation
Jiale Cao, Rao Muhammad Anwer, Jin Xie 0005, Jing Nie 0001, Ai-Ping Yang, Yanwei Pang
Pattern Recognit.7
2024 ESGN: Efficient Stereo Geometry Network for Fast 3D Object Detection
abstract
Fast stereo based 3D object detectors have made great progress recently. However, they suffer from the inferior accuracy. We argue that the main reason is due to the poor geometry-aware feature representation in 3D space. To solve this problem, we propose an efficient stereo geometry network (ESGN). The key in our ESGN is an efficient geometry-aware feature generation (EGFG) module. Our EGFG module first uses a stereo correlation and reprojection module to construct multi-scale stereo volumes in camera frustum space, second employs a multi-scale bird’s eye view (BEV) projection and fusion module to generate multiple geometry-aware features. In these two steps, we adopt deep multi-scale information fusion for discriminative geometry-aware feature generation, without any complex aggregation networks. In addition, we introduce a deep geometry-aware feature distillation scheme to guide stereo feature learning with a LiDAR-based detector. The experiments are performed on the classical KITTI dataset. On KITTI test set, our ESGN outperforms the fast state-of-art-art detector YOLOStereo3D by 5.14% on mAP3d at$62ms$. To the best of our knowledge, our ESGN achieves a best trade-off between accuracy and speed. We hope that our efficient stereo geometry network can provide more possible directions for fast 3D object detection.
Aqi Gao, Yanwei Pang, Jing Nie 0001, Jiale Cao, Yishun Guo, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Few-Shot Action Recognition via Multi-View Representation Learning
abstract
Few-shot action recognition aims to recognize novel action classes with limited labeled samples and has recently received increasing attention. The core objective of few-shot action recognition is to enhance the discriminability of feature representations. In this paper, we propose a novel multi-view representation learning network (MRLN) to model intra-video and inter-video relations for few-shot action recognition. Specifically, we first propose a spatial-aware aggregation refinement module (SARM), which mainly consists of a spatial-aware aggregation sub-module and a spatial-aware refinement sub-module to explore the spatial context of samples at the frame level. Then, we design a temporal-channel enhancement module (TCEM), which can capture the temporal-aware and channel-aware features of samples with the elaborately designed temporal-aware enhancement sub-module and channel-aware enhancement sub-module. Third, we introduce a cross-video relation module (CVRM), which can explore the relations across videos by utilizing the self-attention mechanism. Moreover, we design a prototype-centered mean absolute error loss to improve the feature learning capability of the proposed MRLN. Extensive experiments on four prevalent few-shot action recognition benchmarks show that the proposed MRLN can significantly outperform a variety of state-of-the-art few-shot action recognition methods. Especially, on the 5-way 1-shot setting, our MRLN respectively achieves 75.7%, 86.9%, 65.5% and 45.9% on the Kinetics, UCF101, HMDB51 and SSv2 datasets.
Xiao Wang 0072, Yang Lu 0009, Wanchuan Yu, Yanwei Pang, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.4
2024 NTK-Guided Few-Shot Class Incremental Learning
abstract
The proliferation of Few-Shot Class Incremental Learning (FSCIL) methodologies has highlighted the critical challenge of maintaining robust anti-amnesia capabilities in FSCIL learners. In this paper, we present a novel conceptualization of anti-amnesia in terms of mathematical generalization, leveraging the Neural Tangent Kernel (NTK) perspective. Our method focuses on two key aspects: ensuring optimal NTK convergence and minimizing NTK-related generalization loss, which serve as the theoretical foundation for cross-task generalization. To achieve global NTK convergence, we introduce a principled meta-learning mechanism that guides optimization within an expanded network architecture. Concurrently, to reduce the NTK-related generalization loss, we systematically optimize its constituent factors. Specifically, we initiate self-supervised pre-training on the base session to enhance NTK-related generalization potential. These self-supervised weights are then carefully refined through curricular alignment, followed by the application of dual NTK regularization tailored specifically for both convolutional and linear layers. Through the combined effects of these measures, our network acquires robust NTK properties, ensuring optimal convergence and stability of the NTK matrix and minimizing the NTK-related generalization loss, significantly enhancing its theoretical generalization. On popular FSCIL benchmark datasets, our NTK-FSCIL surpasses contemporary state-of-the-art approaches, elevating end-session accuracy by 2.9% to 9.3%.
Jingren Liu, Zhong Ji, Yanwei Pang, Yunlong Yu 0001
IEEE Trans. Image Process.3
2024 Image Reconstruction for Accelerated MR Scan With Faster Fourier Convolutional Neural Networks
abstract
High quality image reconstruction from undersampled k -space data is key to accelerating MR scanning. Current deep learning methods are limited by the small receptive fields in reconstruction networks, which restrict the exploitation of long-range information, and impede the mitigation of full-image artifacts, particularly in 3D reconstruction tasks. Additionally, the substantial computational demands of 3D reconstruction considerably hinder advancements in related fields. To tackle these challenges, we propose the following: 1) A novel convolution operator named Faster Fourier Convolution (FasterFC), aims at providing an adaptable broad receptive field for spatial domain reconstruction networks with fast computational speed. 2) A split-slice strategy that substantially reduces the computational load of 3D reconstruction, enabling high-resolution, multi-coil, 3D MR image reconstruction while fully utilizing inter-layer and intra-layer information. 3) A single-to-group algorithm that efficiently utilizes scan-specific and data-driven priors to enhance k -space interpolation effects. 4) A multi-stage, multi-coil, 3D fast MRI method, called the faster Fourier convolution based single-to-group network (FAS-Net), comprising a single-to-group k -space interpolation algorithm and a FasterFC-based image domain reconstruction module, significantly minimizes the computational demands of 3D reconstruction through split-slice strategy. Experimental evaluations conducted on the NYU fastMRI and Stanford MRI Data datasets reveal that the FasterFC significantly enhances the quality of both 2D and 3D reconstruction results. Moreover, FAS-Net, characterized as a method that can achieve high-resolution (320, 320, 256), multi-coil, (8 coils), 3D fast MRI, exhibits superior reconstruction performance compared to other state-of-the-art 2D and 3D methods.
Yanwei Pang, Xuebin Sun, Yonghong Hou, Zhenchang Wang, Xuelong Li 0001
IEEE Trans. Image Process.2
2024 Model Attention Expansion for Few-Shot Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) aims at incrementally learning new knowledge from limited training examples without forgetting previous knowledge. However, we observe that existing methods face a challenge known as supervision collapse, where the model disproportionately emphasizes class-specific features of base classes at the detriment of novel class representations, leading to restricted cognitive capabilities. To alleviate this issue, we propose a new framework, Model aTtention Expansion for Few-Shot Class-Incremental Learning (MTE-FSCIL), aimed at expanding the model attention fields to improve transferability without compromising the discriminative capability for base classes. Specifically, the framework adopts a dual-stage training strategy, comprising pre-training and meta-training stages. In the pre-training stage, we present a new regularization technique, named the Reserver (RS) loss, to expand the global perception and reduce over-reliance on class-specific features by amplifying feature map activations. During the meta-training stage, we introduce the Repeller (RP) loss, a novel pair-based loss that promotes variation in representations and improves the model's recognition of sample uniqueness by scattering intra-class samples within the embedding space. Furthermore, we propose a Transformational Adaptation (TA) strategy to enable continuous incorporation of new knowledge from downstream tasks, thus facilitating cross-task knowledge transfer. Extensive experimental results on mini-ImageNet, CIFAR100, and CUB200 datasets demonstrate that our proposed framework consistently outperforms the state-of-the-art methods.
Xuan Wang 0016, Zhong Ji, Yunlong Yu 0001, Yanwei Pang, Jungong Han
IEEE Trans. Image Process.4
2024 USER: Unified Semantic Enhancement With Momentum Contrast for Image-Text Retrieval
abstract
As a fundamental and challenging task in bridging language and vision domains, Image-Text Retrieval (ITR) aims at searching for the target instances that are semantically relevant to the given query from the other modality, and its key challenge is to measure the semantic similarity across different modalities. Although significant progress has been achieved, existing approaches typically suffer from two major limitations: (1) It hurts the accuracy of the representation by directly exploiting the bottom-up attention based region-level features where each region is equally treated. (2) It limits the scale of negative sample pairs by employing the mini-batch based end-to-end training mechanism. To address these limitations, we propose a Unified Semantic Enhancement Momentum Contrastive Learning (USER) method for ITR. Specifically, we delicately design two simple but effective Global representation based Semantic Enhancement (GSE) modules. One learns the global representation via the self-attention algorithm, noted as Self-Guided Enhancement (SGE) module. The other module benefits from the pre-trained CLIP module, which provides a novel scheme to exploit and transfer the knowledge from an off-the-shelf model, noted as CLIP-Guided Enhancement (CGE) module. Moreover, we incorporate the training mechanism of MoCo into ITR, in which two dynamic queues are employed to enrich and enlarge the scale of negative sample pairs. Meanwhile, a Unified Training Objective (UTO) is developed to learn from mini-batch based and dynamic queue based samples. Extensive experiments on the benchmark MSCOCO and Flickr30K datasets demonstrate the superiority of both retrieval accuracy and inference efficiency. For instance, compared with the existing best method NAAF, the metric R@1 of our USER on the MSCOCO 5K Testing set is improved by 5% and 2.4% on caption retrieval and image retrieval without any external knowledge or pre-trained model while enjoying over 60 times faster inference speed. Our source code will be released at https://github.com/zhangy0822/USER.
Yan Zhang 0135, Zhong Ji, Di Wang 0026, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Image Process.4
2024 Multi-Modal Multi-Slice Cooperative Dual-Domain Cascaded De-Aliasing Network for MR Imaging Reconstruction
abstract
Recent advancements in Magnetic Resonance Imaging (MRI) reconstruction techniques aim to accelerate the imaging process. However, these methods still face two key limitations. Firstly, although the same location consistently provides anatomical information across different modalities, such as organs, tissues, or lesions, previous studies have predominantly relied on single-modality information, overlooking the potential advantages of incorporating complementary data from other modalities. Secondly, while adjacent MRI slices often capture the same location or organ with similar anatomical structures, only a few methods consider the information from neighboring slices during the reconstruction process. To address these challenges, we propose aMulti-modalMulti-slice cooperativeDual-domain cascaded de-alising network for MR imagingReconstruction (MMDR). Specifically, we design a multi-slice and multi-modal feature fusion network based on 3D convolution and swin transformer that efficiently extracts multi-modal features from MRI. Then, a dual domain cascaded recurrent network through dense-blocks with large receptive fields for fast MRI reconstruction is explored. Extensive experiments on the IXI datasets were carried out to evaluate the proposed method's robustness across varying network structures, under-sampling rates, and sampling patterns. MMDR demonstrates promising performance across both qualitative and quantitative metrics, particularly with a competitive PSNE of 42.15 and an SSIM of 0.984 for T2 reconstruction using 30% T2WI and PDWI, as well as achieving a PSNE of 39.92 and an SSIM of 0.966 for PDWI reconstruction with 30% PDWI and T2WI.
Xuebin Sun, Yanwei Pang, Caifeng Shan, Shing Shin Cheng
IEEE J. Biomed. Health Informatics2
2024 Toward Generalizable Multispectral Pedestrian Detection
abstract
Multispectral pedestrian detection has achieved great success in past years, which can be used in autonomous driving for intelligent transportation system. Most existing multispectral pedestrian detection approaches are developed on the assumption that training and test data belong to an identical distribution, which does not guarantee a good generalization to cross-domain (unseen) data. In this paper, we aim to develop a generalizable multispectral pedestrian detector, which achieves a favorable performance on both intra-dataset evaluation and cross-dataset evaluation. To achieve this goal, we conduct intra-dataset and cross-dataset experiments using single-modal and multi-modal data. By deep analysis, we find that, compared to visible or multi-modal data, thermal data not only has a best cross-dataset generalization, but also generates high-quality proposals on intra-dataset and cross-dataset evaluations. Inspired by this, we propose a novel thermal-first and fusion-second network (called TFNet) for multispectral pedestrian detection. In our TFNet, we first employ a thermal-based proposal network to extract candidate pedestrian proposals. After that, we design a transformer fusion based head network to further classify/regress these proposals. Experiments are performed on three public datasets. The comprehensive results demonstrate the effectiveness of our proposed TFNet on both intra-dataset and cross-dataset evaluations. We hope that our simple design can promote the future study on generalizable multispectral pedestrian detection.
Fuchen Chu, Jiale Cao, Zhanjie Song, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Intell. Transp. Syst.5
2024 Transformer-Based Stereo-Aware 3D Object Detection From Binocular Images
abstract
Transformers have shown promising progress in various visual object detection tasks, including monocular 2D/3D detection and surround-view 3D detection. More importantly, the attention mechanism in the Transformer model and the 3D information extraction in binocular stereo are both similarity-based. However, directly applying existing Transformer-based detectors to binocular stereo 3D object detection leads to slow convergence and significant precision drops. We argue that a key cause of that defect is that existing Transformers ignore the binocular-stereo-specific image correspondence information. In this paper, we explore the model design of Transformers in binocular 3D object detection, focusing particularly on extracting and encoding task-specific image correspondence information. To achieve this goal, we present TS3D, a Transformer-based Stereo-aware 3D object detector. In the TS3D, a Disparity-Aware Positional Encoding (DAPE) module is proposed to embed the image correspondence information into stereo features. The correspondence is encoded as normalized sub-pixel-level disparity and is used in conjunction with sinusoidal 2D positional encoding to provide the 3D location information of the scene. To enrich multi-scale stereo features, we propose a Stereo Preserving Feature Pyramid Network (SPFPN). The SPFPN is designed to preserve the correspondence information while fusing intra-scale and aggregating cross-scale stereo features. Our proposed TS3D achieves a 41.29% Moderate Car detection average precision on the KITTI test set and takes 88 ms to detect objects from each binocular image pair. It is competitive with advanced counterparts in terms of both precision and inference speed.
Hanqing Sun 0001, Yanwei Pang, Jiale Cao, Jin Xie 0005, Xuelong Li 0001
IEEE Trans. Intell. Transp. Syst.2
2024 Binocular Image Dehazing via a Plain Network Without Disparity Estimation
abstract
Heavy haze leads to severely degraded visual quality for images, and thus the performance of high level image-based tasks such as object detection and semantic segmentation is deteriorated. It is necessary and important to design an effective dehazing method for the computer vision system. It is well known that image haze is a function of depth and binocular images can predict the depth. Existing binocular dehazing methods conduct disparity estimation and dehazing jointly to enhance each other. However, a small error in disparity gives rise to a large variation in depth and in the estimation of haze-free images. To alleviate the problem, we propose a plain binocular image dehazing network in this paper, called BidNet, to dehaze both the left and right images simultaneously. BidNet does not explicitly perform disparity estimation that is time-consuming and well-known to be challenging. Instead, we design a stereo transformation module to mine the relationship and correlation between binocular images, making the best of varying information of cross views. Additionally, we design a Stereo Foggy Cityscapes dataset extended from the Foggy Cityscapes dataset for training the proposed BidNet. Extensive experimental results demonstrate that BidNet significantly outperforms the SOTA dehazing methods on the synthetic stereo foggy datasets as well as in real stereo foggy scenes. Experimental results show that jointly dehazing binocular image pairs is mutually beneficial, which is better than only dehazing left images. Furthermore, when applying BidNet to preprocess foggy inputs, large improvements are obtained in the performance of object detection, instance segmentation, semantic segmentation, and stereo-based 3D object detection.
Jing Nie 0001, Yanwei Pang, Jin Xie 0005, Jungong Han, Xuelong Li 0001
IEEE Trans. Multim.2
2024 DCMSTRD: End-to-end Dense Captioning via Multi-Scale Transformer Decoding
abstract
Dense captioning creates diverse Region of Interests (RoIs) descriptions for complex visual scenes. While promising results have been obtained, several issues persist. In particular: 1) it is hard to find the optimal parameters for artificially designed modules (e.g., non-maximum suppression (NMS)) causing redundancies and fewer interactions to benefit the two sub-tasks of RoI detection and RoI captioning; 2) the absence of a multi-scale decoder in current methods hinders the acquisition of scale-invariant features, thus leading to poor performance. To tackle these limitations, we bypass the artificially designed modules and present an end-to-end dense captioning framework via multi-scale transformer decoding (DCMSTRD). DCMSTRD solves dense captioning by set matching and prediction instead. To further enhance the discriminative quality of the multi-scale representations during caption generation, we introduce a multi-scale module, termed multi-scale language decoder (MSLD). Our proposed method tested on standard datasets achieves a mean Average Precision (mAP) of 16.7% on the challenging VG-COCO dataset, demonstrating its effectiveness against the current methods.
Jungong Han, Kurt Debattista, Yanwei Pang
IEEE Trans. Multim.4
2024 Semantic-Aware Dynamic Generation Networks for Few-Shot Human-Object Interaction Recognition
abstract
Recognizing human-object interaction (HOI) aims at inferring various relationships between actions and objects. Although great progress in HOI has been made, the long-tail problem and combinatorial explosion problem are still practical challenges. To this end, we formulate HOI as a few-shot task to tackle both challenges and design a novel dynamic generation method to address this task. The proposed approach is called semantic-aware dynamic generation networks (SADG-Nets). Specifically, SADG-Net first assigns semantic-aware task representations for different batches of data, which further generates dynamic parameters. It obtains the features that highlight intercategory discriminability and intracategory commonality adaptively. In addition, we also design a dual semantic-aware encoder module (DSAE-Module), that is, verb-aware and noun-aware branches, to yield both action and object prototypes of HOI for each task space, which generalizes to novel combinations by transferring similarities among interactions. Extensive experimental results on two benchmark datasets, that is, humans interacting with common objects (HICO)-FS and trento universal HOI (TUHOI)-FS, illustrate that our SADG-Net achieves superior performance over state-of-the-art approaches, which proves its impressive effectiveness on few-shot HOI recognition.
Zhong Ji, Xiyao Liu 0002, Changxin Gao, Yanwei Pang, Ling Shao 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 HGBER: Heterogeneous Graph Neural Network With Bidirectional Encoding Representation
abstract
Heterogeneous graphs with multiple types of nodes and link relationships are ubiquitous in many real-world applications. Heterogeneous graph neural networks (HGNNs) as an efficient technique have shown superior capacity of dealing with heterogeneous graphs. Existing HGNNs usually define multiple meta-paths in a heterogeneous graph to capture the composite relations and guide neighbor selection. However, these models only consider the simple relationships (i.e., concatenation or linear superposition) between different meta-paths, ignoring more general or complex relationships. In this article, we propose a novel unsupervised framework termed Heterogeneous Graph neural network with bidirectional encoding representation (HGBER) to learn comprehensive node representations. Specifically, the contrastive forward encoding is firstly performed to extract node representations on a set of meta-specific graphs corresponding to meta-paths. We then introduce the reversed encoding for the degradation process from the final node representations to each single meta-specific node representations. Moreover, to learn structure-preserving node representations, we further utilize a self-training module to discover the optimal node distribution through iterative optimization. Extensive experiments on five open public datasets show that the proposed HGBER model outperforms the state-of-the-art HGNNs baselines by 0.8%-8.4% in terms of accuracy on most datasets in various downstream tasks.
Yanbei Liu, Lianxi Fan, Xiao Wang 0017, Zhitao Xiao, Shuai Ma 0001, Yanwei Pang, Jerry Chun-Wei Lin
IEEE Trans. Neural Networks Learn. Syst.6
2024 Integrating Visual Perception With Decision Making in Neuromorphic Fault-Tolerant Quadruplet-Spike Learning Framework
abstract
The brain possesses the remarkable ability to seamlessly integrate perception with decision making within a dynamically changing environment in a fault-tolerant, end-to-end manner. This extraordinary capability offers a compelling solution for brain-inspired intelligence, replete with the advantages of end-to-end decision making: robustness, high accuracy, real-time responsiveness, autonomous intelligence, and a high degree of biological plausibility. Neuromorphic computing stands as a promising avenue for brain-inspired intelligence through the harmonious co-design of algorithms and hardware, aimed at unlocking full potential. This article introduces a comprehensive neuromorphic computing framework for end-to-end intelligence. It introduces the quadruplet spike-timing-dependent plasticity, which serves as a cornerstone for perceptual to decision-making tasks. A fault-tolerant neuromorphic routing strategy is presented to fortify the framework’s robustness. Empirical results underscore its impressive attributes with high accuracy, robustness, fault tolerance, and minimal computational latency when orchestrating end-to-end decision-making alongside visual perception. This study marks a pioneering effort in unified, fault-tolerant neuromorphic framework engineered for brain-inspired end-to-end intelligent tasks, merging visual perception with adaptive decision making. Such an endeavor is profoundly meaningful, as it propels the development of artificial general intelligence, holding vast implications for the field’s advancement.
Shuangming Yang, Yanwei Pang, Yaochu Jin, Bernabé Linares-Barranco
IEEE Trans. Syst. Man Cybern. Syst.3
2024 Multi-feature self-attention super-resolution network
Ai-Ping Yang, Zihao Wei, Jinbin Wang, Jiale Cao, Zhong Ji, Yanwei Pang
Vis. Comput.6
2023 Attentive Alignment Network for Multispectral Pedestrian Detection
abstract
Multispectral pedestrian detection is of great importance in various around-the-clock applications, i.e., self-driving and video surveillance. Fusing the features from RGB images and thermal infrared (TIR) images to explore the complementary information between different modalities is one of the most effective manners to improve multispectral pedestrian detection performance. However, the misalignment between different modalities in spatial dimension and modality reliability would introduce harmful information during feature fusion, limiting the performance of multispectral pedestrian detection. To address the above issues, we propose an attentive alignment network, consisting of an attentive position alignment (APA) module and an attentive modality alignment (AMA) module. Our APA module emphasizes pedestrian regions while aligning the pedestrian regions between different modalities. Our AMA module utilizes a channel-wise attention mechanism with illumination guidance to eliminate the imbalance between different modalities. The experiments are conducted on two widely used multispectral detection datasets, KASIT and CVC-14. Our approach surpasses the current state-of-the-art performance on both datasets.
Nuo Chen 0003, Jin Xie 0005, Jing Nie 0001, Jiale Cao, Yanwei Pang
ACM Multimedia6
2023 DIIK-Net: A full-resolution cross-domain deep interaction convolutional neural network for MR image reconstruction
Yanwei Pang, Jing Nie 0001
Neurocomputing2
2023 EVOLVE: Learning volume-adaptive phases for fast 3D magnetic resonance scan and image reconstruction
Yanwei Pang, Xuebin Sun, Yonghong Hou
Neurocomputing2
2023 Mutual mentor: Online contrastive distillation network for general continual learning
Qiang Wang 0056, Zhong Ji, Jin Li 0054, Yanwei Pang
Neurocomputing4
2023 Spike-driven multi-scale learning with hybrid mechanisms of spiking dendrites
Shuangming Yang, Yanwei Pang, Tao Lei 0003, Yaochu Jin
Neurocomputing2
2023 Self-taught cross-domain few-shot learning with weakly supervised object localization and task-decomposition
Xiyao Liu 0002, Zhong Ji, Yanwei Pang, Zhi Han
Knowl. Based Syst.3
2023 Spatial attention-guided deformable fusion network for salient object detection
Ai-Ping Yang, Simeng Cheng, Jiale Cao, Zhong Ji, Yanwei Pang
Multim. Syst.6
2023 Zero-shot classification with unseen prototype learning
Zhong Ji, Biying Cui, Yunlong Yu 0001, Yanwei Pang, Zhongfei Zhang
Neural Comput. Appl.4
2023 Dual Distillation Discriminator Networks for Domain Adaptive Few-Shot Learning
Xiyao Liu 0002, Zhong Ji, Yanwei Pang, Zhi Han
Neural Networks3
2023 Context and detail interaction network for stereo rain streak and raindrop removal
Jing Nie 0001, Jin Xie 0005, Jiale Cao, Yanwei Pang
Neural Networks4
2023 Visual-quality-driven unsupervised image dehazing
Ai-Ping Yang, Jinbin Wang, Jiale Cao, Zhong Ji, Yanwei Pang
Neural Networks7
2023 COREN: Multi-Modal Co-Occurrence Transformer Reasoning Network for Image-Text Retrieval
Zhong Ji, Yanwei Pang, Zhongfei Zhang
Neural Process. Lett.4
2023 SipMaskv2: Enhanced Fast Image and Video Instance Segmentation
abstract
We propose a fast single-stage method for both image and video instance segmentation, called SipMask, that preserves the instance spatial information by performing multiple sub-region mask predictions. The main module in our method is a light-weight spatial preservation (SP) module that generates a separate set of spatial coefficients for the sub-regions within a bounding-box, enabling a better delineation of spatially adjacent instances. To better correlate mask prediction with object detection, we further propose a mask alignment weighting loss and a feature alignment scheme. In addition, we identify two issues that impede the performance of single-stage instance segmentation and introduce two modules, including a sample selection scheme and an instance refinement module, to address these two issues. Experiments are performed on both image instance segmentation dataset MS COCO and video instance segmentation dataset YouTube-VIS. On MS COCO test-dev set, our method achieves a state-of-the-art performance. In terms of real-time capabilities, it outperforms YOLACT by a gain of 3.0% (mask AP) under the similar settings, while operating at a comparable speed. On YouTube-VIS validation set, our method also achieves promising results. The source code is available at https://github.com/JialeCao001/SipMask.
Jiale Cao, Yanwei Pang, Rao Muhammad Anwer, Hisham Cholakkal, Fahad Shahbaz Khan, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Semantic-aware self-supervised depth estimation for stereo 3D detection
Hanqing Sun 0001, Jiale Cao, Yanwei Pang
Pattern Recognit. Lett.3
2023 G2LP-Net: Global to Local Progressive Video Inpainting Network
abstract
The self-attention based video inpainting methods have achieved promising progress by establishing long-range correlation over the whole video. However, existing methods generally relied on the global self-attention that directly searches missing contents among all reference frames but lacks accurate matching and effective organization on contents, which often blurs the result owing to the loss of local textures. In this paper, we propose a Global-to-Local Progressive Inpainting Network (G2LP-Net) consisting of the following innovative ideas. First, we present a global to local self-attention mechanism by incorporating local self-attention into global self-attention to improve searching efficiency and accuracy, where the self-attention is implemented in multi-scale regions to fully exploit local redundancy for the texture recovery. Second, we propose a progressive video inpainting (PVI) method to organize the generated contents, which completes the target video frames from periphery to core to ensure reliable contents serve first. Last, we develop a window-sliding method for sampling reference frames to obtain rich available information for inpainting. In addition, we release a wire-removal video (WRV) dataset that consists of 150 video clips masked by wires to evaluate the video inpainting on irregularly slender regions. Both quantitative and qualitative experiments on benchmark datasets, DAVIS, YouTube-VOS and our WRV dataset have demonstrated the superiority of our proposed G2LP-Net method.
Zhong Ji, Yimu Su, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 Consensus Knowledge Exploitation for Partial Query Based Image Retrieval
abstract
Partial Query based Image Retrieval (PQIR) enables a search engine to perform an interactive retrieval given by only an initial query and actively provide alternative feedbacks for a user to refine a set of retrieval results. It alleviates the deficiency in practice interactive image retrieval that requires the user to laboriously provide detailed feedbacks, and enables the retrieval on-the-fly with the incomplete initial query. Although significant progress has been made, existing works remain have challenge in actively providing more discriminative feedbacks. To address this challenge, we propose a novel Attributes&Objects-based Consensus Extraction and Representation (AoCer) framework. Specifically, we formulate a simple but effective Attribute&Object Feedback (AOF) paradigm, which employs both attributes and objects as intermediate feedbacks to carry out multiple rounds of interaction. To mine the intrinsic associations among concepts and enhance their feature representations, we further propose an Interventional Consensus Representation Learning (ICRL) module, which mainly constructs an interventional concept graph to yield the Interventional Consensus Representation (ICR). In addition, a Dual-Head Feedback Sampler (DHFS) is developed to sample objects and attributes for conducting the next round retrieval. Extensive experiments demonstrate the superiority of the proposed framework. Our source code will be released athttps://github.com/zhangy0822/AoCer.
Yan Zhang 0135, Zhong Ji, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Dual Contrastive Network for Few-Shot Remote Sensing Image Scene Classification
abstract
Few-shot remote sensing image scene classification (FS-RSISC) aims at classifying remote sensing images with only a few labeled samples. The main challenges lie in small inter-class variances and large intra-class variances, which are the inherent property of remote sensing images. To address these challenges, we propose a transfer-based Dual Contrastive Network (DCN), which incorporates two auxiliary supervised contrastive learning branches during the training process. Specifically, one is a Context-guided Contrastive Learning (CCL) branch and the other is a Detail-guided Contrastive Learning (DCL) branch, which focus on inter-class discriminability and intra-class invariance, respectively. In the CCL branch, we first devise a Condenser Network to capture context features, and then leverage a supervised contrastive learning on top of the obtained context features to facilitate the model to learn more discriminative features. In the DCL branch, a Smelter Network is designed to highlight the significant local detail information. And then we construct a supervised contrastive learning based on the detail feature maps to fully exploit the spatial information in each map, enabling the model to concentrate on invariant detail features. Extensive experiments on four public benchmark remote sensing datasets demonstrate the competitive performance of our proposed DCN.
Zhong Ji, Liyuan Hou 0002, Xuan Wang 0016, Gang Wang 0060, Yanwei Pang
IEEE Trans. Geosci. Remote. Sens.5
2023 Knowledge-Aided Momentum Contrastive Learning for Remote-Sensing Image Text Retrieval
abstract
Remote sensing image-text retrieval (RSITR) has attracted widespread attention due to its great potential for rapid information mining ability on remote sensing images. Although significant progress has been achieved, existing methods typically overlook the challenge posed by the extremely analogous descriptions, where the subtle differences remain largely unexploited or, in some cases, are entirely disregarded. To address the limitation, we propose a Knowledge Aided Momentum Contrastive Learning (KAMCL) method for RSITR. Specifically, we propose a novel Knowledge Aided Learning framework, including knowledge initialization, construction, filtration, and alignment operations, which aims at providing valuable concepts and learning discriminative representations. On this basis, we integrate Momentum Contrastive Learning to promote the capture of key concepts within the representation via expanding the scale of negative sample pairs. Moreover, we design a hierarchical aggregator module to better capture the multi-level information from remote sensing images. Finally, we introduce an innovative two-step training strategy designed to effectively harness the synergy among concepts and leverage their respective functionalities. Extensive experiments conducted on the three public datasets showcase the remarkable performance of our approach in terms of retrieval accuracy and computational efficiency. For instance, compared with the existing state-of-the-art method, our method exhibits notable performance improvements of 2.65% on the RSICD dataset, simultaneously achieving improvements in inference efficiency by 48%. Our source code will be released at https://github.com/mcx-mcx/KAMCL.
Zhong Ji, Changxu Meng, Yan Zhang 0135, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Memorizing Complementation Network for Few-Shot Class-Incremental Learning
abstract
Few-shot Class-Incremental Learning (FSCIL) aims at learning new concepts continually with only a few samples, which is prone to suffer the catastrophic forgetting and overfitting problems. The inaccessibility of old classes and the scarcity of the novel samples make it formidable to realize the trade-off between retaining old knowledge and learning novel concepts. Inspired by that different models memorize different knowledge when learning novel concepts, we propose a Memorizing Complementation Network (MCNet) to ensemble multiple models that complements the different memorized knowledge with each other in novel tasks. Additionally, to update the model with few novel samples, we develop a Prototype Smoothing Hard-mining Triplet (PSHT) loss to push the novel samples away from not only each other in current task but also the old distribution. Extensive experiments on three benchmark datasets, e.g., CIFAR100, miniImageNet and CUB200, have demonstrated the superiority of our proposed method.
Zhong Ji, Zhishen Hou, Xiyao Liu 0002, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Image Process.4
2023 Real-Time Stereo 3D Car Detection With Shape-Aware Non-Uniform Sampling
abstract
Pseudo-LiDAR based stereo 3D detectors have gained popularity due to their high accuracy. However, these methods need dense depth supervision and suffer from inferior speed. To solve these two issues, a recently introduced RTS3D builds an efficient 4D feature-consistency embedding (FCE) space for the object intermediate representation without depth supervision, where FCE space performs uniform sampling to generate feature sampling points, which ignores the importance of different object regions. In this paper, we observe that, compared with the inner region, the outer region of the object plays a more important role for accurate 3D detection. To fully exploit the useful information from the outer region, we propose a novel shape-aware non-uniform sampling strategy. Instead of uniform sampling, our proposed non-uniform sampling strategy performs dense sampling in outer region and sparse sampling in inner region. Therefore, more points are sampled from the outer region and more useful features are extracted for 3D detection. In addition, we design a high-level semantic enhanced FCE module to exploit more contextual information and suppress noise better. As a result, it further improves feature discrimination of each sampling point. Experimental results on the KITTI dataset show the effectiveness of the proposed method. Compared with the baseline RTS3D, our proposed method has 2.57% improvement on car AP$_{3d}$almost without extra network parameters. Moreover, our proposed method outperforms the state-of-the-art methods without extra supervision at a real-time speed.
Aqi Gao, Jiale Cao, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Intell. Transp. Syst.3
2023 Latent Feature Pyramid Network for Object Detection
abstract
Object detection methods based on Convolution Neural Networks (CNN) usually utilize feature pyramid networks to detect objects with various scales. The state-of-the-art feature pyramid networks improve detection accuracy by enhancing multi-level feature representations. Fusing multi-level features is the most effective manner to enhance the feature representations. However, the existing feature pyramid networks usually fuse multi-level features by element-wise operations. It leads to the lack of long-range dependencies in the feature fusion. To address the problem, we propose a simple yet efficient feature pyramid network named latent feature pyramid network (LFPN). LFPN can enhance the feature representations by modeling inner-scale and cross-scale long-range dependencies through conducting inner-scale and cross-scale feature fusion in the latent space. Comprehensive experiments are performed on two challenge object detection datasets: MS COCO and Pascal VOC. The experimental results show consistent improvements on various feature pyramid networks, backbones, and object detectors, which demonstrates the effectiveness and generality of our LFPN.
Jin Xie 0005, Yanwei Pang, Jing Nie 0001, Jiale Cao, Jungong Han
IEEE Trans. Multim.2
2023 Textual Context-Aware Dense Captioning With Diverse Words
abstract
Dense captioning generates more detailed spoken descriptions for complex visual scenes. Despite several promising leads, existing methods still have two broad limitations: 1) The vast majority of prior arts only consider visual contextual clues during captioning but ignore potentially important textual context; 2) current imbalanced learning mechanisms limit the diversity of vocabulary learned from the dictionary, thus giving rise to low language-learning efficiency. To alleviate these gaps, in this paper, we propose an end-to-end enhanced dense captioning architecture, namely Enhanced Transformer Dense Captioner (ETDC), which obtains textual context from surrounding regions and dynamically diversifies the vocabulary bank during captioning. Concretely, we first propose the Textual Context Module (TCM), which is integrated into each self-attention layer of the Transformer decoder, to capture the surrounding textual context. Moreover, we take full advantage of the class information of object context and propose a Dynamic Vocabulary Frequency Histogram (DVFH) re-sampling strategy during training to balance words with different frequencies. The proposed method is tested on the standard dense captioning datasets and surpasses the state-of-the-art methods in terms of mean Average Precision (mAP).
Jungong Han, Kurt Debattista, Yanwei Pang
IEEE Trans. Multim.4
2023 Hierarchical Regression and Classification for Accurate Object Detection
abstract
Accurate object detection requires correct classification and high-quality localization. Currently, most of the single shot detectors (SSDs) conduct simultaneous classification and regression using a fully convolutional network. Despite high efficiency, this structure has some inappropriate designs for accurate object detection. The first one is the mismatch of bounding box classification, where the classification results of the default bounding boxes are improperly treated as the results of the regressed bounding boxes during the inference. The second one is that only one-time regression is not good enough for high-quality object localization. To solve the problem of classification mismatch, we propose a novel reg-offset-cls (ROC) module including three hierarchical steps: the regression of the default bounding box, the prediction of new feature sampling locations, and the classification of the regressed bounding box with more accurate features. For high-quality localization, we stack two ROC modules together. The input of the second ROC module is the output of the first ROC module. In addition, we inject a feature enhanced (FE) module between two stacked ROC modules to extract more contextual information. The experiments on three different datasets (i.e., MS COCO, PASCAL VOC, and UAVDT) are performed to demonstrate the effectiveness and superiority of our method. Without any bells or whistles, our proposed method outperforms state-of-the-art one-stage methods at a real-time speed. The source code is available at https://github.com/JialeCao001/HSD.
Jiale Cao, Yanwei Pang, Jungong Han, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 Complementary Feature Pyramid Network for Object Detection
abstract
The way of constructing a robust feature pyramid is crucial for object detection. However, existing feature pyramid methods, which aggregate multi-level features by using element-wise sum or concatenation, are inefficient to construct a robust feature pyramid. The reason is that these methods cannot be effective in discriminating the relevant semantics of objects. In this article, we propose a Complementary Feature Pyramid Network (CFPN) to aggregate multi-level features selectively and efficiently by exploring complementary information between multi-level features. Specifically, a Spatial Complementary Module (SCM) and a Channel Complementary Module (CCM) are designed and embedded in CFPN to enhance useful information and suppress irrelevant information during feature fusions along spatial and channel dimensions, respectively. CFPN is a generic feature extractor, as evidenced by its seamless integration into single-stage, two-stage, and end-to-end object detectors. Experiments conducted on the COCO and Pascal VOC datasets demonstrate that integrating our CFPN into RetinaNet, Faster RCNN, Cascade RCNN, and Sparse RCNN obtains consistent performance improvements with negligible overheads. Code and models are available at: https://github.com/VIPLab-CQU/CFPN .
Jin Xie 0005, Yanwei Pang, Jing Nie 0001, Jiale Cao, Jungong Han
ACM Trans. Multim. Comput. Commun. Appl.2
2022 PSTR: End-to-End One-Step Person Search With Transformers
abstract
We propose a novel one-step transformer-based person search framework, PSTR, that jointly performs person detection and re-identification (re-id) in a single architecture. PSTR comprises a person search-specialized (PSS) module that contains a detection encoder-decoder for person detection along with a discriminative re-id decoder for person re-id. The discriminative re-id decoder utilizes a multi-level supervision scheme with a shared decoder for discriminative re-id feature learning and also comprises a part attention block to encode relationship between different parts of a person. We further introduce a simple multi-scale scheme to support re-id across person instances at different scales. PSTR jointly achieves the diverse objectives of object-level recognition (detection) and instance-level matching (re-id). To the best of our knowledge, we are the first to propose an end-to-end one-step transformer-based person search framework. Experiments are performed on two popular benchmarks: CUHK-SYSU and PRW. Our extensive ablations reveal the merits of the proposed contributions. Further, the proposed PSTR sets a new state-of-the-art on both benchmarks. On the challenging PRW benchmark, PSTR achieves a mean average precision (mAP) score of 56.5%. The source code is available at https://github.com/JialeCao001/PSTR.
Jiale Cao, Yanwei Pang, Rao Muhammad Anwer, Hisham Cholakkal, Jin Xie 0005, Mubarak Shah, Fahad Shahbaz Khan
CVPR2
2022 Heterogeneous memory enhanced graph reasoning network for cross-modal retrieval
Zhong Ji, Yanwei Pang, Xuelong Li 0001
Sci. China Inf. Sci.4
2022 Teachers cooperation: team-knowledge distillation for multiple cross-domain few-shot learning
Zhong Ji, Jingwei Ni, Xiyao Liu 0002, Yanwei Pang
Frontiers Comput. Sci.4
2022 Multi-stream densely connected network for semantic segmentation
abstract
Abstract Semantic segmentation is a challenging task in computer vision which is widely used in autonomous driving and scene understanding. State‐of‐the‐art semantic segmentation networks, like DeepLab and PSPNet, make full use of multiple feature information to improve spatial resolution. However, the feature resolution in the scale‐axis is not dense enough for practical applications. To tackle this problem, a multi‐stream network is designed with atrous convolutional layers at multiple rates to capture objects and context at multiple scales. Furthermore, intra‐connections and inter‐connections are designed to fuse multi‐scale features densely which produce a feature pyramid with much larger scale diversity and larger receptive field by involving small quantity of computation. The proposed module can be easily used in other methods and it helps to increase the performance. Compared with existing methods, the proposed network, called Multi‐stream Densely Connected Network, reaches competitive results on ADE20K dataset, PASCAL VOC 2012 dataset, and Cityscapes dataset.
Dayu Jia, Jiale Cao, Yanwei Pang
IET Comput. Vis.4
2022 Improving 2D object detection with binocular images for outdoor surveillance
Fuchen Chu, Yanwei Pang, Jiale Cao, Jing Nie 0001, Xuelong Li 0001
Neurocomputing2
2022 PCNet: Paired channel feature volume network for accurate and efficient depth estimation
Dayu Jia, Yanwei Pang, Jiale Cao
Neurocomputing2
2022 Saliency detection network with two-stream encoder and interactive decoder
Ai-Ping Yang, Simeng Cheng, Shangyang Song, Jinbin Wang, Zhong Ji, Yanwei Pang, Jiale Cao
Neurocomputing6
2022 Reinforced pedestrian attribute recognition with group optimization reward
Zhong Ji, Zhenfei Hu, Yanwei Pang
Image Vis. Comput.5
2022 Hierarchical Correlations Replay for Continual Learning
Qiang Wang 0056, Zhong Ji, Yanwei Pang, Zhongfei Zhang
Knowl. Based Syst.4
2022 Non-linear perceptual multi-scale network for single image super-resolution
Ai-Ping Yang, Jinbin Wang, Zhong Ji, Yanwei Pang, Jiale Cao, Zihao Wei
Neural Networks5
2022 From Handcrafted to Deep Features for Pedestrian Detection: A Survey
abstract
Pedestrian detection is an important but challenging problem in computer vision, especially in human-centric tasks. Over the past decade, significant improvement has been witnessed with the help of handcrafted features and deep features. Here we present a comprehensive survey on recent advances in pedestrian detection. First, we provide a detailed review of single-spectral pedestrian detection that includes handcrafted features based methods and deep features based approaches. For handcrafted features based methods, we present an extensive review of approaches and find that handcrafted features with large freedom degrees in shape and space have better performance. In the case of deep features based approaches, we split them into pure CNN based methods and those employing both handcrafted and CNN based features. We give the statistical analysis and tendency of these methods, where feature enhanced, part-aware, and post-processing methods have attracted main attention. In addition to single-spectral pedestrian detection, we also review multi-spectral pedestrian detection, which provides more robust features for illumination variance. Furthermore, we introduce some related datasets and evaluation metrics, and a deep experimental analysis. We conclude this survey by emphasizing open problems that need to be addressed and highlighting various future directions. Researchers can track an up-to-date list at https://github.com/JialeCao001/PedSurvey.
Jiale Cao, Yanwei Pang, Jin Xie 0005, Fahad Shahbaz Khan, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Editorial for the special issue on deep learning for precise and efficient object detection
Yanwei Pang, Jungong Han, Nicola Conci
Pattern Recognit. Lett.1
2022 Stereo Refinement Dehazing Network
abstract
The performance of stereo vision tasks degrades when haze exists in the input stereo image pair. Independently applying single image dehazing algorithm on left and right images is not optimal. To overcome the problem, we propose an effective framework, called SRDNet, for simultaneously dehazing stereo images. The main idea of SRDNet is to make full use of the stereo information from cross views improving dehazing performance. It does not explicitly employ the disparity estimation and the correlation matrix. SRDNet comprises two parts: a weight-sharing coarse dehazing network (WSCDN) and a guided separated refinement network (GSRN). The WSCDN is utilized to predict a coarse dehazed image pair. Then the GSRN is introduced to predict the residues for different views by extracting the fused information of cross views and separating the features of different views with a guided channel and spatial refinement module. The residues are added to the coarse dehazed pair so as to make refinement and remove the remained haze. Experimental results demonstrate that our proposed SRDNet surpasses previous image dehazing methods by a significant margin both quantitatively and qualitatively. Moreover, our SRDNet could be a preprocessing step of the stereo-based 3D object detection and boosts the 3D detection accuracy in hazy scenes.
Jing Nie 0001, Yanwei Pang, Jin Xie 0005, Jungong Han
IEEE Trans. Circuits Syst. Video Technol.2
2022 SMAN: Stacked Multimodal Attention Network for Cross-Modal Image-Text Retrieval
abstract
This article focuses on tackling the task of the cross-modal image-text retrieval which has been an interdisciplinary topic in both computer vision and natural language processing communities. Existing global representation alignment-based methods fail to pinpoint the semantically meaningful portion of images and texts, while the local representation alignment schemes suffer from the huge computational burden for aggregating the similarity of visual fragments and textual words exhaustively. In this article, we propose a stacked multimodal attention network (SMAN) that makes use of the stacked multimodal attention mechanism to exploit the fine-grained interdependencies between image and text, thereby mapping the aggregation of attentive fragments into a common space for measuring cross-modal similarity. Specifically, we sequentially employ intramodal information and multimodal information as guidance to perform multiple-step attention reasoning so that the fine-grained correlation between image and text can be modeled. As a consequence, we are capable of discovering the semantically meaningful visual regions or words in a sentence which contributes to measuring the cross-modal similarity in a more precise manner. Moreover, we present a novel bidirectional ranking loss that enforces the distance among pairwise multimodal instances to be closer. Doing so allows us to make full use of pairwise supervised information to preserve the manifold structure of heterogeneous pairwise data. Extensive experiments on two benchmark datasets demonstrate that our SMAN consistently yields competitive performance compared to state-of-the-art methods.
Zhong Ji, Haoran Wang 0004, Jungong Han, Yanwei Pang
IEEE Trans. Cybern.4
2022 Semantic-Guided Class-Imbalance Learning Model for Zero-Shot Image Classification
abstract
In this article, we focus on the task of zero-shot image classification (ZSIC) that equips a learning system with the ability to recognize visual images from unseen classes. In contrast to the traditional image classification, ZSIC more easily suffers from the class-imbalance issue since it is more concerned with the class-level knowledge transferring capability. In the real world, the sample numbers of different categories generally follow a long-tailed distribution, and the discriminative information in the sample-scarce seen classes is hard to transfer to the related unseen classes in the traditional batch-based training manner, which degrades the overall generalization ability a lot. To alleviate the class-imbalance issue in ZSIC, we propose a sample-balanced training process to encourage all training classes to contribute equally to the learned model. Specifically, we randomly select the same number of images from each class across all training classes to form a training batch to ensure that the sample-scarce classes contribute equally as those classes with sufficient samples during each iteration. Considering that the instances from the same class differ in class representativeness, we further develop an efficient semantic-guided feature fusion model to obtain the discriminative class visual prototype for the following visual-semantic interaction process via distributing different weights to the selected samples based on their class representativeness. Extensive experiments on three imbalanced ZSIC benchmark datasets for both traditional ZSIC and generalized ZSIC tasks demonstrate that our approach achieves promising results, especially for the unseen categories that are closely related to the sample-scarce seen categories. Besides, the experimental results on two class-balanced datasets show that the proposed approach also improves the classification performance against the baseline model.
Zhong Ji, Xuejie Yu, Yunlong Yu 0001, Yanwei Pang, Zhongfei Zhang
IEEE Trans. Cybern.4
2022 DGIG-Net: Dynamic Graph-in-Graph Networks for Few-Shot Human-Object Interaction
abstract
Few-shot learning (FSL) for human-object interaction (HOI) aims at recognizing various relationships between human actions and surrounding objects only from a few samples. It is a challenging vision task, in which the diversity and interactivity of human actions result in great difficulty to learn an adaptive classifier to catch ambiguous interclass information. Therefore, traditional FSL methods usually perform unsatisfactorily in complex HOI scenes. To this end, we propose dynamic graph-in-graph networks (DGIG-Net), a novel graph prototypes framework to learn a dynamic metric space by embedding a visual subgraph to a task-oriented cross-modal graph for few-shot HOI. Specifically, we first build a knowledge reconstruction graph to learn latent representations for HOI categories by reconstructing the relationship among visual features, which generates visual representations under the category distribution of every task. Then, a dynamic relation graph integrates both reconstructible visual nodes and dynamic task-oriented semantic information to explore a graph metric space for HOI class prototypes, which applies the discriminative information from the similarities among actions or objects. We validate DGIG-Net on multiple benchmark datasets, on which it largely outperforms existing FSL approaches and achieves state-of-the-art results.
Xiyao Liu 0002, Zhong Ji, Yanwei Pang, Jungong Han, Xuelong Li 0001
IEEE Trans. Cybern.3
2022 Information Symmetry Matters: A Modal-Alternating Propagation Network for Few-Shot Learning
abstract
Semantic information provides intra-class consistency and inter-class discriminability beyond visual concepts, which has been employed in Few-Shot Learning (FSL) to achieve further gains. However, semantic information is only available for labeled samples but absent for unlabeled samples, in which the embeddings are rectified unilaterally by guiding the few labeled samples with semantics. Therefore, it is inevitable to bring a cross-modal bias between semantic-guided samples and nonsemantic-guided samples, which results in an information asymmetry problem. To address this problem, we propose a Modal-Alternating Propagation Network (MAP-Net) to supplement the absent semantic information of unlabeled samples, which builds information symmetry among all samples in both visual and semantic modalities. Specifically, the MAP-Net transfers the neighbor information by the graph propagation to generate the pseudo-semantics for unlabeled samples guided by the completed visual relationships and rectify the feature embeddings. In addition, due to the large discrepancy between visual and semantic modalities, we design a Relation Guidance (RG) strategy to guide the visual relation vectors via semantics so that the propagated information is more beneficial. Extensive experimental results on three semantic-labeled datasets, i.e., Caltech-UCSD-Birds 200-2011, SUN Attribute Database and Oxford 102 Flower, have demonstrated that our proposed method achieves promising performance and outperforms the state-of-the-art approaches, which indicates the necessity of information symmetry.
Zhong Ji, Zhishen Hou, Xiyao Liu 0002, Yanwei Pang, Jungong Han
IEEE Trans. Image Process.4
2022 CerebelluMorphic: Large-Scale Neuromorphic Model and Architecture for Supervised Motor Learning
abstract
The cerebellum plays a vital role in motor learning and control with supervised learning capability, while neuromorphic engineering devises diverse approaches to high-performance computation inspired by biological neural systems. This article presents a large-scale cerebellar network model for supervised learning, as well as a cerebellum-inspired neuromorphic architecture to map the cerebellar anatomical structure into the large-scale model. Our multinucleus model and its underpinning architecture contain approximately 3.5 million neurons, upscaling state-of-the-art neuromorphic designs by over 34 times. Besides, the proposed model and architecture incorporate 3411k granule cells, introducing a 284 times increase compared to a previous study including only 12k cells. This large scaling induces more biologically plausible cerebellar divergence/convergence ratios, which results in better mimicking biology. In order to verify the functionality of our proposed model and demonstrate its strong biomimicry, a reconfigurable neuromorphic system is used, on which our developed architecture is realized to replicate cerebellar dynamics during the optokinetic response. In addition, our neuromorphic architecture is used to analyze the dynamical synchronization within the Purkinje cells, revealing the effects of firing rates of mossy fibers on the resonance dynamics of Purkinje cells. Our experiments show that real-time operation can be realized, with a system throughput of up to 4.70 times larger than previous works with high synaptic event rate. These results suggest that the proposed work provides both a theoretical basis and a neuromorphic engineering perspective for brain-inspired computing and the further exploration of cerebellar learning.
Shuangming Yang, Jiang Wang 0002, Bin Deng 0001, Yanwei Pang, Mostafa Rahimi Azghadi
IEEE Trans. Neural Networks Learn. Syst.5
2022 Task-Oriented High-Order Context Graph Networks for Few-Shot Human-Object Interaction Recognition
abstract
Few-shot human-object interaction (FS-HOI) recognition aims at inferring new interactions between human actions and surrounding objects merely with a few available instances. It is beneficial to alleviate the long-tail and combinatorial explosion problems in human-object interaction (HOI). Nevertheless, the existing FS-HOI methods only focus on modeling the relationships between labeled samples and unlabeled samples in the Euclidean domain, which neglects the rich relational structures of the visual information among labeled samples and between human actions and objects. Accordingly, we tackle the few-shot HOI task in the non-Euclidean domain and present a graph-based model, namely, task-oriented high-order context graph network (THCG-Net). It contains a task attention module (TA-Module) and a high-order context graph module (HG-Module). In TA-Module, an attention mechanism is designed by utilizing task information to build a task-oriented space, in which the discriminative information for the current task (episode) is captured by embedding the visual features into the task-oriented space. The HG-Module is proposed to construct a task-level graph and takes the context information as high-order knowledge, which provides discriminative guidance for propagating visual information. It captures the discriminability among different categories while highlights the commonality of related categories adaptively, which effectively transfers knowledge to related categories. Extensive experimental results on two benchmark datasets, HICO-FS and TUHOI-FS, are provided. It demonstrates that our THCG-Net significantly outperforms the state-of-the-art approaches, which proves its impressive effectiveness in recognizing various human actions and surrounding objects in few-shot scenarios.
Zhong Ji, Xiyao Liu 0002, Yanwei Pang, Ling Shao 0001, Zhongfei Zhang
IEEE Trans. Syst. Man Cybern. Syst.4
2021 Triple discriminator generative adversarial network for zero-shot image classification
Zhong Ji, Jiangtao Yan, Qiang Wang 0056, Yanwei Pang, Xuelong Li 0001
Sci. China Inf. Sci.4
2021 PSC-Net: learning part spatial co-occurrence for occluded pedestrian detection
Jin Xie 0005, Yanwei Pang, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan, Ling Shao 0001
Sci. China Inf. Sci.2
2021 Multi-label correlation guided feature fusion network for abnormal ECG diagnosis
Zhaoyang Ge, Xiaoheng Jiang, Zhuang Tong, Panpan Feng, Bing Zhou 0003, Mingliang Xu 0001, Zongmin Wang, Yanwei Pang
Knowl. Based Syst.8
2021 A semi-supervised zero-shot image classification method based on soft-target
Zhong Ji, Qiang Wang 0056, Biying Cui, Yanwei Pang, Xianbin Cao 0001, Xuelong Li 0001
Neural Networks4
2021 Cascaded hierarchical atrous spatial pyramid pooling module for semantic segmentation
Xuhang Lian, Yanwei Pang, Jungong Han
Pattern Recognit.2
2021 Efficient Selective Context Network for Accurate Object Detection
abstract
Single-stage detectors have gained great attention due to their high detection accuracy and real-time speed. To detect multi-scale objects, single-stage detectors make scale-aware predictions based on multiple pyramid layers. However, the insufficient context exploration in shallow pyramid layers leads to the detection accuracy of small objects being far from satisfactory. To tackle this problem, we propose a scheme to selectively extract multi-scale context with attention-adaptive weights. Specifically, we propose an efficient selective context network for accurate object detection. It incorporates an enhanced context module and a triple attention module. The enhanced context module consists of multi-branches to extract original-scale, small-scale, and large-scale contextual information. To make full use of this context and filter out noisy information, the triple attention module, which contains global-level, channel-level, and spatial-level attentions, is introduced to carry out selective context fusion. The two modules are easy to implement and can efficiently boost the accuracy of object detection. The performance of our method is validated on two benchmarks: PASCAL VOC and MS COCO. For a 512×512 input, our detector with VGG16 achieves competitive results (80.9 on the Pascal VOC 2012 test set in the case of single-scale inference without MS COCO pre-training). On the MS COCO test-dev set, our detector with ResNet101 outperforms RetinaNet500 by 2.5% AP in terms of overall performance and its speed is 48 milliseconds on a Titan XP GPU. As a result, ESCNet achieves a better trade-off between accuracy and speed.
Jing Nie 0001, Yanwei Pang, Shengjie Zhao 0001, Jungong Han, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Few-Shot Human-Object Interaction Recognition With Semantic-Guided Attentive Prototypes Network
abstract
Extreme instance imbalance among categories and combinatorial explosion make the recognition of Human-Object Interaction (HOI) a challenging task. Few studies have addressed both challenges directly. Motivated by the success of few-shot learning that learns a robust model from a few instances, we formulate HOI as a few-shot task in a meta-learning framework to alleviate the above challenges. Due to the fact that the intrinsical characteristic of HOI is diverse and interactive, we propose a Semantic-guided Attentive Prototypes Network (SAPNet) framework to learn a semantic-guided metric space where HOI recognition can be performed by computing distances to attentive prototypes of each class. Specifically, the model generates attentive prototypes guided by the category names of actions and objects, which highlight the commonalities of images from the same class in HOI. In addition, we design two alternative prototypes calculation methods, i.e., Prototypes Shift (PS) approach and Hallucinatory Graph Prototypes (HGP) approach, which explore to learn a suitable category prototypes representations in HOI. Finally, in order to realize the task of few-shot HOI, we reorganize 2 HOI benchmark datasets with 2 split strategies, i.e., HICO-NN, TUHOI-NN, HICO-NF, and TUHOI-NF. Extensive experimental results on these datasets have demonstrated the effectiveness of our proposed SAPNet approach.
Zhong Ji, Xiyao Liu 0002, Yanwei Pang, Wanli Ouyang, Xuelong Li 0001
IEEE Trans. Image Process.3
2021 Improving Single Shot Object Detection With Feature Scale Unmixing
abstract
Due to the advantages of real-time detection and improved performance, single-shot detectors have gained great attention recently. To solve the complex scale variations, single-shot detectors make scale-aware predictions based on multiple pyramid layers. Typically, small objects are detected on shallow layers while large objects are detected on deep layers. However, the features in the pyramid are not scale-aware enough, which limits the detection performance. Two common problems in single-shot detectors caused by object scale variations can be observed: (1) false negative problem, i.e., small objects are easily missed due to the weak features; (2) part-false positive problem, i.e., the salient part of a large object is sometimes detected as an object. With this observation, a new Neighbor Erasing and Transferring (NET) mechanism is proposed for feature scale-unmixing to explore scale-aware features in this paper. In NET, a Neighbor Erasing Module (NEM) is designed to erase the salient features of large objects and emphasize the features of small objects in shallow layers. A Neighbor Transferring Module (NTM) is introduced to transfer the erased features and highlight large objects in deep layers. With this mechanism, a single-shot network called NETNet is constructed for scale-aware object detection. In addition, we propose to aggregate nearest neighboring pyramid features to enhance our NET. Experiments on MS COCO dataset and UAVDT dataset demonstrate the effectiveness of our method. NETNet obtains 38.5% AP at a speed of 27 FPS and 32.0% AP at a speed of 55 FPS on MS COCO dataset. As a result, NETNet achieves a better trade-off for real-time and accurate object detection.
Yazhao Li, Yanwei Pang, Jiale Cao, Jianbing Shen, Ling Shao 0001
IEEE Trans. Image Process.2
2021 TJU-DHD: A Diverse High-Resolution Dataset for Object Detection
abstract
Vehicles, pedestrians, and riders are the most important and interesting objects for the perception modules of self-driving vehicles and video surveillance. However, the state-of-the-art performance of detecting such important objects (esp. small objects) is far from satisfying the demand of practical systems. Large-scale, rich-diversity, and high-resolution datasets play an important role in developing better object detection methods to satisfy the demand. Existing public large-scale datasets such as MS COCO collected from websites do not focus on the specific scenarios. Moreover, the popular datasets (e.g., KITTI and Citypersons) collected from the specific scenarios are limited in the number of images and instances, the resolution, and the diversity. To attempt to solve the problem, we build a diverse high-resolution dataset (called TJU-DHD). The dataset contains 115354 high-resolution images (52% images have a resolution of 1624×1200 pixels and 48% images have a resolution of at least 2, 560×1.440 pixels) and 709 330 labeled objects in total with a large variance in scale and appearance. Meanwhile, the dataset has a rich diversity in season variance, illumination variance, and weather variance. In addition, a new diverse pedestrian dataset is further built. With the four different detectors (i.e., the one-stage RetinaNet, anchor-free FCOS, two-stage FPN, and Cascade R-CNN), experiments about object detection and pedestrian detection are conducted. We hope that the newly built dataset can help promote the research on object detection and pedestrian detection in these two scenes. The dataset is available at https://github.com/tjubiit/TJU-DHD.
Yanwei Pang, Jiale Cao, Yazhao Li, Jin Xie 0005, Hanqing Sun 0001, Jinfeng Gong
IEEE Trans. Image Process.1
2021 Mask-Guided Attention Network and Occlusion-Sensitive Hard Example Mining for Occluded Pedestrian Detection
abstract
Pedestrian detection relying on deep convolution neural networks has made significant progress. Though promising results have been achieved on standard pedestrians, the performance on heavily occluded pedestrians remains far from satisfactory. The main culprits are intra-class occlusions involving other pedestrians and inter-class occlusions caused by other objects, such as cars and bicycles. These result in a multitude of occlusion patterns. We propose an approach for occluded pedestrian detection with the following contributions. First, we introduce a novel mask-guided attention network that fits naturally into popular pedestrian detection pipelines. Our attention network emphasizes on visible pedestrian regions while suppressing the occluded ones by modulating full body features. Second, we propose the occlusion-sensitive hard example mining method and occlusion-sensitive loss that mines hard samples according to the occlusion level and assigns higher weights to the detection errors occurring at highly occluded pedestrians. Third, we empirically demonstrate that weak box-based segmentation annotations provide reasonable approximation to their dense pixel-wise counterparts. Experiments are performed on CityPersons, Caltech and ETH datasets. Our approach sets a new state-of-the-art on all three datasets. Our approach obtains an absolute gain of 10.3% in log-average miss rate, compared with the best reported results on the heavily occluded HO pedestrian set of the CityPersons test set. Code and models are available at: https://github.com/Leotju/MGAN.
Jin Xie 0005, Yanwei Pang, Muhammad Haris Khan, Rao Muhammad Anwer, Fahad Shahbaz Khan, Ling Shao 0001
IEEE Trans. Image Process.2
2021 Density-Aware Multi-Task Learning for Crowd Counting
abstract
In this paper, we present a method called density-aware convolutional neural network (DensityCNN) to perform the crowd counting task in various crowded scenes. The key idea of the DensityCNN is to utilize high-level semantic information to provide guidance and constraint when generating density maps. To this end, we implement the DensityCNN by adopting a multi-task CNN structure to jointly learn density-level classification and density map estimation. The density-level classification task learns multi-channel semantic features that are aware of the density distributions of the input image. This task is accomplished via our specially designed group-based convolutional structure in a supervised learning manner. In the density map estimation task, these semantic features are deployed together with high-dimension convolutional features to generate density maps with lower count errors. Extensive experiments on four challenging crowd datasets (ShanghaiTech, UCF_CC_50, UCF-QNCF, and WorldExpo'10) and one vehicle dataset TRANCOS demonstrate the effectiveness of the proposed method.
Xiaoheng Jiang, Li Zhang 0072, Tianzhu Zhang 0001, Pei Lv, Bing Zhou 0003, Yanwei Pang, Mingliang Xu 0001, Changsheng Xu
IEEE Trans. Multim.6
2021 Deep Attentive Video Summarization With Distribution Consistency Learning
abstract
This article studies supervised video summarization by formulating it into a sequence-to-sequence learning framework, in which the input and output are sequences of original video frames and their predicted importance scores, respectively. Two critical issues are addressed in this article: short-term contextual attention insufficiency and distribution inconsistency. The former lies in the insufficiency of capturing the short-term contextual attention information within the video sequence itself since the existing approaches focus a lot on the long-term encoder-decoder attention. The latter refers to the distributions of predicted importance score sequence and the ground-truth sequence is inconsistent, which may lead to a suboptimal solution. To better mitigate the first issue, we incorporate a self-attention mechanism in the encoder to highlight the important keyframes in a short-term context. The proposed approach alongside the encoder-decoder attention constitutes our deep attentive models for video summarization. For the second one, we propose a distribution consistency learning method by employing a simple yet effective regularization loss term, which seeks a consistent distribution for the two sequences. Our final approach is dubbed as Attentive and Distribution consistent video Summarization (ADSum). Extensive experiments on benchmark data sets demonstrate the superiority of the proposed ADSum approach against state-of-the-art approaches.
Zhong Ji, Yanwei Pang, Xi Li 0001, Jungong Han
IEEE Trans. Neural Networks Learn. Syst.3
2020 SGAP-Net: Semantic-Guided Attentive Prototypes Network for Few-Shot Human-Object Interaction Recognition
abstract
Extreme instance imbalance among categories and combinatorial explosion make the recognition of Human-Object Interaction (HOI) a challenging task. Few studies have addressed both challenges directly. Motivated by the success of few-shot learning that learns a robust model from a few instances, we formulate HOI as a few-shot task in a meta-learning framework to alleviate the above challenges. Due to the fact that the intrinsic characteristic of HOI is diverse and interactive, we propose a Semantic-Guided Attentive Prototypes Network (SGAP-Net) to learn a semantic-guided metric space where HOI recognition can be performed by computing distances to attentive prototypes of each class. Specifically, the model generates attentive prototypes guided by the category names of actions and objects, which highlight the commonalities of images from the same class in HOI. In addition, we design a novel decision method to alleviate the biases produced by different patterns of the same action in HOI. Finally, in order to realize the task of few-shot HOI, we reorganize two HOI benchmark datasets, i.e., HICO-FS and TUHOI-FS, to realize the task of few-shot HOI. Extensive experimental results on both datasets have demonstrated the effectiveness of our proposed SGAP-Net approach.
Zhong Ji, Xiyao Liu 0002, Yanwei Pang, Xuelong Li 0001
AAAI3
2020 D2Det: Towards High Quality Object Detection and Instance Segmentation
abstract
We propose a novel two-stage detection method, D2Det, that collectively addresses both precise localization and accurate classification. For precise localization, we introduce a dense local regression that predicts multiple dense box offsets for an object proposal. Different from traditional regression and keypoint-based localization employed in two-stage detectors, our dense local regression is not limited to a quantized set of keypoints within a fixed region and has the ability to regress position-sensitive real number dense offsets, leading to more precise localization. The dense local regression is further improved by a binary overlap prediction strategy that reduces the influence of background region on the final box regression. For accurate classification, we introduce a discriminative RoI pooling scheme that samples from various sub-regions of a proposal and performs adaptive weighting to obtain discriminative features. On MS COCO test-dev, our D2Det outperforms existing two-stage methods, with a single-model performance of 45.4 AP, using ResNet101 backbone. When using multi-scale training and inference, D2Det obtains AP of 50.1. In addition to detection, we adapt D2Det for instance segmentation, achieving a mask AP of 40.2 with a two-fold speedup, compared to the state-of-the-art. We also demonstrate the effectiveness of our D2Det on airborne sensors by performing experiments for object detection in UAV images (UAVDT dataset) and instance segmentation in satellite images (iSAID dataset). Source code is available at https://github.com/JialeCao001/D2Det.
Jiale Cao, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan, Yanwei Pang, Ling Shao 0001
CVPR5
2020 Attention Scaling for Crowd Counting
abstract
Convolutional Neural Network (CNN) based methods generally take crowd counting as a regression task by outputting crowd densities. They learn the mapping between image contents and crowd density distributions. Though having achieved promising results, these data-driven counting networks are prone to overestimate or underestimate people counts of regions with different density patterns, which degrades the whole count accuracy. To overcome this problem, we propose an approach to alleviate the counting performance differences in different regions. Specifically, our approach consists of two networks named Density Attention Network (DANet) and Attention Scaling Network (ASNet). DANet provides ASNet with attention masks related to regions of different density levels. ASNet first generates density maps and scaling factors and then multiplies them by attention masks to output separate attention-based density maps. These density maps are summed to give the final density map. The attention scaling factors help attenuate the estimation errors in different regions. Furthermore, we present a novel Adaptive Pyramid Loss (APLoss) to hierarchically calculate the estimation losses of sub-regions, which alleviates the training bias. Extensive experiments on four challenging datasets (ShanghaiTech Part A, UCF_CC_50, UCF-QNRF, and WorldExpo'10) demonstrate the superiority of the proposed approach.
Xiaoheng Jiang, Li Zhang 0072, Mingliang Xu 0001, Tianzhu Zhang 0001, Pei Lv, Bing Zhou 0003, Xin Yang 0011, Yanwei Pang
CVPR8
2020 NETNet: Neighbor Erasing and Transferring Network for Better Single Shot Object Detection
abstract
Due to the advantages of real-time detection and improved performance, single-shot detectors have gained great attention recently. To solve the complex scale variations, single-shot detectors make scale-aware predictions based on multiple pyramid layers. However, the features in the pyramid are not scale-aware enough, which limits the detection performance. Two common problems in single-shot detectors caused by object scale variations can be observed: (1) small objects are easily missed; (2) the salient part of a large object is sometimes detected as an object. With this observation, we propose a new Neighbor Erasing and Transferring (NET) mechanism to reconfigure the pyramid features and explore scale-aware features. In NET, a Neighbor Erasing Module (NEM) is designed to erase the salient features of large objects and emphasize the features of small objects in shallow layers. A Neighbor Transferring Module (NTM) is introduced to transfer the erased features and highlight large objects in deep layers. With this mechanism, a single-shot network called NETNet is constructed for scale-aware object detection. In addition, we propose to aggregate nearest neighboring pyramid features to enhance our NET. NETNet achieves 38.5% AP at a speed of 27 FPS and 32.0% AP at a speed of 55 FPS on MS COCO dataset. As a result, NETNet achieves a better trade-off for real-time and accurate object detection.
Yazhao Li, Yanwei Pang, Jianbing Shen, Jiale Cao, Ling Shao 0001
CVPR2
2020 BidNet: Binocular Image Dehazing Without Explicit Disparity Estimation
abstract
Heavy haze results in severe image degradation and thus hampers the performance of visual perception, object detection, etc. On the assumption that dehazed binocular images are superior to the hazy ones for stereo vision tasks such as 3D object detection and according to the fact that image haze is a function of depth, this paper proposes a Binocular image dehazing Network (BidNet) aiming at dehazing both the left and right images of binocular images within the deep learning framework. Existing binocular dehazing methods rely on simultaneously dehazing and estimating disparity, whereas BidNet does not need to explicitly perform time-consuming and well-known challenging disparity estimation. Note that a small error in disparity gives rise to a large variation in depth and in estimation of haze-free image. The relationship and correlation between binocular images are explored and encoded by the proposed Stereo Transformation Module (STM). Jointly dehazing binocular image pairs is mutually beneficial, which is better than only dehazing left images. We extend the Foggy Cityscapes dataset to a Stereo Foggy Cityscapes dataset with binocular foggy image pairs. Experimental results demonstrate that BidNet significantly outperforms state-of-the-art dehazing methods in both subjective and objective assessments.
Yanwei Pang, Jing Nie 0001, Jin Xie 0005, Jungong Han, Xuelong Li 0001
CVPR1
2020 Hierarchical Human Parsing With Typed Part-Relation Reasoning
abstract
Human parsing is for pixel-wise human semantic understanding. As human bodies are underlying hierarchically structured, how to model human structures is the central theme in this task. Focusing on this, we seek to simultaneously exploit the representational capacity of deep graph networks and the hierarchical human structures. In particular, we provide following two contributions. First, three kinds of part relations, i.e., decomposition, composition, and dependency, are, for the first time, completely and precisely described by three distinct relation networks. This is in stark contrast to previous parsers, which only focus on a portion of the relations and adopt a type-agnostic relation modeling strategy. More expressive relation information can be captured by explicitly imposing the parameters in the relation networks to satisfy the specific characteristics of different relations. Second, previous parsers largely ignore the need for an approximation algorithm over the loopy human hierarchy, while we instead address an iterative reasoning process, by assimilating generic message-passing networks with their edge-typed, convolutional counterparts. With these efforts, our parser lays the foundation for more sophisticated and flexible human relation patterns of reasoning. Comprehensive experiments on five datasets demonstrate that our parser sets a new state-of-the-art on each.
Wenguan Wang, Hailong Zhu, Jifeng Dai, Yanwei Pang, Jianbing Shen, Ling Shao 0001
CVPR4
2020 SipMask: Spatial Information Preservation for Fast Image and Video Instance Segmentation
Jiale Cao, Rao Muhammad Anwer, Hisham Cholakkal, Fahad Shahbaz Khan, Yanwei Pang, Ling Shao 0001
ECCV (14)5
2020 Consensus-Aware Visual-Semantic Embedding for Image-Text Matching
Haoran Wang 0004, Zhong Ji, Yanwei Pang
ECCV (24)4
2020 Count- and Similarity-Aware R-CNN for Pedestrian Detection
Jin Xie 0005, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan, Yanwei Pang, Ling Shao 0001, Mubarak Shah
ECCV (17)5
2020 Learning Local Structure of Representative Points for Point Cloud Classification and Semantic Segmentation
abstract
Directly processing large point clouds is inefficient. State-of-the-art frameworks hierarchically employ Farthest Point Sampling (FPS) to down-sample the data for point cloud classification, semantic segmentation, etc. However, using only geometric information, FPS neglects the importance of semantic information of each point. To solve this problem, we propose a learning-based block, named Representative Points Block (RPB), to select the most representative points of an irregular point cloud according to the task. RPB takes the information of semantic interest of each point into account, and preserves the structure of the point cloud. We construct our network termed RP-Net by performing feature extraction on hierarchical representative points for point cloud classification and semantic segmentation. With further observation that local shape of representative points is different, a graph-based method is used to explore the features of representative points. Experimental results on challenging benchmarks demonstrate that RPB is more efficient and effective than FPS and RP-Net achieves state-of-the-art performance.
Yanwei Pang, Yuefeng Wu, Yazhao Li
ICASSP2
2020 Special focus on deep learning for computer vision
Xiang Bai, Yanwei Pang, Guofeng Zhang 0001
Sci. China Inf. Sci.2
2020 Preserving details in semantics-aware context for scene parsing
Shuai Ma 0001, Yanwei Pang, Ling Shao 0001
Sci. China Inf. Sci.2
2020 CGNet: cross-guidance network for semantic segmentation
Yanwei Pang
Sci. China Inf. Sci.2
2020 Deep attentive and semantic preserving video summarization
Zhong Ji, Fang Jiao, Yanwei Pang, Ling Shao 0001
Neurocomputing3
2020 Dual triplet network for image zero-shot learning
Zhong Ji, Yanwei Pang, Ling Shao 0001
Neurocomputing3
2020 Toward improving ECG biometric identification using cascaded convolutional neural networks
Yazhao Li, Yanwei Pang, Kongqiao Wang, Xuelong Li 0001
Neurocomputing2
2020 Semantic segmentation with hybrid pyramid pooling and stacked pyramid structure
Xuhang Lian, Yanwei Pang, Jungong Han
Neurocomputing2
2020 Lightweight group convolutional network for single image super-resolution
Ai-Ping Yang, Bingwang Yang, Zhong Ji, Yanwei Pang, Ling Shao 0001
Inf. Sci.4
2020 Multi-layer Attention Based CNN for Target-Dependent Sentiment Classification
Suqi Zhang, Xinyun Xu, Yanwei Pang, Jungong Han
Neural Process. Lett.3
2020 Stacked squeeze-and-excitation recurrent residual network for visual-semantic matching
Haoran Wang 0004, Zhong Ji, Zhigang Lin, Yanwei Pang, Xuelong Li 0001
Pattern Recognit.4
2020 Improved prototypical networks for few-Shot learning
Zhong Ji, Xingliang Chai, Yunlong Yu 0001, Yanwei Pang, Zhongfei Zhang
Pattern Recognit. Lett.4
2020 Pedestrian attribute recognition based on multiple time steps attention
Zhong Ji, Zhenfei Hu, Erlu He, Jungong Han, Yanwei Pang
Pattern Recognit. Lett.5
2020 Cross-modal guidance based auto-encoder for multi-video summarization
Zhong Ji, Yanwei Pang, Xuelong Li 0001
Pattern Recognit. Lett.3
2020 High-Level Semantic Networks for Multi-Scale Object Detection
abstract
To better solve scale variance problem, deep multi-scale methods usually detect objects of different scales by different in-network layers. However, the semantic levels of features from different layers are usually inconsistent. In this paper, we propose a multi-branch and high-level semantic network by gradually splitting a base network into multiple different branches. As a result, the different branches have same depth and the output features of different branches have similarly high-level semantics. Due to the difference of receptive fields, the different branches are suitable to detect objects of different scales. Meanwhile, the multi-branch network does not introduce additional parameters by sharing the convolutional weights of different branches. To further improve detection performance, skip-layer connections are used to add context to the branch of relatively small receptive field, and dilated convolution is incorporated to enlarge the resolutions of output feature maps. When they are embedded into Faster RCNN architecture, the weighted scores of proposal generation network and proposal classification network are further proposed. Experiments on three pedestrian datasets (i.e., the KITTI dataset, the Caltech dataset, and the Citypersons dataset), one face dataset (i.e., the WIDER FACE dataset), and two general object datasets (i.e., the COCO benchmark and the PASCAL VOC dataset) demonstrate the effectiveness and generality of proposed method. On these datasets, our method achieves state-of-the-art performance.
Jiale Cao, Yanwei Pang, Shengjie Zhao 0001, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 Video Summarization With Attention-Based Encoder-Decoder Networks
abstract
This paper addresses the problem of supervised video summarization by formulating it as a sequence-to-sequence learning problem, where the input is a sequence of original video frames, and the output is a keyshot sequence. Our key idea is to learn a deep summarization network with attention mechanism to mimic the way of selecting the keyshots of human. To this end, we propose a novel video summarization framework named attentive encoder-decoder networks for video summarization (AVS), in which the encoder uses a bidirectional long short-term memory (BiLSTM) to encode the contextual information among the input video frames. As for the decoder, two attention-based LSTM networks are explored by using additive and multiplicative objective functions, respectively. Extensive experiments are conducted on two video summarization benchmark datasets, i.e., SumMe and TVSum. The results demonstrate the superiority of the proposed AVS-based approaches against the state-of-the-art approaches, with remarkable improvements on both datasets.
Zhong Ji, Kailin Xiong, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 Spectral Clustering by Joint Spectral Embedding and Spectral Rotation
abstract
Spectral clustering is an important clustering method widely used for pattern recognition and image segmentation. Classical spectral clustering algorithms consist of two separate stages: 1) solving a relaxed continuous optimization problem to obtain a real matrix followed by 2) applying K -means or spectral rotation to round the real matrix (i.e., continuous clustering result) into a binary matrix called the cluster indicator matrix. Such a separate scheme is not guaranteed to achieve jointly optimal result because of the loss of useful information. To obtain a better clustering result, in this paper, we propose a joint model to simultaneously compute the optimal real matrix and binary matrix. The existing joint model adopts an orthonormal real matrix to approximate the orthogonal but nonorthonormal cluster indicator matrix. It is noted that only in a very special case (i.e., all clusters have the same number of samples), the cluster indicator matrix is an orthonormal matrix multiplied by a real number. The error of approximating a nonorthonormal matrix is inevitably large. To overcome the drawback, we propose replacing the nonorthonormal cluster indicator matrix with a scaled cluster indicator matrix which is an orthonormal matrix. Our method is capable of obtaining better performance because it is easy to minimize the difference between two orthonormal matrices. Experimental results on benchmark datasets demonstrate the effectiveness of the proposed method (called JSESR).
Yanwei Pang, Jin Xie 0005, Feiping Nie 0001, Xuelong Li 0001
IEEE Trans. Cybern.1
2020 Taking a Look at Small-Scale Pedestrians and Occluded Pedestrians
abstract
Small-scale pedestrian detection and occluded pedestrian detection are two challenging tasks. However, most state-of-the-art methods merely handle one single task each time, thus giving rise to relatively poor performance when the two tasks, in practice, are required simultaneously. In this paper, it is found that small-scale pedestrian detection and occluded pedestrian detection actually have a common problem, i.e., an inaccurate location problem. Therefore, solving this problem enables to improve the performance of both tasks. To this end, we pay more attention to the predicted bounding box with worse location precision and extract more contextual information around objects, where two modules (i.e., location bootstrap and semantic transition) are proposed. The location bootstrap is used to reweight regression loss, where the loss of the predicted bounding box far from the corresponding ground-truth is upweighted and the loss of the predicted bounding box near the corresponding ground-truth is downweighted. Additionally, the semantic transition adds more contextual information and relieves semantic inconsistency of the skip-layer fusion. Since the location bootstrap is not used at the test stage and the semantic transition is lightweight, the proposed method does not add many extra computational costs during inference. Experiments on the challenging CityPersons and Caltech datasets show that the proposed method outperforms the state-of-the-art methods on the small-scale pedestrians and occluded pedestrians (e.g., 5.20% and 4.73% improvements on the Caltech).
Jiale Cao, Yanwei Pang, Jungong Han, Bolin Gao, Xuelong Li 0001
IEEE Trans. Image Process.2
2020 Attribute-Guided Network for Cross-Modal Zero-Shot Hashing
abstract
Zero-shot hashing (ZSH) aims at learning a hashing model that is trained only by instances from seen categories but can generate well to those of unseen categories. Typically, it is achieved by utilizing a semantic embedding space to transfer knowledge from seen domain to unseen domain. Existing efforts mainly focus on single-modal retrieval task, especially image-based image retrieval (IBIR). However, as a highlighted research topic in the field of hashing, cross-modal retrieval is more common in real-world applications. To address the cross-modal ZSH (CMZSH) retrieval task, we propose a novel attribute-guided network (AgNet), which can perform not only IBIR but also text-based image retrieval (TBIR). In particular, AgNet aligns different modal data into a semantically rich attribute space, which bridges the gap caused by modality heterogeneity and zero-shot setting. We also design an effective strategy that exploits the attribute to guide the generation of hash codes for image and text within the same network. Extensive experimental results on three benchmark data sets (AwA, SUN, and ImageNet) demonstrate the superiority of AgNet on both cross-modal and single-modal zero-shot image retrieval tasks.
Zhong Ji, Yunlong Yu 0001, Yanwei Pang, Jungong Han
IEEE Trans. Neural Networks Learn. Syst.4
2020 Learning Multi-Level Density Maps for Crowd Counting
abstract
People in crowd scenes often exhibit the characteristic of imbalanced distribution. On the one hand, people size varies largely due to the camera perspective. People far away from the camera look smaller and are likely to occlude each other, whereas people near to the camera look larger and are relatively sparse. On the other hand, the number of people also varies greatly in the same or different scenes. This article aims to develop a novel model that can accurately estimate the crowd count from a given scene with imbalanced people distribution. To this end, we have proposed an effective multi-level convolutional neural network (MLCNN) architecture that first adaptively learns multi-level density maps and then fuses them to predict the final output. Density map of each level focuses on dealing with people of certain sizes. As a result, the fusion of multi-level density maps is able to tackle the large variation in people size. In addition, we introduce a new loss function named balanced loss (BL) to impose relatively BL feedback during training, which helps further improve the performance of the proposed network. Furthermore, we introduce a new data set including 1111 images with a total of 49 061 head annotations. MLCNN is easy to train with only one end-to-end training stage. Experimental results demonstrate that our MLCNN achieves state-of-the-art performance. In particular, our MLCNN reaches a mean absolute error (MAE) of 242.4 on the UCF_CC_50 data set, which is 37.2 lower than the second-best result.
Xiaoheng Jiang, Li Zhang 0072, Pei Lv, Yibo Guo, Ruijie Zhu 0001, Yanwei Pang, Xi Li 0001, Bing Zhou 0003, Mingliang Xu 0001
IEEE Trans. Neural Networks Learn. Syst.7
2019 Triply Supervised Decoder Networks for Joint Detection and Segmentation
abstract
Joint object detection and semantic segmentation is essential in many fields such as self-driving cars. An initial attempt towards this goal is to simply share a single network for multi-task learning. We argue that it does not make full use of the fact that detection and segmentation are mutually beneficial. In this paper, we propose a framework called TripleNet to deeply boost these two tasks. On the one hand, to deeply join the two tasks at different scales, triple supervisions including detection-oriented supervision and class-aware/agnostic segmentation supervisions are imposed on each layer of the decoder. Class-agnostic segmentation provides an objectness prior to detection and segmentation. On the other hand, to further intercross the two tasks and refine the features in each scale, two light-weight modules (i.e., the inner-connected module and the attention skip-layer fusion) are incorporated. Because segmentation supervision on each decoder layer are not performed at the test stage and two added modules are light-weight, the proposed TripleNet can run at a real-time speed (16 fps). Experiments on the VOC 2007/2012 and COCO datasets show that TripleNet outperforms all the other one-stage methods on both two tasks (e.g., 81.9% mAP and 83.3% mIoU on VOC 2012, and 37.1% mAP and 59.6% mIoU on COCO) by a single network.
Jiale Cao, Yanwei Pang, Xuelong Li 0001
CVPR2
2019 Efficient Featurized Image Pyramid Network for Single Shot Detector
abstract
Single-stage object detectors have recently gained popularity due to their combined advantage of high detection accuracy and real-time speed. However, while promising results have been achieved by these detectors on standard-sized objects, their performance on small objects is far from satisfactory. To detect very small/large objects, classical pyramid representation can be exploited, where an image pyramid is used to build a feature pyramid (featurized image pyramid), enabling detection across a range of scales. Existing single-stage detectors avoid such a featurized image pyramid representation due to its memory and time complexity. In this paper, we introduce a light-weight architecture to efficiently produce featurized image pyramid in a single-stage detection framework. The resulting multi-scale features are then injected into the prediction layers of the detector using an attention module. The performance of our detector is validated on two benchmarks: PASCAL VOC and MS COCO. For a 300×300 input, our detector operates at 111 frames per second (FPS) on a Titan X GPU, providing state-of-the-art detection accuracy on PASCAL VOC 2007 testset. On the MS COCO testset, our detector achieves state-of-the-art results surpassing all existing single-stage methods in the case of single-scale inference.
Yanwei Pang, Tiancai Wang, Rao Muhammad Anwer, Fahad Shahbaz Khan, Ling Shao 0001
CVPR1
2019 Hierarchical Shot Detector
abstract
Single shot detector simultaneously predicts object categories and regression offsets of the default boxes. Despite of high efficiency, this structure has some inappropriate designs: (1) The classification result of the default box is improperly assigned to that of the regressed box during inference, (2) Only regression once is not good enough for accurate object detection. To solve the first problem, a novel reg-offset-cls (ROC) module is proposed. It contains three hierarchical steps: box regression, the feature sampling location predication, and the regressed box classification with the features of offset locations. To further solve the second problem, a hierarchical shot detector (HSD) is proposed, which stacks two ROC modules and one feature enhanced module. The second ROC treats the regressed boxes and the feature sampling locations of features in the first ROC as the inputs. Meanwhile, the feature enhanced module injected between two ROCs aims to extract the local and non-local context. Experiments on the MS COCO and PASCAL VOC datasets demonstrate the superiority of proposed HSD. Without the bells or whistles, HSD outperforms all one-stage methods at real-time speed.
Jiale Cao, Yanwei Pang, Jungong Han, Xuelong Li 0001
ICCV2
2019 Saliency-Guided Attention Network for Image-Sentence Matching
abstract
This paper studies the task of matching image and sentence, where learning appropriate representations to bridge the semantic gap between image contents and language appears to be the main challenge. Unlike previous approaches that predominantly deploy symmetrical architecture to represent both modalities, we introduce a Saliency-guided Attention Network (SAN) that is characterized by building an asymmetrical link between vision and language to efficiently learn a fine-grained cross-modal correlation. The proposed SAN mainly includes three components: saliency detector, Saliency-weighted Visual Attention (SVA) module, and Saliency-guided Textual Attention (STA) module. Concretely, the saliency detector provides the visual saliency information to drive both two attention modules. Taking advantage of the saliency information, SVA is able to learn more discriminative visual features. By fusing the visual information from SVA and intra-modal information as a multi-modal guidance, STA affords us powerful textual representations that are synchronized with visual clues. Extensive experiments demonstrate SAN can improve the state-of-the-art results on the benchmark Flickr30K and MSCOCO datasets by a large margin.
Zhong Ji, Haoran Wang 0004, Jungong Han, Yanwei Pang
ICCV4
2019 Enriched Feature Guided Refinement Network for Object Detection
abstract
We propose a single-stage detection framework that jointly tackles the problem of multi-scale object detection and class imbalance. Rather than designing deeper networks, we introduce a simple yet effective feature enrichment scheme to produce multi-scale contextual features. We further introduce a cascaded refinement scheme which first instills multi-scale contextual features into the prediction layers of the single-stage detector in order to enrich their discriminative power for multi-scale detection. Second, the cascaded refinement scheme counters the class imbalance problem by refining the anchors and enriched features to improve classification and regression. Experiments are performed on two benchmarks: PASCAL VOC and MS COCO. For a 320×320 input on the MS COCO test-dev, our detector achieves state-of-the-art single-stage detection accuracy with a COCO AP of 33.2 in the case of single-scale inference, while operating at 21 milliseconds on a Titan XP GPU. For a 512×512 input on the MS COCO test-dev, our approach obtains an absolute gain of 1.6% in terms of COCO AP, compared to the best reported single-stage results[5]. Source code and models are available at: https://github.com/Ranchentx/EFGRNet.
Jing Nie 0001, Rao Muhammad Anwer, Hisham Cholakkal, Fahad Shahbaz Khan, Yanwei Pang, Ling Shao 0001
ICCV5
2019 Towards Bridging Semantic Gap to Improve Semantic Segmentation
abstract
Aggregating multi-level features is essential for capturing multi-scale context information for precise scene semantic segmentation. However, the improvement by directly fusing shallow features and deep features becomes limited as the semantic gap between them increases. To solve this problem, we explore two strategies for robust feature fusion. One is enhancing shallow features using a semantic enhancement module (SeEM) to alleviate the semantic gap between shallow features and deep features. The other strategy is feature attention, which involves discovering complementary information (i.e., boundary information) from low-level features to enhance high-level features for precise segmentation. By embedding these two strategies, we construct a parallel feature pyramid towards improving multi-level feature fusion. A Semantic Enhanced Network called SeENet is constructed with the parallel pyramid to implement precise segmentation. Experiments on three benchmark datasets demonstrate the effectiveness of our method for robust multi-level feature aggregation. As a result, our SeENet has achieved better performance than other state-of-the-art methods for semantic segmentation.
Yanwei Pang, Yazhao Li, Jianbing Shen, Ling Shao 0001
ICCV1
2019 Mask-Guided Attention Network for Occluded Pedestrian Detection
abstract
Pedestrian detection relying on deep convolution neural networks has made significant progress. Though promising results have been achieved on standard pedestrians, the performance on heavily occluded pedestrians remains far from satisfactory. The main culprits are intra-class occlusions involving other pedestrians and inter-class occlusions caused by other objects, such as cars and bicycles. These results in a multitude of occlusion patterns. We propose an approach for occluded pedestrian detection with the following contributions. First, we introduce a novel mask-guided attention network that fits naturally into popular pedestrian detection pipelines. Our attention network emphasizes on visible pedestrian regions while suppressing the occluded ones by modulating full body features. Second, we empirically demonstrate that coarse-level segmentation annotations provide reasonable approximation to their dense pixel-wise counterparts. Experiments are performed on CityPersons and Caltech datasets. Our approach sets a new state-of-the-art on both datasets. Our approach obtains an absolute gain of 9.5% in log-average miss rate, compared to the best reported results [32] on the heavily occluded HO pedestrian set of CityPersons test set. Further, on the HO pedestrian set of Caltech dataset, our method achieves an absolute gain of 5.0% in log-average miss rate, compared to the best reported results [13]. Code and models are available at: https://github.com/Leotju/MGAN.
Yanwei Pang, Jin Xie 0005, Muhammad Haris Khan, Rao Muhammad Anwer, Fahad Shahbaz Khan, Ling Shao 0001
ICCV1
2019 Learning Rich Features at High-Speed for Single-Shot Object Detection
abstract
Single-stage object detection methods have received significant attention recently due to their characteristic realtime capabilities and high detection accuracies. Generally, most existing single-stage detectors follow two common practices: they employ a network backbone that is pretrained on ImageNet for the classification task and use a top-down feature pyramid representation for handling scale variations. Contrary to common pre-training strategy, recent works have demonstrated the benefits of training from scratch to reduce the task gap between classification and localization, especially at high overlap thresholds. However, detection models trained from scratch require significantly longer training time compared to their typical finetuning based counterparts. We introduce a single-stage detection framework that combines the advantages of both fine-tuning pretrained models and training from scratch. Our framework constitutes a standard network that uses a pre-trained backbone and a parallel light-weight auxiliary network trained from scratch. Further, we argue that the commonly used top-down pyramid representation only focuses on passing high-level semantics from the top layers to bottom layers. We introduce a bi-directional network that efficiently circulates both low-/mid-level and high-level semantic information in the detection framework. Experiments are performed on MS COCO and UAVDT datasets. Compared to the baseline, our detector achieives an absolute gain of 7.4% and 4.2% in average precision (AP) on MS COCO and UAVDT datasets, respectively using VGG backbone. For a 300×300 input on the MS COCO test set, our detector with ResNet backbone surpasses existing single-stage detection methods for single-scale inference achieving 34.3 AP, while operating at an inference time of 19 milliseconds on a single Titan X GPU. Code is avail- able at https://github.com/vaesl/LRF-Net.
Tiancai Wang, Rao Muhammad Anwer, Hisham Cholakkal, Fahad Shahbaz Khan, Yanwei Pang, Ling Shao 0001
ICCV5
2019 Deep Contextual Attention for Human-Object Interaction Detection
abstract
Human-object interaction detection is an important and relatively new class of visual relationship detection tasks, essential for deeper scene understanding. Most existing approaches decompose the problem into object localization and interaction recognition. Despite showing progress, these approaches only rely on the appearances of humans and objects and overlook the available context information, crucial for capturing subtle interactions between them. We propose a contextual attention framework for human-object interaction detection. Our approach leverages context by learning contextually-aware appearance features for human and object instances. The proposed attention module then adaptively selects relevant instance-centric context information to highlight image regions likely to contain human-object interactions. Experiments are performed on three benchmarks: V-COCO, HICO-DET and HCVRD. Our approach outperforms the state-of-the-art on all datasets. On the V-COCO dataset, our method achieves a relative gain of 4.4% in terms of role mean average precision (mAP role ), compared to the existing best approach.
Tiancai Wang, Rao Muhammad Anwer, Muhammad Haris Khan, Fahad Shahbaz Khan, Yanwei Pang, Ling Shao 0001, Jorma Laaksonen
ICCV5
2019 Learning Compositional Neural Information Fusion for Human Parsing
abstract
This work proposes to combine neural networks with the compositional hierarchy of human bodies for efficient and complete human parsing. We formulate the approach as a neural information fusion framework. Our model assembles the information from three inference processes over the hierarchy: direct inference (directly predicting each part of a human body using image information), bottom-up inference (assembling knowledge from constituent parts), and top-down inference (leveraging context from parent nodes). The bottom-up and top-down inferences explicitly model the compositional and decompositional relations in human bodies, respectively. In addition, the fusion of multi-source information is conditioned on the inputs, i.e., by estimating and considering the confidence of the sources. The whole model is end-to-end differentiable, explicitly modeling information flows and structures. Our approach is extensively evaluated on four popular datasets, outperforming the state-of-the-arts in all cases, with a fast processing speed of 23fps. Our code and results have been released to help ease future research in this direction.
Wenguan Wang, Siyuan Qi, Jianbing Shen, Yanwei Pang, Ling Shao 0001
ICCV5
2019 Complementary Features with Reasonable Receptive Field for Road Scene 3D Object Detection
abstract
Accurate and efficient 3D object detection is of great importance for autonomous driving and robot perception. There are two problems in lidar based 3D object detection networks. (1) Semantic information (e.g., class label of each point) and spatial details are not fully explored for feature extraction. (2) The variance of object sizes represented by point cloud is much smaller than those represented by 2D images. But existing methods do not make use of this property and the receptive field sizes generally mismatch the physical sizes of road scene objects. Based on these two aspects, we propose Complementary Features with Reasonable receptive field networks (CFRNet). CFRNet first exploits Complementary Feature Extractor to learn semantic and positional features, then utilizes a RPN (Region Proposal Networks) with reasonable receptive field to collect correlated context in road scene. Experimental results on KITTI benchmark show the effectiveness of our method. Moreover, our method achieves state-of-the-art performance at a high inference speed.
Yuefeng Wu, Yanwei Pang, Bolin Gao, Jungong Han
ICIP2
2019 Dual-Path in Dual-Path Network for Single Image Dehazing
abstract
Recently, deep learning-based single image dehazing method has been a popular approach to tackle dehazing. However, the existing dehazing approaches are performed directly on the original hazy image, which easily results in image blurring and noise amplifying. To address this issue, the paper proposes a DPDP-Net (Dual-Path in Dual-Path network) framework by employing a hierarchical dual path network. Specifically, the first-level dual-path network consists of a Dehazing Network and a Denoising Network, where the Dehazing Network is responsible for haze removal in the structural layer, and the Denoising Network deals with noise in the textural layer, respectively. And the second-level dual-path network lies in the Dehazing Network, which has an AL-Net (Atmospheric Light Network) and a TM-Net (Transmission Map Network), respectively. Concretely, the AL-Net aims to train the non-uniform atmospheric light, while the TM-Net aims to train the transmission map that reflects the visibility of the image. The final dehazing image is obtained by nonlinearly fusing the output of the Denoising Network and the Dehazing Network. Extensive experiments demonstrate that our proposed DPDP-Net achieves competitive performance against the state-of-the-art methods on both synthetic and real-world images.
Ai-Ping Yang, Zhong Ji, Yanwei Pang, Ling Shao 0001
IJCAI4
2019 ET-Net: A Generic Edge-aTtention Guidance Network for Medical Image Segmentation
Huazhu Fu, Hang Dai, Jianbing Shen, Yanwei Pang, Ling Shao 0001
MICCAI (1)5
2019 Small and Dense Commodity Object Detection with Multi-Scale Receptive Field Attention
abstract
Small and dense commodity object detection is highly valued to the applications in practical scenario. Unlike existing approaches mostly focus on detecting generic objects, this paper studies the problem of specific commodity detection, which is characterized by searching for small and dense instances with similar appearances. Since there is no available dataset or benchmark specialized for exploring this issue, we release a Small and Dense Object Dataset of Milk Tea (SDOD-MT) for promoting the research. Besides, our main solutions for mitigating the detection performance drop caused by the existence of small and dense objects can be concluded as two items. First, for the sake of highlighting the information of positive objects in the feature map, we propose a Multi-Scale Receptive Field (MSRF) attention to generate an attention map to weight the importance on each location of the image feature. Second, for eliminating the negative impact for detection performance brought by the issue of sample imbalance, we present a new loss function named ω-focal loss, which significantly improves the detection accuracy of the categories with few objects. Incorporating these two components into an end-to-end deep architecture, we propose a one-stage detecting framework, dubbed CommodityNet. Extensive experimental results on SDODMT demonstrate that the proposed approach achieves a superior performance on small dense object detection.
Zhong Ji, Qiankun Kong, Haoran Wang 0004, Yanwei Pang
ACM Multimedia4
2019 Special focus on deep learning for computer vision
Yanwei Pang, Xiang Bai, Guofeng Zhang 0001
Sci. China Inf. Sci.1
2019 Class-specific synthesized dictionary model for Zero-Shot Learning
Zhong Ji, Junyue Wang, Yunlong Yu 0001, Yanwei Pang, Jungong Han
Neurocomputing4
2019 Multi-video summarization with query-dependent weighted archetypal analysis
Zhong Ji, Yanwei Pang, Xuelong Li 0001
Neurocomputing3
2019 Query-aware sparse coding for web multi-video summarization
Zhong Ji, Yaru Ma, Yanwei Pang, Xuelong Li 0001
Inf. Sci.3
2019 Visual Haze Removal by a Unified Generative Adversarial Network
abstract
Existence of haze significantly degrades visual quality and hence negatively affects the performance of visual surveillance, video analysis, and human–machine interaction. To remove haze from a visual signal, in this paper, we propose a generative adversarial network for visual haze removal called HRGAN. HRGAN consists of a generator network and a discriminator network. A unified network jointly estimating transmission maps, atmospheric light, and haze-free images (called UNTA) is proposed as the generator network of HRGAN. Instead of being optimized by minimizing the pixel-wise loss, HRGAN is optimized by minimizing a novel loss function consisting of pixel-wise loss, perceptual loss, and adversarial loss produced by a discriminator network. Classical model-based image dehazing algorithms consist of three separate stages: 1) estimating transmission map; 2) estimating atmospheric light; and 3) restoring haze-free image by using an atmospheric scattering model to process the transmission map and atmospheric light. Such a separate scheme is not guaranteed to achieve optimal results. On the contrary, UNTA performs transmission map estimation and atmospheric light estimation simultaneously to obtain joint optimal solutions. The experimental results on both synthetic and real-world image databases demonstrate that HRGAN outperforms the state-of-the-art algorithms in terms of both effectiveness and efficiency.
Yanwei Pang, Jin Xie 0005, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2019 JCS-Net: Joint Classification and Super-Resolution Network for Small-Scale Pedestrian Detection in Surveillance Images
abstract
While convolutional neural network (CNN)-based pedestrian detection methods have proven to be successful in various applications, detecting small-scale pedestrians from surveillance images is still challenging. The major reason is that the small-scale pedestrians lack much detailed information compared to the large-scale pedestrians. To solve this problem, we propose to utilize the relationship between the large-scale pedestrians and the corresponding small-scale pedestrians to help recover the detailed information of the small-scale pedestrians, thus improving the performance of detecting small-scale pedestrians. Specifically, a unified network (called JCS-Net) is proposed for small-scale pedestrian detection, which integrates the classification task and the super-resolution task in a unified framework. As a result, the super-resolution and classification are fully engaged, and the super-resolution sub-network can recover some useful detailed information for the subsequent classification. Based on HOG+LUV and JCS-Net, multi-layer channel features (MCF) are constructed to train the detector. The experimental results on the Caltech pedestrian dataset and the KITTI benchmark demonstrate the effectiveness of the proposed method. To further enhance the detection, multi-scale MCF based on JCS-Net for pedestrian detection is also proposed, which achieves the state-of-the-art performance.
Yanwei Pang, Jiale Cao, Jian Wang 0087, Jungong Han
IEEE Trans. Inf. Forensics Secur.1
2019 Simultaneously Learning Neighborship and Projection Matrix for Supervised Dimensionality Reduction
abstract
Explicitly or implicitly, most dimensionality reduction methods need to determine which samples are neighbors and the similarities between the neighbors in the original high-dimensional space. The projection matrix is then learnt on the assumption that the neighborhood information, e.g., the similarities, are known and fixed prior to learning. However, it is difficult to precisely measure the intrinsic similarities of samples in high-dimensional space because of the curse of dimensionality. Consequently, the neighbors selected according to such similarities and the projection matrix obtained according to such similarities and the corresponding neighbors might not be optimal in the sense of classification and generalization. To overcome this drawback, in this paper, we propose to let the similarities and neighbors be variables and model these in a low-dimensional space. Both the optimal similarity and projection matrix are obtained by minimizing a unified objective function. Nonnegative and sum-to-one constraints on the similarity are adopted. Instead of empirically setting the regularization parameter, we treat it as a variable to be optimized. It is interesting that the optimal regularization parameter is adaptive to the neighbors in a low-dimensional space and has an intuitive meaning. Experimental results on the YALE B, COIL-100, and MNIST data sets demonstrate the effectiveness of the proposed method.
Yanwei Pang, Feiping Nie 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 Stacked Semantics-Guided Attention Model for Fine-Grained Zero-Shot Learning
abstract
Zero-Shot Learning (ZSL) is generally achieved via aligning the semantic relationships between the visual features and the corresponding class semantic descriptions. However, using the global features to represent fine-grained images may lead to sub-optimal results since they neglect the discriminative differences of local regions. Besides, different regions contain distinct discriminative information. The important regions should contribute more to the prediction. To this end, we propose a novel stacked semantics-guided attention (S2GA) model to obtain semantic relevant features by using individual class semantic features to progressively guide the visual features to generate an attention map for weighting the importance of different local regions. Feeding both the integrated visual features and the class semantic features into a multi-class classification architecture, the proposed framework can be trained end-to-end. Extensive experimental results on CUB and NABird datasets show that the proposed approach has a consistent improvement on both fine-grained zero-shot classification and retrieval tasks.
Yunlong Yu 0001, Zhong Ji, Yanwei Fu 0001, Jichang Guo, Yanwei Pang, Zhongfei Zhang
NeurIPS5
2018 GlanceNets - efficient convolutional neural networks with adaptive hard example mining
Hanqing Sun 0001, Yanwei Pang
Sci. China Inf. Sci.2
2018 Randomly translational activation inspired by the input distributions of ReLU
Jiale Cao, Yanwei Pang, Xuelong Li 0001, Jingkun Liang
Neurocomputing2
2018 Semantic softmax loss for zero-shot learning
Zhong Ji, Yunlong Yu 0001, Jichang Guo, Yanwei Pang
Neurocomputing5
2018 Deep neural networks with Elastic Rectified Linear Units for object recognition
Xiaoheng Jiang, Yanwei Pang, Xuelong Li 0001, Yinghong Xie
Neurocomputing2
2018 Patient-specific ECG classification by deeper CNN from generic to dedicated
Yazhao Li, Yanwei Pang, Jian Wang 0087, Xuelong Li 0001
Neurocomputing2
2018 Learning intensity and detail mapping parameters for dehazing
Xuhang Lian, Yanwei Pang, Ai-Ping Yang
Multim. Tools Appl.2
2018 Fusion-Attention Network for person search with free-form natural language
Zhong Ji, Shengjia Li, Yanwei Pang
Pattern Recognit. Lett.3
2018 LightenNet: A Convolutional Neural Network for weakly illuminated image enhancement
Chongyi Li, Jichang Guo, Fatih Porikli, Yanwei Pang
Pattern Recognit. Lett.4
2018 Hypergraph dominant set based multi-video summarization
Zhong Ji, Yanwei Pang, Xuelong Li 0001
Signal Process.3
2018 Incremental Learning With Saliency Map for Moving Object Detection
abstract
Moving object detection is a key to intelligent video analysis. On the one hand, what moves are not only interesting objects but also noise and cluttered background. On the other hand, moving objects without rich texture are prone to not be detected. Therefore, there are undesirable false alarms and missed alarms in the results of many algorithms of moving object detection. To reduce the false alarms and missed alarms, in this paper we propose to incorporate a saliency map into an incremental subspace analysis framework in which the saliency map makes the estimated background have less of a chance than the foreground (i.e., moving objects) to contain salient objects. The proposed objective function systematically takes into account the properties of sparsity, low rank, connectivity, and saliency. An alternative minimization algorithm is proposed to seek the optimal solutions. The experimental results on both the Perception Test Images Sequences data set and Wallflower data set demonstrate that the proposed method is effective in reducing false alarms and missed alarms.
Yanwei Pang, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2018 Cascaded Subpatch Networks for Effective CNNs
abstract
Conventional convolutional neural networks use either a linear or a nonlinear filter to extract features from an image patch (region) of spatial size (typically, is small and is equal to , e.g., is 5 or 7). Generally, the size of the filter is equal to the size of the input patch. We argue that the representational ability of equal-size strategy is not strong enough. To overcome the drawback, we propose to use subpatch filter whose spatial size is smaller than . The proposed subpatch filter consists of two subsequent filters. The first one is a linear filter of spatial size and is aimed at extracting features from spatial domain. The second one is of spatial size and is used for strengthening the connection between different input feature channels and for reducing the number of parameters. The subpatch filter convolves with the input patch and the resulting network is called a subpatch network. Taking the output of one subpatch network as input, we further repeat constructing subpatch networks until the output contains only one neuron in spatial domain. These subpatch networks form a new network called the cascaded subpatch network (CSNet). The feature layer generated by CSNet is called the csconv layer. For the whole input image, we construct a deep neural network by stacking a sequence of csconv layers. Experimental results on five benchmark data sets demonstrate the effectiveness and compactness of the proposed CSNet. For example, our CSNet reaches a test error of 5.68% on the CIFAR10 data set without model averaging. To the best of our knowledge, this is the best result ever obtained on the CIFAR10 data set.
Xiaoheng Jiang, Yanwei Pang, Manli Sun, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.2
2018 Convolution in Convolution for Network in Network
abstract
Network in network (NiN) is an effective instance and an important extension of deep convolutional neural network consisting of alternating convolutional layers and pooling layers. Instead of using a linear filter for convolution, NiN utilizes shallow multilayer perceptron (MLP), a nonlinear function, to replace the linear filter. Because of the powerfulness of MLP and convolutions in spatial domain, NiN has stronger ability of feature representation and hence results in better recognition performance. However, MLP itself consists of fully connected layers that give rise to a large number of parameters. In this paper, we propose to replace dense shallow MLP with sparse shallow MLP. One or more layers of the sparse shallow MLP are sparely connected in the channel dimension or channel-spatial domain. The proposed method is implemented by applying unshared convolution across the channel dimension and applying shared convolution across the spatial dimension in some computational layers. The proposed method is called convolution in convolution (CiC). The experimental results on the CIFAR10 data set, augmented CIFAR10 data set, and CIFAR100 data set demonstrate the effectiveness of the proposed CiC method.
Yanwei Pang, Manli Sun, Xiaoheng Jiang, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 Transductive Zero-Shot Learning With Adaptive Structural Embedding
abstract
Zero-shot learning (ZSL) endows the computer vision system with the inferential capability to recognize new categories that have never seen before. Two fundamental challenges in it are visual-semantic embedding and domain adaptation in cross-modality learning and unseen class prediction steps, respectively. This paper presents two corresponding methods named Adaptive STructural Embedding (ASTE) and Self-PAced Selective Strategy (SPASS) for both challenges. Specifically, ASTE formulates the visual-semantic interactions in a latent structural support vector machine framework by adaptively adjusting the slack variables to embody different reliablenesses among training instances. To alleviate the domain shift problem in ZSL, SPASS borrows the idea from self-paced learning by iteratively selecting the unseen instances from reliable to less reliable to gradually adapt the knowledge from the seen domain to the unseen domain. Consequently, by combining SPASS and ASTE, we present a self-paced Transductive ASTE (TASTE) method to progressively reinforce the classification capacity. Extensive experiments on three benchmark data sets (i.e., AwA, CUB, and aPY) demonstrate the superiorities of ASTE and TASTE. Furthermore, we also propose a fast training (FT) strategy to improve the efficiency of most existing ZSL methods. The FT strategy is surprisingly simple and general enough, which speeds up the training time of most existing ZSL methods by 4~300 times while holding the previous performance.
Yunlong Yu 0001, Zhong Ji, Jichang Guo, Yanwei Pang
IEEE Trans. Neural Networks Learn. Syst.4
2017 ECG Waveform Extraction from Paper Records
Jian Wang 0087, Yanwei Pang
ICIG (2)2
2017 Deep pedestrian attribute recognition based on LSTM
abstract
Automatically recognizing attributes such as gender, age, footwear and clothing style from pedestrian images at far distance is an important task in surveillance scenarios. However, the appearance diversity and ambiguity in these images make it a challenging task. This paper presents an end-to-end Neural Pedestrian Attribute Recognition (Neural PAR) model to address these challenges. Rather than taking it as a recognition problem like previous methods, Neural PAR formulates it as an end-to-end image to attribute description problem. To this end, the training images and their corresponding attributes are used as inputs. Specifically, the attributes are concatenated into different attribute descriptions to well contextualize the potential relationships among them. Then, a neural network model is trained based on CNN and LSTM to learn the complex relations between visual features and their corresponding attributes. Extensive experiments show that the proposed Neural PAR significantly outperforms the state-of-the-art methods on the benchmark PETA dataset.
Zhong Ji, Weixiong Zheng, Yanwei Pang
ICIP3
2017 Learning Pooling for Convolutional Neural Network
Manli Sun, Zhanjie Song, Xiaoheng Jiang, Yanwei Pang
Neurocomputing5
2017 Zero-shot learning with regularized cross-modality ranking
Yunlong Yu 0001, Zhong Ji, Jichang Guo, Yanwei Pang
Neurocomputing4
2017 Manifold regularized cross-modal embedding for zero-shot learning
Zhong Ji, Yunlong Yu 0001, Yanwei Pang, Jichang Guo, Zhongfei Zhang
Inf. Sci.3
2017 Zero-shot learning with Multi-Battery Factor Analysis
Zhong Ji, Yunlong Yu 0001, Yanwei Pang, Zhongfei Zhang
Signal Process.3
2017 Learning Sampling Distributions for Efficient Object Detection
abstract
Object detection is an important task in computer vision and machine intelligence systems. Multistage particle windows (MPW), proposed by Gualdi et al., is an algorithm of fast and accurate object detection. By sampling particle windows (PWs) from a proposal distribution (PD), MPW avoids exhaustively scanning the image. Despite its success, it is unknown how to determine the number of stages and the number of PWs in each stage. Moreover, it has to generate too many PWs in the initialization step and it unnecessarily regenerates too many PWs around object-like regions. In this paper, we attempt to solve the problems of MPW. An important fact we used is that there is a large probability for a randomly generated PW not to contain the object because the object is a sparse event relative to the huge number of candidate windows. Therefore, we design a PD so as to efficiently reject the huge number of nonobject windows. Specifically, we propose the concepts of rejection, acceptance, and ambiguity windows and regions. Then, the concepts are used to form and update a dented uniform distribution and a dented Gaussian distribution. This contrasts to MPW which utilizes only on region of support. The PD of MPW is acceptance-oriented whereas the PD of our method (called iPW) is rejection-oriented. Experimental results on human and face detection demonstrate the efficiency and the effectiveness of the iPW algorithm. The source code is publicly accessible.
Yanwei Pang, Jiale Cao, Xuelong Li 0001
IEEE Trans. Cybern.1
2017 Cascade Learning by Optimally Partitioning
abstract
Cascaded AdaBoost classifier is a well-known efficient object detection algorithm. The cascade structure has many parameters to be determined. Most of existing cascade learning algorithms are designed by assigning detection rate and false positive rate to each stage either dynamically or statically. Their objective functions are not directly related to minimum computation cost. These algorithms are not guaranteed to have optimal solution in the sense of minimizing computation cost. On the assumption that a strong classifier is given, in this paper, we propose an optimal cascade learning algorithm (iCascade) which iteratively partitions the strong classifiers into two parts until predefined number of stages are generated. iCascade searches the optimal partition point of each stage by directly minimizing the computation cost of the cascade. Theorems are provided to guarantee the existence of the unique optimal solution. Theorems are also given for the proposed efficient algorithm of searching optimal parameters . Once a new stage is added, the parameter for each stage decreases gradually as iteration proceeds, which we call decreasing phenomenon. Moreover, with the goal of minimizing computation cost, we develop an effective algorithm for setting the optimal threshold of each stage. In addition, we prove in theory why more new weak classifiers in the current stage are required compared to that of the previous stage. Experimental results on face detection and pedestrian detection demonstrate the effectiveness and efficiency of the proposed algorithm.
Yanwei Pang, Jiale Cao, Xuelong Li 0001
IEEE Trans. Cybern.1
2017 Learning Multilayer Channel Features for Pedestrian Detection
abstract
Pedestrian detection based on the combination of convolutional neural network (CNN) and traditional handcrafted features (i.e., HOG+LUV) has achieved great success. In general, HOG+LUV are used to generate the candidate proposals and then CNN classifies these proposals. Despite its success, there is still room for improvement. For example, CNN classifies these proposals by the fully connected layer features, while proposal scores and the features in the inner-layers of CNN are ignored. In this paper, we propose a unifying framework called multi-layer channel features (MCF) to overcome the drawback. It first integrates HOG+LUV with each layer of CNN into a multi-layer image channels. Based on the multi-layer image channels, a multi-stage cascade AdaBoost is then learned. The weak classifiers in each stage of the multi-stage cascade are learned from the image channels of corresponding layer. Experiments on Caltech data set, INRIA data set, ETH data set, TUD-Brussels data set, and KITTI data set are conducted. With more abundant features, an MCF achieves the state of the art on Caltech pedestrian data set (i.e., 10.40% miss rate). Using new and accurate annotations, an MCF achieves 7.98% miss rate. As many non-pedestrian detection windows can be quickly rejected by the first few stages, it accelerates detection speed by 1.43 times. By eliminating the highly overlapped detection windows with lower scores after the first stage, it is 4.07 times faster than negligible performance loss.
Jiale Cao, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Image Process.2
2016 Pedestrian Detection Inspired by Appearance Constancy and Shape Symmetry
abstract
The discrimination and simplicity of features are very important for effective and efficient pedestrian detection. However, most state-of-the-art methods are unable to achieve good tradeoff between accuracy and efficiency. Inspired by some simple inherent attributes of pedestrians (i.e., appearance constancy and shape symmetry), we propose two new types of non-neighboring features (NNF): side-inner difference features (SIDF) and symmetrical similarity features (SSF). SIDF can characterize the difference between the background and pedestrian and the difference between the pedestrian contour and its inner part. SSF can capture the symmetrical similarity of pedestrian shape. However, it's difficult for neighboring features to have such above characterization abilities. Finally, we propose to combine both non-neighboring and neighboring features for pedestrian detection. It's found that nonneighboring features can further decrease the average miss rate by 4.44%. Experimental results on INRIA and Caltech pedestrian datasets demonstrate the effectiveness and efficiency of the proposed method. Compared to the state-of the-art methods without using CNN, our method achieves the best detection performance on Caltech, outperforming the second best method (i.e., Checkerboards) by 1.63%.
Jiale Cao, Yanwei Pang, Xuelong Li 0001
CVPR2
2016 Single underwater image restoration by blue-green channels dehazing and red channel correction
abstract
Restoring underwater image from a single image is know to be ill-posed, and some assumptions made in previous methods are not suitable for many situations. In this paper, we propose a method based on blue-green channels dehazing and red channel correction for underwater image restoration. Firstly, blue-green channels are recovered via dehazing algorithm based on an extension and modification of Dark Channel Prior algorithm. Then, red channel is corrected following the Gray-World assumption theory. Finally, in order to resolve the problem which some recovered image regions may look too dim or too bright, an adaptive exposure map is built. Qualitative analysis demonstrates that our method significantly improves visibility and contrast, and reduces the effects of light absorption and scattering. For quantitative analysis, our results obtain best values in terms of entropy, local feature points and average gradient, which outperform three existing physical model available methods.
Chongyi Li, Jichang Quo, Yanwei Pang, Shanji Chen, Jian Wang 0087
ICASSP3
2016 Underwater image restoration based on minimum information loss principle and optical properties of underwater imaging
abstract
Restoring underwater image from a single image is known to be an ill-posed problem. Some assumptions made in previous methods are not suitable in many situations. In this paper, an effective method is proposed to restore underwater images. Using the quad-tree subdivision and graph-based segmentation, the global background light can be robustly estimated. The medium transmission map is estimated based on minimum information loss principle and optical properties of underwater imaging. Qualitative experiments show that our results are characterized by relatively genuine color, natural appearance, and improved contrast and visibility. Quantitative comparisons demonstrate that the proposed method can achieve better quality of underwater images when compared with several other methods.
Chongyi Li, Jichang Guo, Shanji Chen, Yibin Tang, Yanwei Pang, Jian Wang 0087
ICIP5
2016 Enhancement for Dust-Sand Storm Images
Jian Wang 0087, Yanwei Pang, Changshu Liu
MMM (1)2
2016 Speed up deep neural network based pedestrian detection by sharing features across multi-scale models
abstract
Deep neural networks (DNNs) have now demonstrated state-of-the-art detection performance on pedestrian datasets. However, because of their high computational complexity, detection efficiency is still a frustrating problem even with the help of Graphics Processing Units (GPUs). To improve detection efficiency, this paper proposes to share features across a group of DNNs that correspond to pedestrian models of different sizes. By sharing features, the computational burden for extracting features from an image pyramid can be significantly reduced. Simultaneously, we can detect pedestrians of several different scales on one single layer of an image pyramid. Furthermore, the improvement of detection efficiency is achieved with negligible loss of detection accuracy. Experimental results demonstrate the robustness and efficiency of the proposed algorithm.
Xiaoheng Jiang, Yanwei Pang, Xuelong Li 0001
Neurocomputing2
2016 Special issue on dimensionality reduction for visual big data
Yanwei Pang, Ling Shao 0001
Neurocomputing1
2016 Relevance and irrelevance graph based marginal Fisher analysis for image search reranking
Zhong Ji, Yanwei Pang, Yuan Yuan 0001
Signal Process.2
2016 Classifying Discriminative Features for Blur Detection
abstract
Blur detection in a single image is challenging especially when the blur is spatially-varying. Developing discriminative blur features is an open problem. In this paper, we propose a new kernel-specific feature vector consisting of the information of a blur kernel and the information of an image patch. Specifically, the kernel specific-feature is composed of the multiplication of the variance of filtered kernel and the variance of filtered patch gradients. The feature origins from a blur-classification theorem and its discrimination can also be intuitively explained. To make the kernel-specific features useful for real applications, we build a pool of kernels consisting of motion-blur kernels, defocus-blur (out-of-focus) kernels, and their combinations. By extracting such features followed by the classifiers, the proposed algorithm outperforms the state-of-the-art blur detection method. Experimental results on public databases demonstrate the effectiveness of the proposed method.
Yanwei Pang, Hailong Zhu, Xuelong Li 0001
IEEE Trans. Cybern.1
2016 Pedestrian Detection Inspired by Appearance Constancy and Shape Symmetry
abstract
Most state-of-the-art methods in pedestrian detection are unable to achieve a good trade-off between accuracy and efficiency. For example, ACF has a fast speed but a relatively low detection rate, while checkerboards have a high detection rate but a slow speed. Inspired by some simple inherent attributes of pedestrians (i.e., appearance constancy and shape symmetry), we propose two new types of non-neighboring features: side-inner difference features (SIDF) and symmetrical similarity features (SSFs). SIDF can characterize the difference between the background and pedestrian and the difference between the pedestrian contour and its inner part. SSF can capture the symmetrical similarity of pedestrian shape. However, it is difficult for neighboring features to have such above characterization abilities. Finally, we propose to combine both non-neighboring features and neighboring features for pedestrian detection. It is found that non-neighboring features can further decrease the log-average miss rate by 4.44%. The relationship between our proposed method and some state-of-the-art methods is also given. Experimental results on INRIA, Caltech, and KITTI data sets demonstrate the effectiveness and efficiency of the proposed method. Compared with the state-of-the-art methods without using CNN, our method achieves the best detection performance on Caltech, outperforming the second best method (i.e., checkerboards) by 2.27%. Using the new annotations of Caltech, it can achieve 11.87% miss rate, which outperforms other methods.
Jiale Cao, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Image Process.2
2016 Underwater Image Enhancement by Dehazing With Minimum Information Loss and Histogram Distribution Prior
abstract
Images captured under water are usually degraded due to the effects of absorption and scattering. Degraded underwater images show some limitations when they are used for display and analysis. For example, underwater images with low contrast and color cast decrease the accuracy rate of underwater object detection and marine biology recognition. To overcome those limitations, a systematic underwater image enhancement method, which includes an underwater image dehazing algorithm and a contrast enhancement algorithm, is proposed. Built on a minimum information loss principle, an effective underwater image dehazing algorithm is proposed to restore the visibility, color, and natural appearance of underwater images. A simple yet effective contrast enhancement algorithm is proposed based on a kind of histogram distribution prior, which increases the contrast and brightness of underwater images. The proposed method can yield two versions of enhanced output. One version with relatively genuine color and natural appearance is suitable for display. The other version with high contrast and brightness can be used for extracting more valuable information and unveiling more details. Simulation experiment, qualitative and quantitative comparisons, as well as color accuracy and application tests are conducted to evaluate the performance of the proposed method. Extensive experiments demonstrate that the proposed method achieves better visual quality, more valuable information, and more accurate color restoration than several state-of-the-art methods, even for underwater images taken under several challenging scenes.
Chongyi Li, Jichang Guo, Runmin Cong, Yanwei Pang, Bo Wang 0070
IEEE Trans. Image Process.4
2015 Interactive Head 3D Reconstruction Based Combine of Key Points and Voxel
Yanwei Pang, Changshu Liu
ICIG (2)1
2015 Semi-supervised LPP algorithms for learning-to-rank-based visual search reranking
Zhong Ji, Yanwei Pang, Huanfen Zhang
Inf. Sci.2
2015 Flexible sliding windows with adaptive pixel strides
Xiaoheng Jiang, Yanwei Pang, Xuelong Li 0001
Signal Process.2
2015 Efficient object detection by prediction in 3D space
Yanwei Pang, Xiaoheng Jiang, Xuelong Li 0001
Signal Process.1
2015 Truncation Error Analysis on Reconstruction of Signal From Unsymmetrical Local Average Sampling
abstract
The classical Shannon sampling theorem is suitable for reconstructing a band-limited signal from its sampled values taken at regular instances with equal step by using the well-known sinc function. However, due to the inertia of the measurement apparatus, it is impossible to measure the value of a signal precisely at such discrete time. In practice, only unsymmetrically local averages of signal near the regular instances can be measured and used as the inputs for a signal reconstruction method. In addition, when implemented in hardware, the traditional sinc function cannot be directly used for signal reconstruction. We propose using the Taylor expansion of sinc function to reconstruct signal sampled from unsymmetrically local averages and give the upper bound of the reconstruction error (i.e., truncation error). The convergency of the reconstruction method is also presented.
Yanwei Pang, Zhanjie Song, Xuelong Li 0001
IEEE Trans. Cybern.1
2015 Relevance Preserving Projection and Ranking for Web Image Search Reranking
abstract
An image search reranking (ISR) technique aims at refining text-based search results by mining images' visual content. Feature extraction and ranking function design are two key steps in ISR. Inspired by the idea of hypersphere in one-class classification, this paper proposes a feature extraction algorithm named hypersphere-based relevance preserving projection (HRPP) and a ranking function called hypersphere-based rank (H-Rank). Specifically, an HRPP is a spectral embedding algorithm to transform an original high-dimensional feature space into an intrinsically low-dimensional hypersphere space by preserving the manifold structure and a relevance relationship among the images. An H-Rank is a simple but effective ranking algorithm to sort the images by their distances to the hypersphere center. Moreover, to capture the user's intent with minimum human interaction, a reversed k-nearest neighbor (KNN) algorithm is proposed, which harvests enough pseudorelevant images by requiring that the user gives only one click on the initially searched images. The HRPP method with reversed KNN is named one-click-based HRPP (OC-HRPP). Finally, an OC-HRPP algorithm and the H-Rank algorithm form a new ISR method, H-reranking. Extensive experimental results on three large real-world data sets show that the proposed algorithms are effective. Moreover, the fact that only one relevant image is required to be labeled makes it has a strong practical significance.
Zhong Ji, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Image Process.2
2014 Weighted Deformable Part Model for Robust Human Detection
Tianshuo Li, Yanwei Pang, Changshu Liu
ICIC (1)2
2014 Frequency Domain Directional Filtering Based Rain Streaks Removal from a Single Color Image
Changbo Liu, Yanwei Pang, Jian Wang 0087, Ai-Ping Yang
ICIC (1)2
2014 Distributed Object Detection With Linear SVMs
abstract
In vision and learning, low computational complexity and high generalization are two important goals for video object detection. Low computational complexity here means not only fast speed but also less energy consumption. The sliding window object detection method with linear support vector machines (SVMs) is a general object detection framework. The computational cost is herein mainly paid in complex feature extraction and innerproduct-based classification. This paper first develops a distributed object detection framework (DOD) by making the best use of spatial-temporal correlation, where the process of feature extraction and classification is distributed in the current frame and several previous frames. In each framework, only subfeature vectors are extracted and the response of partial linear classifier (i.e., subdecision value) is computed. To reduce the dimension of traditional block-based histograms of oriented gradients (BHOG) feature vector, this paper proposes a cell-based HOG (CHOG) algorithm, where the features in one cell are not shared with overlapping blocks. Using CHOG as feature descriptor, we develop CHOG-DOD as an instance of DOD framework. Experimental results on detection of hand, face, and pedestrian in video show the superiority of the proposed method.
Yanwei Pang, Yuan Yuan 0001, Kongqiao Wang
IEEE Trans. Cybern.1
2014 Learning Regularized LDA by Clustering
abstract
As a supervised dimensionality reduction technique, linear discriminant analysis has a serious overfitting problem when the number of training samples per class is small. The main reason is that the between- and within-class scatter matrices computed from the limited number of training samples deviate greatly from the underlying ones. To overcome the problem without increasing the number of training samples, we propose making use of the structure of the given training data to regularize the between- and within-class scatter matrices by between- and within-cluster scatter matrices, respectively, and simultaneously. The within- and between-cluster matrices are computed from unsupervised clustered data. The within-cluster scatter matrix contributes to encoding the possible variations in intraclasses and the between-cluster scatter matrix is useful for separating extra classes. The contributions are inversely proportional to the number of training samples per class. The advantages of the proposed method become more remarkable as the number of training samples per class decreases. Experimental results on the AR and Feret face databases demonstrate the effectiveness of the proposed method.
Yanwei Pang, Yuan Yuan 0001
IEEE Trans. Neural Networks Learn. Syst.1
2013 Efficient 2D-to-3D Correspondence Filtering for Scalable 3D Object Recognition
abstract
3D model-based object recognition has been a noticeable research trend in recent years. Common methods find 2D-to-3D correspondences and make recognition decisions by pose estimation, whose efficiency usually suffers from noisy correspondences caused by the increasing number of target objects. To overcome this scalability bottleneck, we propose an efficient 2D-to-3D correspondence filtering approach, which combines a light-weight neighborhood-based step with a finer-grained pairwise step to remove spurious correspondences based on 2D/3D geometric cues. On a dataset of 300 3D objects, our solution achieves ~10 times speed improvement over the baseline, with a comparable recognition accuracy. A parallel implementation on a quad-core CPU can run at ~3fps for 1280×720 images.
Rui Cai 0002, Zhiwei Li 0006, Lei Zhang 0001, Yanwei Pang, Feng Wu 0001, Yong Rui
CVPR5
2013 Image Search Reranking with Semi-supervised LPP and Ranking SVM
Zhong Ji, Yanru Yu, Yuting Su 0001, Yanwei Pang
MMM (1)4
2013 Robust probabilistic tensor analysis for time-variant collaborative filtering
Zhao Ma, Yanwei Pang, Yuan Yuan 0001
Neurocomputing3
2013 Special issue on image feature detection and description
Yanwei Pang, Xianbin Cao 0001, Lei Zhang 0001, Amir Hussein
Neurocomputing1
2013 Sparse representations for image and video analysis
Jinhui Tang 0001, Shuicheng Yan, John Wright 0001, Qi Tian 0001, Yanwei Pang, Edwige E. Pissaloux
J. Vis. Commun. Image Represent.5
2013 Rank canonical correlation analysis and its application in visual search reranking
Zhong Ji, Peiguang Jing, Yuting Su 0001, Yanwei Pang
Signal Process.4
2013 Energy-saving object detection by efficiently rejecting a set of neighboring sub-images
Yanwei Pang, Yuan Yuan 0001, Kongqiao Wang
Signal Process.2
2013 Ranking Graph Embedding for Learning to Rerank
abstract
Dimensionality reduction is a key step to improving the generalization ability of reranking in image search. However, existing dimensionality reduction methods are typically designed for classification, clustering, and visualization, rather than for the task of learning to rank. Without using of ranking information such as relevance degree labels, direct utilization of conventional dimensionality reduction methods in ranking tasks generally cannot achieve the best performance. In this paper, we show that introducing ranking information into dimensionality reduction significantly increases the performance of image search reranking. The proposed method transforms graph embedding, a general framework of dimensionality reduction, into ranking graph embedding (RANGE) by modeling the global structure and the local relationships in and between different relevance degree sets, respectively. The proposed method also defines three types of edge weight assignment between two nodes: binary, reconstruction, and global. In addition, a novel principal components analysis based similarity calculation method is presented in the stage of global graph construction. Extensive experimental results on the MSRA-MM database demonstrate the effectiveness and superiority of the proposed RANGE method and the image search reranking framework.
Yanwei Pang, Zhong Ji, Peiguang Jing, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2012 3D visual phrases for landmark recognition
abstract
In this paper, we study the problem of landmark recognition and propose to leverage 3D visual phrases to improve the performance. A 3D visual phrase is a triangular facet on the surface of a reconstructed 3D landmark model. In contrast to existing 2D visual phrases which are mainly based on co-occurrence statistics in 2D image planes, such 3D visual phrases explicitly characterize the spatial structure of a 3D object (landmark), and are highly robust to projective transformations due to viewpoint changes. We present an effective solution to discover, describe, and detect 3D visual phrases. The experiments on 10 landmarks have achieved promising results, which demonstrate that our approach provides a good balance between precision and recall of landmark recognition while reducing the dependence on post-verification to reject false positives.
Rui Cai 0002, Zhiwei Li 0006, Lei Zhang 0001, Yanwei Pang, Feng Wu 0001
CVPR5
2012 Incremental threshold learning for classifier selection
Yanwei Pang, Junping Deng, Yuan Yuan 0001
Neurocomputing1
2012 Fully affine invariant SURF for image matching
Yanwei Pang, Yuan Yuan 0001
Neurocomputing1
2012 Scale invariant image matching using triplewise constraint and weighted voting
Yanwei Pang, Mianyou Shang, Yuan Yuan 0001
Neurocomputing1
2012 Learning optimal spatial filters by discriminant analysis for brain-computer-interface
Yanwei Pang, Yuan Yuan 0001, Kongqiao Wang
Neurocomputing1
2012 Efficient image matching using weighted voting
Yuan Yuan 0001, Yanwei Pang, Kongqiao Wang, Mianyou Shang
Pattern Recognit. Lett.2
2012 An Improved Nyquist-Shannon Irregular Sampling Theorem From Local Averages
abstract
The Nyquist–Shannon sampling theorem is on the reconstruction of a band-limited signal from its uniformly sampled samples. The higher the signal bandwidth gets, the more challenging the uniform sampling may become. To deal with this problem, signal reconstruction from local averages has been studied in the literature. In this paper, we obtain an improved Nyquist–Shannon sampling theorem from general local averages. In practice, the measurement apparatus gives a weighted average over an asymmetrical interval. As a special case, for local averages from symmetrical interval, we show that the sampling rate is much lower than that of a result by Gröchenig. Moreover, we obtain two exact dual frames from local averages, one of which improves a result by Sun and Zhou. At the end of this paper, as an example application of local average sampling, we consider a reconstruction algorithm: the piecewise linear approximations.
Zhanjie Song, Yanwei Pang, Chunping Hou, Xuelong Li 0001
IEEE Trans. Inf. Theory3
2012 Robust CoHOG Feature Extraction in Human-Centered Image/Video Management System
abstract
Many human-centered image and video management systems depend on robust human detection. To extract robust features for human detection, this paper investigates the following shortcomings of co-occurrence histograms of oriented gradients (CoHOGs) which significantly limit its advantages: 1) The magnitudes of the gradients are discarded, and only the orientations are used; 2) the gradients are not smoothed, and thus, aliasing effect exists; and 3) the dimensionality of the CoHOG feature vector is very large (e.g., 200,000). To deal with these problems, in this paper, we propose a framework that performs the following: 1) utilizes a novel gradient decomposition and combination strategy to make full use of the information of gradients; (2) adopts a two-stage gradient smoothing scheme to perform efficient gradient interpolation; and (3) employs incremental principal component analysis to reduce the large dimensionality of the CoHOG features. Experimental results on the two different human databases demonstrate the effectiveness of the proposed method.
Yanwei Pang, Yuan Yuan 0001, Kongqiao Wang
IEEE Trans. Syst. Man Cybern. Part B1
2011 Diversifying the Image Relevance Reranking with Absorbing Random Walks
abstract
Image visual reranking holds the simple search mechanism preferred by typical users, and exploits the visual information and image analysis methods in another way. Therefore, it integrates characteristics of real-time and accuracy, and has great importance to establish practical image search system. A novel reranking method named DIRRA is proposed in this paper, in which absorbing random walks is utilized to enhance the diversity as well as relevance of the initial search results. Four kinds of image visual features are extracted firstly, and then a graph is built, where nodes are images and edges are the similarities between images. Next, the first item is decided by teleporting random walks on the graph, and the other items are decided by absorbing random walks on the graph at last. Experiments are performed on a web image database including 10 queries, which prove the reranking results are both diverse and relevant, and practical to improve user's satisfaction in web search.
Zhong Ji, Yuting Su 0001, Yanwei Pang, Xiaojie Qu
ICIG3
2011 Robust Sparse Tensor Decomposition by Probabilistic Latent Semantic Analysis
abstract
Movie recommendation system is becoming more and more popular in recent years. As a result, it is becoming increasingly important to develop machine learning algorithm on partially-observed matrix to predict users' preferences on missing data. Motivated by the user ratings prediction problem, we propose a novel robust tensor probabilistic latent semantic analysis (RT-pLSA) algorithm that not only takes time variable into account, but also uses the periodic property of data in time attribute. Different from the previous algorithms of predicting missing values on two-dimensional sparse matrix, we formulize the prediction problem as a probabilistic tensor factorization problem with periodicity constraint on time coordinate. Furthermore, we apply the Tsallis divergence error measure in the context of RT-pLSA tensor decomposition that is able to robustly predict the latent variable in the presence of noise. Our experimental results on two benchmark movie rating dataset: Netflix and Movie lens, show a good predictive accuracy of the model.
Yanwei Pang, Zhao Ma, Yuan Yuan 0001
ICIG1
2011 Integrating kAS and SIFT-like Descriptor for Image Description
abstract
Shape-descriptor (e.g. Adjacent Contour Segments, i.e. kAS) and key point-descriptor (e.g. Scale Invariant Feature Transform, i.e. SIFT) are widely used for computer vision. However, few works principally integrate shape-descriptor and key point-descriptor to describe the content of an image. On one hand, in some cases the degree of locality of keying-descriptor is too high to capture semantic characteristics of an object. On the other hand, though the shape has higher semantic level than key point, it contains no texture information because only the information of contour/edge is used. To make full use of the information of both shape and key point for generate robust and distinctive features, in this paper we propose an algorithm to integrate shape and key point descriptor. Specifically, we employ kAS to extract useful shape information. Then key points of a kAS shape are defined at which we propose to extract SIFT-like features. Experimental results on image matching demonstrate the effectiveness of the proposed algorithm.
Mianyou Shang, Yanwei Pang, Yuan Yuan 0001
ICIG3
2011 Multimodal learning for multi-label image classification
abstract
We tackle the challenge of web image classification using additional tags information. Unlike traditional methods that only use the combination of several low-level features, we try to use semantic concepts to represent images and corresponding tags. At first, we extract the latent topic information by probabilistic latent semantic analysis (pLSA) algorithm, and then use multi-label multiple kernel learning to combine visual and textual features to make a better image classification. In our experiments on PASCAL VOC'07 set and MIR Flickr set, we demonstrate the benefit of using multimodal feature to improve image classification. Specifically, we discover that on the issue of image classification, utilizing latent semantic feature to represent images and associated tags can obtain better classification results than other ways that integrating several low-level features.
Yanwei Pang, Zhao Ma, Yuan Yuan 0001, Xuelong Li 0001, Kongqiao Wang
ICIP1
2011 From one tree to a forest: a unified solution for structured web data extraction
abstract
Structured data, in the form of entities and associated attributes, has been a rich web resource for search engines and knowledge databases. To efficiently extract structured data from enormous websites in various verticals (e.g., books, restaurants), much research effort has been attracted, but most existing approaches either require considerable human effort or rely on strong features that lack of flexibility. We consider an ambitious scenario -- can we build a system that (1) is general enough to handle any vertical without re-implementation and (2) requires only one labeled example site from each vertical for training to automatically deal with other sites in the same vertical? In this paper, we propose a unified solution to demonstrate the feasibility of this scenario. Specifically, we design a set of weak but general features to characterize vertical knowledge (including attribute-specific semantics and inter-attribute layout relationships). Such features can be adopted in various verticals without redesign; meanwhile, they are weak enough to avoid overfitting of the learnt knowledge to seed sites. Given a new unseen site, the learnt knowledge is first applied to identify page-level candidate attribute values, while inevitably involve false positives. To remove noise, site-level information of the new site is then exploited to boost up the true values. The site-level information is derived in an unsupervised manner, without harm to the applicability of the solution. Promising experimental performance on 80 websites in 8 distinct verticals demonstrated the feasibility and flexibility of the proposed solution.
Rui Cai 0002, Yanwei Pang, Lei Zhang 0001
SIGIR3
2011 Summarizing tourist destinations by mining user-generated travelogues and photos
Yanwei Pang, Yuan Yuan 0001, Tanji Hu, Rui Cai 0002, Lei Zhang 0001
Comput. Vis. Image Underst.1
2011 Travelogue Enriching and Scenic Spot Overview Based on Textual and Visual Topic Models
abstract
We consider the problem of enriching the travelogue associated with a small number (even one) of images with more web images. Images associated with the travelogue always consist of the content and the style of textual information. Relying on this assumption, in this paper, we present a framework of travelogue enriching, exploiting both textual and visual information generated by different users. The framework aims to select the most relevant images from automatically collected candidate image set to enrich the given travelogue, and form a comprehensive overview of the scenic spot. To do these, we propose to build two-layer probabilistic models, i.e. a text-layer model and image-layer models, on offline collected travelogues and images. Each topic (e.g. Sea, Mountain, Historical Sites) in the text-layer model is followed by an image-layer model with sub-topics learnt (e.g. the topic of sea is with the sub-topic like beach, tree, sunrise and sunset). Based on the model, we develop strategies to enrich travelogues in the following steps: (1) remove noisy names of scenic spots from travelogues; (2) generate queries to automatically gather candidate image set; (3) select images to enrich the travelogue; and (4) choose images to portray the visual content of a scenic spot. Experimental results on Chinese travelogues demonstrate the potential of the proposed approach on tasks of travelogue enrichment and the corresponding scenic spot illustration.
Yanwei Pang, Yuan Yuan 0001, Xuelong Li 0001
Int. J. Pattern Recognit. Artif. Intell.1
2011 Efficient HOG human detection
Yanwei Pang, Yuan Yuan 0001, Xuelong Li 0001
Signal Process.1
2010 Photo2Trip: generating travel routes from geo-tagged photos for trip planning
abstract
Travel route planning is an important step for a tourist to prepare his/her trip. As a common scenario, a tourist usually asks the following questions when he/she is planning his/her trip in an unfamiliar place: 1) Are there any travel route suggestions for a one-day or three-day trip in Beijing? 2) What is the most popular travel path within the Forbidden City? To facilitate a tourist's trip planning, in this paper, we target at solving the problem of automatic travel route planning. We propose to leverage existing travel clues recovered from 20 million geo-tagged photos collected from www.panoramio.com to suggest customized travel route plans according to users' preferences. As the footprints of tourists at memorable destinations, the geo-tagged photos could be naturally used to discover the travel paths within a destination (attractions/landmarks) and travel routes between destinations. Based on the information discovered from geo-tagged photos, we can provide a customized trip plan for a tourist, i.e., the popular destinations to visit, the visiting order of destinations, the time arrangement in each destination, and the typical travel path within each destination. Users are also enabled to specify personal preference such as visiting location, visiting time/season, travel duration, and destination style in an interactive manner to guide the system. Owning to 20 million geo-tagged photos and 200,000 travelogues, an online system has been developed to help users plan travel routes for over 30,000 attractions/landmarks in more than 100 countries and territories. Experimental results show the intelligence and effectiveness of the proposed framework.
Changhu Wang, Jiang-Ming Yang, Yanwei Pang, Lei Zhang 0001
ACM Multimedia4
2010 Equip tourists with knowledge mined from travelogues
abstract
With the prosperity of tourism and Web 2.0 technologies, more and more people have willingness to share their travel experiences on the Web (e.g., weblogs, forums, or Web 2.0 communities). These so-called travelogues contain rich information, particularly including location-representative knowledge such as attractions (e.g., Golden Gate Bridge), styles (e.g., beach, history), and activities (e.g., diving, surfing). The location-representative information in travelogues can greatly facilitate other tourists' trip planning, if it can be correctly extracted and summarized. However, since most travelogues are unstructured and contain much noise, it is difficult for common users to utilize such knowledge effectively. In this paper, to mine location-representative knowledge from a large collection of travelogues, we propose a probabilistic topic model, named as Location-Topic model. This model has the advantages of (1) differentiability between two kinds of topics, i.e., local topics which characterize locations and global topics which represent other common themes shared by various locations, and (2) representation of locations in the local topic space to encode both location-representative knowledge and similarities between locations. Some novel applications are developed based on the proposed model, including (1) destination recommendation for on flexible queries, (2) characteristic summarization for a given destination with representative tags and snippets, and (3) identification of informative parts of a travelogue and enriching such highlights with related images. Based on a large collection of travelogues, the proposed framework is evaluated using both objective and subjective evaluation methods and shows promising results.
Rui Cai 0002, Changhu Wang, Rong Xiao 0003, Jiang-Ming Yang, Yanwei Pang, Lei Zhang 0001
WWW6
2010 Outlier-resisting graph embedding
Yanwei Pang, Yuan Yuan 0001
Neurocomputing1
2010 Robust Tensor Analysis With L1-Norm
abstract
Tensor analysis plays an important role in modern image and vision computing problems. Most of the existing tensor analysis approaches are based on the Frobenius norm, which makes them sensitive to outliers. In this paper, we propose L1-norm-based tensor analysis (TPCA-L1), which is robust to outliers. Experimental results upon face and other datasets demonstrate the advantages of the proposed approach.
Yanwei Pang, Xuelong Li 0001, Yuan Yuan 0001
IEEE Trans. Circuits Syst. Video Technol.1
2010 Footwear for Gender Recognition
abstract
Conventionally, biometrics resources, such as face, gait silhouette, footprint, and pressure, have been utilized in gender recognition systems. However, the acquisition and processing time of these biometrics data makes the analysis difficult. This letter demonstrates for the first time how effective the footwear appearance is for gender recognition as a biometrics resource. A footwear database is also established with reprehensive shoes (footwears). Preliminary experimental results suggest that footwear appearance is a promising resource for gender recognition. Moreover, it also has the potential to be used jointly with other developed biometrics resources to boost performance.
Yuan Yuan 0001, Yanwei Pang, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2010 Deterministic Column-Based Matrix Decomposition
abstract
In this paper, we propose a deterministic column-based matrix decomposition method. Conventional column-based matrix decomposition (CX) computes the columns by randomly sampling columns of the data matrix. Instead, the newly proposed method (termed as CX_D) selects columns in a deterministic manner, which well approximates singular value decomposition. The experimental results well demonstrate the power and the advantages of the proposed method upon three real-world data sets.
Xuelong Li 0001, Yanwei Pang
IEEE Trans. Knowl. Data Eng.2
2010 L1-Norm-Based 2DPCA
abstract
In this paper, we first present a simple but effective L1-norm-based two-dimensional principal component analysis (2DPCA). Traditional L2-norm-based least squares criterion is sensitive to outliers, while the newly proposed L1-norm 2DPCA is robust. Experimental results demonstrate its advantages.
Xuelong Li 0001, Yanwei Pang, Yuan Yuan 0001
IEEE Trans. Syst. Man Cybern. Part B2
2009 A fast feature extraction method
abstract
A fast subspace analysis and feature extraction algorithm is proposed which is based on fast Haar transform and integral vector. In rapid object detection and conventional binary subspace learning, Haar-like functions have been frequently used but true Haar functions are seldom employed. In this paper we have shown that true Haar functions can be successfully used to accelerate subspace analysis and feature extraction. Both the training and testing speed of the proposed method is higher than conventional algorithms. Experimental results on face database demonstrated its effectiveness.
Yanwei Pang, Xuelong Li 0001, Yuan Yuan 0001, Dacheng Tao
ICASSP2
2009 Generating location overviews with images and tags by mining user-generated travelogues
abstract
Automatically generating location overviews in the form of both visual and textual descriptions is highly desired for online services such as travel planning, to provide attractive and comprehensive outlines of travel destinations. Actually, user-generated content (e.g., travelogues) on the Web provides abundant information to various aspects (e.g., landmarks, styles, activities) of most locations in the world. To leverage the experience shared by Web users, in this paper we propose a location overview generation approach, which first mines location-representative tags from travelogues and then uses such tags to retrieve web images. The learnt tags and retrieved images are finally presented via a novel user interface which provides an informative overview for a given location. Experimental results based on 23,756 travelogues and evaluation over 20 locations show promising results on both travelogue mining and location overview generation.
Rui Cai 0002, Xin-Jing Wang, Jiang-Ming Yang, Yanwei Pang, Lei Zhang 0001
ACM Multimedia5
2009 Scene segmentation based on IPCA for visual surveillance
Yuan Yuan 0001, Yanwei Pang, Xuelong Li 0001
Neurocomputing2
2009 Binary Sparse Nonnegative Matrix Factorization
abstract
This paper presents a fast part-based subspace selection algorithm, termed the binary sparse nonnegative matrix factorization (B-SNMF). Both the training process and the testing process of B-SNMF are much faster than those of binary principal component analysis (B-PCA). Besides, B-SNMF is more robust to occlusions in images. Experimental results on face images demonstrate the effectiveness and the efficiency of the proposed B-SNMF.
Yuan Yuan 0001, Xuelong Li 0001, Yanwei Pang, Dacheng Tao
IEEE Trans. Circuits Syst. Video Technol.3
2009 Fast Haar transform based feature extraction for face representation and recognition
abstract
Subspace learning is the process of finding a proper feature subspace and then projecting high-dimensional data onto the learned low-dimensional subspace. The projection operation requires many floating-point multiplications and additions, which makes the projection process computationally expensive. To tackle this problem, this paper proposes twosimple-but-effectivefast subspace learning and image projection methods, fast Haar transform (FHT) based principal component analysis and FHT based spectral regression discriminant analysis. The advantages of these two methods result from employing both the FHT for subspace learning and the integral vector for feature extraction. Experimental results on three face databases demonstrated their effectiveness and efficiency.
Yanwei Pang, Xuelong Li 0001, Yuan Yuan 0001, Dacheng Tao
IEEE Trans. Inf. Forensics Secur.1
2009 Iterative Subspace Analysis Based on Feature Line Distance
abstract
Nearest feature line-based subspace analysis is first proposed in this paper. Compared with conventional methods, the newly proposed one brings better generalization performance and incremental analysis. The projection point and feature line distance are expressed as a function of a subspace, which is obtained by minimizing the mean square feature line distance. Moreover, by adopting stochastic approximation rule to minimize the objective function in a gradient manner, the new method can be performed in an incremental mode, which makes it working well upon future data. Experimental results on the FERET face database and the UCI satellite image database demonstrate the effectiveness.
Yanwei Pang, Yuan Yuan 0001, Xuelong Li 0001
IEEE Trans. Image Process.1
2008 Discriminant adaptive edge weights for graph embedding
abstract
Many linear dimensionality reduction (LDR) methods, such as PCA and LDA, can be reformulated in the framework of graph embedding (GE). In this framework, those LDR methods are differentiated by values of edge weights of a graph. This paper first proposes a linear dimensionality reduction method, which assigns edges with discriminant adaptive weights. Specifically, we compute a local decision hyper-plane by using support vector machine (SVM). Then edge weighs corresponding to the local region are expressed as a function of the angle between the direction of the edges and the normal vector of the hyper-plane. Experimental results demonstrate the advantages of this proposed method.
Yuan Yuan 0001, Yanwei Pang
ICASSP2
2008 Boosting simple projections for multi-class dimensionality reduction
abstract
This paper presents a novel method for dimensionality reduction and for multi-class classification tasks. This method iteratively selects a series of simple but effective 1D subspaces, and then combines the corresponding 1D projections by Adaboost.M2. Its major advantages are: 1) it does not impose specific assumptions on data distribution; 2) it minimizes Bayes error estimation in low-dimensional space; 3) it simplifies existing subspace-based methods to eigenvalue decomposition problem; and 4) each of the 1D subspaces (with associated nearest neighbor classifier) has different emphasis - measured by weighted training error. Experiments on both synthetic and real-world data demonstrate the effectiveness of the proposed method.
Yuan Yuan 0001, Yanwei Pang
SMC2
2008 Gabor-Based Region Covariance Matrices for Face Recognition
abstract
This paper presents a new method for human face recognition by utilizing Gabor-based region covariance matrices as face descriptors. Both pixel locations and Gabor coefficients are employed to form the covariance matrices. Experimental results demonstrate the advantages of this proposed method.
Yanwei Pang, Yuan Yuan 0001, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2008 Binary Two-Dimensional PCA
abstract
Fast training and testing procedures are crucial in biometrics recognition research. Conventional algorithms, e.g., principal component analysis (PCA), fail to efficiently work on large-scale and high-resolution image data sets. By incorporating merits from both two-dimensional PCA (2DPCA)-based image decomposition and fast numerical calculations based on Haarlike bases, this technical correspondence first proposes binary 2DPCA (B-2DPCA). Empirical studies demonstrated the advantages of B-2DPCA compared with 2DPCA and binary PCA.
Yanwei Pang, Dacheng Tao, Yuan Yuan 0001, Xuelong Li 0001
IEEE Trans. Syst. Man Cybern. Part B1
2008 Effective Feature Extraction in High-Dimensional Space
abstract
This correspondence first kernalizes the region covariance matrix and formalizes the similarity metric using four block matrices. The effectiveness of the proposed methods is proven with experiments on face recognition.
Yanwei Pang, Yuan Yuan 0001, Xuelong Li 0001
IEEE Trans. Syst. Man Cybern. Part B1
2006 Regularized Local Discrimimant Embedding
abstract
Recently, Chen et al. CVPR 2005) proposed a new manifold embedding method, Local Discriminant Embedding LDE), which utilizes the neighbor and class relations of data to construct the embedding for classification. While having powerful classification ability, LDE suffers from small size sample problem, which leads to unstably numerical computation. To deal with this problem, we propose to a method of regularized LDE RLDE) by imposing additional regularizing constraints on LDE. Experimental results show the effectiveness of the proposed method.
Yanwei Pang, Nenghai Yu
ICASSP (3)1
2006 A new nonlinear feature extraction method for face recognition
Yanwei Pang, Zhengkai Liu, Nenghai Yu
Neurocomputing1
2005 Neighborhood Preserving Projections (NPP): A Novel Linear Dimension Reduction Method
Yanwei Pang, Lei Zhang 0001, Zhengkai Liu, Nenghai Yu, Houqiang Li
ICIC (1)1
2004 Fusion of SVD and LDA for face recognition
abstract
A face recognition method based on the fusion of linear discriminant analysis (LDA) and singular value decomposition (SVD) is presented. In theory, fusion of different data or classifiers can achieve better performance when they are independent of each other or they can overcome shortcomings of each other. As one of the subspace methods, LDA-based method has a drawback that LDA is sensitive (variant) to translation, rotation and other geometric transforms. SVD-based method, as an algebraic feature extraction approach, has the merit of invariance to translation, rotation and mirror transforms. By combining these two methods, it is expected that better recognition performance can be obtained. Experiment results on ORL face database show the effectiveness of the proposed method.
Yanwei Pang, Nenghai Yu, Rong Zhang 0004, Jiawei Rong, Zhengkai Liu
ICIP1