EDBT 2026 Demo / reviewers in the wild / expert
Yanning Zhang 0001
dblp:14/6655 · also Yan-Ning Zhang 0001
· DBLP profile ↗
540ranked-venue papers
9as first author
308since 2021 · last 2026
0000-0002-2977-8057ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 285 · 3 first-author · 164 since 2021Artificial intelligence and machine learning · 239 · 6 first-author · 137 since 2021Applied, interdisciplinary, general and emerging computing · 83 · 1 first-author · 50 since 2021Computer networks · 9 · 3 since 2021Databases, data management, data science and information retrieval · 8 · 4 since 2021Human-computer interaction and ubiquitous computing · 3Security and privacy · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SOMA: Feature Gradient Enhanced Affine-Flow Matching for SAR-Optical RegistrationabstractAchieving pixel-level registration between SAR and optical images remains a challenging task due to their fundamentally different imaging mechanisms and visual characteristics. Although deep learning has achieved great success in many cross-modal tasks, its performance on SAR-Optical registration tasks is still unsatisfactory. Gradient-based information has traditionally played a crucial role in handcrafted descriptors by highlighting structural differences. However, such gradient cues have not been effectively leveraged in deep learning frameworks for SAR-Optical image matching. To address this gap, we propose SOMA, a dense registration framework that integrates structural gradient priors into deep features and refines alignment through a hybrid matching strategy. Specifically, we introduce the Feature Gradient Enhancer (FGE), which embeds multi-scale, multi-directional gradient filters into the feature space using attention and reconstruction mechanisms to boost feature distinctiveness. Furthermore, we propose the Global-Local Affine-Flow Matcher (GLAM), which combines affine transformation and flow-based refinement within a coarse-to-fine architecture to ensure both structural consistency and local accuracy. Experimental results demonstrate that SOMA significantly improves registration precision, increasing the CMR@1px by 12.29% on the SEN1-2 dataset and 18.50% on the GFGE_SO dataset. In addition, SOMA exhibits strong robustness and generalizes well across diverse scenes and resolutions. Tao Zhuo, Xiuwei Zhang 0001, Hanlin Yin, Wencong Wu, Yanning Zhang 0001 |
AAAI | 6 |
| 2026 | YOLO-IOD: Towards Real Time Incremental Object DetectionabstractCurrent methodologies for incremental object detection (IOD) primarily rely on Faster R-CNN or DETR series detectors; however, these approaches do not accommodate the real-time YOLO detection frameworks. In this paper, we first identify three primary types of knowledge conflicts that contribute to catastrophic forgetting in YOLO-based incremental detectors: foreground-background confusion, parameter interference, and misaligned knowledge distillation. Subsequently, we introduce YOLO-IOD, a real-time Incremental Object Detection (IOD) framework that is constructed upon the pretrained YOLO-World model, facilitating incremental learning via a stage-wise parameter-efficient finetuning process. Specifically, YOLO-IOD encompasses three principal components: 1) Conflict-Aware Pseudo-Label Refinement (CPR), which mitigates the foreground-background confusion by leveraging the confidence levels of pseudo labels and identifying potential objects relevant to future tasks. 2) Importance-based Kernel Selection (IKS), which identifies and updates the pivotal convolution kernels pertinent to the current task during the current learning stage. 3)Cross-Stage Asymmetric Knowledge Distillation (CAKD), which addresses the misaligned knowledge distillation conflict by transmitting the features of the student target detector through the detection heads of both the previous and current teacher detectors, thereby facilitating asymmetric distillation between existing and newly introduced categories. We further introduce LoCo COCO, a more realistic benchmark that eliminates data leakage across stages. Experiments on both conventional and LoCo COCO benchmarks show that YOLO-IOD achieves superior performance with minimal forgetting. Shizhou Zhang, Xueqiang Lv, Yinghui Xing, Qirui Wu, Di Xu 0010, Chen Zhao 0009, Yanning Zhang 0001 |
AAAI | 7 |
| 2026 | Coordinated B diffusion Gaussian distribution AlGaN/GaN HEMT device by quasi-van der Waals epitaxy
Yanning Zhang 0001, Haidi Wu, Xinchen Ji, Zhichun Yang, Xinbo Zhang, Ling Bai, Juncheng Zheng, Yue Hao 0001, Jincheng Zhang 0001 |
Sci. China Inf. Sci. | 1 |
| 2026 | Alternating exposure control network for real-world environments
Chenyuan Zhao, Yu Zhu 0004, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | MRFMA: A hybrid paradigm integrating multi-receptive field network with mediator attention for 3D multi-organ segmentation
Hengfei Cui, Jiatong Li 0006, Dianrong Du, Yanning Zhang 0001, Yong Xia 0001 |
Expert Syst. Appl. | 4 |
| 2026 | History-aware adaptive teacher for cross-domain object detection
Yaoqi Hu, Axi Niu, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001 |
Expert Syst. Appl. | 7 |
| 2026 | Inter-view dual-domain guided stable diffusion for real-world stereo image super-resolution
Yu Zhu 0004, Axi Niu, Jinqiu Sun, Yanning Zhang 0001 |
Expert Syst. Appl. | 5 |
| 2026 | Single-frame supervision for temporal video anomaly grounding
Yuzhou Long, Peng Wu 0015, Guansong Pang, Peng Wang 0015, Yanning Zhang 0001 |
Neurocomputing | 6 |
| 2026 | Semi-Supervised VQA Multi-Modal Explanation via Self-Critical LearningabstractVQA explanation task aims to explain the decision-making process of VQA models in a way that is easily understandable to humans. Existing methods mostly use visual location or natural language explanation approaches to generate corresponding rationales. Although significant progress has been made, these frameworks are bottlenecked by the following challenges: 1) Uni-modal paradigm inevitably leads to semantic ambiguity of explanations. 2) The reasoning process cannot be faithfully responded to and suffers from logical inconsistency. 3) Human-annotated explanations are expensive and time-consuming to collect. In this paper, we introduce a new Semi-supervised VQA Multi-modal Explanation (SME) method via self-critical learning, which addresses the above challenges by leveraging both visual and textual explanations to comprehensively reveal the inference process of the model. Meanwhile, in order to improve the logical consistency between answers and rationales, we design a novel self-critical strategy to evaluate candidate explanations based on answer reward scores. More importantly, our method can benefit from a tremendous amount of samples without human-annotated explanations with semi-supervised learning. Extensive automatic measures and human evaluations all show the effectiveness of our method. Finally, the framework achieves a new state-of-the-art performance on the three VQA explanation datasets. Wei Suo, Ji Ma 0008, Mengyang Sun, Hanwang Zhang, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | A Hierarchical Prior Mining Approach for Non-Local Multi-View StereoabstractAs a fundamental problem in computer vision, multi-view stereo (MVS) aims at recovering the 3D geometry of the target from a set of 2D images. However, the reconstructed quality is significantly impacted by the presence of low-textured areas. In this paper, we propose a Hierarchical Prior Mining (HPM) framework for non-local multi-view stereo. Different from most existing works dedicated to focusing on local information and only using a single prior, HPM captures non-local structural cues and leverages multi-source priors for geometry recovery. Based on the framework, we first propose HPM-MVS, which obtains precise initial hypotheses through non-local operations, simultaneously constructing a better planar prior model in an HPM framework to further facilitate hypothesis generation. In addition, we futher propose HPM-MVS++, which excavates the structured region information of images and spatial geometric relationships of hypotheses as prior knowledge. Then, it incorporates them into probabilistic graphical models, ultimately deducing two novel multi-view matching costs. This significantly enhances the robustness to challenging situations and improves the completeness of the reconstruction. Experimental results on the ETH3D and Tanks & Temples have verified the superior performance and strong generalization capability of our approach. Jiaqi Yang 0002, Yanan He, Chunlin Ren, Qingshan Xu 0001, Siwen Quan, Xiyu Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | VPT-NSP2++: Importance-Aware Visual Prompt Tuning in Null Space for Continual LearningabstractContinual learning (CL) enables AI models to adapt to evolving environments while mitigating catastrophic forgetting, which is a critical capability for dynamic real-world applications. With the growing popularity of pre-trained Vision Transformer (ViT) models and visual prompt tuning (VPT) technique in CL, this work explores a CL method on top of the ViT-based foundation model, through VPT mechanism with theoretical guarantees. Inspired by the orthogonal projection method, we aim to leverage this approach for VPT to enhance CL performance, particularly in long-term scenarios. However, since the orthogonal projection is originally designed for linear operations in CNNs, applying it to ViTs poses challenges induced by the non-linear self-attention mechanism and the distribution drift within LayerNorm. To address these issues, we deduced two orthogonality conditions to achieve the prompt gradient orthogonal projection, which provide a theoretical guarantee of maintaining stability. Considering the strict orthogonal constraints can diminish model capacity and reduce plasticity, we further propose an importance-aware orthogonal regularization framework. By applying varying degrees of orthogonal constraints to different parameters based on their importance to old and new tasks, the framework adaptively enhances model capacity and thereby promotes long-sequence CL while improving the stability-plasticity trade-off. To implement the proposed approach, a null-space-based approximation solution is employed to efficiently achieve the prompt gradient orthogonal projection. Extensive experiments on various class-incremental learning benchmarks demonstrate that our method achieves state-of-the-art performance across diverse CL scenarios. Shizhou Zhang, Yue Lu 0008, De Cheng, Yinghui Xing, Nannan Wang 0001, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | AMG-Net: A multitask network with adaptive mutual guidance for Semantic Change Detection
Yuduo Bian, Wei Wei 0008, Chen Ding 0002, Lei Zhang 0038, Jiangbin Zheng 0001, Yanning Zhang 0001 |
Pattern Recognit. | 6 |
| 2026 | Task-Adapter++: Task-specific adaptation with order-aware alignment for few-shot action recognition
Congqi Cao, Peiheng Han, Yueran Zhang, Yating Yu, Qinyi Lv, Lingtong Min, Yanning Zhang 0001 |
Pattern Recognit. | 7 |
| 2026 | Nearest-neighbor class prototype prompt and simulated logits for continual learning
Yue Lu 0008, Shizhou Zhang, Yinghui Xing, Guoqiang Liang 0001, Yanning Zhang 0001 |
Pattern Recognit. | 6 |
| 2026 | Push the limit of scene text recognition using character and text length guided text super-resolution
Jiangtao Nie, Boxiong Wu, Wenyu Peng, Wei Wei 0008, Lei Zhang 0054, Chen Ding 0002, Yanning Zhang 0001 |
Pattern Recognit. | 7 |
| 2026 | AdaPrompt-IR: Adaptive learning to perceive degradation semantic and prompting for all-in-one image restoration
Wei Sun 0036, Qianzhou Wang, Qingsen Yan, Yanning Zhang 0001 |
Pattern Recognit. | 6 |
| 2026 | Category text-guided RGBT tracking with shared-specific feature representation
Wei Wei 0008, Haolie Wang, Yuduo Bian, Haijiao Xing, Chen Ding 0002, Lei Zhang 0054, Tao Zhou 0009, Jiangbin Zheng 0001, Yanning Zhang 0001 |
Pattern Recognit. | 9 |
| 2026 | Lightweight modal-guided cross-attention fusion network for visible-infrared object detection
Wencong Wu, Hongxi Zhang, Xiuwei Zhang 0001, Hanlin Yin, Yanning Zhang 0001 |
Pattern Recognit. | 5 |
| 2026 | DT-RSRGAN: An one-off domain translation generative model for real image super-resolution
Shaolin Su, Yu Zhu 0004, Lingmei Zhang, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001 |
Pattern Recognit. | 7 |
| 2026 | Do it yourself dynamic single image super resolution network via ODE
Xiao Zhang 0058, Zhen Zhang 0008, Wei Wei 0008, Lei Zhang 0054, Yanning Zhang 0001 |
Pattern Recognit. | 5 |
| 2026 | VoMarkSplat: Robust watermarking for 3D Gaussian splatting with patch and multi-convolutional voting
Tianyu Xiong, Rui Li 0013, Jiaqi Yang 0002, Yanning Zhang 0001 |
Pattern Recognit. Lett. | 5 |
| 2026 | Hyperspectral Image Compression With Spectral-Spatial Coupling and Group-Wise Context ModelingabstractThe rich spectral information within hyperspectral images (HSIs) results in large data volumes. Thus finding a compact representation for HSIs while maintaining reconstruction quality is a fundamental task for numerous applications. Though the existing learning-based compression methods and context models have shown strong rate-distortion (RD) performance, these methods only pay their attention on spatial redundancy without considering the spectral redundancy of HSIs, which thus impedes further improvement of their performance on HSI. Moreover, the strictly sequential autoregressive nature of context models leads to inefficiency, further limiting their practical applications. In this paper, leveraging the spectral priors unique to HSIs, we propose a hybrid Transformer-CNN architecture to find compact latent representations of HSIs. In specific, we construct Spectral-Spatial Coupling Transformer Group (SSCTG) to cooperatively extract spatial and spectral features of HSIs. Additionally, we propose Group-wise Context Model (GCM) to further enhance the parallel processing capability of autoregression within context models, significantly improving the coding efficiency. Extensive experiments demonstrate the effectiveness of the proposed method, achieving superior RD performance compared to state-of-the-art methods while maintaining high efficiency of codecs. Wei Wei 0008, Shuyi Zhao, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | High Dynamic Range Imaging via Spatial-Frequency Interactionabstractpublicly available.High Dynamic Range (HDR) imaging aims to reconstruct scenes with a wide range of luminance by fusing multi-exposure Low Dynamic Range (LDR) images. In dynamic scenes with pronounced foreground motion or camera jitter, especially under challenging conditions including extremely low or high luminance, widespread saturation, and substantial motion, existing approaches often encounter ghosting artifacts, spatial misalignment, and degradation of fine structural details. Traditional techniques based on handcrafted priors struggle to generalize to complex motion patterns, while most deep learning-based methods operate exclusively in the spatial domain, limiting their ability to capture global contextual cues and restore high-frequency structures that are better represented in the frequency domain. To address these challenges, we introduce a Dual-Domain Parallel Fusion Network with Prompt Refinement (DDPF-PR), which jointly leverages spatial and frequency-domain features for enhanced HDR reconstruction. Specifically, the proposed framework consists of a Bi-Domain Interaction Module(BDIM), which integrates spatial features for local detail and frequency features for global structure to suppress ghosting artifacts caused by motion. In addition, a Prompt Refinement Module(PRM) is designed to recover fine details in degraded regions such as saturated or misaligned areas by adaptively generating structural cues. Extensive experiments demonstrate that DDPF-PR consistently outperforms state-of-the-art methods across multiple benchmarks in both qualitative and quantitative evaluations. The code will be made publicly available. Weiyu Zhou, Yongqing Yang, Tao Hu 0013, Pu Hui, Yu Cao 0016, Qingsen Yan, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | AxisPose: Model-Free Matching-Free Single-Shot 6D Object Pose Estimation via Axis GenerationabstractObject pose estimation is a fundamental task in computer vision and plays an important role in various applications such as robotics, augmented reality, and autonomous manipulation. Existing studies often demand complex inputs or depend on correspondence-based matching between 2D image features and 3D object representations. While effective, these methods rely strongly on explicit appearance matching, often requiring multi-view inputs, depth sensors, or CAD models, which limits their scalability and robustness. Building on top of the pioneering generative studies, we propose AxisPose, a model-free, matching-free, and single-view 6D pose estimation framework that departs from conventional correspondence-based paradigms. Unlike existing methods, AxisPose directly infers a pose representation by learning a latent distribution of object orientation axes through a diffusion model. Specifically, AxisPose introduces an Axis Generation Module (AGM) that progressively denoises tri-axial orientation fields guided by geometric consistency constraints, and a Triaxial Back-projection Module (TBM) to recover the final 6D pose from the generated orientation axes without relying on explicit 2D- 2D/3D correspondences. AxisPose achieves strong cross-instance generalization, enabling a single model to handle multiple object categories without retraining. Extensive experiments on LINEMOD and YCB-Video datasets demonstrate that AxisPose improves the Average Distance Deviation score from 0.733 to 0.814 over the strong baseline NOPE, using only a single RGB input. The code is available at https://github.com/pubyLu/AxisPose/tree/main. Yang Zou 0004, Zhaoshuai Qi, Weipeng Sun, Xingyuan Li 0005, Jiaqi Yang 0002, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2026 | Boosting HDR Image Reconstruction via Semantic Knowledge TransferabstractRecovering High Dynamic Range (HDR) images from multiple Standard Dynamic Range (SDR) images becomes challenging when the SDR images exhibit noticeable degradation and missing content. Leveraging scene-specific semantic priors offers a promising solution for restoring heavily degraded regions. However, these priors are typically extracted from sRGB SDR images, the domain/format gap poses a significant challenge when applying it to HDR imaging. To address this issue, we propose a general framework that transfers semantic knowledge derived from SDR domain via self-distillation to boost existing HDR reconstruction. Specifically, the proposed framework first introduces the Semantic Priors Guided Reconstruction Model (SPGRM), which leverages SDR image semantic knowledge to address ill-posed problems in the initial HDR reconstruction results. Subsequently, we leverage a self-distillation mechanism that constrains the color and content information with semantic knowledge, aligning the external outputs between the baseline and SPGRM. Furthermore, to transfer the semantic knowledge of the internal features, we utilize a Semantic Knowledge Alignment Module (SKAM) to fill the missing semantic contents with the complementary masks. Extensive experiments demonstrate that our framework significantly boosts HDR imaging quality for existing methods without altering the network architecture. Tao Hu 0013, Longyao Wu, Wei Dong 0010, Peng Wu 0015, Jinqiu Sun, Xiaogang Xu 0002, Qingsen Yan, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 8 |
| 2026 | Ghost-Free HDR Imaging via Latent Low-Frequency Priors and Deformable Attention AlignmentabstractRecovering ghost-free High Dynamic Range (HDR) images from multiple Low Dynamic Range (LDR) images becomes challenging when the LDR images exhibit saturation and significant motion. Recent Diffusion Models (DMs) have been introduced in HDR imaging field, showing promising performance, particularly in achieving visually perceptible better results compared to previous DNN-based methods. However, DMs require extensive iterations with large models to estimate entire images, resulting in inefficiency that hinders their practical application. To address this challenge, we propose the Low-Frequency aware Diffusion (LF-Diff) model for ghost-free HDR imaging. The key idea of LF-Diff is implementing the DMs in a highly compacted latent space and integrating it into a regression-based model to enhance the details of reconstructed images. Specifically, as low-frequency information is closely related to human visual perception we propose to utilize DMs to create compact low-frequency priors for the reconstruction process. These priors are integrated into a carefully designed Dynamic HDR Reconstruction Network (DHRNet), which employs a regression-based approach to produce high-quality HDR images. Furthermore, we introduce the Attention-guided Deformable Alignment Module (ADAM) that utilizes correlation-driven feature matching to learn deformable receptive fields for self-attention, enabling efficient pre-alignment of LDR images by focusing on salient regions. Extensive experiments on synthetic and real-world benchmark datasets demonstrate that our LF-Diff performs favorably against several state-of-the-art methods and is $10\times $ faster than previous DM-based methods. Tao Hu 0013, Qingsen Yan, Wei Dong 0010, Peng Wu 0015, Yuankai Qi, Weisi Lin, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 8 |
| 2026 | Pseudo Sentences Evaluation and Quality-Aware Robust Learning for Unsupervised Text-Based Person SearchabstractUnsupervised Text-Based Person Search (TBPS) eliminates the need for costly manual sentence annotations by generating pseudo sentences via Multi-modal Large Language Models (MLLMs). However, these pseudo sentences often face the quality defect issues, resulting in semantic misalignment across modalities, which will hinder discriminative representation learning. To address this problem, we propose the PSE-QRL (Pseudo Sentences Evaluation and Quality-aware Robust Learning), a unified framework that enhances robustness to pseudo sentences for unsupervised TBPS. The PSE-QRL dynamically couples an evolving TBPS model with MLLMs to assess pseudo sentences' reliability, and adaptively leverages high-quality ones during training. It consists of three key components: 1) Multi-granularity Sentence Augmentation, for enriching pseudo sentences with multiple granularities to broaden the diversity of image-sentence pairs; 2) Hybrid Quality Evaluation, to combine MLLM's cross-modal reasoning knowledge with TBPS model's person-specific distinguishing capabilities for effective sentence quality assessment; and 3) Quality-aware Robust Learning, for selecting and re-weighting samples based on quality scores to emphasize reliable sentence annotations while suppressing low-quality ones. Extensive experiments on CUHK-PEDES, ICFG-PEDES, and RSTPReid benchmarks demonstrate the effectiveness of PSE-QRL for improving learning robustness, achieving state-of-the-art (SOTA) retrieval performance for unsupervised TBPS. Kai Niu 0002, Xinyue Song, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | Less Is More: Infrared and Visible Images Fusion via Semantic-Guided Mixture of Multi-Feature ExpertsabstractInfrared (IR) and visible image fusion (IVIF) has become prevalent in recent years. By leveraging the complementary characteristics of infrared and visible images, we can obtain visually-appealing fused images, which further facilitate subsequent scene understanding and object detection from day to night. Integrating complementary information while simultaneously eliminating redundancy is a crucial challenge in fusion. Most of available deep learning based methods, after being trained, execute static inference on all pairs of infrared and visible images. They struggle to effectively handle redundancy of modality across diverse scenarios, resulting in superfluous information such as thermal noise in infrared images and artifacts in visible images. In this paper, we propose an IVIF method based on a semantic-guided mixture of multi-feature experts, where multiple types of features are extracted, each assigned to a dedicated expert network specialized in processing a specific type of features. Through an expert routing mechanism, these experts are chosen dynamically, ensuring that the most significant features of each image modality are routed to a specific group of experts. In order to align fusion task with subsequent semantic segmentation task, we introduce a segmentation head to semantically guide the selection of the complementary features. Extensive experiments on five infrared and visible image fusion and segmentation benchmarks demonstrate the effectiveness of our method, both for image fusion and subsequent semantic segmentation tasks. The code will be available at https://github.com/ZhilongNiu/SD-MoMFE. Yinghui Xing, Zhilong Niu, Shizhou Zhang, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | Decoupling Target Semantics via Text-Anchored Visual Contrast for Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning (SSL) provides an effective means of reducing reliance on large-scale annotated datasets by leveraging unlabeled data. However, existing SSL methods often struggle with semantic ambiguity, especially under limited supervision. Recent studies have incorporated textual information to provide contextual guidance, yet most focus on feature fusion rather than emphasizing target semantics critical for segmentation. In this paper, we proposed a novel Text-anchored Visual Decoupling (TeViD) framework for semi-supervised medical image segmentation. TeViD is built upon a teacher-student architecture with a dual-decoder design that explicitly disentangles target and background representations using both labeled and unlabeled data. For unlabeled data, a reversed cross-supervision mechanism is introduced to enhance decoder diversity and semantic separation. Furthermore, two contrastive learning objectives are proposed: a teacher-guided visual contrastive loss and a text-anchored contrastive loss, both designed to reinforce semantic disentanglement from visual and textual perspectives. Extensive experiments on five public datasets (covering X-ray, pathology, ultrasound, MRI, and CT) demonstrate that TeViD consistently outperforms both standard SSL and text-enhanced SSL methods, achieving average improvements of 5.72% in Dice and 8.15% in mIoU over the second-best competitor. The code is available at: https://github.com/jgfiuuuu/TeViD. Qingjie Zeng, Xinke Ma, Zilin Lu, Mengkang Lu, Yanning Zhang 0001, Yong Xia 0001 |
IEEE Trans. Image Process. | 7 |
| 2026 | Meta-Exploiting Complementary Semantic Consistency for Cross-Domain Few-Shot Learning PromotionabstractMeta-learning has emerged as an effective solver for cross-domain few-shot learning (CD-FSL) tasks. Despite achieving obvious progress recently, the typical episodic learning paradigm often causes the feature embedding model collapsing into the simplicity bias pitfall, viz., the model tends to prioritize some shortcut patterns (e.g., color, style, background) that are only sufficient to distinguish categories in source domain, while fail to generalize across domains. To mitigate this problem, we present a novel meta-learning framework which emphasizes meta-exploiting inductive bias to alleviate simplicity bias for CD-FSL promotion, and mainly contributes in the following four aspects. 1) We establish a novel inductive bias for CD-FSL, termed complementary semantic consistency (CSC). The rationale behind lies in that forcing the semantic consistency between two complementary feature learning schemes is beneficial to distill cross-domain transferable features. 2) We establish a solid theoretical foundation, supported by rigorous mathematical proofs and key lemmas, which demonstrates that CSC establishes a tighter generalization bound and facilitates the learning of domain-invariant features. 3) Inspired by CSC, we propose a general meta-learning framework, which implements complementary feature embedding models using parallel networks with the same architecture but different input forms, and introduce proper knowledge distillation losses to encourage the semantic consistency between different branches during meta-training. This framework can be seamlessly integrated with any complementary feature learning schemes. 4) To clarify this point, we instantiate two effective meta-learners based on the proposed framework. The former establishes a two-branch network that simultaneously classifies both the query image and its random local crops. The latter decomposes the query image into high-frequency and low-frequency components, which are then integrated into a parallel feature embedding network for category prediction, analogous to the original query image. Subsequently, a KL divergence based knowledge distillation loss is separately leveraged to force the prediction consistency between the complementary branches (e.g., local-global, spatial-frequency) during meta-training. By doing these, both learners are able to distill cross-domain transferable features with better generalization performance. Empirical results on diverse benchmarks consistently affirm the proposed framework's advantages, while additional analysis provides compelling support for our key claims. Fei Zhou 0008, Lei Zhang 0054, Wei Wei 0008, Chen Ding 0002, Guosheng Lin, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 7 |
| 2026 | From Few to More: Scribble-Based Medical Image Segmentation via Masked Context Modeling and Continuous Pseudo LabelsabstractScribble-based weakly supervised segmentation methods have shown promising results in medical image segmentation, significantly reducing annotation costs. However, existing approaches often rely on auxiliary tasks to enforce semantic consistency and use hard pseudo labels for supervision, overlooking the unique challenges faced by models trained with sparse annotations. These models must predict pixel-wise segmentation maps from limited data, making it crucial to handle varying levels of annotation richness effectively. In this paper, we propose MaCo, a weakly supervised model designed for medical image segmentation, based on the principle of "from few to more." MaCo leverages Masked Context Modeling (MCM) and Continuous Pseudo Labels (CPL). MCM employs an attention-based masking strategy to perturb the input image, ensuring that the model's predictions align with those of the original image. CPL converts scribble annotations into continuous pixel-wise labels by applying an exponential decay function to distance maps, producing confidence maps that represent the likelihood of each pixel belonging to a specific category, rather than relying on hard pseudo labels. We evaluate MaCo on three public datasets, comparing it with other weakly supervised methods. Our results show that MaCo outperforms competing methods across all datasets, establishing a new record in weakly supervised medical image segmentation. Zhisong Wang, Yiwen Ye, Ziyang Chen 0003, Minglei Shu, Yanning Zhang 0001, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2026 | Harnessing Text Insights With Visual Alignment for Medical Image SegmentationabstractPre-trained vision-language models (VLMs) and language models (LMs) have recently garnered significant attention due to their remarkable ability to represent textual concepts, opening up new avenues in vision tasks. In medical image segmentation, efforts are being made to integrate text and image data using VLMs and LMs. However, current text-enhanced approaches face several challenges. First, using separate pre-trained vision and text models to encode image and text data can result in semantic shifts. Second, while VLMs can establish the correspondence between visual and textual features when pre-trained on paired image-text data, this alignment often deteriorates during segmentation tasks due to misalignment between the text and vision components in ongoing learning. In this paper, we propose TeViA, a novel approach that seamlessly integrates with various vision and text models, irrespective of their pre-training relationships. This integration is achieved through a segmentation-specific text-to-vision alignment design, ensuring both information gain and semantic consistency. Specifically, for each training data, a foreground visual representation is extracted from the segmentation head and used to supervise projection layers, thereby adjusting the textual features to better contribute to the segmentation task. Additionally, a historic visual prototype is created by aggregating target semantics from all training data and is updated using a momentum-based manner. This prototype aims to enhance the visual representation of each data instance by establishing feature-level connections, which in turn refines the textual features. The superiority of TeViA is validated on five public datasets, exhibiting over 6% Dice improvements compared to vision-only methods. Code is available at: https://github.com/jgfiuuuu/TeViA. Qingjie Zeng, Zilin Lu, Yutong Xie 0001, Zhiyong Wang 0001, Yanning Zhang 0001, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2026 | Deep Learning for Video Anomaly Detection: A ReviewabstractVideo anomaly detection (VAD) aims to discover behaviors or events deviating from the normality in videos. As a long-standing task in the field of computer vision, VAD has witnessed much good progress. In the era of deep learning, with the explosion of architectures of continuously growing capability and capacity, a great variety of deep learning-based methods are constantly emerging for the VAD task, greatly improving the generalization ability of detection algorithms and broadening the application scenarios. Therefore, such a multitude of methods and a large body of literature make a comprehensive survey a pressing necessity. In this article, we present an extensive and comprehensive research review, covering the spectrum of five different categories, namely, semi-supervised, weakly supervised, fully supervised, unsupervised, and open-set supervised VAD, and we also delve into the latest VAD works based on pretrained large models and open-world learning, remedying the limitations of past reviews in terms of only focusing on semi-supervised VAD and small model-based methods. For the VAD task with different levels of supervision, we construct a well-organized taxonomy, profoundly discuss the characteristics of different types of methods, and show their performance comparisons. In addition, this review involves the public datasets, open-source codes, and evaluation metrics covering all the aforementioned VAD tasks. Finally, we provide several important research directions for the VAD community. Additional details of the survey are available on the project homepage: https://github.com/Roc-Ng/DeepVAD. Peng Wu 0015, Chengyu Pan, Guansong Pang, Qingsen Yan, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2026 | STPrompt\(\boldsymbol{++}\): Prompting Vision-Language Models for Weakly Supervised Video Anomaly Detection and Fine-Grained LocalizationabstractTraditional weakly supervised video anomaly detection (WSVAD) tasks typically rely on coarse-grained frame-level labels for training. Although this approach reduces annotation costs, it results in weak semantic understanding and spatial localization capabilities due to the absence of fine-grained annotations, hindering precise pixel-level anomaly detection and localization. Thanks to the success of vision-language models (VLMs), e.g., CLIP, recent approaches leveraging large VLMs focus on exploiting their strong semantic understanding capabilities, but they typically feed only keyframes or short video segments into the models, without supplying sufficient prior contextual information (e.g., contextual frames around anomalies, zoomed-in anomaly regions, and detailed anomaly descriptions), which restricts the models’ capability for fine-grained anomaly understanding and precise localization. More recently, a few methods leveraging VLMs, attempt to achieve training-free spatial anomaly localization by fusing patch-level visual features with simple textual features. However, these methods employ simplistic textual descriptions, lacking deep semantic comprehension of anomalies, leading to coarse localization results with significant irrelevant background noise. To address these issues, we propose STPrompt \(++\) , a novel weakly supervised spatio-temporal video anomaly detection and localization method based on VLMs. In our work, we systematically leverage preliminary coarse localization regions derived from anomaly scores as spatial priors, together with contextual frames around keyframes, zoomed-in views of suspected anomalous regions, and refined textual descriptions of anomalies. This comprehensive prompting mechanism guides the VLMs toward deep semantic comprehension of video anomalies, enabling accurate pixel-level spatial localization. The proposed STPrompt \(++\) requires no additional training and significantly enhances the precision of anomaly understanding and localization through a carefully designed multi-round and multi-modal prompting mechanism. Extensive experiments on two widely used WSVAD benchmarks, UCF-Crime and UBnormal, show that our method achieves state-of-the-art spatial localization performance and competitive temporal anomaly detection results. Notably, on UCF-Crime dataset, our approach improves spatial localization accuracy (in TIoU) by 5.61% over the current best method (from 23.90% to 29.51%), underscoring its superior capabilities in precise anomaly localization and semantic understanding. Peng Wu 0015, Chengyu Pan, Guansong Pang, Xiangteng He, Zhiwei Yang 0013, Peng Wang 0015, Yanning Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | MAC++: Going Further with Maximal Cliques for 3D RegistrationabstractMaximal cliques (MAC) represent a novel state-of-theart approach for 3D registration from correspondences, however, it still suffers from extremely severe outliers. In this paper, we introduce a robust learning-free estimator called MAC++, exploring maximal cliques for$3 D$registration from the following two perspectives: 1)$A$novel hypothesis generation method utilizing putative seeds through voting to guide the construction of maximal clique pools, effectively preserving more potential correct hypotheses. 2) A progressive hypothesis evaluation method that continuously reduces the solution space in a “global-clusters-cluster-individual” manner rather than traditional one-shot techniques, greatly alleviating the issue of missing good hypotheses. Experiments conducted on U3M, 3DMatch/3DLoMatch, and KITTI-LC datasets show the new state-of-the-art performance of MAC++. MAC++ demonstrates the capability to handle extremely low inlier ratio data where MAC fails (e.g., showing 27.1%/30.6% registration recall improvements on 3DMatch/3DLoMatch with$<1 \%$inliers). Xiyu Zhang 0001, Yanning Zhang 0001, Jiaqi Yang 0002 |
3DV | 2 |
| 2025 | Training Consistent Mixture-of-Experts-Based Prompt Generator for Continual LearningabstractVisual prompt tuning-based continual learning (CL) methods have shown promising performance in exemplar-free scenarios, where their key component can be viewed as a prompt generator. Existing approaches generally rely on freezing old prompts, slow updating and task discrimination for prompt generators to preserve stability and minimize forgetting. In contrast, we introduce a novel approach that trains a consistent prompt generator to ensure stability during CL. Consistency means that for any instance from an old task, its corresponding instance-ware prompt generated by the prompt generator remains consistent even as the generator continually updates in a new task. This ensures that the representation of a specific instance remains stable across tasks and thereby prevents forgetting. We employ a mixture of experts (MoE) as the prompt generator, which contains a router and multiple experts. By deriving conditions sufficient to achieve the consistency for the MoE prompt generator, we demonstrate that: during training in a new task, if the router and experts update in the directions orthogonal to the subspaces spanned by old input features and gating vectors, respectively, the consistency can be theoretically guaranteed. To implement this orthogonality, we project parameter gradients to those orthogonal directions using the orthogonal projection matrices computed via the null space method. Extensive experiments on four class-incremental learning benchmarks validate the effectiveness and superiority of our approach. Yue Lu 0008, Shizhou Zhang, De Cheng, Guoqiang Liang 0001, Yinghui Xing, Nannan Wang 0001, Yanning Zhang 0001 |
AAAI | 7 |
| 2025 | VarCMP: Adapting Cross-Modal Pre-Training Models for Video Anomaly RetrievalabstractVideo anomaly retrieval (VAR) aims to retrieve pertinent abnormal or normal videos from collections of untrimmed and long videos through cross-modal requires such as textual descriptions and synchronized audios. Cross-modal pre-training (CMP) models, by pre-training on large-scale cross-modal pairs, e.g., image and text, can learn the rich associations between different modalities, and this cross-modal association capability gives CMP an advantage in conventional retrieval tasks. Inspired by this, how to utilize the robust cross-modal association capabilities of CMP in VAR to search crucial visual component from these untrimmed and long videos becomes a critical research problem. Therefore, this paper proposes a VAR method based on CMP models, named VarCMP. First, a unified hierarchical alignment strategy is proposed to constrain the semantic and spatial consistency between video and text, as well as the semantic, temporal, and spatial consistency between video and audio. It fully leverages the efficient cross-modal association capabilities of CMP models by considering cross-modal similarities at multiple granularities, enabling VarCMP to achieve effective all-round information matching for both video-text and video-audio VAR tasks. Moreover, to further solve the problem of untrimmed and long video alignment, an anomaly-biased weighting is devised in the fine-grained alignment, which identifies key segments in untrimmed long videos using anomaly priors, giving them more attention, thereby discarding irrelevant segment information, and achieving more accurate matching with cross-modal queries. Extensive experiments demonstrates high efficacy of VarCMP in both video-text and video-audio VAR tasks, achieving significant improvements on both text-video (UCFCrime-AR) and audio-video (XDViolence-AR) datasets against the best competitors by 5.0% and 5.3% R@1. Peng Wu 0015, Wanshun Su, Xiangteng He, Peng Wang 0015, Yanning Zhang 0001 |
AAAI | 5 |
| 2025 | Low-Biased General Annotated Dataset GenerationabstractPre-training backbone networks on a general annotated dataset (e.g., ImageNet) that comprises numerous manually collected images with category annotations has proven to be indispensable for enhancing the generalization capacity of downstream visual tasks. However, those manually collected images often exhibit bias, which is non-transferable across either categories or domains, thus causing the model’s generalization capacity degeneration. To mitigate this problem, we present a low-biased general annotated dataset generation framework (lbGen). Instead of expensive manual collection, we aim at directly generating low-biased images with category annotations. To achieve this goal, we propose to leverage the advantage of a multimodal foundation model (e.g., CLIP), in terms of aligning images in a low-biased semantic space defined by language. Specifically, we develop a bi-level semantic alignment loss, which not only forces all generated images to be consistent with the semantic distribution of all categories belonging to the target dataset in an adversarial learning manner, but also requires each generated image to match the semantic description of its category name. In addition, we further cast an existing image quality scoring model into a quality assurance loss to preserve the quality of the generated image. By leveraging these two loss functions, we can obtain a low-biased image generation model by simply fine-tuning a pre-trained diffusion model using only all category names in the target dataset as input. Experimental results confirm that, compared with the manually labeled dataset or other synthetic datasets, the utilization of our generated low-biased dataset leads to stable generalization capacity enhancement of different backbone networks across various tasks, especially in tasks where the manually labeled samples are scarce. Code is available at: https://github.com/vvvvvjdy/lbGen Dengyang Jiang, Haoyu Wang 0016, Lei Zhang 0054, Wei Wei 0008, Guang Dai, Yanning Zhang 0001 |
CVPR | 8 |
| 2025 | Sparse2DGS: Geometry-Prioritized Gaussian Splatting for Surface Reconstruction from Sparse ViewsabstractWe present a Gaussian Splatting method for surface reconstruction using sparse input views. Previous methods relying on dense views struggle with extremely sparse Structure-from-Motion points for initialization. While learning-based Multi-view Stereo (MVS) provides dense 3D points, directly combining it with Gaussian Splatting leads to suboptimal results due to the ill-posed nature of sparse-view geometric optimization. We propose Sparse2DGS, an MVS-initialized Gaussian Splatting pipeline for complete and accurate reconstruction. Our key insight is to incorporate the geometric-prioritized enhancement schemes, allowing for direct and robust geometric learning under illposed conditions. Sparse2DGS outperforms existing methods by notable margins while being 2× faster than the NeRF-based fine-tuning approach. Code is available at https://github.com/Wuuu3511/Sparse2DGS. Rui Li 0013, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001 |
CVPR | 6 |
| 2025 | Dual-Granularity Semantic Guided Sparse Routing Diffusion Model for General PansharpeningabstractPansharpening aims at integrating complementary information from panchromatic and multispectral images. Available deep-learning based pansharpening methods typically perform exceptionally with particular satellite datasets. At the same time, it has been observed that these models also exhibit scene dependence, for example, if the majority of the training samples come from the urban scenes, the model’s performance may decline in the river scene. To address the domain gap produced by varying satellite sensors and distinct scenes, we propose a dual-granularity semantic guided sparse routing diffusion model for general pansharpening. By utilizing the large Vision-Language Models (VLMs) in the field of geoscience, e.g, GeoChat, we introduce the dual granularity semantics to generate dynamic sparse routing scores for adaptation of different satellite sensors and scenes. This scene-level and region-level dual-granularity semantic information serves as guidance for dynamically activating specialized experts within the diffusion model. Extensive experiments on WorldView-3, QuickBird, and GaoFen-2 datasets show the effectiveness of our proposed method. Notably, the proposed method outperforms the comparison approaches in adapting to new satellite sensors and scenes. The codes are available at https://github.com/codgodtao/SGDiff. Yinghui Xing, Litao Qu, Shizhou Zhang, Di Xu 0010, Yingkun Yang, Yanning Zhang 0001 |
CVPR | 6 |
| 2025 | HVI: A New Color Space for Low-light Image EnhancementabstractLow-Light Image Enhancement (LLIE) is a crucial computer vision task that aims to restore detailed visual information from corrupted low-light images. Many existing LLIE methods are based on standard RGB (sRGB) space, which often produce color bias and brightness artifacts due to inherent high color sensitivity in sRGB. While converting the images using Hue, Saturation and Value (HSV) color space helps resolve the brightness issue, it introduces significant red and black noise artifacts. To address this issue, we propose a new color space for LLIE, namely Horizontal/Vertical-Intensity (HVI), defined by polarized HS maps and learnable intensity. The former enforces small distances for red coordinates to remove the red artifacts, while the latter compresses the low-light regions to remove the black artifacts. To fully leverage the chromatic and intensity information, a novel Color and Intensity Decoupling Network (CIDNet) is further introduced to learn accurate photometric mapping function under different lighting conditions in the HVI space. Comprehensive results from benchmark and ablation experiments show that the proposed HVI color space with CIDNet outperforms the state-of-the-art methods on 10 datasets. The code is available at https://github.com/Fediory/HVI-CIDNet. Qingsen Yan, Yixu Feng, Guansong Pang, Kangbiao Shi, Peng Wu 0015, Wei Dong 0010, Jinqiu Sun, Yanning Zhang 0001 |
CVPR | 9 |
| 2025 | Revisiting Generative Replay for Class Incremental Object DetectionabstractGenerative replay has gained significant attention in class-incremental learning; however, its application to Class Incremental Object Detection (CIOD) remains limited due to the challenges in generating complex images with precise spatial arrangements. In this study, motivated by the observation that the forgetting of prior knowledge is predominantly present in the classification sub-task as opposed to the localization sub-task, we revisit the generative replay method for class incremental object detection. Our method utilize a standard Stable Diffusion model to generate image-level replay data for all old and new tasks. Accordingly, the old detector and a stage-wise detector are conducted on the synthetic images respectively to determine the bounding box positions through pseudo-labeling. Furthermore, we propose to use a Similarity-based Cross Sampling mechanism to select valuable confusing data between old and new tasks to more effectively mitigate catastrophic forgetting and reduce the false alarm rate for the new task. Finally, all synthetic and real data are integrated for current-stage detector training, where the images generated for previous tasks are highly beneficial in minimizing the forgetting of existing knowledge, while those synthesized for the new task can help bridge the domain gap between real and synthetic images. We conducted extensive experiments on PASCAL VOC 2007 and MS COCO benchmark datasets in multiple settings to showcase the efficacy of our proposed approach, which achieves state-of-the-art results. The code is available at https://github.com/qiangzailv/RGR-IOD. Shizhou Zhang, Xueqiang Lv, Yinghui Xing, Qirui Wu, Di Xu 0010, Yanning Zhang 0001 |
CVPR | 6 |
| 2025 | Gradient Decomposition and Alignment for Incremental Object Detection
Wenlong Luo, Shizhou Zhang, De Cheng, Yinghui Xing, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001 |
ICCV | 7 |
| 2025 | Pruning All-Rounder: Rethinking and Improving Inference Efficiency for Large Vision Language ModelsabstractAlthough Large Vision-Language Models (LVLMs) have achieved impressive results, their high computational costs pose a significant barrier to wide application. To enhance inference efficiency, most existing approaches can be categorized as parameter-dependent or token-dependent strategies to reduce computational demands. However, parameter-dependent methods require retraining LVLMs to recover performance while token-dependent strategies struggle to consistently select the most relevant tokens. In this paper, we systematically analyze the above challenges and provide a series of valuable insights for inference acceleration. Based on these findings, we propose a novel framework, the Pruning All-Rounder (PAR). Different from previous works, PAR develops a meta-router to adaptively organize pruning flows across both tokens and layers. With a self-supervised learning manner, our method achieves a superior balance between performance and efficiency. Notably, PAR is highly flexible, offering multiple pruning versions to address a range of acceleration scenarios. The code for this work is publicly available at https://github.com/ASGO-MM/Pruning-All-Rounder. Wei Suo, Ji Ma 0008, Mengyang Sun, Lin Wu 0001, Peng Wang 0015, Yanning Zhang 0001 |
ICCV | 6 |
| 2025 | Learning to Generalize Without Bias for Open-Vocabulary Action RecognitionabstractLeveraging the effective visual-text alignment and static generalizability from CLIP, recent video learners adopt CLIP initialization with further regularization or recombination for generalization in open-vocabulary action recognition in-context. However, due to the static bias of CLIP, such video learners tend to overfit on shortcut static features, thereby compromising their generalizability, especially to novel out-of-context actions. To address this issue, we introduce Open-MeDe, a novel Meta-optimization framework with static Debiasing for Open-vocabulary action recognition. From a fresh perspective of generalization, Open-MeDe adopts a meta-learning approach to improve known-to-open generalizing and image-to-video debiasing in a cost-effective manner. Specifically, Open-MeDe introduces a cross-batch meta-optimization scheme that explicitly encourages video learners to quickly generalize to arbitrary subsequent data via virtual evaluation, steering a smoother optimization landscape. In effect, the free of CLIP regularization during optimization implicitly mitigates the inherent static bias of the video meta-learner. We further apply self-ensemble over the optimization trajectory to obtain generic optimal parameters that can achieve robust generalization to both in-context and out-of-context novel data. Extensive evaluations show that Open-MeDe not only surpasses state-of-the-art regularization methods tailored for in-context open-vocabulary action recognition but also substantially excels in out-of-context scenarios.Code is released at https://github.com/Mia-YatingYu/Open-MeDe. Yating Yu, Congqi Cao, Yifan Zhang 0001, Yanning Zhang 0001 |
ICCV | 4 |
| 2025 | Autoregressive Denoising Score Matching Is a Good Video Anomaly DetectorabstractVideo anomaly detection (VAD) is an important computer vision problem. Thanks to the mode coverage capabilities of generative models, the likelihood-based paradigm is catching growing interest, as it can model normal distribution and detect out-of-distribution anomalies. However, these likelihood-based methods are blind to the anomalies located in local modes near the learned distribution. To handle these ``unseen" anomalies, we dive into three gaps uniquely existing in VAD regarding scene, motion and appearance. Specifically, we first build a noise-conditioned score transformer for denoising score matching. Then, we introduce a scene-dependent and motion-aware score function by embedding the scene condition of input sequences into our model and assigning motion weights based on the difference between key frames of input sequences. Next, to solve the problem of blindness in principle, we integrate unaffected visual information via a novel autoregressive denoising score matching mechanism for inference. Through autoregressively injecting intensifying Gaussian noise into the denoised data and estimating the corresponding score function, we compare the denoised data with the original data to get a difference and aggregate it with the score function for an enhanced appearance perception and accumulate the abnormal context. With all three gaps considered, we can compute a more comprehensive anomaly indicator. Experiments on three popular VAD benchmarks demonstrate the state-of-the-art performance of our method. Hanwen Zhang 0017, Congqi Cao, Qinyi Lv, Lingtong Min, Yanning Zhang 0001 |
ICCV | 5 |
| 2025 | HyperGCT: A Dynamic Hyper-GNN-Learned Geometric Constraint for 3D RegistrationabstractGeometric constraints between feature matches are critical in 3D point cloud registration problems. Existing approaches typically model unordered matches as a consistency graph and sample consistent matches to generate hypotheses. However, explicit graph construction introduces noise, posing great challenges for handcrafted geometric constraints to render consistency. To overcome this, we propose HyperGCT, a flexible dynamic Hyper-GNN-learned geometric ConstrainT that leverages high-order consistency among 3D correspondences. To our knowledge, HyperGCT is the first method that mines robust geometric constraints from dynamic hypergraphs for 3D registration. By dynamically optimizing the hypergraph through vertex and edge feature aggregation, HyperGCT effectively captures the correlations among correspondences, leading to accurate hypothesis generation. Extensive experiments on 3DMatch, 3DLoMatch, KITTI-LC, and ETH show that HyperGCT achieves state-of-the-art performance. Furthermore, HyperGCT is robust to graph noise, demonstrating a significant advantage in terms of generalization. Xiyu Zhang 0001, Jiayi Ma 0001, Zhaoshuai Qi, Fei Hui, Jiaqi Yang 0002, Yanning Zhang 0001 |
ICCV | 8 |
| 2025 | Towards Effective Foundation Model Adaptation for Extreme Cross-Domain Few-Shot Learning
Fei Zhou 0008, Lei Zhang 0038, Wei Wei 0008, Chen Ding 0002, Guosheng Lin, Yanning Zhang 0001 |
ICCV | 7 |
| 2025 | Demystifying Catastrophic Forgetting in Two-Stage Incremental Object DetectorabstractCatastrophic forgetting is a critical chanllenge for incremental object detection (IOD). Most existing methods treat the detector monolithically, relying on instance replay or knowledge distillation without analyzing component-specific forgetting. Through dissection of Faster R-CNN, we reveal a key insight: Catastrophic forgetting is predominantly localized to the RoI Head classifier, while regressors retain robustness across incremental stages. This finding challenges conventional assumptions, motivating us to develop a framework termed NSGP-RePRE. Regional Prototype Replay (RePRE) mitigates classifier forgetting via replay of two types of prototypes: coarse prototypes represent class-wise semantic centers of RoI features, while fine-grained prototypes model intra-class variations. Null Space Gradient Projection (NSGP) is further introduced to eliminate prototype-feature misalignment by updating the feature extractor in directions orthogonal to subspace of old inputs via gradient projection, aligning RePRE with incremental learning dynamics. Our simple yet effective design allows NSGP-RePRE to achieve state-of-the-art performance on the Pascal VOC and MS COCO datasets under various settings. Our work not only advances IOD methodology but also provide pivotal insights for catastrophic forgetting mitigation in IOD. Code will be available soon. Qirui Wu, Shizhou Zhang, De Cheng, Yinghui Xing, Di Xu 0010, Peng Wang 0015, Yanning Zhang 0001 |
ICML | 7 |
| 2025 | Prompt-Free Conditional Diffusion for Multi-object Image AugmentationabstractDiffusion model has underpinned much recent advances of dataset augmentation in various computer vision tasks. However, when involving generating multi-object images as real scenarios, most existing methods either rely entirely on text condition, resulting in a deviation between the generated objects and the original data, or rely too much on the original images, resulting in a lack of diversity in the generated images, which is of limited help to downstream tasks. To mitigate both problems with one stone, we propose a prompt-free conditional diffusion framework for multi-object image augmentation. Specifically, we introduce a local-global semantic fusion strategy to extract semantics from images to replace text, and inject knowledge into the diffusion model through LoRA to alleviate the category deviation between the original model and the target dataset. In addition, we design a reward model based counting loss to assist the traditional reconstruction loss for model training. By constraining the object counts of each category instead of pixel-by-pixel constraints, bridging the quantity deviation between the generated data and the original data while improving the diversity of the generated data. Experimental results demonstrate the superiority of the proposed method over several representative state-of-the-art baselines and showcase strong downstream task gain and out-of-domain generalization capabilities. Code is available at \href{https://github.com/00why00/PFCD}{here}. Haoyu Wang 0016, Lei Zhang 0054, Wei Wei 0008, Chen Ding 0002, Yanning Zhang 0001 |
IJCAI | 5 |
| 2025 | Boosting Multi-Modal Alignment: Geometric Feature Separation for Class Incremental LearningabstractClass Incremental Learning (CIL) aims to continually learn new classes from a stream of data without forgetting previously learned ones. Recent approaches have leveraged pre-trained models (PTMs) to improve performance, especially vision-language models, which offer better generalization than models trained solely on visual data. Many of these methods rely on simple language templates to generate class representations, which then serve as classifiers. However, due to differences between the pre-training data and downstream tasks, these textual features can become too similar for certain classes, leading to prediction errors. To address this issue, we propose a method that optimizes the geometric structure of both visual and textual features across different classes. Inspired by neural collapse theory, we introduce a multi-modal alignment strategy: for each class, a reference vector is chosen from a simplex Equiangular Tight Frame, and both the visual and textual features of the class are aligned with this vector. To better capture intra-class variations, we also construct multiple visual prototypes for each class. A multi-prototype supervised contrastive loss is then employed to pull an image feature toward the closest matching prototype of its true class and push it away from prototypes of other classes. We evaluate our approach on five widely used CIL benchmarks. The results show that our method achieves state-of-the-art performance, demonstrating its effectiveness in addressing the challenges of class incremental learning. Our code is available at https://github.com/qcNPU/NCSCMP. Guoqiang Liang 0001, De Cheng, Shizhou Zhang, Yanning Zhang 0001 |
ACM Multimedia | 5 |
| 2025 | Test-Time Adaptation for Text-Based Person SearchabstractText-based person search (TBPS), aiming to retrieve target pedestrian images with natural language descriptions, has seen significant progress in recent years. However, severe domain shift remains a key challenge in this field, causing source-domain-trained models to degrade significantly when applied to an unseen target domain. To address this, we propose the Identity-preserving Cross-modal Alignment and Adaptation (ICAA) model, a novel test-time adaptation framework for TBPS that enables seamless domain adaptation using only unlabeled target samples. Our method tackles two key challenges: 1) Cross-modal domain-shift misalignment: textual and visual modalities exhibit inconsistent distributional shifts across domains. To this end, our Cross-Modal Alignment adaptation (CMA) module identifies pseudo-positive image-text pairs and minimizes their matching discrepancies in the target domain, adapting to new cross-modal distribution relationships. 2) Identity semantic absence: crucial identity annotations are usually unavailable in both target text and image data. To mitigate this, we introduce the Identity-Preserving Dynamic adaptation (IPD) module, which dynamically associates image-text pairs with potential identity prototypes to enhance identity consistency in cross-modal alignment during adaptation. Our method is simple yet effective, establishing new state-of-the-art cross-domain results for TBPS on three public benchmarks, i.e., CUHK-PEDES, ICFG-PEDES, and RSTPReid. Kai Niu 0002, Liucun Shi, Qinzi Zhao, Yanning Zhang 0001 |
ACM Multimedia | 6 |
| 2025 | Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers
Ji Ma 0008, Wei Suo, Peng Wang 0015, Yanning Zhang 0001 |
ACM Multimedia | 4 |
| 2025 | Amplitude-aware Domain Style Replay for Lifelong Person Re-identificationabstractLifelong Person Re-identification (LReID) focuses on continuously adapting to new domains over time while preserving knowledge from previously seen domains, particularly under the domain incremental learning setting. The major challenge of LReID is catastrophic forgetting, typically caused by large domain shifts during training. To address this, we propose a novel Amplitude-aware Domain Style Replay (ADSR) framework, which introduces a Fourier-based Style Transfer (FST) mechanism to generate synthetic data that reflects the style of previously encountered domains. These proxy images help retain prior knowledge without the need to store actual past data. Our method transfers stylistic information-mainly encoded in the amplitude spectrum-from old domains to new ones, creating old-stylized images that preserve the content of new domain data while adopting the visual style of earlier domains. To further boost generalization, we design a Self-Stylization Normalization (SSN) module that adapts the current domain's style distribution, making the model more robust to stylistic variations. Additionally, we introduce a Multi-Granularity Transfer (MGT) module that uses K-Means clustering to extract multiple representative style features from each domain, enabling compact yet comprehensive storage and replay of domain-specific information. Extensive experiments on multiple LReID benchmarks show that ADSR achieves superior performance over existing approaches, effectively reducing forgetting and improving cross-domain generalization. Our code is available at https://github.com/cclong8/MM2025-ADSR. De Cheng, Shizhou Zhang, Yinghui Xing, Di Xu 0010, Yanning Zhang 0001 |
ACM Multimedia | 6 |
| 2025 | Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models
Mingyu Fu, Wei Suo, Ji Ma 0008, Lin Wu 0001, Peng Wang 0015, Yanning Zhang 0001 |
ACM Multimedia | 6 |
| 2025 | Text-Visual Semantic Constrained AI-Generated Image Quality Assessment
Qingsen Yan, Haojian Huang, Peng Wu 0015, Haokui Zhang, Yanning Zhang 0001 |
ACM Multimedia | 6 |
| 2025 | PoseCrafter: Extreme Pose Estimation with Hybrid Video SynthesisabstractPairwise camera pose estimation from sparsely overlapping image pairs remains a critical and unsolved challenge in 3D vision.
Most existing methods struggle with image pairs that have small or no overlap. Recent approaches attempt to address this by synthesizing intermediate frames using video interpolation and selecting key frames via a self-consistency score. However, the generated frames are often blurry due to small overlap inputs, and the selection strategies are slow and not explicitly aligned with pose estimation.
To solve these cases, we propose Hybrid Video Generation (HVG) to synthesize clearer intermediate frames by coupling a video interpolation model with a pose-conditioned novel view synthesis model, where we also propose a Feature Matching Selector (FMS) based on feature correspondence to select intermediate frames appropriate for pose estimation from the synthesized results. Extensive experiments on Cambridge Landmarks, ScanNet, DL3DV-10K, and NAVI demonstrate that, compared to existing SOTA methods, PoseCrafter can obviously enhance the pose estimation performances, especially on examples with small or no overlap. Qing Mao, Tianxin Huang, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001, Gim Hee Lee |
NeurIPS | 5 |
| 2025 | Draw Sketch, Draw Flesh: Whole-Body Computed Tomography from Any X-Ray Views
Yongsheng Pan, Yiwen Ye, Yanning Zhang 0001, Yong Xia 0001, Dinggang Shen |
Int. J. Comput. Vis. | 3 |
| 2025 | Audio-visual correspondences based joint learning for instrumental playing source separation
Peng Zhang 0005, Siliang Wang, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001 |
Neurocomputing | 6 |
| 2025 | Enhancing the noise robustness of sparse-form patches for image denoising
Liping Qi, Yu Zhu 0004, Wei Sun 0036, Axi Niu, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001 |
Knowl. Based Syst. | 8 |
| 2025 | Generalized pixel-aware deep function-mixture network for effective spectral super-resolution
Jiangtao Nie, Lei Zhang 0054, Chongxing Song, Zhiqiang Lang, Weixin Ren, Wei Wei 0008, Chen Ding 0002, Yanning Zhang 0001 |
Knowl. Based Syst. | 8 |
| 2025 | Multi-level cross-knowledge fusion with edge guidance for camouflaged object detection
Wei Sun 0036, Qianzhou Wang, Yulong Tian, Xianguang Kong, Yizhuo Dong, Yanning Zhang 0001 |
Knowl. Based Syst. | 7 |
| 2025 | CLIP-guided continual novel class discovery
Qingsen Yan, Yiting Yang, Yutong Dai 0001, Katarzyna Wiltos, Marcin Wozniak, Wei Dong 0010, Yanning Zhang 0001 |
Knowl. Based Syst. | 8 |
| 2025 | Unleashing the potential of open-set noisy samples against label noise for medical image classification
Zehui Liao, Shishuai Hu, Yanning Zhang 0001, Yong Xia 0001 |
Medical Image Anal. | 3 |
| 2025 | A multi-scale feature cross-dimensional interaction network for stereo image super-resolution
Yu Zhu 0004, Shengjun Peng, Axi Niu, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001 |
Multim. Syst. | 7 |
| 2025 | Modeling optical imaging pipeline and learning contrastive-based representation for hybrid-corrupted image restoration
Chenyuan Zhao, Yu Zhu 0004, Qingsen Yan, Jinqiu Sun, Axi Niu, Yanning Zhang 0001 |
Multim. Syst. | 6 |
| 2025 | SAR remote sensing image segmentation based on feature enhancement
Wei Wei 0008, Yanyu Ye, Guochao Chen, Yanning Zhang 0001 |
Neural Networks | 7 |
| 2025 | Scene-Dependent Prediction in Latent Space for Video Anomaly Detection and AnticipationabstractVideo anomaly detection (VAD) plays a crucial role in intelligent surveillance. However, an essential type of anomaly named scene-dependent anomaly is overlooked. Moreover, the task of video anomaly anticipation (VAA) also deserves attention. To fill these gaps, we build a comprehensive dataset named NWPU Campus, which is the largest semi-supervised VAD dataset and the first dataset for scene-dependent VAD and VAA. Meanwhile, we introduce a novel forward-backward framework for scene-dependent VAD and VAA, in which the forward network individually solves the VAD and jointly solves the VAA with the backward network. Particularly, we propose a scene-dependent generative model in latent space for the forward and backward networks. First, we propose a hierarchical variational auto-encoder to extract scene-generic features. Next, we design a score-based diffusion model in latent space to refine these features more compact for the task and generate scene-dependent features with a scene information auto-encoder, modeling the relationships between video events and scenes. Finally, we develop a temporal loss from key frames to constrain the motion consistency of video clips. Extensive experiments demonstrate that our method can handle both scene-dependent anomaly detection and anticipation well, achieving state-of-the-art performance on ShanghaiTech, CUHK Avenue, and the proposed NWPU Campus datasets. Congqi Cao, Hanwen Zhang 0017, Yue Lu 0008, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Learning Dual-Stream Conditional Concepts in Compositional Zero-Shot LearningabstractCompositional Zero-Shot Learning (CZSL) aims to recognize unseen compositional concepts composed of seen single concepts. One of the problems of CZSL is to model attributes interacting with objects and objects interacting with attributes. In this work, we focus on this problem and propose Dual-Stream Conditional Network (DSCNet) that learns dual-stream conditional concepts as a solution, where the conditional visual and semantic embeddings of attributes and objects are learned. First, we argue that the condition of the attribute or object is supposed to contain the recognized object and input image, or the recognized attribute and input image. Next, for each concept which can either be an attribute or object, in the semantic stream, we propose to encode the recognized object or attribute semantic features and the input image visual features as the encoded condition, which is then injected into all concept semantic embeddings by a semantic cross encoder to acquire conditional semantic embeddings. In the visual stream, the conditional attribute or object visual embeddings are acquired by injecting the semantic features of the recognized object or attribute into the mapped attribute or object visual features. Experimental results on CZSL benchmarks demonstrate the superiority of our proposed method. Qingsheng Wang, Lingqiao Liu, Chenchen Jing, Peng Wang 0015, Yanning Zhang 0001, Chunhua Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | A masking, linkage and guidance framework for online class incremental learning
Guoqiang Liang 0001, Zhaojie Chen, Shibin Su, Shizhou Zhang, Yanning Zhang 0001 |
Pattern Recognit. | 5 |
| 2025 | CHA: Conditional Hyper-Adapter method for detecting human-object interaction
Mengyang Sun, Wei Suo, Peng Wang 0015, Yanning Zhang 0001 |
Pattern Recognit. | 5 |
| 2025 | Multi-scale feature extraction and fusion with attention interaction for RGB-T tracking
Haijiao Xing, Wei Wei 0008, Lei Zhang 0054, Yanning Zhang 0001 |
Pattern Recognit. | 4 |
| 2025 | Domain consistency learning for continual test-time adaptation in image semantic segmentation
Yanyu Ye, Wei Wei 0008, Lei Zhang 0054, Chen Ding 0002, Yanning Zhang 0001 |
Pattern Recognit. | 5 |
| 2025 | Adapt Anything: Tailor Any Image Classifier Across Domains and Categories Using Text-to-Image Diffusion ModelsabstractWe study a novel problem in this paper, that is, if a modern text-to-image diffusion model can tailor any image classifier across domains and categories. Existing domain adaption works exploit both source and target data for domain alignment so as to transfer the knowledge from the labeled source data to the unlabeled target data. However, as the development of text-to-image diffusion models, we wonder if the high-fidelity synthetic data can serve as a surrogate of the source data in real world. In this way, we do not need to collect and annotate the source data for each image classification task in a one-for-one manner. Instead, we utilize only one off-the-shelf text-to-image model to synthesize images with labels derived from text prompts, and then leverage them as a bridge to dig out the knowledge from the task-agnostic text-to-image generator to the task-oriented image classifier via domain adaptation. Such a one-for-all adaptation paradigm allows us to adapt anything in the world using only one text-to-image generator as well as any unlabeled target data. Extensive experiments validate the feasibility of this idea, which even surprisingly surpasses the state-of-the-art domain adaptation works using the source data collected and annotated in real world. Weijie Chen 0006, Haoyu Wang 0016, Shicai Yang, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Luojun Lin, Di Xie, Yueting Zhuang |
IEEE Trans. Big Data | 6 |
| 2025 | LINR: A Plug-and-Play Local Implicit Neural Representation Module for Visual Object TrackingabstractCurrent one-stream trackers suffer from limitations in distinguishing targets from complex backgrounds owing to their uniform token division strategy. By treating all regions equally, these methods allocate inadequate attention to crucial target details while overemphasizing redundant background information. Consequently, their performance deteriorates significantly in scenarios involving similar distractors or background clutter. In this work, we propose a Local Implicit Neural Representation (LINR) module specifically designed for local fine-grained object modeling. It consists of two key modules: (1) Local Window Selection: Leveraging template-guided CNN-based cross-correlation, it accurately identify crucial target-relevant regions, reducing background information redundant and computation burden. (2) INR-based Window Refinement: Using implicit neural networks, it optimizes token density and spatial continuity to improve local fine-grained instance-level representations, facilitating the discriminative ability between the target and the background. Moreover, the LINR module exhibits three remarkable advantages as a generalized enhancement for visual tracking. Firstly, it is plug-and-play, seamlessly integrating into existing one-stream trackers, both non-real-time and real-time ones, without architectural modifications, achieving significant performance improvements. Secondly, it is highly portable since it does not introduce new loss functions, additional training strategies or data. Thirdly, it is efficiency-friendly, having minimal impact on model parameters and tracking speed,e.g., AQATrack-LINR increases only 1.9% of the parameters and reduces the tracking speed by only 6fps. We incorporate the LINR module into two non-real-time trackers, OSTrack based on ViT-B and AQATrack based on HiViT-B, and one real-time tracker, FERMT based on ViT-tiny, respectively. The resultant OSTrack-LINR, AQATrack-LINR, and FERMT-LINR achieve state-of-the-art performance across seven widely utilized datasets, such as TrackingNet, LaSOT, and NFS30. The source code is available at https://github.com/Xiaochen918/LINR. Guancheng Jia, Yufei Zha, Peng Zhang 0005, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Lightweight Image Deblurring via Recurrent Gated Attention and Efficient DecouplingabstractIn recent years, deep learning has been significantly advancing the field of image deblurring. However, existing deep learning models usually rely on overloaded large kernel convolutions or overweighted attention modules. This leads to a heavy computational burden and restricts real applications. To address this issue, we propose a lightweight deblurring network, termed RGE-Net. Our RGE-Net possesses two novel features: 1) We propose a recurrent path into the convolutions to ensure each kernel weight can learn better and stronger feature information, thus increasing the parameter efficiency and reducing the parameters. Furthermore, we propose gated attention to suppress incorrect features flowing in the recurrent path, thus improving performance. 2) We decouple the kernels into spatial and channel components to reduce learning difficulty by reducing parameters and then perform an attention mechanism to obtain significant performance. Extensive experiments on benchmark datasets demonstrate the superiority of RGE-Net over state-of-the-art deblurring models in terms of both effectiveness and efficiency. Shilin Ye, Geng Chen 0001, Meklit Mesfin Atlaw, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Visual Object Tracking With Multi-Frame Distractor SuppressionabstractWith the rapid development of CNN or Transformer, the present mainstream approaches regard an image patch as the reference of the target to perform tracking, which is known as template matching-based trackers. However, most existing template matching-based trackers only consider the per-frame localization accuracy, neglecting the potential distractor (similar object) dependencies among multiple video frames, which poses a fundamental challenge in template matching-based tracking. In this work, we propose a novel comprehensive framework with multi-frame distractor suppression for visual object tracking (MFDSTrack), which explicitly models the temporal history of both the target object and potential distractors. Specifically, we utilize a universal target candidate generation module to detect target candidates (both target and distractors), providing a holistic view of the scene. In addition, a temporal and distractor-aware association module is designed to suppress multi-frame distractors by adopting a simple encoder-decoder Transformer architecture. The encoder accepts inputs of target candidates’ history, while the decoder takes current target candidate queries and the output of the encoder as inputs to associate current target candidate queries with historical trajectories. We extensively evaluate our trackers, MFDSTrack-SD, MFDSTrack-OS, MFDSTrack-GRM, and MFDSTrack-LT on the LaSOT,${\mathrm {LaSOT}}_{ext}$, TrackingNet, GOT-10k, UAV123, NFS, and OTB100 benchmark. Extensive experiments show that our methods outperform previous state-of-the-art trackers on seven tracking benchmarks. Mingyu Cai, Zhixuan Bai, Tao Zhuo, Hongming Zhang 0002, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Multi-Granularity Language-Guided Training for Multi-Object TrackingabstractMost existing multi-object tracking methods typically learn visual tracking features via maximizing dis-similarities of different instances and minimizing similarities of the same instance. While such a feature learning scheme achieves promising performance, learning discriminative features solely based on visual information is challenging especially in case of environmental interference such as occlusion, blur and domain variance. In this work, we argue that multi-modal language-driven features provide complementary information to classical visual features, thereby aiding in improving the robustness to such environmental interference. To this end, we propose a new multi-object tracking framework, named LG-MOT, that explicitly leverages language information at different levels of granularity (scene-and instance-level) and combines it with standard visual features to obtain discriminative representations. To develop LG-MOT, we annotate existing MOT datasets with scene-and instance-level language descriptions. We then encode both scene-and instance-level language information into high-dimensional embeddings, which are utilized to guide the visual features during training. At inference, our LG-MOT uses the standard visual features without relying on annotated language descriptions. Extensive experiments on three benchmarks, MOT17, DanceTrack and SportsMOT, reveal the merits of the proposed contributions leading to state-of-the-art performance. On the DanceTrack test set, our LG-MOT achieves an absolute gain of 2.2% in terms of target object association (IDF1 score), compared to the baseline using only visual features. Further, our LG-MOT exhibits strong cross-domain generalizability. Source code and pre-trained models are available at https://github.com/WesLee88524/LG-MOT. Jiale Cao, Muzammal Naseer, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001, Fahad Shahbaz Khan |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Learning From Multi-Perception Features for Real-Word Image Super-ResolutionabstractActual image super-resolution is an extremely challenging task due to complex degradations existing in the image. To solve this problem, two dominant methodologies have emerged: degradation-estimation-based Addressing actual image super-resolution remains a formidable challenge due to the intricate degradations present in images. Two primary methodologies have emerged: degradation-estimation-based and blind-based methods. The former often struggle to accurately estimate degradation, limiting their effectiveness on real low-resolution images. Conversely, blind-based methods rely on a single perceptual perspective, constraining their adaptability to diverse perceptual characteristics. In response to these challenges, we present MPF-Net, a novel super-resolution approach aimed at enhancing real-world image super-resolution tasks by enabling the model to learn multiple perceptual features from input images. Our method features a Multi-Perception Feature Extraction module (MPFE) designed to extract diverse perceptual details, complemented by Cross-Perception Blocks (CPB) facilitating the fusion of this information for efficient super-resolution reconstruction. Additionally, we introduce a contrastive regularization term (CR) to enhance the model’s learning by leveraging newly generated HR and LR images as positive and negative samples. Experimental results on challenging real-world SR datasets demonstrate the superiority of our approach over existing state-of-the-art methods, both qualitatively and quantitatively. Axi Niu, Kang Zhang 0008, Trung X. Pham, Jinqiu Sun, In-So Kweon, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Pseudo Labeling Methods for Semi-Supervised Semantic Segmentation: A Review and Future PerspectivesabstractSemantic segmentation is a fundamental task in computer vision and finds extensive applications in scene understanding, medical image analysis, and remote sensing. With the advent of deep learning, significant advancements have been made in segmentation tasks. However, deep learning models require a substantial amount of labeled data for training, and accurately annotating datasets is labor-intensive and costly. Recently, numerous studies have explored the semantic segmentation task through the lens of semi-supervised learning, with the pseudo-labeling (PL) method emerging as a straightforward and widely applicable approach. This paper provides a comprehensive review and analysis of various PL methods and their applications in semi-supervised semantic segmentation (SSSS) from multiple angles. Initially, it captures the essence of individual model self-training and the collaborative training of multiple models from a model-centric viewpoint. Next, it explores strategies for refining or dismissing unreliable methods. Then, it categorizes techniques for addressing noisy PL data and inspects improvements in PL methods from the perspective of data augmentation. It further provides insights into optimization strategies. Furthermore, it examines PL methods from an application-oriented standpoint, such as in medical image segmentation and remote sensing image segmentation. Lastly, this paper evaluates the performance of cutting-edge methods on public datasets and concludes by discussing the challenges and potential directions for future research. Lingyan Ran, Guoqiang Liang 0001, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Flexible Temperature Parallel Distillation for Dense Object Detection: Make Response-Based Knowledge Distillation Great AgainabstractFeature-based approaches have been the focal point of previous research on knowledge distillation (KD) for dense object detection. These methods employ feature imitation and result in competitive performance. Despite being able to achieve comparable performance in image recognition, response-based KD methods can not reach the same level in dense object detection. Inspired by improving distillation performance from two key aspects: where to distill and how to distill, in this paper, a parallel distillation (PD) is introduced to fully utilize the sophisticated detection head and transfer all the output responses from the teacher to the student efficiently. In particular, the proposed PD takes an important consideration of the specific location of distillation, which is crucial for effective knowledge transfer. Regarding the discrepancies in output responses between the localization branch and the classification branch, we propose a novel Dynamic Localization Temperature (DLT) module to enhance the precision of distilling localization information. As for the classification branch, a Classification Temperature-Free (CTF) module is also designed to increase the robustness of distillation in heterogeneous networks. By incorporating the DLT and CTF into the PD framework to avoid setting temperature values manually, the Flexible Temperature Parallel Distillation (FTPD) is proposed to achieve a state-of-the-art (SOTA) performance, which can also be further combined with mainstream feature-based methods for better results. In terms of accuracy and robustness with extensive experiments, the proposed FTPD outperforms other KD methods in the task of dense object detection. Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Efficient Image Enhancement With a Diffusion-Based Frequency PriorabstractDue to the lack of appropriate priors, generating the content of dark regions remains a challenge in low-light image enhancement tasks. Currently, diffusion models employ robust image generation capabilities for enhancing low-light images. However, diffusion models require multiple iterations at the image feature level to generate details and content, which limits the speed. Moreover, the diffusion-based methods tend to generate unexpected artifacts in the degraded regions. To address these issues, we propose a Frequency Priors-guided Image Enhancement (FPIE) network, including a frequency prior generation network and an image restoration network. FPIE significantly accelerates inference by learning abstract prior with frequency domain constraints. Concretely, to learn compacted priors at the frequency domain, we introduce a joint training approach for the prior generation and restoration models to constrain the distribution of priors. Furthermore, to better utilize frequency-domain features for enhancing the network’s generation capabilities, a wavelet-based transformer block is introduced to produce intricate details and avoid the artifacts of the output. Extensive experimental results on the commonly used benchmarks demonstrate that our approach achieves state-of-the-art performances and well generalization to real-world images. Qingsen Yan, Tao Hu 0013, Peng Wu 0015, Duwei Dai, Shuhang Gu, Wei Dong 0010, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | From Dynamic to Static: Stepwisely Generate HDR Image for Ghost RemovalabstractGenerating high-quality high dynamic range (HDR) images in dynamic scenes is particularly challenging due to the influence of large motion. Despite the effectiveness of existing deep learning methods, they still suffer from ghosting artifacts when saturation and motion coexist. Inspired by fusion on static scenes, we propose an inpainting and fusion strategy to enhance the quality of the generated HDR images. The proposed method consists of pseudo-static LDR generation and detail-guided HDR generation, which creates pseudo-static images and then generates ghost-free HDR images. Specifically, the pseudo-static LDR generation network utilizes semantic information to identify the motion regions, and employs a diffusion model-based inpainting approach to produce pseudo-static LDR images that closely resemble real scenes. In the detail-guided HDR generation network, we employ a detail enhancement module to refine diverse high-frequency features with detailed information extracted from pseudo-static LDR images, which effectively enhances the visual quality. Extensive experiments on four public datasets demonstrate the superiority of the proposed method, both quantitatively and qualitatively. Qingsen Yan, Kangzhen Yang, Tao Hu 0013, Genggeng Chen, Kexin Dai, Peng Wu 0015, Wenqi Ren, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | Diffusion-Augmented Cross-Domain Prototypical Knowledge Distillation for Few-Shot Learning in Hyperspectral Image ClassificationabstractCross-domain few-shot learning (FSL) has demonstrated remarkable new classes recognition capabilities in hyperspectral image classification tasks. However, existing domain adaptation methods face two critical challenges in the cross-domain feature alignment process: first, the domain shift leads to misaligned feature transfer and diminished classification accuracy; second, the intra-class feature dispersion and inter-class boundary blurring in few-shot tasks result in degraded classification performance for novel classes. Moreover, the impact of redundant and noisy data on model discriminability is rarely considered in existing approaches. To solve these issues, this article proposes a cross-domain FSL hyperspectral image classification method based on diffusion-augmented prototype knowledge distillation (DAPKD-CFSL). Firstly, we introduce a diffusion-augmented unsupervised domain adaptation pre-training (DA-PT) framework to address the domain shift by performing a domain-adversarial denoising and reconstruction task using visible source data and masked target data. Second, our dual-branch spatial-spectral attention (DB-SSA) captures global and local spectral-spatial dependencies to enhance feature representation. Then, the proposed global-local prototype knowledge distillation (GL-PKD) performs global prototype alignment while conducting local contrastive learning, addressing feature dispersion and boundary ambiguity. Finally, a dynamic learning strategy prioritizes feature alignment early and gradually strengthens classification supervision through adaptive loss weights, and incorporates an SNR-enhanced loss to effectively mitigate noise interference. The experimental results on three HSI datasets demonstrate the superiority and effectiveness of the proposed DAPKD-CFSL. Chen Ding 0002, Sirui Zheng, Mengmeng Zheng, Yizhou Dong, Wenqiang Hua, Wei Wei 0008, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | TAP-Track: Generalizable Spacecraft Pose Tracking by Tracking Any PointsabstractRecent learning-based spacecraft pose tracking methods have demonstrated impressive improvement in estimation accuracy and potential scalability to complex space environment. However, most of them still rely on the detection of discriminative keypoints on a known 3D model, limiting the generalization to unknown spacecraft. To this end, we propose, to the best of our knowledge, the first generalizable spacecraft pose tracking method. Instead of requiring a known model, we only assume the existence of at least one planar structure, e.g. solar panels, which holds for most satellites in general scenes. Additionally, the proposed method tracks any points on the plane across multi-frame followed by a re-projection error minimization, rather than detecting keypoints between image pairs, allowing robust capture of “long-term” temporal information among frames even for textureless surfaces without sufficient keypoints. Moreover, we also constructed the first large-scale dataset G-SPET for generalizable spacecraft pose estimation and tracking. It covers 174 satellites with diversity structures and rich annotations, increasing the number of targets in previous datasets by almost two orders of magnitude. Extensive evaluations on the proposed dataset have demonstrated the superiority of our method over state-of-the-art methods. The code and dataset will be made publicly available soon. Zhaoshuai Qi, Pulin Chen, Huilin Fan, Yu Zhu 0004, Jiaqi Yang 0002, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | IrregFusion: A Generalized Framework for Hyperspectral Image Fusion Across Diverse Spectral DataabstractFusing a low-resolution (LR) hyperspectral image (HSI) with a high-resolution (HR) multispectral image (MSI) has emerged as a promising strategy for reconstructing high-quality HSIs that combine rich spectral and fine spatial information. However, most existing HSI fusion methods operate under the restrictive assumption that the LR HSI and HR MSI are spatially aligned and fully consistent on the field-of-view (FoV), which significantly limits their applicability in real-world scenarios when such alignment is unavailable. To overcome these limitations, we propose IrregFusion, a generalized HSI fusion framework capable of handling both FoV-consistent and inconsistent fusion scenarios. Specifically, IrregFusion incorporates a Transformer-based reconstruction module that captures both intra- and inter-modal correlations between the diverse spectra data and the MSI, enhancing the model’s ability to perceive and reconstruct non-local spectral–spatial structures. To further address the challenges posed by FoV inconsistencies, we introduce a spectral propagation strategy that diffuses observed spectral information into adjacent spectral-blank regions, thereby easing the reconstruction of missing spectral content. Additionally, a self-supervised adaptation mechanism is integrated into the framework, enabling robust spectral–spatial representation learning and enhancing generalization across diverse and challenging conditions. Extensive experiments conducted on benchmark datasets demonstrate that IrregFusion effectively addresses the challenges of diverse spectral data fusion and consistently outperforms state-of-the-art methods in both reconstruction accuracy and visual fidelity. The source code will be released in https://github.com/JiangtaoNie/IrregFusion.git. Jiangtao Nie, Wei Wei 0008, Lei Zhang 0054, Chen Ding 0002, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | AdaSemiCD: An Adaptive Semi-Supervised Change Detection Method Based on Pseudo-Label EvaluationabstractChange detection (CD) is an essential field in remote sensing, with a primary focus on identifying areas of change in bitemporal image pairs captured at varying intervals of the same region. The data annotation process for CD tasks is both time-consuming and labor-intensive. To better utilize the scarce labeled data and abundant unlabeled data, we introduce an adaptive semi-supervised learning (SSL) method, AdaSemiCD, to improve pseudo-label usage and optimize the training process. Initially, due to the extreme class imbalance inherent in CD, the model is more inclined to focus on the background class, and it is easy to confuse the boundary of the target object. Considering these two points, we develop a measurable evaluation metric for pseudo-labels that enhances the representation of information entropy by class rebalancing and amplification of ambiguous areas, assigning greater weights to prospective change objects. Subsequently, to enhance the reliability of sample wise pseudo-labels, we introduce the AdaFusion module, to dynamically identify the most uncertain region and substitute it with more trustworthy content. Lastly, to ensure better training stability, we introduce the AdaEMA module, which updates the teacher model using only batches of trusted samples. Experimental results on ten public CD datasets validate the efficacy and generalizability of our proposed adaptive training framework. Lingyan Ran, Wen Dongcheng, Tao Zhuo, Shizhou Zhang, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | River Ice Fine-Grained Segmentation: A GF-2 Satellite Image Dataset and Deep Learning BenchmarkabstractSemantic segmentation of river ice image serves as a critical technological foundation for hydrological monitoring and ice flood early warning system. Current publicly available river ice datasets predominantly utilize UAV-captured image and ground-based photographic observations. To address the limitations of spatial coverage in existing datasets, we present NWPU_YRCC_GFICE - a satellite remote sensing dataset constructed from multi-spectral GF-2 satellite images. The dataset innovatively categorizes river ice into six fine-grained classes across freeze-thaw cycles and covers river ice data from Yellow River (Ningxia-Inner Mongolia section) spanning the past 10 years. We further establish a comprehensive deep learning benchmark, which evaluates 33 state-of-the-art segmentation models and two improved segmentation models based on YOLO and Segformer architecture, separately. Experiments are conducted on the NWPU_YRCC_GFICE dataset and three public river ice datasets (NWPU_YRCC_EX, NWPU_YRCC2, and Alberta river ice segmentation dataset). The proposed models exhibit excellent performance, surpassing the state-of-the-art methods. The presented NWPU_YRCC_GFICE dataset and benchmark enriches the river ice dataset and favors in promoting fine-grained river ice segmentation research from satellite view. Our dataset and code is available at https://github.com/ASGOLabMultisourceCooperationGroup/NWPU_YRCC_GFICE. Chenxu Wei, Haohao Zhou, Omirzhan Taukebayev, Wencong Wu, Amirkhan Temirbayev, Lingyan Ran, Hanlin Yin, Peng Wang 0015, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 14 |
| 2025 | Effective Road Segmentation With Selective State-Space Model and Frequency Feature CompensationabstractRoad segmentation from high-resolution remote sensing imagery is critical for tasks such as autonomous driving, urban planning, and geographic information systems. However, challenges such as intensity nonuniformity, pixel ambiguity, and the visual similarity between roads and natural features make accurate segmentation difficult. In this article, we propose a road segmentation framework built upon the Mamba architecture, integrating a novel frequency feature compensation (FFC) approach to improve segmentation performance. Specifically, we introduce a progressive FFC method, leveraging wavelet decomposition to capture fine-grained details by separating features into high- and low-frequency components. Multistage features extracted from the Mamba backbone are decomposed using this approach and progressively integrated to compensate for the essential details for accurate road segmentation. We also introduce a wavelet loss (WL) to improve the model’s ability to capture fine structural variations in the frequency domain. Furthermore, we develop a spatial perception Mamba block (SPMB) to enhance the capture of spatial relationships. By seamlessly integrating global context and local structures with the selective state-space model and FFC, our framework significantly boosts road segmentation accuracy. Extensive experiments on three publicly available road segmentation datasets demonstrate that our method achieves state-of-the-art performance, surpassing existing approaches in segmenting complex roads. Ting Liu 0012, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Generalizable Person Re-Identification From a 3D Perspective: Addressing Unpredictable Viewpoint ChangesabstractMost existing Domain Generalizable Person Re-identification (DG-ReID) methods focus on addressing style disparities between domains but often overlook the impact of unpredictable camera view changes, which we have identified as a significant factor responsible for poor generalization performance. To address this issue, we propose a novel approach from a 3D perspective, utilizing a customized 2D-to-3D reconstruction model to convert images captured from arbitrary camera views into canonical view images. However, merely applying a 3D reconstruction model in isolation may not result in improved DG-ReID performance, as reconstruction quality can be influenced by multiple factors, such as insufficient image resolution, extreme viewpoint, and environmental variations. These factors may lead to error accumulation and the loss of critical discriminative clues in the reconstructed results. To address this difficulty, we propose fusing the canonical view image with the original image using a transformer-based module. The transformer’s cross-attention mechanism is ideal for aligning and fusing the key semantic clues of the original image with the canonical view image, compensating for reconstruction errors. We demonstrate the effectiveness of our method through extensive experiments in various evaluation settings, achieving superior DG-ReID performance compared to existing approaches. Our approach addresses the impact of unpredictable camera view changes and provides a new perspective for designing DG-ReID methods. Bingliang Jiao, Lingqiao Liu, Liying Gao, Dapeng Oliver Wu, Guosheng Lin, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | Contrastive Neuron Pruning for Backdoor DefenseabstractRecent studies have revealed that deep neural networks (DNNs) are susceptible to backdoor attacks, in which attackers insert a pre-defined backdoor into a DNN model by poisoning a few training samples. A small subset of neurons in DNN is responsible for activating this backdoor and pruning these backdoor-associated neurons has been shown to mitigate the impact of such attacks. Current neuron pruning techniques often face challenges in accurately identifying these critical neurons, and they typically depend on the availability of labeled clean data, which is not always feasible. To address these challenges, we propose a novel defense strategy called Contrastive Neuron Pruning (CNP). This approach is based on the observation that poisoned samples tend to cluster together and are distinguishable from benign samples in the feature space of a backdoored model. Given a backdoored model, we initially apply a reversed trigger to benign samples, generating multiple positive (benign-benign) and negative (benign-poisoned) feature pairs from the backdoored model. We then employ contrastive learning on these pairs to improve the separation between benign and poisoned features. Subsequently, we identify and prune neurons in the Batch Normalization layers that show significant response differences to the generated pairs. By removing these backdoor-associated neurons, CNP effectively defends against backdoor attacks while requiring the pruning of only about 1% of the total neurons. Comprehensive experiments conducted on various benchmarks validate the efficacy of CNP, demonstrating its robustness and effectiveness in mitigating backdoor attacks compared to existing methods. Benteng Ma, Dongnan Liu, Yanning Zhang 0001, Tom Weidong Cai, Yong Xia 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Enhancing Feature Learning With Hard Samples in Mutual Learning for Online Class Incremental LearningabstractOnline Class-Incremental Learning (OCIL) aims to solve the problem of incrementally learning new classes from a non-i.i.d. and single-pass data stream. Compared to the offline setting, OCIL is much closer to a live learning experience requiring higher model update frequency at less computational budget. Due to its one-epoch training constraint, the model is likely to learn non-essential features and encounter the under-fitting issue, which severely affects the model's stability. In this paper, we investigate how to use hard samples to improve data variability, eventually enhancing feature learning and addressing the under-fitting problem. Specifically, by introducing a scoring function assessing the sample value, we build an OCIL formulation that simultaneously generates high-value samples and optimizes the OCIL model, improving generalization ability within the constraint of single-epoch training. Empirically, we found that strong data augmentation is a simple but effective way to generate a higher proportion of high-score samples. To make the most of these augmented samples, we design an OCIL model based on mutual learning with two networks of identical structures. Moreover, a collaborative learning mechanism is developed by aligning the features and class probabilities from the two networks to promote their interaction. Extensive experiments on three widely used datasets for OCIL have demonstrated the effectiveness of our method, obtaining superior performance to state-of-the-art methods. The code is available at https://github.com/susususushi/SDA-MCL. Guoqiang Liang 0001, Shibin Su, De Cheng, Shizhou Zhang, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | A Perception CNN for Facial Expression RecognitionabstractConvolutional neural networks (CNNs) can automatically learn data patterns to express face images for facial expression recognition (FER). However, they may ignore effect of facial segmentation of FER. In this paper, we propose a perception CNN for FER as well as PCNN. Firstly, PCNN can use five parallel networks to simultaneously learn local facial features based on eyes, cheeks and mouth to realize the sensitive capture of the subtle changes in FER. Secondly, we utilize a multi-domain interaction mechanism to register and fuse between local sense organ features and global facial structural features to better express face images for FER. Finally, we design a two-phase loss function to restrict accuracy of obtained sense information and reconstructed face images to guarantee performance of obtained PCNN in FER. Experimental results show that our PCNN achieves superior results on several lab and real-world FER benchmarks: CK+, JAFFE, FER2013, FERPlus, RAF-DB and Occlusion and Pose Variant Dataset. Its code is available at https://github.com/hellloxiaotian/PCNN. Chunwei Tian, Jingyuan Xie, Lingjun Li, Wangmeng Zuo, Yanning Zhang 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Prompt-Based Modality Alignment for Effective Multi-Modal Object Re-IdentificationabstractA critical challenge for multi-modal Object Re-Identification (ReID) is the effective aggregation of complementary information to mitigate illumination issues. State-of-the-art methods typically employ complex and highly-coupled architectures, which unavoidably result in heavy computational costs. Moreover, the significant distribution gap among different image spectra hinders the joint representation of multi-modal features. In this paper, we propose a framework named as PromptMA to establish effective communication channels between different modality paths, thereby aggregating modal complementary information and bridging the distribution gap. Specifically, we inject a series of learnable multi-modal prompts into the Image Encoder and introduce a prompt exchange mechanism to enable the prompts to alternately interact with different modal token embeddings, thus capturing and distributing multi-modal features effectively. Building on top of the multi-modal prompts, we further propose Prompt-based Token Selection (PBTS) and Prompt-based Modality Fusion (PBMF) modules to achieve effective multi-modal feature fusion while minimizing background interference. Additionally, due to the flexibility of our prompt exchange mechanism, our method is well-suited to handle scenarios with missing modalities. Extensive evaluations are conducted on four widely used benchmark datasets and the experimental results demonstrate that our method achieves state-of-the-art performances, surpassing the current benchmarks by over 15% on the challenging MSVR310 dataset and by 6% on the RGBNT201. The code is available at https://github.com/FHR-L/PromptMA. Shizhou Zhang, Wenlong Luo, De Cheng, Yinghui Xing, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 7 |
| 2025 | P2TC: A Lightweight Pyramid Pooling Transformer-CNN Network for Accurate 3D Whole Heart SegmentationabstractCardiovascular disease is a leading global cause of death, requiring accurate heart segmentation for diagnosis and surgical planning. Deep learning methods have been demonstrated to achieve superior performances in cardiac structures segmentation. However, there are still limitations in 3D whole heart segmentation, such as inadequate spatial context modeling, difficulty in capturing long-distance dependencies, high computational complexity, and limited representation of local high-level semantic information. To tackle the above problems, we propose a lightweight Pyramid Pooling Transformer-CNN (P2TC) network for accurate 3D whole heart segmentation. The proposed architecture comprises a dual encoder-decoder structure with a 3D pyramid pooling Transformer for multi-scale information fusion and a lightweight large-kernel Convolutional Neural Network (CNN) for local feature extraction. The decoder has two branches for precise segmentation and contextual residual handling. The first branch is used to generate segmentation masks for pixel-level classification based on the features extracted by the encoder to achieve accurate segmentation of cardiac structures. The second branch highlights contextual residuals across slices, enabling the network to better handle variations and boundaries. Extensive experimental results on the Multi-Modality Whole Heart Segmentation (MM-WHS) 2017 challenge dataset demonstrate that P2TC outperforms the most advanced methods, achieving the Dice scores of 92.6% and 88.1% in Computed Tomography (CT) and Magnetic Resonance Imaging (MRI) modalities respectively, which surpasses the baseline model by 1.5% and 1.7%, and achieves state-of-the-art segmentation results. Hengfei Cui, Yifan Wang 0033, Yan Li 0129, Yanning Zhang 0001, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Hierarchical Grafting Network With Structural Alignment for Ultra-High Resolution Image Segmentation
Ting Liu 0012, Shikui Wei, Yanning Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | MCINet: Fusing Low-Light Visible-Infrared Image via Max-Merge Complementary InformationabstractFusing complementary information in the visible-infrared image offers a promising approach to enhance the performance of downstream computer vision tasks (e.g., object detection, segmentation etc) in complicated imaging conditions (e.g., low-illumination). However, due to the robust imaging capacity of the infrared sensor in complicated imaging conditions, most existing methods primarily rely on the salient object intensity information in the infrared modality for fusion, while the visible information (e.g., color, texture etc) is not adequately utilized, and thus limit their generalization capacity in downstream computer vision tasks. In this study, we present a novel image fusion framework, i.e.,MCInet, which attempts toMaximize and merge theComplementaryInformation across visible-infrared modalities for more informative image fusion. To this end, we first introduce the modality-specific processing module into the fusion framework to improve the information representation of each modality image. For visible images, a pre-trained low-light enhance module is adopted to enhance its color and texture information. In addition, for infrared images, a nonlinear mapping module is constructed to suppress the excessive salient object intensity information of infrared modality. Then we establish a reusable MCI block that embeds a cross-image mutual information minimization scheme into an input-aware fusion module. This empowers us to dynamically maximize and merge the complementary information between two input images according to their feature representation. In addition, we introduce a cycle reconstruction loss to self-supervised regularize the fusion results for further enhancement. Experiments on image fusion, object detection, and segmentation tasks demonstrate that the proposed framework can produce more informative fusion results and exhibit better performance in downstream computer vision tasks. Jiangtao Nie, Boxiong Wu, Wei Wei 0008, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Frequency-Guided Spatial Adaptation for Camouflaged Object DetectionabstractCamouflaged object detection (COD) aims to segment camouflaged objects which exhibit very similar patterns with the surrounding environment. Recent research works have shown that enhancing the feature representation via the frequency information can greatly alleviate the ambiguity problem between the foreground objects and the background. With the emergence of vision foundation models, like InternImage, Segment Anything Model etc, adapting the pretrained model on COD tasks with a lightweight adapter module shows a novel and promising research direction. Existing adapter modules mainly care about the feature adaptation in the spatial domain. In this paper, we propose a novel frequency-guided spatial adaptation method for COD task. Specifically, we transform the input features of the adapter into frequency domain. By grouping and interacting with frequency components located within non overlapping circles in the spectrogram, different frequency components are dynamically enhanced or weakened, making the intensity of image details and contour features adaptively adjusted. At the same time, the features that are conducive to distinguishing object and background are highlighted, indirectly implying the position and shape of camouflaged object. We conduct extensive experiments on four widely adopted benchmark datasets and the proposed method outperforms 26 state-of-the-art methods with large margins. Code will be released. Shizhou Zhang, Dexuan Kong, Yinghui Xing, Yue Lu 0008, Lingyan Ran, Guoqiang Liang 0001, Hexu Wang, Yanning Zhang 0001 |
IEEE Trans. Multim. | 8 |
| 2024 | VadCLIP: Adapting Vision-Language Models for Weakly Supervised Video Anomaly DetectionabstractThe recent contrastive language-image pre-training (CLIP) model has shown great success in a wide range of image-level tasks, revealing remarkable ability for learning powerful visual representations with rich semantics. An open and worthwhile problem is efficiently adapting such a strong model to the video domain and designing a robust video anomaly detector. In this work, we propose VadCLIP, a new paradigm for weakly supervised video anomaly detection (WSVAD) by leveraging the frozen CLIP model directly without any pre-training and fine-tuning process. Unlike current works that directly feed extracted features into the weakly supervised classifier for frame-level binary classification, VadCLIP makes full use of fine-grained associations between vision and language on the strength of CLIP and involves dual branch. One branch simply utilizes visual features for coarse-grained binary classification, while the other fully leverages the fine-grained language-image alignment. With the benefit of dual branch, VadCLIP achieves both coarse-grained and fine-grained video anomaly detection by transferring pre-trained knowledge from CLIP to WSVAD task. We conduct extensive experiments on two commonly-used benchmarks, demonstrating that VadCLIP achieves the best performance on both coarse-grained and fine-grained WSVAD, surpassing the state-of-the-art methods by a large margin. Specifically, VadCLIP achieves 84.51% AP and 88.02% AUC on XD-Violence and UCF-Crime, respectively. Code and features are released at https://github.com/nwpu-zxr/VadCLIP. Peng Wu 0015, Xuerong Zhou, Guansong Pang, Lingru Zhou, Qingsen Yan, Peng Wang 0015, Yanning Zhang 0001 |
AAAI | 7 |
| 2024 | GSDD: Generative Space Dataset Distillation for Image Super-resolutionabstractSingle image super-resolution (SISR), especially in the real world, usually builds a large amount of LR-HR image pairs to learn representations that contain rich textural and structural information. However, relying on massive data for model training not only reduces training efficiency, but also causes heavy data storage burdens. In this paper, we attempt a pioneering study on dataset distillation (DD) for SISR problems to explore how data could be slimmed and compressed for the task. Unlike previous coreset selection methods which select a few typical examples directly from the original data, we remove the limitation that the selected data cannot be further edited, and propose to synthesize and optimize samples to preserve more task-useful representations. Concretely, by utilizing pre-trained GANs as a suitable approximation of realistic data distribution, we propose GSDD, which distills data in a latent generative space based on GAN-inversion techniques. By optimizing them to match with the practical data distribution in an informative feature space, the distilled data could then be synthesized. Experimental results demonstrate that when trained with our distilled data, GSDD can achieve comparable performance to the state-of-the-art (SOTA) SISR algorithms, while a nearly ×8 increase in training efficiency and a saving of almost 93.2% data storage space can be realized. Further experiments on challenging real-world data also demonstrate the promising generalization ability of GSDD. Shaolin Su, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001 |
AAAI | 5 |
| 2024 | Generating Content for HDR Deghosting from Frequency ViewabstractRecovering ghost-free High Dynamic Range (HDR) images from multiple Low Dynamic Range (LDR) images becomes challenging when the LDR images exhibit saturation and significant motion. Recent Diffusion Models (DMs) have been introduced in HDR imaging field, demonstrating promising performance, particularly in achieving visually perceptible results compared to previous DNN-based methods. However, DMs require extensive iterations with large models to estimate entire images, resulting in inefficiency that hinders their practical application. To address this challenge, we propose the Low-Frequency aware Diffusion (LF-Diff) model for ghost-free HDR imaging. The key idea of LF-Diff is implementing the DMs in a highly compacted latent space and integrating it into a regression-based model to enhance the details of reconstructed images. Specifically, as low-frequency information is closely related to human visual perception we propose to utilize DMs to create compact low-frequency priors for the reconstruction process. In addition, to take full advantage of the above low-frequency priors, the Dynamic HDR Reconstruction Network (DHRNet) is carried out in a regression-based manner to obtain final HDR images. Extensive experiments conducted on synthetic and real-world benchmark datasets demonstrate that our LF-Diff performs favorably against several state-of-the-art methods and is 10x faster than previous DM-based methods. Tao Hu 0013, Qingsen Yan, Yuankai Qi, Yanning Zhang 0001 |
CVPR | 4 |
| 2024 | GoMVS: Geometrically Consistent Cost Aggregation for Multi-View StereoabstractMatching cost aggregation plays a fundamental role in learning-based multi-view stereo networks. However, di-rectly aggregating adjacent costs can lead to suboptimal results due to local geometric inconsistency. Related meth-ods either seek selective aggregation or improve aggregated depth in the 2D space, both are unable to handle geomet-ric inconsistency in the cost volume effectively. In this pa-per, we propose GoMVS to aggregate geometrically consis-tent costs, yielding better utilization of adjacent geometries. More specifically, we correspond and propagate adjacent costs to the reference pixel by leveraging the local geomet-ric smoothness in conjunction with surface normals. We achieve this by the geometric consistent propagation (GCP) module. It computes the correspondence from the adjacent depth hypothesis space to the reference depth space using surface normals, then uses the correspondence to propa-gate adjacent costs to the reference geometry, followed by a convolution for aggregation. Our method achieves new state-of-the-art performance on DTU, Tanks & Temple, and ETH3D datasets. Notably, our method ranks 1st on the Tanks & Temple Advanced benchmark. Code is available at https://github.com/Wuuu3511IGoMVS. Rui Li 0013, Haofei Xu, Wenxun Zhao, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001 |
CVPR | 7 |
| 2024 | Open-Vocabulary Video Anomaly DetectionabstractCurrent video anomaly detection (VAD) approaches with weak supervisions are inherently limited to a closed-set setting and may struggle in open-world applications where there can be anomaly categories in the test data unseen during training. A few recent studies attempt to tackle a more realistic setting, open-set VAD, which aims to de-tect unseen anomalies given seen anomalies and normal videos. However, such a setting focuses on predicting frame anomaly scores, having no ability to recognize the specific categories of anomalies, despite the fact that this ability is essential for building more informed video surveillance systems. This paper takes a step further and explores open-vocabulary video anomaly detection (OVVAD), in which we aim to leverage pretrained large models to detect and cate-gorize seen and unseen anomalies. To this end, we propose a model that decouples OVVAD into two mutually comple-mentary tasks - class-agnostic detection and class-specific classification - and jointly optimizes both tasks. Particu-larly, we devise a semantic knowledge injection module to introduce semantic knowledge from large language models for the detection task, and design a novel anomaly synthesis module to generate pseudo unseen anomaly videos with the help of large vision generation models for the classification task. These semantic knowledge and synthesis anomalies substantially extend our model's capability in detecting and categorizing a variety of seen and unseen anomalies. Exten-sive experiments on three widely-used benchmarks demonstrate our model achieves state-of-the-art performance on OVVAD task. Peng Wu 0015, Xuerong Zhou, Guansong Pang, Jing Liu 0006, Peng Wang 0015, Yanning Zhang 0001 |
CVPR | 7 |
| 2024 | Rethinking and Improving Visual Prompt Selection for In-Context Learning Segmentation
Wei Suo, Lanqing Lai, Mengyang Sun, Hanwang Zhang, Peng Wang 0015, Yanning Zhang 0001 |
ECCV (46) | 6 |
| 2024 | 3D Single-Object Tracking in Point Clouds with High Temporal Variation
Qiao Wu, Kun Sun 0002, Pei An, Mathieu Salzmann, Yanning Zhang 0001, Jiaqi Yang 0002 |
ECCV (7) | 5 |
| 2024 | Cross-Platform Video Person ReID: A New Benchmark Dataset and Adaptation Approach
Shizhou Zhang, Wenlong Luo, De Cheng, Qingchun Yang, Lingyan Ran, Yinghui Xing, Yanning Zhang 0001 |
ECCV (27) | 7 |
| 2024 | Multiple Object Tracking Based on Occlusion-Aware Embedding Consistency LearningabstractThe Joint Detection and Embedding (JDE) framework has achieved remarkable progress for multiple object tracking. Existing methods often employ extracted embeddings to re-establish associations between new detections and previously disrupted tracks. However, the reliability of embeddings diminishes when the region of the occluded object frequently contains adjacent objects or clutters, especially in scenarios with severe occlusion. To alleviate this problem, we propose a novel multiple object tracking method based on visual embedding consistency, mainly including: 1) Occlusion Prediction Module (OPM) and 2) Occlusion-Aware Association Module (OAAM). The OPM predicts occlusion information for each true detection, facilitating the selection of valid samples for consistency learning of the track’s visual embedding. The OAAM leverages occlusion cues and visual embeddings to generate two separate embeddings for each track, guaranteeing consistency in both unoccluded and occluded detections. By integrating these two modules, our method is capable of addressing track interruptions caused by occlusion in online tracking scenarios. Extensive experimental results demonstrate that our approach achieves promising performance levels in both unoccluded and occluded tracking scenarios. Yaoqi Hu, Axi Niu, Yu Zhu 0004, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001 |
ICASSP | 6 |
| 2024 | Diffevent: Event Residual Diffusion for Image DeblurringabstractTraditional frame-based cameras inevitably suffer from non-uniform blur in real-world scenarios. Event cameras that record the intensity changes with high temporal resolution provide an effective solution for image deblurring. In this paper, we formulate the event-based image deblurring as an image generation problem by designing diffusion priors for the image and residual. Specifically, we propose an alternative diffusion sampling framework to jointly estimate clear and residual images to ensure the quality of the final result. In addition, to further enhance the subtle details, a pseudoinverse guidance module is leveraged to guide the prediction closer to the input with event data. Note that the proposed method can effectively handle the real unknown degradation without kernel estimation. The experiments on the benchmark event datasets demonstrate the effectiveness of our method. Jiumei He, Qingsen Yan, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001 |
ICASSP | 6 |
| 2024 | Complementary Fusion Network Based on Frequency Hybrid Attention for PansharpeningabstractPansharpening is a feasible way to obtain the high-resolution (HR) multispectral (MS) images by using panchromatic (PAN) images to sharpen low-resolution MS images. Despite its great advances, most existing pansharpening methods neglect the importance of integrating local and non-local characteristics of images, resulting in the imbalance of spatial and spectral distribution. In this paper, we propose a complementary fusion network (CFNet) based on frequency hybrid attention mechanism for pansharpening. By introducing the frequency transformation and the deformable cross-attention, our model takes image-wide receptive field into consideration to explore global feature learning. Combined with the convolutional layers with local receptive field, CFNet can well capture local and non-local features. Experimental results demonstrate that the proposed method outperforms the comparison methods in terms of visual and quantitative qualities. Yinghui Xing, Litao Qu, Kai Zhang 0010, Yan Zhang 0127, Xiuwei Zhang 0001, Yanning Zhang 0001 |
ICASSP | 6 |
| 2024 | Edge-Guided Detector-Free Network for Robust and Accurate Visible-Thermal Image MatchingabstractRecent detector-free models strive to leverage both local and global context for image matching, showcasing enhanced robustness, particularly in scenarios with weak-textured scenes. Despite these advancements, automatically establishing feature correspondences between visible and thermal images still introduces additional challenges. Differences in radiation and geometry between these modalities often result in degraded performance for the majority of existing methods. To this end, we propose edge-guided detector-free model termed EDMatcher for visible-thermal image matching. Besides local and global context in the images, EDMatcher also leverages modality-robust structural information in image edges, which demonstrates promising robustness to images with distinct modalities. Moreover, an edge-masked ground-truth matrix generation strategy is introduced during the training, which helps EDMatcher to further focus on more salient regions while leaving out texture-less regions, leading to more efficient learning. Extensive experiments show that EDMatcher has strong generalization and achieves excellent matching performances. Zhaoshuai Qi, Xiuwei Zhang 0001, Tao Zhuo, Yanning Zhang 0001 |
ICME | 6 |
| 2024 | Dual Supervised Contrastive Learning Based on Perturbation Uncertainty for Online Class Incremental Learning
Shibin Su, Zhaojie Chen, Guoqiang Liang 0001, Shizhou Zhang, Yanning Zhang 0001 |
ICPR (9) | 5 |
| 2024 | RGB-T Object Detection via Group Shuffled Multi-receptive Attention and Multi-modal Supervision
Jinzhong Wang, Xuetao Tian, Shun Dai, Tao Zhuo, Haorui Zeng, Hongjuan Liu, Xiuwei Zhang 0001, Yanning Zhang 0001 |
ICPR (17) | 9 |
| 2024 | Dual-Branch Task Residual Enhancement with Parameter-Free Attention for Zero-Shot Multi-label Image Recognition
Shizhou Zhang, Kairui Dang, De Cheng, Yinghui Xing, Qirui Wu, Dexuan Kong, Yanning Zhang 0001 |
ICPR (22) | 7 |
| 2024 | CR-SSRNet: Cross-Sensor Robust Spectral Super-Resolution Network Guided by Cognition FeaturesabstractSpectral Super-Resolution (SSR) aims at reconstructing a latent hyperspectral images (HSI) from a RGB image. Recent progress mainly focused on building a deep spectral super-resolution networkto directly map the input RGB image to the corresponding HSI. Their pleasing performance depends on the assumption that the spectral response function determined by the RGB sensor is consistent across training and test data. However, in practice, the training and test data are inevitably captured by different RGB sensors, thus resulting in obvious performance drop when using these networks. To mitigate this problem, we present a novel cognitive feature guided cross-sensor robust spectral super-resolution network. In a specific, a U-shape multi-scale network is first established to learn the deep mapping between input RGB image and the latent HSI. Then, a large-scale foundation cognitive model is introduced to extract multi-level cross-sensor invariant cognitive features from the input RGB. Moreover, these features are separately adapted and injected into different decoder blocks in the U-shape spectral super-resolution network. By doing these, the proposed network learns to appropriately guide the coarse-to-fine spectral reconstruction process using multilevel cognitive features, and thus shows better generalization performance in the cross-sensor SSR tasks. Experiments on two benchmark datasets demonstrate the superiority of the proposed method over several state-of-the-art baselines. Weixin Ren, Ruiling Liu, Lei Zhang 0054, Wei Wei 0008, Chen Ding 0002, Yanning Zhang 0001 |
IGARSS | 6 |
| 2024 | C3L: Content Correlated Vision-Language Instruction Tuning Data Generation via Contrastive Learning
Ji Ma 0008, Wei Suo, Peng Wang 0015, Yanning Zhang 0001 |
IJCAI | 4 |
| 2024 | Task-Adapter: Task-specific Adaptation of Image Models for Few-shot Action Recognition
Congqi Cao, Yueran Zhang, Yating Yu, Qinyi Lv, Lingtong Min, Yanning Zhang 0001 |
ACM Multimedia | 6 |
| 2024 | One-shot In-context Part SegmentationabstractIn this paper, we present the One-shot In-context Part Segmentation (OIParts) framework, designed to tackle the challenges of part segmentation by leveraging visual foundation models (VFMs). Existing training-based one-shot part segmentation methods that utilize VFMs encounter difficulties when faced with scenarios where the one-shot image and test image exhibit significant variance in appearance and perspective, or when the object in the test image is partially visible. We argue that training on the one-shot example often leads to overfitting, thereby compromising the model's generalization capability. Our framework offers a novel approach to part segmentation that is training-free, flexible, and data-efficient, requiring only a single in-context example for precise segmentation with superior generalization ability. By thoroughly exploring the complementary strengths of VFMs, specifically DINOv2 and Stable Diffusion, we introduce an adaptive channel selection approach by minimizing the intra-class distance for better exploiting these two features, thereby enhancing the discriminatory power of the extracted features for the fine-grained parts. We have achieved remarkable segmentation performance across diverse object categories. The OIParts framework not only eliminates the need for extensive labeled data but also demonstrates superior generalization ability. Through comprehensive experimentation on three benchmark datasets, we have demonstrated the superiority of our proposed method over existing part segmentation approaches in one-shot settings. Zhenqi Dai, Ting Liu 0012, Xingxing Zhang 0001, Yunchao Wei, Yanning Zhang 0001 |
ACM Multimedia | 5 |
| 2024 | Sustainable Self-evolution Adversarial TrainingabstractWith the wide application of deep neural network models in various computer vision tasks, there has been a proliferation of adversarial example generation strategies aimed at deeply exploring model security. However, existing adversarial training defense models, which rely on single or limited types of attacks under a one-time learning process, struggle to adapt to the dynamic and evolving nature of attack methods. Therefore, to achieve defense performance improvements for models in long-term applications, we propose a novel Sustainable Self-Evolution Adversarial Training (SSEAT) framework. Specifically, we introduce a continual adversarial defense pipeline to realize learning from various kinds of adversarial examples across multiple stages. Additionally, to address the issue of model catastrophic forgetting caused by continual learning from ongoing novel attacks, we propose an adversarial data replay module to better select more diverse and key relearning data. Furthermore, we design a consistency regularization strategy to encourage current defense models to learn more from previously trained ones, guiding them to retain more past knowledge and maintain accuracy on clean samples. Extensive experiments have been conducted to verify the efficacy of the proposed SSEAT defense method, which demonstrates superior defense performance and classification accuracy compared to competitors. Wenxuan Wang 0003, Chenglei Wang, Huihui Qi, Menghao Ye, Xuelin Qian, Peng Wang 0015, Yanning Zhang 0001 |
ACM Multimedia | 7 |
| 2024 | Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal PromptsabstractCurrent weakly supervised video anomaly detection (WSVAD) task aims to achieve frame-level anomalous event detection with only coarse video-level annotations available. Existing works typically involve extracting global features from full-resolution video frames and training frame-level classifiers to detect anomalies in the temporal dimension. However, most anomalous events tend to occur in localized spatial regions rather than the entire video frames, which implies existing frame-level feature based works may be misled by the dominant background information and lack the interpretation of the detected anomalies. To address this dilemma, this paper introduces a novel method called STPrompt that learns spatio-temporal prompt embeddings for weakly supervised video anomaly detection and localization (WSVADL) based on pre-trained vision-language models (VLMs). Our proposed method employs a two-stream network structure, with one stream focusing on the temporal dimension and the other primarily on the spatial dimension. By leveraging the learned knowledge from pre-trained VLMs and incorporating natural motion priors from raw videos, our model learns prompt embeddings that are aligned with spatio-temporal regions of videos (e.g., patches of individual frames) for identify specific local regions of anomalies, enabling accurate video anomaly detection while mitigating the influence of background information. Without relying on detailed spatio-temporal annotations or auxiliary object detection/tracking, our method achieves state-of-the-art performance on three public benchmarks for the WSVADL task. Peng Wu 0015, Xuerong Zhou, Guansong Pang, Zhiwei Yang 0013, Qingsen Yan, Peng Wang 0015, Yanning Zhang 0001 |
ACM Multimedia | 7 |
| 2024 | A Plug-and-Play Method for Rare Human-Object Interactions Detection by Bridging Domain GapabstractHuman-object interactions (HOI) detection aims at capturing human-object pairs in images and corresponding actions. It is an important step toward high-level visual reasoning and scene understanding. However, due to the natural bias from the real world, existing methods mostly struggle with rare human-object pairs and lead to sub-optimal results. Recently, with the development of the generative model, a straightforward approach is to construct a more balanced dataset based on a group of supplementary samples. Unfortunately, there is a significant domain gap between the generated data and the original data, and simply merging the generated images into the original dataset cannot significantly boost the performance. To alleviate the above problem, we present a novel model-agnostic framework called Context-Enhanced Feature Alignment (CEFA) module, which can effectively align the generated data with the original data at the feature level and bridge the domain gap. Specifically, CEFA consists of a feature alignment module and a context enhancement module. On one hand, considering the crucial role of human-object pairs information in HOI tasks, the feature alignment module aligns the human-object pairs by aggregating instance information. On the other hand, to mitigate the issue of losing important context information caused by the traditional discriminator-style alignment method, we employ a context-enhanced image reconstruction module to improve the model's learning ability of contextual cues. Extensive experiments have shown that our method can serve as a plug-and-play module to improve the detection performance of HOI models on rare categories. Wei Suo, Peng Wang 0015, Yanning Zhang 0001 |
ACM Multimedia | 4 |
| 2024 | Visual Prompt Tuning in Null Space for Continual LearningabstractExisting prompt-tuning methods have demonstrated impressive performances in continual learning (CL), by selecting and updating relevant prompts in the vision-transformer models. On the contrary, this paper aims to learn each task by tuning the prompts in the direction orthogonal to the subspace spanned by previous tasks' features, so as to ensure no interference on tasks that have been learned to overcome catastrophic forgetting in CL. However, different from the orthogonal projection in the traditional CNN architecture, the prompt gradient orthogonal projection in the ViT architecture shows completely different and greater challenges, i.e., 1) the high-order and non-linear self-attention operation; 2) the drift of prompt distribution brought by the LayerNorm in the transformer block. Theoretically, we have finally deduced two consistency conditions to achieve the prompt gradient orthogonal projection, which provide a theoretical guarantee of eliminating interference on previously learned knowledge via the self-attention mechanism in visual prompt tuning. In practice, an effective null-space-based approximation solution has been proposed to implement the prompt gradient orthogonal projection. Extensive experimental results demonstrate the effectiveness of anti-forgetting on four class-incremental benchmarks with diverse pre-trained baseline models, and our approach achieves superior performances to state-of-the-art methods. Our code is available at https://github.com/zugexiaodui/VPTinNSforCL Yue Lu 0008, Shizhou Zhang, De Cheng, Yinghui Xing, Nannan Wang 0001, Peng Wang 0015, Yanning Zhang 0001 |
NeurIPS | 7 |
| 2024 | Meta-Exploiting Frequency Prior for Cross-Domain Few-Shot LearningabstractMeta-learning offers a promising avenue for few-shot learning (FSL), enabling models to glean a generalizable feature embedding through episodic training on synthetic FSL tasks in a source domain. Yet, in practical scenarios where the target task diverges from that in the source domain, meta-learning based method is susceptible to over-fitting. To overcome this, we introduce a novel framework, Meta-Exploiting Frequency Prior for Cross-Domain Few-Shot Learning, which is crafted to comprehensively exploit the cross-domain transferable image prior that each image can be decomposed into complementary low-frequency content details and high-frequency robust structural characteristics. Motivated by this insight, we propose to decompose each query image into its high-frequency and low-frequency components, and parallel incorporate them into the feature embedding network to enhance the final category prediction. More importantly, we introduce a feature reconstruction prior and a prediction consistency prior to separately encourage the consistency of the intermediate feature as well as the final category prediction between the original query image and its decomposed frequency components. This allows for collectively guiding the network's meta-learning process with the aim of learning generalizable image feature embeddings, while not introducing any extra computational cost in the inference phase. Our framework establishes new state-of-the-art results on multiple cross-domain few-shot learning benchmarks. Fei Zhou 0008, Peng Wang 0023, Lei Zhang 0054, Zhenghua Chen, Wei Wei 0008, Chen Ding 0002, Guosheng Lin, Yanning Zhang 0001 |
NeurIPS | 8 |
| 2024 | Take a prior from other tasks for severe blur removal
Yu Zhu 0004, Danna Xue, Qingsen Yan, Jinqiu Sun, Sung-Eui Yoon, Yanning Zhang 0001 |
Comput. Vis. Image Underst. | 7 |
| 2024 | Image deraining via invertible disentangled representations
Xueling Chen, Wei Sun 0036, Yanning Zhang 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | An Adaptive Correlation Filtering Method for Text-Based Person Search
Mengyang Sun, Wei Suo, Peng Wang 0015, Kai Niu 0002, Le Liu 0008, Guosheng Lin, Yanning Zhang 0001, Qi Wu 0001 |
Int. J. Comput. Vis. | 7 |
| 2024 | How does Layer Normalization improve Batch Normalization in self-supervised sound source localization?
Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001 |
Neurocomputing | 6 |
| 2024 | FDE-Net: A memory-efficiency densely connected network inspired from fractional-order differential equations for single image super-resolution
Xiao Zhang 0058, Lei Zhang 0054, Wei Wei 0008, Chunna Tian, Yanning Zhang 0001 |
Neurocomputing | 6 |
| 2024 | Dynamic center point learning for multiple object tracking under Severe occlusions
Yaoqi Hu, Axi Niu, Jinqiu Sun, Yu Zhu 0004, Qingsen Yan, Wei Dong 0010, Marcin Wozniak, Yanning Zhang 0001 |
Knowl. Based Syst. | 8 |
| 2024 | DPGS: Cross-cooperation guided dynamic points generation for scene text spotting
Wei Sun 0036, Qianzhou Wang, Xueling Chen, Qingsen Yan, Yanning Zhang 0001 |
Knowl. Based Syst. | 6 |
| 2024 | GRAN: ghost residual attention network for single image super resolution
Axi Niu, Yu Zhu 0004, Jinqiu Sun, Qingsen Yan, Yanning Zhang 0001 |
Multim. Tools Appl. | 6 |
| 2024 | Scale-aware local difference attention on pyramidal features for crowd counting
Qian Zhang 0046, Shizhou Zhang, Xinyao Liu, Yanning Zhang 0001 |
Multim. Tools Appl. | 4 |
| 2024 | Unsupervised Test-Time Adaptation Learning for Effective Hyperspectral Image Super-Resolution With Unknown DegenerationabstractFusing a low-resolution hyperspectral image (HSI) with a high-resolution (HR) multi-spectral image has provided an effective way for HSI super-resolution (SR). The key lies on inferring the posteriori of the latent (i.e., HR) HSI using an appropriate image prior and the likelihood determined by the degeneration between the latent HSI and the observed images. However, in scenarios with complex imaging environments and various imaging scenes, the prior of HSIs can be prohibitively complicated and the degeneration is often unknown, which causes it difficult to accurately infer the posteriori of each latent HSI. To tackle this problem, we present an unsupervised test-time adaptation learning (UTAL) framework for HSI SR under unknown degeneration. Instead of directly modeling the complicated image prior, it first implicitly learns a content-agnostic prior shared across different images through supervisedly pre-training a mutual-guiding fusion module on extensive synthetic data. Then, it adapts the shared prior to those private characteristics in the latent HSI for posteriori inference through unsupervisedly learning a self-guiding adaptation module and a degeneration estimation network on two observed images in the test phase. Such a two-stage learning scheme models the complicated image prior in a divide-and-conquer manner, which eases the modeling difficulty and improves the prior accuracy. Moreover, the unknown degeneration can be estimated properly. Both of these two advantages empower us to accurately infer the posteriori of the latent HSI, thereby increasing the generalization performance in real applications. Additionally, in order to further mitigate the over-fitting in coping with more challenging cases (e.g., degenerations in both spectral and spatial domains are unknown) and speed up, we propose to meta-train UTAL on extensive synthetic SR tasks and solve it using an alternative optimization strategy such that UTAL learns to produce good generalization performance in real challenging cases with a small number of gradient descent steps. To verify the efficacy of UTAL, we evaluate it on HSI SR tasks with different unknown degenerations as well as some other HSI restoration tasks (e.g., compressive sensing), and report strong results superior to that of existing competitors. Lei Zhang 0054, Jiangtao Nie, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Mutual Voting for Ranking 3D CorrespondencesabstractConsistent correspondences between point clouds are vital to 3D vision tasks such as registration and recognition. In this paper, we present a mutual voting method for ranking 3D correspondences. The key insight is to achieve reliable scoring results for correspondences by refining both voters and candidates in a mutual voting scheme. First, a graph is constructed for the initial correspondence set with the pairwise compatibility constraint. Second, nodal clustering coefficients are introduced to preliminarily remove a portion of outliers and speed up the following voting process. Third, we model nodes and edges in the graph as candidates and voters, respectively. Mutual voting is then performed in the graph to score correspondences. Finally, the correspondences are ranked based on the voting scores and top-ranked ones are identified as inliers. Feature matching, 3D point cloud registration, and 3D object recognition experiments on various datasets with different nuisances and modalities verify that MV is robust to heavy outliers under different challenging settings, and can significantly boost 3D point cloud registration and 3D object recognition performance. Jiaqi Yang 0002, Xiyu Zhang 0001, Shichao Fan, Chunlin Ren, Yanning Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | MAC: Maximal Cliques for 3D RegistrationabstractThis paper presents a 3D registration method with maximal cliques (MAC) for 3D point cloud registration (PCR). The key insight is to loosen the previous maximum clique constraint and mine more local consensus information in a graph for accurate pose hypotheses generation: 1) A compatibility graph is constructed to render the affinity relationship between initial correspondences. 2) We search for maximal cliques in the graph, each representing a consensus set. 3) Transformation hypotheses are computed for the selected cliques by the SVD algorithm and the best hypothesis is used to perform registration. In addition, we present a variant of MAC if given overlap prior, called MAC-OP. Overlap prior further enhances MAC from many technical aspects, such as graph construction with re-weighted nodes, hypotheses generation from cliques with additional constraints, and hypothesis evaluation with overlap-aware weights. Extensive experiments demonstrate that both MAC and MAC-OP effectively increase registration recall, outperform various state-of-the-art methods, and boost the performance of deep-learned methods. For instance, MAC combined with GeoTransformer achieves a state-of-the-art registration recall of [Formula: see text] on 3DMatch / 3DLoMatch. We perform synthetic experiments on 3DMatch-LIR / 3DLoMatch-LIR, a dataset with extremely low inlier ratios for 3D registration in ultra-challenging cases. Jiaqi Yang 0002, Xiyu Zhang 0001, Peng Wang 0015, Yulan Guo, Kun Sun 0002, Qiao Wu, Shikun Zhang, Yanning Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | Momentum recursive DARTS
Benteng Ma, Yanning Zhang 0001, Yong Xia 0001 |
Pattern Recognit. | 2 |
| 2024 | Closed-loop unified knowledge distillation for dense object detection
Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001 |
Pattern Recognit. | 6 |
| 2024 | KGSR: A kernel guided network for real-world blind super-resolution
Qingsen Yan, Axi Niu, Wei Dong 0010, Marcin Wozniak, Yanning Zhang 0001 |
Pattern Recognit. | 6 |
| 2024 | Uncertainty estimation in HDR imaging with Bayesian neural networks
Qingsen Yan, Haishen Wang, Yuhang Liu 0002, Wei Dong 0010, Marcin Wozniak, Yanning Zhang 0001 |
Pattern Recognit. | 7 |
| 2024 | Meta-collaborative comparison for effective cross-domain few-shot learning
Fei Zhou 0008, Peng Wang 0023, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
Pattern Recognit. | 5 |
| 2024 | Palette-Based Color Harmonization via Color NamingabstractColor harmony refers to combinations of colors that look pleasing together. We present a novel strategy to harmonize an image's colors using color-palette manipulation and color naming. Palette-based color manipulation is a method that extracts a few colors to represent the image. Modifying the palette colors modifies the color appearance of the image. A color-naming model is a mechanism to categorize colors into a fixed number of basic color terms. Working from a color-naming model, we derive a set ofprototype colorsand demonstrate that mapping an image's extracted color palette to the nearest prototype colors effectively harmonizes the image's colors. This straightforward approach yields visually compelling, outperforming more complex color harmony methods. Danna Xue, Javier Vazquez-Corral, Luis Herranz, Yanning Zhang 0001, Michael S. Brown |
IEEE Signal Process. Lett. | 4 |
| 2024 | VS-TransGRU: A Novel Transformer-GRU-Based Framework Enhanced by Visual-Semantic Fusion for Egocentric Action AnticipationabstractEgocentric action anticipation is a challenging task that aims to make advanced predictions of future actions from current and historical observations in the first-person view. Most existing methods focus on improving the model architecture and loss function based on the visual input and recurrent neural network to boost the anticipation performance. However, these methods, which merely consider visual information and rely on a single network architecture, gradually reach a performance plateau. In order to fully understand what has been observed and capture the dependencies between current observations and future actions well enough, we propose a novel visual-semantic fusion enhanced and Transformer-GRU-based action anticipation framework in this paper. Firstly, high-level semantic information is introduced to improve the performance of action anticipation for the first time. We propose to use the semantic features generated based on the class labels or directly from the visual observations to augment the original visual features. Secondly, to take advantage of both the parallel and autoregressive models, we design a Transformer-based encoder for long-term sequential modeling and a GRU-based decoder for flexible iteration decoding. This hybrid architecture allows for better performance with fewer parameters and computations. Thirdly, an effective visual-semantic fusion module is proposed to make up for the semantic gap and fully utilize the complementarity of different modalities. Extensive experiments on two large-scale first-person view datasets and two third-person datasets validate the effectiveness of our proposed method, which achieves new state-of-the-art performance, outperforming previous approaches by a large margin. The code will be released after acceptance athttps://github.com/sunze992/VS-TransGRU. Congqi Cao, Ze Sun, Qinyi Lv, Lingtong Min, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Co-Occurrence Matters: Learning Action Relation for Temporal Action LocalizationabstractTemporal action localization (TAL) is a prevailing task due to its great application potential. Existing works in this field mainly suffer from two weaknesses: (1) They often neglect the multi-label case and only focus on temporal modeling. (2) They ignore the semantic information in class labels and only use the visual information. To solve these problems, we propose a novel Co-Occurrence Relation Module (CORM) that explicitly models the co-occurrence relationship between actions. Besides the visual information, it further utilizes the semantic embeddings of class labels to model the co-occurrence relationship. The CORM works in a plug-and-play manner and can be easily incorporated with the existing sequence models. By considering both visual and semantic co-occurrence, our method achieves high multi-label relationship modeling capacity. Meanwhile, existing datasets in TAL always focus on low-semantic atomic actions. Thus we construct a challenging multi-label dataset UCF-Crime-TAL that focuses on high-semantic actions by annotating the UCF-Crime dataset at frame level and considering the semantic overlap of different events. Extensive experiments on two commonly used TAL datasets, i.e., MultiTHUMOS and TSU, and our newly proposed UCF-Crime-TAL demenstrate the effectiveness of the proposed CORM, which achieves state-of-the-art performance on these datasets. Congqi Cao, Yueran Zhang, Yue Lu 0008, Xin Zhang 0168, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Contrastive Pedestrian Attentive and Correlation Learning Network for Occluded Person Re-IdentificationabstractOccluded person Re-identification (ReID) aims to match occluded and holistic pedestrian images across different camera views. This task presents two primary challenges. First, it is crucial to accurately capture pedestrian foregrounds from seriously occluded person images. Second, a noticeable information asymmetry exists between the partial body in occluded images and the complete body in corresponding holistic images, which could cause the ReID model to underestimate their similarities. To address these challenges, we introduce a contrastive pedestrian attentive and correlation learning (CpaCol) model. Within CpaCol, we first design a Contrastive Pedestrian Attention (ContrastAttn) module to capture pedestrian foregrounds from occluded images. In this process, we notice that most existing attention-based methods only supervise the final predictions with identity loss yet neglect its causality with the generated attention maps, which could mislead the model to capture some salient yet pedestrian-irrelevant noises as discriminative clues. To rectify this, we integrate contrastive learning into our ContrastAttn module to guide it to learn the semantic divergence between pedestrian foregrounds and noises, thereby capturing pedestrian foregrounds more accurately. Besides, we propose a correlation learning module, where we tailor an effective dense feature correlation learning tool, 4D convolution, to enable it to adapt to pedestrian images and capture corresponding clues between comparing images. By focusing more on corresponding clues, our model could avoid overemphasizing the inherent information asymmetry between occluded and holistic images, thereby improving re-identification. Empowered by these modules, our CpaCol achieves state-of-the-art performance on three relevant ReID settings,i.e., occluded, partial, and holistic ReID. Our code is available in https://github.com/nwpugaoliying/CpaCol. Liying Gao, Bingliang Jiao, Yuzhou Long, Kai Niu 0002, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Vehicle Re-Identification in Aerial Images and Videos: Dataset and ApproachabstractIn this work, we propose a large-scale dataset, VRAI, and an effective Orientation Adaptive and Salience Attentive (OASA) Network for vehicle re-identification (ReID) in aerial imagery. The VRAI dataset includes two subsets: VRAI-Image, which contains over 137,000 images of 13,000 vehicle instances, and VRAI-Video, which comprises more than 14,000 video trajectories of 7,000 identities. To our best knowledge, this is the largest dataset for UAV-based vehicle ReID, and the first dataset proposed for video-based ReID under UAV views. Based on the VRAI dataset, we design an OASA network to address two crucial challenges of vehicle ReID in aerial imagery. Firstly, the significant vehicle orientation variations in aerial images could cause great vehicle pattern deformations, making it difficult to identify vehicles across UAV views. To overcome this challenge, in our OASA, an orientation adaptive dynamic convolution module is designed, which constructs customized kernels for each vehicle instance to extract orientation-invariant features. Besides, the unique vertical view and long focal length of the UAV platform often render many salient vehicle attributes, such as logos and license plates, invisible, which brings a great challenge to ReID models to extract distinguishable vehicle features. To address this issue, in the OASA, we design a transformer-based salience attentive module (Trans-Attn) that guides the model to focus on subtle yet discriminative clues of vehicle instances in aerial imagery. Through extensive experiments, both of our designed modules are verified effective. Besides, our OASA model outperforms state-of-the-art algorithms both on our VRAI dataset and other surveillance-based datasets. Our VRAI dataset is available in https://github.com/JiaoBL1234/VRAI-Dataset. Bingliang Jiao, Lu Yang 0016, Liying Gao, Peng Wang 0015, Shizhou Zhang, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | New Insights on Relieving Task-Recency Bias for Online Class Incremental LearningabstractTo imitate the ability of keeping learning of human, continual learning which can learn from a never-ending data stream has attracted more interests recently. In all settings, the online class incremental learning (OCIL), where incoming samples from data stream can be used only once, is more challenging and can be encountered more frequently in real world. Actually, all continual learning models face a stability-plasticity dilemma, where the stability means the ability to preserve old knowledge while the plasticity denotes the ability to incorporate new knowledge. Although replay-based methods have shown exceptional promise, most of them concentrate on the strategy for updating and retrieving memory to keep stability at the expense of plasticity. To strike a preferable trade-off between stability and plasticity, we propose an Adaptive Focus Shifting algorithm (AFS), which dynamically adjusts focus to ambiguous samples and non-target logits in model learning. Through a deep analysis of the task-recency bias caused by class imbalance, we propose a revised focal loss to mainly keep stability. By utilizing a new weight function, the revised focal loss will pay more attention to current ambiguous samples, which are the potentially valuable samples to make model progress quickly. To promote plasticity, we introduce a virtual knowledge distillation. By designing a virtual teacher, it assigns more attention to non-target classes, which can surmount overconfidence and encourage model to focus on inter-class information. Extensive experiments on three popular datasets for OCIL have shown the effectiveness of AFS. The code will be available at https://github.com/czjghost/AFS. Guoqiang Liang 0001, Zhaojie Chen, Zhaoqiang Chen, Shiyu Ji, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | An Overview of Text-Based Person Search: Recent Advances and Future DirectionsabstractDue to the practical significance in smart video surveillance systems, Text-Based Person Search (TBPS) has been one of the research hotspots recently, which refers to searching for the interested pedestrian images given natural language sentences. To help researchers quickly grasp the developments of this important task, we comprehensively summarize the recent research advances of TBPS from two perspectives,i.e., Feature Extraction (FE) and Semantic Alignments (SA). Specifically, the FE mainly consists of pre-processing approaches and end-to-end frameworks, and the SA could be briefly divided into cross-modal attention mechanism, non-attention alignments, training objectives, and generative approaches. Afterwards, we elaborate four widely-used benchmarks and also the evaluation criterion for TBPS. And comparisons and analyses among the state-of-the-art (SOTA) solutions are provided based on these large-scale benchmarks. At last, we point out some future research directions that need to be further addressed, which will greatly facilitate the practical applications of TBPS. Kai Niu 0002, Yanyi Liu, Yuzhou Long, Yan Huang 0008, Liang Wang 0001, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | A Minimal Solution for Sphere-Based Camera-Projector Pair CalibrationabstractWe propose a minimal solution for sphere-based camera-projector pair (CPP) calibration. Previous works often treated the camera and projector calibration as two independent problems, which exploit only intra-view information from geometric properties of sphere dual image formation and hence require at least three spheres for CPP calibration. However, other than intra-view information, we observe that inter-view information between camera and projector provides additional constraints. Combining these two kinds of information yields a minimal solution for CPP calibration, where only a single sphere is required. Extensive experiments have verified the effectiveness of proposed minimal solver, which demonstrates higher flexibility and comparable accuracy to the state-of-the-art methods. Moreover, the achieved flexibility allows high-quality 3D reconstruction with an uncalibrated CPP, given only a single sphere in the scene. Zhaoshuai Qi, Jingqi Pang, Yifeng Hao, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Non-Exemplar Class-Incremental Learning by Random Auxiliary Classes Augmentation and Mixed FeaturesabstractNon-exemplar class-incremental learning refers to continual classifying of new and old classes without storing samples of old classes. Since only new class samples are available, catastrophic forgetting of old knowledge often occurs. In this paper, we propose an effective non-exemplar method called RAMF consisting of Random Auxiliary classes augmentation and Mixed Features. On the one hand, we design a novel random auxiliary classes augmentation method, where one augmentation is randomly selected from three augmentations and applied to inputs to generate augmented samples and extra class labels. By extending the data and label space, the model can learn more diverse and transferable representations, which can prevent the model from being biased towards learning task-specific features and facilitate the transfer among different tasks. In a word, when learning new tasks, the random auxiliary class augmentation will reduce the change of feature space and improve model generalization. On the other hand, we propose to replace the new features with mixed features for model optimization since only using new features will largely affect the previous representation embedded in the old feature space. Instead, by mixing new and old features, the cosine similarity is improved by reducing the angle between the current and old features, which allows for better stability over long-term incremental learning without increasing the computational complexity. We have conducted extensive experiments on three benchmarks CIFAR-100, Tiny-ImageNet and ImageNet-Subset, where our method outperforms the state-of-the-art non-exemplar methods and is comparable to high-performance replay-based methods. Guoqiang Liang 0001, Zhaojie Chen, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | A Self-Supervised CNN for Image Watermark RemovalabstractPopular convolutional neural networks mainly use paired images in a supervised way for image watermark removal. However, watermarked images do not have reference images in the real world, which results in poor robustness of image watermark removal techniques. In this paper, we propose a self-supervised convolutional neural network (CNN) in image watermark removal (SWCNN). SWCNN uses a self-supervised way to construct reference watermarked images rather than given paired training samples, according to watermark distribution. A heterogeneous U-Net architecture is used to extract more complementary structural information via simple components for image watermark removal. Taking into account texture information, a mixed loss is exploited to improve visual effects of image watermark removal. Besides, a watermark dataset is conducted. Experimental results show that the proposed SWCNN is superior to popular CNNs in image watermark removal. Chunwei Tian, Menghua Zheng, Tiancai Jiao, Wangmeng Zuo, Yanning Zhang 0001, Chia-Wen Lin |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Perceptive Self-Supervised Learning Network for Noisy Image Watermark RemovalabstractPopular methods usually use a degradation model in a supervised way to learn a watermark removal model. However, it is true that reference images are difficult to obtain in the real world, as well as collected images by cameras suffer from noise. To overcome these drawbacks, we propose a perceptive self-supervised learning network for noisy image watermark removal (PSLNet) in this paper. PSLNet depends on a parallel network to remove noise and watermarks. The upper network uses task decomposition ideas to remove noise and watermarks in sequence. The lower network utilizes the degradation model idea to simultaneously remove noise and watermarks. Specifically, mentioned paired watermark images are obtained in a self-supervised way, and paired noisy images (i.e., noisy and reference images) are obtained in a supervised way. To enhance the clarity of obtained images, interacting two sub-networks and fusing obtained clean images are used to improve the effects of image watermark removal in terms of structural information and pixel enhancement. Taking into texture information account, a mixed loss uses obtained images and features to achieve a robust model of noisy image watermark removal. Comprehensive experiments show that our proposed method is very effective in comparison with popular convolutional neural networks (CNNs) for noisy image watermark removal. Codes can be obtained at https://github.com/hellloxiaotian/PSLNet. Chunwei Tian, Menghua Zheng, Bo Li 0004, Yanning Zhang 0001, Shichao Zhang 0001, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Fine-Granularity Alignment for Text-Based Person Retrieval Via Semantics-Centric Visual DivisionabstractText-based Person Retrieval aims to search the target pedestrian image from video surveillance or a large image database with a text description. Previous works have recognized the significance of mining local information in images and descriptions and performing fine-grained alignment. These approaches adopt hard division or auxiliary networks for locating local visual regions. However, the two existing ways are not flexible enough for various images and may even bring noise. Meanwhile, the Vision-Language Pre-training models like CLIP exhibit strong generalization and zero-shot abilities, which provide an available way to this issue. In this paper, we propose a novel Fine-Granularity Alignment model with Semantics-Centric Visual Division (SCVD). Our method contains a Semantics Deconstructor (SD), a Cross-modal Guided Interaction (CGI) module, and a Dynamic Focus Alignment (DFA) module. The SD aims to extract fine-grained semantic prompts from the raw description which is easy-understand for CLIP. In CGI, we propose a Text-Guided Visual Localization (TVL) module to generate local visual representations according to the semantic prompts and a Vision-Guided Semantics Reconstruction (VSR) module to integrate the prompts into the textual representation. The DFA is used finally to align vision-text fine-grained information. The extensive experiments demonstrate that our proposed framework significantly outperforms current state-of-the-art methods in terms of Rank@1 metric on three benchmarks by an absolute gain of 6.56%, 8.93%, and 11.53%, respectively. Our code is available in https://github.com/tujun233/SCVD.git. Zhimin Wei, Peng Wu 0015, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Adjustable Visible and Infrared Image FusionabstractThe visible and infrared image fusion (VIF) method aims to utilize the complementary information between these two modalities to synthesize a new image containing richer information. Although it has been extensively studied, the synthesized image that has the best visual results is difficult to reach consensus since users have different opinions. To address this problem, we propose an adjustable VIF framework termed AdjFusion, which introduces a global controlling coefficient into VIF to enforce it can interact with users. Within AdjFusion, a semantic-aware modulation module is proposed to transform the global controlling coefficient into a semantic-aware controlling coefficient, which provides pixel-wise guidance for AdjFusion considering both interactivity and semantic information within visible and infrared images. In addition, the introduced global controlling coefficient not only can be utilized as an external interface for interaction with users but also can be easily customized by the downstream tasks (e.g., VIF-based detection and segmentation), which can help to select the best fusion result for the downstream tasks. Taking advantage of this, we further propose a lightweight adaptation module for AdjFusion to learn the global controlling coefficient to be suitable for the downstream tasks better. Experimental results demonstrate the proposed AdjFusion can 1) provide ways to dynamically synthesize images to meet the diverse demands of users; and 2) outperform the previous state-of-the-art methods on both VIF-based detection and segmentation tasks, with the constructed lightweight adaptation method. Our code will be released after accepted athttps://github.com/BearTo2/AdjFusion. Boxiong Wu, Jiangtao Nie, Wei Wei 0008, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Toward High-Quality HDR Deghosting With Conditional Diffusion ModelsabstractHigh Dynamic Range (HDR) images can be recovered from several Low Dynamic Range (LDR) images by existing Deep Neural Networks (DNNs) techniques. Despite the remarkable progress, DNN-based methods still generate ghosting artifacts when LDR images have saturation and large motion, which hinders potential applications in real-world scenarios. To address this challenge, we formulate the HDR deghosting problem as an image generation that leverages LDR features as the diffusion model’s condition, consisting of the feature condition generator and the noise predictor. Feature condition generator employs attention and Domain Feature Alignment (DFA) layer to transform the intermediate features to avoid ghosting artifacts. With the learned features as conditions, the noise predictor leverages a stochastic iterative denoising process for diffusion models to generate an HDR image by steering the sampling process. Furthermore, to mitigate semantic confusion caused by the saturation problem of LDR images, we design a sliding window noise estimator to sample smooth noise in a patch-based manner. In addition, an image space loss is proposed to avoid the color distortion of the estimated HDR results. We empirically evaluate our model on benchmark datasets for HDR imaging. The results demonstrate that our approach achieves state-of-the-art performances and well generalization to real-world images. Qingsen Yan, Tao Hu 0013, Hao Tang 0005, Yu Zhu 0004, Wei Dong 0010, Luc Van Gool, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2024 | Toward Meta-Shape-Based Multi-View 3D Point Cloud Registration: An EvaluationabstractReducing cumulative registration error is critical to accurate 3D multi-view registration. Meta-shape based methods optimize rigid transformations of point clouds by iteratively registering each point cloud with a meta-shape, which remain popular solutions to 3D multi-view registration. However, the merits and demerits of existing meta-shape based methods remain unclear. Moreover, we argue that simpler meta-shape based solutions can achieve even better performance. To this end, we evaluate seven representative meta-shape based methods in this work, including four existing ones and three modified ones, in order to investigate the problem of defining a good meta-shape. In particular, we first abstract the main steps of considered methods. Then, experiments on both object and scene datasets with real and synthetic cumulative registration errors are deployed for an in-depth evaluation. Finally, based on the experimental outcomes, we give a discussion on the advantages and limitations of meta-shape based methods. We demonstrate prior works have used unnecessarily complicated techniques for cumulative error elimination and our slightly modified simpler solutions can achieve competitive performance on experimental datasets. Shikun Zhang, Jiaqi Yang 0002, Zhaoshuai Qi, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Arbitrary-Scale Hyperspectral Image Super-Resolution From a Fusion Perspective With Spatial PriorsabstractHigh-resolution hyperspectral image (HR HSI) plays a crucial role in remote sensing applications. The single HSI super-resolution (SR) method aims to obtain an HR HSI in the spatial domain from its low-resolution (LR) counterpart. Although it has been widely studied, the performance of the existing HSI SR method is still limited because the HSI data structure itself cannot provide sufficient spatial information for reconstruction, especially with a large SR factor. In this study, we cast single HSI SR as a task fusing LR HSI with its spectral response RGB image, from which the prevalent extra high-resolution RGB images can be introduced to provide sufficient and high-quality spatial prior information for HSI SR even with a large SR factor. Within this framework, we further propose an HSI arbitrary-scale SR method, which naturally incorporates such a spatial prior in both feature extraction and local implicit image function (LIIF). Extensive experiments on two benchmark remote sensing HSI datasets, showcasing the exceptional SR performance of our proposed method. The proposed SPG-ASSR method outperforms state-of-the-art (SOTA) approaches, demonstrating its effectiveness and practical applicability. Guochao Chen, Jiangtao Nie, Wei Wei 0008, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | GLGAT-CFSL: Global-Local Graph Attention Network-Based Cross-Domain Few-Shot Learning for Hyperspectral Image ClassificationabstractFew-shot learning (FSL) is an effective approach to address the issue of limited labeled data in hyperspectral image classification (HSIC). However, it overlooks the domain shift between the source domain (SD) and the target domain (TD) in cross-domain tasks. Most existing domain adaptation (DA) methods alleviate the domain shift problem to some extent, but DA methods based on traditional convolutional operators overlook the nonlocal spatial relationships in HSI, while methods based on graph neural networks (GNNs), although effective in leveraging nonlocal spatial information for domain alignment, overly emphasize global relationships, which is disadvantageous for pixel-level classification in HSI. To solve these issues, this article proposes a novel globalp-local graph attention network-based cross-domain FSL (GLGAT-CFSL), which comprehensively reduces domain shift through global-to-local domain alignment. It has the following advantages: 1) an innovative dynamic triplet graph attention network is devised to identify nonlocal spatial relationships in HSI for global graph alignment (GGA) while also addressing common overfitting and oversmoothing issues in GNNs; 2) an ingenious local similarity learning (LSL) strategy is designed after global domain alignment, utilizing intradomain connectivity structures and interdomain node similarities for local DA, promoting cross-domain information propagation and more comprehensive reduction of domain shift; and 3) we propose a novel triaxial dynamic convolutional neural network (TDCNN) as the feature extractor, promoting cross-dimensional interaction between spectral and spatial dimensions, establishing a more generalizable and rich feature representation between the SD and the TD. The experimental results on three HSI datasets demonstrate the superiority and effectiveness of the proposed GLGAT-CFSL. Chen Ding 0002, Zhicong Deng, Yaoyang Xu, Mengmeng Zheng, Lei Zhang 0054, Yu Cao 0016, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2024 | Integrating Prototype Learning With Graph Convolution Network for Effective Active Hyperspectral Image ClassificationabstractIn recent years, active learning (AL) methods have provided a feasible approach to alleviate the problem of limited labeled samples in deep learning projects. Existing AL algorithms generally tend to select sample without labeled, whose category is difficult to distinguish. However, the sample in the category center is difficult to determine in AL operations, resulting in inaccurate category measuring and inaccurate sample selection. In addition, hyperspectral images (HSIs) have rich spectral reflective bands with strong correlations, which leads to the phenomenon that the spatial distribution between different categories in HSIs characterizes staggered distribution, which undoubtedly influences the HSI classification effect. In this article, we propose a new AL method (called PLGCN) which combines prototype learning (PL) and graph convolution network (GCN) to solve few-shot HSI classification tasks, and this method can add into existing deep learning-based HSI classification models. It includes two advantages: 1) the prototype of each category is iteratively updated to ensure the optimality of prototype in each sampling stage and 2) the spatial distribution of unlabeled samples is extracted via graph convolution neural network in order to obtain the better features in new space for easier discriminating. Experimental results on three commonly used benchmark HSI datasets demonstrate the effectiveness of the PLGCN in HSI classification tasks with limited labeled samples. Chen Ding 0002, Mengmeng Zheng, Sirui Zheng, Yaoyang Xu, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | DDF: A Novel Dual-Domain Image Fusion Strategy for Remote Sensing Image Semantic Segmentation With Unsupervised Domain AdaptationabstractThe semantic segmentation of remote sensing (RS) images is a challenging and hot issue due to the large amount of unlabeled data and domain variation. Unsupervised domain adaptation (UDA) has proven to be advantageous in leveraging unlabeled information from the target domain. However, traditional approaches of independently fine-tuning UDA models in the source and target domains have a limited effect on the result. In this article, we propose a hybrid training strategy that boosts self-training methods with domain fusion images. First, we introduce a novel dual-domain image fusion (DDF) strategy to effectively utilize the original image, the style-transferred image, and the intermediate-domain information. Second, to further refine the precision of pseudolabels, we present a region-specific reweighting strategy that assigns different weights to pseudolabel regions based on their spatial context. Finally, we conduct a series of extensive benchmark experiments and ablation studies on the ISPRS Vaihingen and Potsdam datasets. These results show the efficiency of our approach and establish a practical basis for implementing semantic segmentation in remote sensors. Lingyan Ran, Lushuang Wang, Tao Zhuo, Yinghui Xing, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Hierarchical Shared Architecture Search for Real-Time Semantic Segmentation of Remote Sensing ImagesabstractReal-time semantic segmentation of remote-sensing images demands a trade-off between speed and accuracy, which makes it challenging. Apart from manually designed networks, researchers seek to adopt neural architecture search (NAS) to discover a real-time semantic segmentation model with optimal performance automatically. Most existing NAS methods stack up no more than two types of searched cells, omitting the characteristics of resolution variation. This paper proposes the Hierarchical shared Architecture Search (HAS) method to automatically build a real-time semantic segmentation model for remote sensing images. Our model contains a lightweight backbone and a multi-scale feature fusion module. The lightweight backbone is carefully designed with low computational cost. The multi-scale feature fusion module is searched using the NAS method, where only the blocks from the same layer share identical cells. Extensive experiments reveal that our searched real-time semantic segmentation model of remote sensing images achieves the state-of-the-art trade-off between accuracy and speed. Specifically, on the LoveDA, Potsdam, and Vaihingen datasets, the searched network achieves 54.5% mIoU, 87.8% mIoU, and 84.1% mIoU, respectively, with an inference speed of 132.7 FPS. Besides, our searched network achieves 72.6% mIoU at 164.0 FPS on the CityScapes dataset and 72.3% mIoU at 186.4 FPS on the CamVid dataset. Wenna Wang, Lingyan Ran, Hanlin Yin, Mingjun Sun, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Empower Generalizability for Pansharpening Through Text-Modulated Diffusion ModelabstractPansharpening is crucial to remote sensing applications by fusing high-resolution (HR) panchromatic (PAN) images with low-resolution multispectral (LRMS) images to generate HR multispectral (HRMS) images. Recently, diffusion probabilistic models (DPMs) have provided high-quality results than regression-based methods when trained on specific pairwise data for their specific purpose. However, their performance degrades when applied to a new satellite dataset, which represents different imaging properties and spectral ranges, limiting the generalization ability of them. For better generalizability of pansharpening, in this article, we propose a text-modulated diffusion model (TMDiff) for unified pansharpening of different satellites. TMDiff takes a text-modulated 3-D UNet (TM3DU) as denoising network to gradually recover HRMS through iterative refinement over multiple time steps. By introducing satellite’s physical properties as text prompts, TM3DU is able to learn meta-knowledge across different satellites and thus can sharpen LRMS images with diverse spatial and spectral attributes. Extensive experiments on various satellite datasets demonstrate the state-of-the-art performance of our model in both qualitative and quantitative metrics. Furthermore, our model exhibits superior generalization ability to unseen datasets, highlighting its practical significance. Code is available athttps://github.com/codgodtao/TMDiff. Yinghui Xing, Litao Qu, Shizhou Zhang, Jiapeng Feng, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Improving Reliability of Heterogeneous Change Detection by Sample Synthesis and Knowledge TransferabstractDetecting changes in heterogeneous images without the supervision of changed label is a challenging yet critical task for quick responding natural disaster relief. Nevertheless, most of available unsupervised heterogeneous change detection methods strong rely on the quality of pseudo labels, and they suffer from performance degradation, even irreversible model collapse, when encounter the low-quality pseudo labels, leading to unreliable detection results. In order to improve the reliability of unsupervised heterogeneous change detection, in this paper, we propose a novel change detection paradigm based on sample synthesis and knowledge transfer. We address the issue of label reliability by artificially creating a changed region and assigning labels rather than constructing pseudo labels. These constructed labels guide the network in automatically learning the correspondence between heterogeneous images, confirming the reliability of changed regions. Moreover, an augmentation with synthetic samples on real samples makes it possible to generate more transferable samples while reducing the domain gap coarsely. A dual-branch joint training with feature contrastive learning is further developed to transfer the knowledge of changes from the synthetic sample domain to real sample domain. Experimental results on five public datasets demonstrate that our proposed method has superior performance when compared with available state-of-the-art methods. Our code is available at https://github.com/zhangqiiii/SS-KT. Yinghui Xing, Lingyan Ran, Xiuwei Zhang 0001, Hanlin Yin, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | SCAFNet: Semantic-Guided Cascade Adaptive Fusion Network for Infrared Small Target DetectionabstractInfrared small target detection is a crucial component of infrared target tracking and search. It is challenging due to the complex backgrounds, low contrast between targets and backgrounds, and the small, dim nature of the targets. Therefore, effectively representing the targets and enhancing the distinction between targets and backgrounds is essential. Existing deep-learning (DL)-based methods struggle to capture the subtle details of weak targets, neglecting the complementary characteristics of multilevel features, which leads to inaccurate localization of targets. In this article, we propose a semantic-guided cascade adaptive fusion network (SCAFNet) to address these challenges. To improve the representation of small targets in the deeper layers, we introduce a multiresolution auxiliary enhancement (MAE) encoder to progressively enhance detailed information within the deep features. After extracting multiscale features, an adaptive fusion (AdaFus) decoder is proposed to fuse them. It has a semantic-guided cascade fusion (SGCF) module to integrate feature maps at three different resolutions. Specifically, SGCF first employs rich semantic features from the high-level feature map to guide the spatial distribution of the low-level feature maps, thereby improving the distinction between the target and the background. Then, AdaFus weights are generated to guide the fusion process, ensuring that the final feature map combines rich semantic information with precise spatial details. Furthermore, we perform long-distance modeling on the feature map to achieve detailed reconstruction, which aids in restoring the shape information of the target. The effectiveness of our method is validated through experiments on various public infrared small target detection datasets. Shizhou Zhang, Yinghui Xing, Liangkui Lin, Xiaoting Su, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Remote Sensing Image Semantic Change Detection Boosted by Semi-Supervised Contrastive Learning of Semantic SegmentationabstractSemantic change detection (SCD) is a challenging task in remote sensing image (RSI) interpretation, which adopts multitemporal images to detect, locate, and analyze pixel-level land-cover “from-to” changes. In SCD, the severe class imbalance problem and the occurrence of confusing categories are very typical, making it challenging to accurately distinguish the easily confused categories with limited semantic context information. However, previous works did not address these issues in depth. This article proposes a novel SCD method named semi-supervised contrastive learning (SSCLNet), in which a simple and effective SCD network is designed as a strong baseline, and a semi-supervised contrastive learning module of semantic segmentation (SS) is presented to enhance the distinguishability of categories. Our baseline extracts semantic context through high-resolution network (HRNet), gets change information simply through an absolute difference, and then directly performs SCD based on the fusion of semantic context and change information. To utilize the semantic context information of the unlabeled non-changed regions, we employ a self-training (ST) method for semi-supervised SS. To learn distinguishable feature representations for easily confused categories, we present contrastive learning with an adaptive sampling strategy for SS. It selects challenging negative samples for each category from the other categories that exhibit similar features or attributes. The sampling space includes both the labeled changed samples and the non-changed samples predicted by ST. The comprehensive experiments on the SECOND and the Landsat-SCD dataset demonstrate that the proposed SSCLNet achieves the state-of-the-art (SOTA) performance, with a significant improvement of 2.07% and 4.15% in the score value, respectively. Xiuwei Zhang 0001, Yizhe Yang, Lingyan Ran, Kangwei Wang, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2024 | Context Recovery and Knowledge Retrieval: A Novel Two-Stream Framework for Video Anomaly DetectionabstractVideo anomaly detection aims to find the events in a video that do not conform to the expected behavior. The prevalent methods mainly detect anomalies by snippet reconstruction or future frame prediction error. However, the error is highly dependent on the local context of the current snippet and lacks the understanding of normality. To address this issue, we propose to detect anomalous events not only by the local context, but also according to the consistency between the testing event and the knowledge about normality from the training data. Concretely, we propose a novel two-stream framework based on context recovery and knowledge retrieval, where the two streams can complement each other. For the context recovery stream, we propose a spatiotemporal U-Net which can fully utilize the motion information to predict the future frame. Furthermore, we propose a maximum local error mechanism to alleviate the problem of large recovery errors caused by complex foreground objects. For the knowledge retrieval stream, we propose an improved learnable locality-sensitive hashing, which optimizes hash functions via a Siamese network and a mutual difference loss. The knowledge about normality is encoded and stored in hash tables, and the distance between the testing event and the knowledge representation is used to reveal the probability of anomaly. Finally, we fuse the anomaly scores from the two streams to detect anomalies. Extensive experiments demonstrate the effectiveness and complementarity of the two streams, whereby the proposed two-stream framework achieves state-of-the-art performance on ShanghaiTech, Avenue and Corridor datasets among the methods without object detection. Even if compared with the methods using object detection, our method reaches competitive or better performance on the ShanghaiTech, Avenue, and Ped2 datasets. Congqi Cao, Yue Lu 0008, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Toward Accurate Human Parsing Through Edge Guided DiffusionabstractExisting human parsing frameworks commonly employ joint learning of semantic edge detection and human parsing to facilitate the localization around boundary regions. Nevertheless, the parsing prediction within the interior of the part contour may still exhibit inconsistencies due to the inherent ambiguity of fine-grained semantics. In contrast, binary edge detection does not suffer from such fine-grained semantic ambiguity, leading to a typical failure case where misclassification occurs inner the part contour while the semantic edge is accurately detected. To address these challenges, we develop a novel diffusion scheme that incorporates guidance from the detected semantic edge to mitigate this problem by propagating corrected classified semantics into the misclassified regions. Building upon this diffusion scheme, we present an Edge Guided Diffusion Network (EGDNet) for human parsing, which can progressively refine the parsing predictions to enhance the accuracy and coherence of human parsing results. Moreover, we design a horizontal-vertical aggregation to exploit inherent correlations among body parts along both the horizontal and vertical axes, which aims at enhancing the initial parsing results. Extensive experimental evaluations on various challenging datasets demonstrate the effectiveness of the proposed EGDNet. Remarkably, our EGDNet shows impressive performances on six benchmark datasets, including four human body parsing datasets (LIP, CIHP, ATR, and PASCAL-Person-Part), and two human face parsing datasets (CelebAMask-HQ and LaPa). Ting Liu 0012, Hongkun Zhu, Yunchao Wei, Shikui Wei, Yao Zhao 0001, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 6 |
| 2024 | Comprehensive Attribute Prediction Learning for Person Search by LanguageabstractPerson search by language refers to searching for the interested pedestrian images given natural language sentences, which requires capturing fine-grained differences to accurately distinguish different pedestrians, while still far from being well addressed by most of the current solutions. In this paper, we propose the Comprehensive Attribute Prediction Learning (CAPL) method, which explicitly carries out attribute prediction learning, for improving the modeling capabilities of fine-grained semantic attributes and obtaining more discriminative visual and textual representations. First, we construct the semantic ATTribute Vocabulary (ATT-Vocab) based on sentence analysis. Second, the complementary context-wise and attribute-wise attribute predictions are simultaneously conducted to better model the high-frequency in-vocab attributes in our In-vocab Attribute Prediction (IAP) module. Third, to additionally consider the out-of-vocab semantics, we present the Attribute Completeness Learning (ACL) module for better capturing the low-frequency attributes outside the ATT-Vocab, obtaining more comprehensive representations. Combining the IAP and ACL modules together, our CAPL method has obtained the currently state-of-the-art retrieval performance on two widely-used benchmarks, i.e., CUHK-PEDES and ICFG-PEDES datasets. Extensive experiments and analyses have been carried out to validate the effectiveness and generalization capacities of our CAPL method. Kai Niu 0002, Linjiang Huang, Yuzhou Long, Yan Huang 0008, Liang Wang 0001, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 6 |
| 2024 | Toward Video Anomaly Retrieval From Video Anomaly Detection: New Benchmarks and ModelabstractVideo anomaly detection (VAD) has been paid increasing attention due to its potential applications, its current dominant tasks focus on online detecting anomalies, which can be roughly interpreted as the binary or multiple event classification. However, such a setup that builds relationships between complicated anomalous events and single labels, e.g., "vandalism", is superficial, since single labels are deficient to characterize anomalous events. In reality, users tend to search a specific video rather than a series of approximate videos. Therefore, retrieving anomalous events using detailed descriptions is practical and positive but few researches focus on this. In this context, we propose a novel task called Video Anomaly Retrieval (VAR), which aims to pragmatically retrieve relevant anomalous videos by cross-modalities, e.g., language descriptions and synchronous audios. Unlike the current video retrieval where videos are assumed to be temporally well-trimmed with short duration, VAR is devised to retrieve long untrimmed videos which may be partially relevant to the given query. To achieve this, we present two large-scale VAR benchmarks and design a model called Anomaly-Led Alignment Network (ALAN) for VAR. In ALAN, we propose an anomaly-led sampling to focus on key segments in long untrimmed videos. Then, we introduce an efficient pretext task to enhance semantic associations between video-text fine-grained representations. Besides, we leverage two complementary alignments to further match cross-modal contents. Experimental results on two benchmarks reveal the challenges of VAR task and also demonstrate the advantages of our tailored method. Captions are publicly released at https://github.com/Roc-Ng/VAR. Peng Wu 0015, Jing Liu 0006, Xiangteng He, Yuxin Peng 0001, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 6 |
| 2024 | CrossDiff: Exploring Self-SupervisedRepresentation of Pansharpening via Cross-Predictive Diffusion ModelabstractFusion of a panchromatic (PAN) image and corresponding multispectral (MS) image is also known as pansharpening, which aims to combine abundant spatial details of PAN and spectral information of MS images. Due to the absence of high-resolution MS images, available deep-learning-based methods usually follow the paradigm of training at reduced resolution and testing at both reduced and full resolution. When taking original MS and PAN images as inputs, they always obtain sub-optimal results due to the scale variation. In this paper, we propose to explore the self-supervised representation for pansharpening by designing a cross-predictive diffusion model, named CrossDiff. It has two-stage training. In the first stage, we introduce a cross-predictive pretext task to pre-train the UNet structure based on conditional Denoising Diffusion Probabilistic Model (DDPM). While in the second stage, the encoders of the UNets are frozen to directly extract spatial and spectral features from PAN and MS images, and only the fusion head is trained to adapt for pansharpening task. Extensive experiments show the effectiveness and superiority of the proposed model compared with state-of-the-art supervised and unsupervised methods. Besides, the cross-sensor experiments also verify the generalization ability of proposed self-supervised representation learners for other satellite datasets. Code is available at https://github.com/codgodtao/CrossDiff. Yinghui Xing, Litao Qu, Shizhou Zhang, Kai Zhang 0010, Yanning Zhang 0001, Lorenzo Bruzzone |
IEEE Trans. Image Process. | 5 |
| 2024 | MS-DETR: Multispectral Pedestrian Detection Transformer With Loosely Coupled Fusion and Modality-Balanced OptimizationabstractMultispectral pedestrian detection is an important task for many around-the-clock applications, since the visible and thermal modalities can provide complementary information especially under low light conditions. Due to the presence of two modalities, misalignment and modality imbalance are the most significant issues in multispectral pedestrian detection. In this paper, we propose MultiSpectral pedestrian DEtection TRansformer (MS-DETR) to fix above issues. MS-DETR consists of two modality-specific backbones and Transformer encoders, followed by a multi-modal Transformer decoder, and the visible and thermal features are fused in the multi-modal Transformer decoder. To well resist the misalignment between multi-modal images, we design a loosely coupled fusion strategy by sparsely sampling some keypoints from multi-modal features independently and fusing them with adaptively learned attention weights. Moreover, based on the insight that not only different modalities, but also different pedestrian instances tend to have different confidence scores to final detection, we further propose an instance-aware modality-balanced optimization strategy, which preserves visible and thermal decoder branches and aligns their predicted slots through an instance-wise dynamic loss. Our end-to-end MS-DETR shows superior performance on the challenging KAIST, CVC-14 and LLVIP benchmark datasets. The source code is available athttps://github.com/YinghuiXing/MS-DETR. Yinghui Xing, Song Wang 0002, Shizhou Zhang, Guoqiang Liang 0001, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | Human Cognition-Based Consistency Inference Networks for Multi-Modal Fake News DetectionabstractThe existing models for multi-modal fake news detection focus mainly on capturing common similar semantics between different modalities to improve detection performance. However, they ignore the extraction of inconsistent features between these modalities. The intuitive cognition way people identify a piece of fake news is generally to discover if there are inconsistent semantics among news content itself and its comments, which could be abstracted as “comparing news image-text consistency - finding valuable comments - reasoning in-/consistency between news and comments”. Inspired by the cognitive process, we propose Human Cognition-based Consistency Inference Networks (HCCIN) to comprehensively explore consistent and inconsistent semantics for multi-modal fake news detection. Specifically, we first design cross-modal alignment layer to learn consistent semantics between textual and visual information within the multi-modal news, and then the comment clue discovery layer is devoted to ascertaining the most-concerned semantics by audiences between comments. Finally, we develop collaborative inference layer to drive news consistent semantics and the most-concerned semantics to reason and discover consistent and inconsistent information between them. Experiments on three public datasets, including Weibo, Twitter, and PHEME, reveal the superiority of our HCCIN. Lianwei Wu, Pusheng Liu, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | Fusion-Embedding Siamese Network for Light Field Salient Object DetectionabstractLight field salient object detection (SOD) has shown remarkable success and gained considerable attention from the computer vision community. Existing methods usually employ a single-/two-stream network to detect saliency. However, these methods can only handle up to two different modalities at a time, preventing them from being able to fully explore the rich information in multi-modal light field derived data. To address this, we propose the first joint multi-modal learning framework, called FES-Net, for light field SOD, which can take rich inputs not limited to two modalities. Specifically, we propose an attention-aware adaptation module to first transform the multi-modal inputs for use in our joint learning framework. The transformed inputs are then fed to a Siamese network along with multiple embedded feature fusion modules to extract informative multi-modal features. Finally, we predict saliency maps from the high-level extracted features using a saliency decoder module. Our joint multi-modal learning framework effectively resolves the limitations of existing methods, providing efficient and effective multi-modal learning that can fully explore the valuable information in light field data for accurate saliency detection. Furthermore, we improve the performance by introducing the Transformer as our backbone network. To the best of our knowledge, the improved version of our model, called FES-Trans, is the first attempt to address the challenging light field SOD with the powerful Transformer technique. Extensive experiments on benchmark datasets demonstrate that our models are superior light field SOD approaches and outperform cutting-edge models remarkably. Geng Chen 0001, Huazhu Fu, Tao Zhou 0002, Guobao Xiao, Keren Fu, Yong Xia 0001, Yanning Zhang 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | Going the Extra Mile in Face Image Quality Assessment: A Novel Database and ModelabstractAn accurate computational model for image quality assessment (IQA) benefits many vision applications, such as image filtering, image processing, and image generation. Although the study of face images is an important subfield in computer vision research, the lack of face IQA data and models limits the precision of current IQA metrics on face image processing tasks such as face superresolution, face enhancement, and face editing. To narrow this gap, in this article, we first introduce the largest annotated IQA database developed to date, which contains 20,000 human faces – an order of magnitude larger than all existing rated datasets of faces – of diverse individuals in highly varied circumstances. Based on the database, we further propose a novel deep learning model to accurately predict face image quality, which, for the first time, explores the use of generative priors for IQA. By taking advantage of rich statistics encoded in well pretrained off-the-shelf generative models, we obtain generative prior information and use it as latent references to facilitate blind IQA. The experimental results demonstrate both the value of the proposed dataset for face IQA and the superior performance of the proposed model. Shaolin Su, Hanhe Lin, Vlad Hosu, Oliver Wiedemann, Jinqiu Sun, Yu Zhu 0004, Hantao Liu, Yanning Zhang 0001, Dietmar Saupe |
IEEE Trans. Multim. | 8 |
| 2024 | Dual Modality Prompt Tuning for Vision-Language Pre-Trained ModelabstractWith the emergence of large pretrained vison-language models such as CLIP, transferable representations can be adapted to a wide range of downstream tasks via prompt tuning. Prompt tuning probes for beneficial information for downstream tasks from the general knowledge stored in the pretrained model. A recently proposed method named Context Optimization (CoOp) introduces a set of learnable vectors as text prompts from the language side. However, tuning the text prompt alone can only adjust the synthesized “classifier”, while the computed visual features of the image encoder cannot be affected, thus leading to suboptimal solutions. In this article, we propose a novel dual-modality prompt tuning (DPT) paradigm through learning text and visual prompts simultaneously. To make the final image feature concentrate more on the target visual concept, a class-aware visual prompt tuning (CAVPT) scheme is further proposed in our DPT. In this scheme, the class-aware visual prompt is generated dynamically by performing the cross attention between text prompt features and image patch token embeddings to encode both the downstream task-related information and visual instance information. Extensive experimental results on 11 datasets demonstrate the effectiveness and generalization ability of the proposed method. Yinghui Xing, Qirui Wu, De Cheng, Shizhou Zhang, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | Human-Centric Behavior Description in Videos: New Benchmark and ModelabstractIn the domain of video surveillance, describing the behavior of each individual within the video is becoming increasingly essential, especially in complex scenarios with multiple individuals present. This is because describing each individual's behavior provides more detailed situational analysis, enabling accurate assessment and response to potential risks, ensuring the safety and harmony of public places. Currently, video-level captioning datasets cannot provide fine-grained descriptions for each individual's specific behavior. However, mere descriptions at the video-level fail to provide an in-depth interpretation of individual behaviors, making it challenging to accurately determine the specific identity of each individual. To address this challenge, we construct a human-centric video surveillance captioning dataset, which provides detailed descriptions of the dynamic behaviors of 7,820 individuals. Specifically, we have labeled several aspects of each person, such as location, clothing, and interactions with other elements in the scene, and these people are distributed across 1,012 videos. Based on this dataset, we can link individuals to their respective behaviors, allowing for further analysis of each person's behavior in surveillance videos. Besides the dataset, we propose a novel video captioning approach that can describe individual behavior in detail on a person-level basis, achieving state-of-the-art results. Lingru Zhou, Yiqi Gao, Manqing Zhang, Peng Wu 0015, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | A Heterogeneous Group CNN for Image Super-ResolutionabstractConvolutional neural networks (CNNs) have obtained remarkable performance via deep architectures. However, these CNNs often achieve poor robustness for image super-resolution (SR) under complex scenes. In this article, we present a heterogeneous group SR CNN (HGSRCNN) via leveraging structure information of different types to obtain a high-quality image. Specifically, each heterogeneous group block (HGB) of HGSRCNN uses a heterogeneous architecture containing a symmetric group convolutional block and a complementary convolutional block in a parallel way to enhance the internal and external relations of different channels for facilitating richer low-frequency structure information of different types. To prevent the appearance of obtained redundant features, a refinement block (RB) with signal enhancements in a serial way is designed to filter useless information. To prevent the loss of original information, a multilevel enhancement mechanism guides a CNN to achieve a symmetric architecture for promoting expressive ability of HGSRCNN. Besides, a parallel upsampling mechanism is developed to train a blind SR model. Extensive experiments illustrate that the proposed HGSRCNN has obtained excellent SR performance in terms of both quantitative and qualitative analysis. Codes can be accessed at https://github.com/hellloxiaotian/HGSRCNN. Chunwei Tian, Yanning Zhang 0001, Wangmeng Zuo, Chia-Wen Lin, David Zhang 0001, Yixuan Yuan |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Progressive Neighborhood Aggregation for Semantic Segmentation RefinementabstractMulti-scale features from backbone networks have been widely applied to recover object details in segmentation tasks. Generally, the multi-level features are fused in a certain manner for further pixel-level dense prediction. Whereas, the spatial structure information is not fully explored, that is similar nearby pixels can be used to complement each other. In this paper, we investigate a progressive neighborhood aggregation (PNA) framework to refine the semantic segmentation prediction, resulting in an end-to-end solution that can perform the coarse prediction and refinement in a unified network. Specifically, we first present a neighborhood aggregation module, the neighborhood similarity matrices for each pixel are estimated on multi-scale features, which are further used to progressively aggregate the high-level feature for recovering the spatial structure. In addition, to further integrate the high-resolution details into the aggregated feature, we apply a self-aggregation module on the low-level features to emphasize important semantic information for complementing losing spatial details. Extensive experiments on five segmentation datasets, including Pascal VOC 2012, CityScapes, COCO-Stuff 10k, DeepGlobe, and Trans10k, demonstrate that the proposed framework can be cascaded into existing segmentation models providing consistent improvements. In particular, our method achieves new state-of-the-art performances on two challenging datasets, DeepGlobe and Trans10k. The code is available at https://github.com/liutinglt/PNA. Ting Liu 0012, Yunchao Wei, Yanning Zhang 0001 |
AAAI | 3 |
| 2023 | See How You Read? Multi-Reading Habits Fusion Reasoning for Multi-Modal Fake News DetectionabstractThe existing approaches based on different neural networks automatically capture and fuse the multimodal semantics of news, which have achieved great success for fake news detection. However, they still suffer from the limitations of both shallow fusion of multimodal features and less attention to the inconsistency between different modalities. To overcome them, we propose multi-reading habits fusion reasoning networks (MRHFR) for multi-modal fake news detection. In MRHFR, inspired by people's different reading habits for multimodal news, we summarize three basic cognitive reading habits and put forward cognition-aware fusion layer to learn the dependencies between multimodal features of news, so as to deepen their semantic-level integration. To explore the inconsistency of different modalities of news, we develop coherence constraint reasoning layer from two perspectives, which first measures the semantic consistency between the comments and different modal features of the news, and then probes the semantic deviation caused by unimodal features to the multimodal news content through constraint strategy. Experiments on two public datasets not only demonstrate that MRHFR not only achieves the excellent performance but also provides a new paradigm for capturing inconsistencies between multi-modal news. Lianwei Wu, Pusheng Liu, Yanning Zhang 0001 |
AAAI | 3 |
| 2023 | Stop-Gradient Softmax Loss for Deep Metric LearningabstractDeep metric learning aims to learn a feature space that models the similarity between images, and feature normalization is a critical step for boosting performance. However directly optimizing L2-normalized softmax loss cause the network to fail to converge. Therefore some SOTA approaches appends a scale layer after the inner product to relieve the convergence problem, but it incurs a new problem that it's difficult to learn the best scaling parameters. In this letter, we look into the characteristic of softmax-based approaches and propose a novel learning objective function Stop-Gradient Softmax Loss (SGSL) to solve the convergence problem in softmax-based deep metric learning with L2-normalization. In addition, we found a useful trick named Remove the last BN-ReLU (RBR). It removes the last BN-ReLU in the backbone to reduce the learning burden of the model. Experimental results on four fine-grained image retrieval benchmarks show that our proposed approach outperforms most existing approaches, i.e., our approach achieves 75.9% on CUB-200-2011, 94.7% on CARS196 and 83.1% on SOP which outperforms other approaches at least 1.7%, 2.9% and 1.7% on Recall@1. Lu Yang 0016, Peng Wang 0015, Yanning Zhang 0001 |
AAAI | 3 |
| 2023 | A New Comprehensive Benchmark for Semi-supervised Video Anomaly Detection and AnticipationabstractSemi-supervised video anomaly detection (VAD) is a critical task in the intelligent surveillance system. However, an essential type of anomaly in VAD named scene-dependent anomaly has not received the attention of researchers. Moreover, there is no research investigating anomaly anticipation, a more significant task for preventing the occurrence of anomalous events. To this end, we propose a new comprehensive dataset, NWPU Campus, containing 43 scenes, 28 classes of abnormal events, and 16 hours of videos. At present, it is the largest semi-supervised VAD dataset with the largest number of scenes and classes of anomalies, the longest duration, and the only one considering the scene-dependent anomaly. Meanwhile, it is also the first dataset proposed for video anomaly anticipation. We further propose a novel model capable of detecting and anticipating anomalous events simultaneously. Compared with 7 outstanding VAD algorithms in recent years, our method can cope with scene-dependent anomaly detection and anomaly anticipation both well, achieving state-of-the-art performance on ShanghaiTech, CUHK Avenue, IITB Corridor and the newly proposed NWPU Campus datasets consistently. Our dataset and code is available at: https://campusvad.github.io. Congqi Cao, Yue Lu 0008, Peng Wang 0015, Yanning Zhang 0001 |
CVPR | 4 |
| 2023 | Learning to Fuse Monocular and Multi-view Cues for Multi-frame Depth Estimation in Dynamic ScenesabstractMulti-frame depth estimation generally achieves high accuracy relying on the multi-view geometric consistency. When applied in dynamic scenes, e.g., autonomous driving, this consistency is usually violated in the dynamic areas, leading to corrupted estimations. Many multi-frame methods handle dynamic areas by identifying them with explicit masks and compensating the multi-view cues with monocular cues represented as local monocular depth or features. The improvements are limited due to the uncontrolled quality of the masks and the underutilized benefits of the fusion of the two types of cues. In this paper, we propose a novel method to learn to fuse the multi-view and monocular cues encoded as volumes without needing the heuristically crafted masks. As unveiled in our analyses, the multiview cues capture more accurate geometric information in static areas, and the monocular cues capture more useful contexts in dynamic areas. To let the geometric perception learned from multi-view cues in static areas propagate to the monocular representation in dynamic areas and let monocular cues enhance the representation of multi-view cost volume, we propose a cross-cue fusion (CCF) module, which includes the cross-cue attention (CCA) to encode the spatially non-local relative intra-relations from each source to enhance the representation of the other. Experiments on real-world datasets prove the significant effectiveness and generalization ability of the proposed method. Rui Li 0013, Dong Gong, Wei Yin 0006, Hao Chen 0041, Yu Zhu 0004, Xiaozhi Chen, Jinqiu Sun, Yanning Zhang 0001 |
CVPR | 9 |
| 2023 | S3C: Semi-Supervised VQA Natural Language Explanation via Self-Critical LearningabstractVQA Natural Language Explanation (VQA-NLE) task aims to explain the decision-making process of VQA models in natural language. Unlike traditional attention or gradient analysis, free-text rationales can be easier to understand and gain users' trust. Existing methods mostly use post-hoc or selfrationalization models to obtain a plausible explanation. However, these frameworks are bottle-necked by the following challenges: 1) the reasoning process cannot be faithfully responded to and suffer from the problem of logical inconsistency. 2) Human-annotated explanations are expensive and time-consuming to collect. In this paper, we propose a new Semi-Supervised VQA-NLE via Self-Critical Learning (S3C), which evaluates the candidate explanations by answering rewards to improve the logical consistency between answers and rationales. With a semi-supervised learning framework, the S3C can benefit from a tremendous amount of samples without human-annotated explanations. A large number of automatic measures and human evaluations all show the effectiveness of our method. Meanwhile, the framework achieves a new state-of-the-art performance on the two VQA-NLE datasets. Wei Suo, Mengyang Sun, Weisong Liu, Yiqi Gao, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001 |
CVPR | 6 |
| 2023 | Glocal Energy-based Learning for Few-Shot Open-Set RecognitionabstractFew-shot open-set recognition (FSOR) is a challenging task of great practical value. It aims to categorize a sample to one of the predefined, closed-set classes illustrated by few examples while being able to reject the sample from unknown classes. In this work, we approach the FSOR task by proposing a novel energy-based hybrid model. The model is composed of two branches, where a classification branch learns a metric to classify a sample to one of closed-set classes and the energy branch explicitly estimates the open-set probability. To achieve holistic detection of open-set samples, our model leverages both class-wise and pixel-wise features to learn a glocal energy-based score, in which a global energy score is learned using the class-wise features, while a local energy score is learned using the pixel-wise features. The model is enforced to assign large energy scores to samples that are deviated from the few-shot examples in either the class-wise features or the pixel-wise features, and to assign small energy scores otherwise. Experiments on three standard FSOR datasets show the superior performance of our model.11Code is available at https://github.com/00why00/Glocal Haoyu Wang 0016, Guansong Pang, Peng Wang 0023, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
CVPR | 6 |
| 2023 | A Unified HDR Imaging Method with Pixel and Patch LevelabstractMapping Low Dynamic Range (LDR) images with different exposures to High Dynamic Range (HDR) remains nontrivial and challenging on dynamic scenes due to ghosting caused by object motion or camera jitting. With the success of Deep Neural Networks (DNNs), several DNNs-based methods have been proposed to alleviate ghosting, they cannot generate approving results when motion and saturation occur. To generate visually pleasing HDR images in various cases, we propose a hybrid HDR deghosting network, called HyHDRNet, to learn the complicated relationship between reference and non-reference images. The proposed HyHDRNet consists of a content alignment subnetwork and a Transformer-based fusion subnetwork. Specifically, to effectively avoid ghosting from the source, the content alignment subnetwork uses patch aggregation and ghost attention to integrate similar content from other non-reference images with patch level and suppress undesired components with pixel level. To achieve mutual guidance between patch-level and pixel-level, we leverage a gating module to sufficiently swap useful information both in ghosted and saturated regions. Furthermore, to obtain a high-quality HDR image, the Transformer-based fusion subnetwork uses a Residual Deformable Transformer Block (RDTB) to adaptively merge information for different exposed regions. We examined the proposed method on four widely used public HDR image deghosting datasets. Experiments demonstrate that HyHDRNet outperforms state-of-the-art methods both quantitatively and qualitatively, achieving appealing HDR visualization with unified textures and colors. Qingsen Yan, Weiye Chen, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001 |
CVPR | 6 |
| 2023 | SMAE: Few-shot Learning for HDR Deghosting with Saturation-Aware Masked AutoencodersabstractGenerating a high-quality High Dynamic Range (HDR) image from dynamic scenes has recently been extensively studied by exploiting Deep Neural Networks (DNNs). Most DNNs-based methods require a large amount of training data with ground truth, requiring tedious and time-consuming work. Few-shot HDR imaging aims to generate satisfactory images with limited data. However, it is difficult for modern DNNs to avoid overfitting when trained on only a few images. In this work, we propose a novel semi-supervised approach to realize few-shot HDR imaging via two stages of training, called SSHDR. Unlikely previous methods, directly recovering content and removing ghosts simultaneously, which is hard to achieve optimum, we first generate content of saturated regions with a self-supervised mechanism and then address ghosts via an iterative semi-supervised learning framework. Concretely, considering that saturated regions can be regarded as masking Low Dynamic Range (LDR) input regions, we design a Saturated Mask AutoEncoder (SMAE) to learn a robust feature representation and reconstruct a non-saturated HDR image. We also propose an adaptive pseudo-label selection strategy to pick high-quality HDR pseudo-labels in the second stage to avoid the effect of mislabeled samples. Experiments demonstrate that SSHDR outperforms state-of-the-art methods quantitatively and qualitatively within and across different datasets, achieving appealing HDR visualization with few labeled samples. Qingsen Yan, Weiye Chen, Hao Tang 0005, Yu Zhu 0004, Jinqiu Sun, Luc Van Gool, Yanning Zhang 0001 |
CVPR | 8 |
| 2023 | 3D Registration with Maximal CliquesabstractAs a fundamental problem in computer vision, 3D point cloud registration (PCR) aims to seek the optimal pose to align a point cloud pair. In this paper, we present a 3D registration method with maximal cliques (MAC). The key insight is to loosen the previous maximum clique constraint, and mine more local consensus information in a graph for accurate pose hypotheses generation: 1) A compatibility graph is constructed to render the affinity relationship between initial correspondences. 2) We search for maximal cliques in the graph, each of which represents a consensus set. We perform node-guided clique selection then, where each node corresponds to the maximal clique with the greatest graph weight. 3) Transformation hypotheses are computed for the selected cliques by the SVD algorithm and the best hypothesis is used to perform registration. Extensive experiments on U3M, 3DMatch, 3DLoMatch and KITTI demonstrate that MAC effectively increases registration accuracy, outperforms various state-of-the-art methods and boosts the performance of deep-learned methods. MAC combined with deep-learned methods achieves state-of-the-art registration recall of 95.7% /78.9% on 3DMatch /3DLoMatch. Xiyu Zhang 0001, Jiaqi Yang 0002, Shikun Zhang, Yanning Zhang 0001 |
CVPR | 4 |
| 2023 | Revisiting Prototypical Network for Cross Domain Few-Shot LearningabstractPrototypical Network is a popular few-shot solver that aims at establishing a feature metric generalizable to novel few-shot classification (FSC) tasks using deep neural networks. However, its performance drops dramatically when generalizing to the FSC tasks in new domains. In this study, we revisit this problem and argue that the devil lies in the simplicity bias pitfall in neural networks. In specific, the network tends to focus on some biased shortcut features (e.g., color, shape, etc.) that are exclusively sufficient to distinguish very few classes in the meta-training tasks within a pre-defined domain, but fail to generalize across domains as some desirable semantic features. To mitigate this problem, we propose a Local-global Distillation Prototypical Network (LDP-net). Different from the standard Prototypical Network, we establish a two-branch network to classify the query image and its random local crops, respectively. Then, knowledge distillation is conducted among these two branches to enforce their class affiliation consistency. The rationale behind is that since such global-local semantic relationship is expected to hold regardless of data domains, the local-global distillation is beneficial to exploit some cross-domain transferable semantic features for feature metric establishment. Moreover, such local-global semantic consistency is further enforced among different images of the same class to reduce the intra-class semantic variation of the resultant feature. In addition, we propose to update the local branch as Exponential Moving Average (EMA) over training episodes, which makes it possible to better distill cross-episode knowledge and further enhance the generalization performance. Experiments on eight cross-domain FSC benchmarks empirically clarify our argument and show the state-of-the-art results of LDP-net. Code is available in https://github.com/NWPUZhoufei/LDP-Net Fei Zhou 0008, Peng Wang 0023, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
CVPR | 5 |
| 2023 | Burst Perception-Distortion Tradeoff: Analysis and EvaluationabstractBurst image restoration attempts to effectively utilize the complementary cues appearing in sequential images to produce a high-quality image. Most current methods use all the available images to obtain the reconstructed image. However, using more images for burst restoration is not always the best option regarding reconstruction quality and efficiency, as the images acquired by handheld imaging devices suffer from degradation and misalignment caused by the camera noise and shake. In this paper, we extend the perception-distortion tradeoff theory by introducing multiple-frame information. We propose the area of the unattainable region as a new metric for perception-distortion tradeoff evaluation and comparison. Based on this metric, we analyse the performance of burst restoration from the perspective of the perception-distortion tradeoff under both aligned bursts and misaligned bursts situations. Our analysis reveals the importance of inter-frame alignment for burst restoration and shows that the optimal burst length for the restoration model depends both on the degree of degradation and misalignment. Danna Xue, Luis Herranz, Javier Vazquez-Corral, Yanning Zhang 0001 |
ICASSP | 4 |
| 2023 | Boosting No-Reference Super-Resolution Image Quality Assessment with Knowledge Distillation and ExtensionabstractDeep learning (DL) based image super-resolution (SR) tech-niques have been well investigated for recent years. However, studies dedicated to SR image quality assessment (SR-IQA) have not been fully developed, which is even more difficult if pristine high-resolution (HR) images are lacking as a reference. Due to the challenge, existing widely used no-reference (NR) SR-IQA metrics (e.g., PI, NIQE, and Ma) are still far from meeting the practical requirements of providing accurate estimations which align well with human mean opinion scores (MOS). To this end, we propose a novel Knowledge Extension Super-Resolution Image Quality Assessment (KE-SR-IQA) framework to predict SR image quality by leveraging a semi-supervised knowledge distillation (KD) strategy. Concretely, we first employ a well-trained full-reference (FR) SR-IQA model as the teacher, then we perform knowledge extension (KE) by additional pseudo-labeled data to further distill a NR-student for promoting the prediction accuracy. Extensive experiments on several benchmarks validate the ef-fectiveness of our approach. Shaolin Su, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001 |
ICASSP | 5 |
| 2023 | AerialVLN: Vision-and-Language Navigation for UAVsabstractRecently emerged Vision-and-Language Navigation (VLN) tasks have drawn significant attention in both computer vision and natural language processing communities. Existing VLN tasks are built for agents that navigate on the ground, either indoors or outdoors. However, many tasks require intelligent agents to carry out in the sky, such as UAV-based goods delivery, traffic/security patrol, and scenery tour, to name a few. Navigating in the sky is more complicated than on the ground because agents need to consider the flying height and more complex spatial relationship reasoning. To fill this gap and facilitate research in this field, we propose a new task named AerialVLN, which is UAV-based and towards outdoor environments. We develop a 3D simulator rendered by near-realistic pictures of 25 city-level scenarios. Our simulator supports continuous navigation, environment extension and configuration. We also proposed an extended baseline model based on the widely-used cross-modal-alignment (CMA) navigation methods. We find that there is still a significant gap between the baseline model and human performance, which suggests AerialVLN is a new challenging task. Dataset and code is available at https://github.com/AirVLN/AirVLN. Yuankai Qi, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001 |
ICCV | 5 |
| 2023 | MixCycle: Mixup Assisted Semi-Supervised 3D Single Object Tracking with Cycle Consistencyabstract3D single object tracking (SOT) is an indispensable part of automated driving. Existing approaches rely heavily on large, densely labeled datasets. However, annotating point clouds is both costly and time-consuming. Inspired by the great success of cycle tracking in unsupervised 2D SOT, we introduce the first semi-supervised approach to 3D SOT. Specifically, we introduce two cycle-consistency strategies for supervision: 1) Self tracking cycles, which leverage labels to help the model converge better in the early stages of training; 2) forward-backward cycles, which strengthen the tracker’s robustness to motion variations and the template noise caused by the template update strategy. Furthermore, we propose a data augmentation strategy named SOTMixup to improve the tracker’s robustness to point cloud diversity. SOTMixup generates training samples by sampling points in two point clouds with a mixing rate and assigns a reasonable loss weight for training according to the mixing rate. The resulting MixCycle approach generalizes to appearance matching-based trackers. On the KITTI benchmark, based on the P2B tracker [16], MixCycle trained with 10% labels outperforms P2B trained with 100% labels, and achieves a 28.4% precision improvement when using 1% labels. Our code will be released at https://github.com/Mumuqiao/MixCycle. Qiao Wu, Jiaqi Yang 0002, Kun Sun 0002, Chu'ai Zhang, Yanning Zhang 0001, Mathieu Salzmann |
ICCV | 5 |
| 2023 | WSAD-Net: Weakly Supervised Anomaly Detection in Untrimmed Surveillance Videos
Peng Wu 0015, Yanning Zhang 0001 |
ICIG (5) | 2 |
| 2023 | CDPMSR: Conditional Diffusion Probabilistic Models for Single Image Super-ResolutionabstractDiffusion probabilistic models (DPM) have been widely adopted in image-to-image translation to generate high-quality images. Prior attempts at applying the DPM to image super-resolution (SR) have shown that iteratively refining a pure Gaussian noise with a conditional image using a U-Net trained on denoising at various-level noises can help obtain a satisfied high-resolution image for the low-resolution one. To further improve the performance and simplify current DPM-based super-resolution methods, we propose a simple but non-trivial DPM-based super-resolution post-process framework, i.e., cDPMSR. After applying a pre-trained SR model on the to-be-test LR image to provide the conditional input, we adapt the standard DPM to conduct conditional image generation and perform super-resolution through a deterministic iterative denoising process. Our method surpasses prior attempts on both qualitative and quantitative results and can generate more photo-realistic counterparts for the low-resolution images with various benchmark datasets including Set5, Set14, Urban100, BSD100, and Manga109. Code will be published after accepted. Axi Niu, Kang Zhang 0008, Trung X. Pham, Jinqiu Sun, Yu Zhu 0004, In-So Kweon, Yanning Zhang 0001 |
ICIP | 7 |
| 2023 | Weakly Supervised Video Anomaly Detection Based on Cross-Batch Clustering GuidanceabstractWeakly supervised video anomaly detection (WSVAD) is a challenging task since only video-level labels are available for training. In previous studies, the discriminative power of the learned features is not strong enough, and the data imbalance resulting from the mini-batch training strategy is ignored. To address these two issues, we propose a novel WSVAD method based on cross-batch clustering guidance. To enhance the discriminative power of features, we propose a batch clustering based loss to encourage a clustering branch to generate distinct normal and abnormal clusters based on a batch of data. Meanwhile, we design a cross-batch learning strategy by introducing clustering results from previous minibatches to reduce the impact of data imbalance. In addition, we propose to generate more accurate segment-level anomaly scores based on batch clustering guidance to further improve the performance of WSVAD. Extensive experiments on two public datasets demonstrate the effectiveness of our approach. Congqi Cao, Xin Zhang 0168, Shizhou Zhang, Peng Wang 0015, Yanning Zhang 0001 |
ICME | 5 |
| 2023 | Wavelet Transform Based Network for Spectral Super-ResolutionabstractSpectral super-resolution (SSR) aims at reconstructing a hyperspectral image (HSI) from an observed RGB image through interpolation in the spectral domain. Recent progress mainly focus on establishing various deep interpolation networks to directly exploit the spatial-spectral information of the RGB image for SSR. However, few of them pay attention on its frequency information, which proves to be orthogonal to the spatial-spectral information and also crucial for SSR, and thus their generalization performance can be further improved. To mitigate this problem, in this study we proposes a wavelet transform based network (WTNet) for SSR. Different from existing SSR networks in image-domain, the Haar wavelet transform is employed to decompose the input RGB image into four different frequency bands. Moreover, a multi-scale convolution and self-attention based feature extraction block and a cross-attention based band interaction block are constructed to separately exploit the statistics within each band as well as the inter-band frequency correlation. By doing these, the proposed WTNet is able to sufficiently exploit the frequency information of the input RGB image for accurate SSR. Experimental results on two datasets demonstrate the efficacy and superior SSR performance of the proposed WTNet. Weixin Ren, Qianyue Duan, Tiange Huang, Lei Zhang 0054, Wei Wei 0008, Chen Ding 0002, Yanning Zhang 0001 |
IGARSS | 7 |
| 2023 | SSML-QNet: Scale-Separative Metric Learning Quadruplet Network for Multi-modal Image Patch MatchingabstractMulti-modal image matching is very challenging due to the significant diversities in visual appearance of different modal images. Typically, the existing well-performed methods mainly focus on learning invariant and discriminative features for measuring the relation between multi-modal image pairs. However, these methods often take the features as a whole and largely overlook the fact that different scale features for a same image pair may have different similarity, which may lead to sub-optimal results only. In this work, we propose a Scale-Separative Metric Learning Quadruplet network (SSML-QNet) for multi-modal image patch matching. Specifically, SSML-QNet can extract both relevant and irrelevant features of imaging modality with the proposed quadruplet network architecture. Then, the proposed Scale-Separative Metric Learning module separately encodes the similarity of different scale features with the pyramid structure. And for each scale, cross-modal consistent features are extracted and measured by coordinate and channel-wise attention sequentially. This makes our network robust to appearance divergence caused by different imaging mechanism. Experiments on the benchmark dataset (VIS-NIR, VIS-LWIR, Optical-SAR, and Brown) have verified that the proposed SSML-QNet is able to outperform other state-of-the-art methods. Furthermore, the cross-dataset transferring experiments on these four datasets also have shown that the proposed method has powerful ability of cross-dataset transferring. Xiuwei Zhang 0001, Hanlin Yin, Yinghui Xing, Yanning Zhang 0001 |
IJCAI | 7 |
| 2023 | Dichotomous Image Segmentation with Frequency PriorsabstractDichotomous image segmentation (DIS) has a wide range of real-world applications and gained increasing research attention in recent years. In this paper, we propose to tackle DIS with informative frequency priors. Our model, called FP-DIS, stems from the fact that prior knowledge in the frequency domain can provide valuable cues to identify fine-grained object boundaries. Specifically, we propose a frequency prior generator to jointly utilize a fixed filter and learnable filters to extract informative frequency priors. Before embedding the frequency priors into the network, we first harmonize the multi-scale side-out features to reduce their heterogeneity. This is achieved by our feature harmonization module, which is based on a gating mechanism to harmonize the grouped features. Finally, we propose a frequency prior embedding module to embed the frequency priors into multi-scale features through an adaptive modulation strategy. Extensive experiments on the benchmark dataset, DIS5K, demonstrate that our FP-DIS outperforms state-of-the-art methods by a large margin in terms of key evaluation metrics. Bo Dong 0001, Yuanfeng Wu, Wentao Zhu 0002, Geng Chen 0001, Yanning Zhang 0001 |
IJCAI | 6 |
| 2023 | Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source LocalizationabstractSelf-supervised sound source localization is usually challenged by the modality inconsistency. In recent studies, contrastive learning based strategies have shown promising to establish such a consistent correspondence between audio and sound sources in visual scenarios. Unfortunately, the insufficient attention to the heterogeneity influence in the different modality features still limits this scheme to be further improved, which also becomes the motivation of our work. In this study, an Induction Network is proposed to bridge the modality gap more effectively. By decoupling the gradients of visual and audio modalities, the discriminative visual representations of sound sources can be learned with the designed Induction Vector in a bootstrap manner, which also enables the audio modality to be aligned with the visual modality consistently. In addition to a visual weighted contrastive loss, an adaptive threshold selection strategy is introduced to enhance the robustness of the Induction Network. Substantial experiments conducted on SoundNet-Flickr and VGG-Sound Source datasets have demonstrated a superior performance compared to other state-of-the-art works in different challenging scenarios. The code is available at https://github.com/Tahy1/AVIN. Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001 |
ACM Multimedia | 6 |
| 2023 | Automatic Network Architecture Search for RGB-D Semantic SegmentationabstractRecent RGB-D semantic segmentation networks are usually manually designed. However, due to limited human efforts and time costs, their performance might be inferior for complex scenarios. To address this issue, we propose the first Neural Architecture Search (NAS) method that designs the network automatically. Specifically, the target network consists of an encoder and a decoder. The encoder is designed with two independent branches, where each branch specializes in extracting features from RGB and depth images, respectively. The decoder fuses the features and generates the final segmentation result. Besides, for automatic network design, we design a grid-like network-level search space combined with a hierarchical cell-level search space. By further developing an effective gradient-based search strategy, the network structure with hierarchical cell architectures is discovered. Extensive results on two datasets show that the proposed method outperforms the state-of-the-art approaches, which achieves a mIoU score of 55.1% on the NYU-Depth v2 dataset and 50.3% on the SUN-RGBD dataset. Wenna Wang, Tao Zhuo, Xiuwei Zhang 0001, Mingjun Sun, Hanlin Yin, Yinghui Xing, Yanning Zhang 0001 |
ACM Multimedia | 7 |
| 2023 | Ground-to-Aerial Person Search: Benchmark Dataset and ApproachabstractIn this work, we construct a large-scale dataset for Ground-to-Aerial Person Search, named G2APS, which contains 31,770 images of 260,559 annotated bounding boxes for 2,644 identities appearing in both of the UAVs and ground surveillance cameras. To our knowledge, this is the first dataset for cross-platform intelligent surveillance applications, where the UAVs could work as a powerful complement for the ground surveillance cameras. To more realistically simulate the actual cross-platform Ground-to-Aerial surveillance scenarios, the surveillance cameras are fixed about 2 meters above the ground, while the UAVs capture videos of persons at different location, with a variety of view-angles, flight attitudes and flight modes. Therefore, the dataset has the following unique characteristics: 1) drastic view-angle changes between query and gallery person images from cross-platform cameras; 2) diverse resolutions, poses and views of the person images under 9 rich real-world scenarios. On basis of the G2APS benchmark dataset, we demonstrate detailed analysis about current two-step and end-to-end person search methods, and further propose a simple yet effective knowledge distillation scheme on the head of the ReID network, which achieves state-of-the-art performances on both of the G2APS and the previous two public person search datasets, i.e., PRW and CUHK-SYSU. The dataset and source code available on https://github.com/yqc123456/HKD_for_person_search. Shizhou Zhang, Qingchun Yang, De Cheng, Yinghui Xing, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001 |
ACM Multimedia | 7 |
| 2023 | All-in-one Multi-degradation Image Restoration Network via Hierarchical Degradation RepresentationabstractThe aim of image restoration is to recover high-quality images from distorted ones. However, current methods usually focus on a single task (e.g., denoising, deblurring or super-resolution) which cannot address the needs of real-world multi-task processing, especially on mobile devices. Thus, developing an all-in-one method that can restore images from various unknown distortions is a significant challenge. Previous works have employed contrastive learning to learn the degradation representation from observed images, but this often leads to representation drift caused by deficient positive and negative pairs. To address this issue, we propose a novel All-in-one Multi-degradation Image Restoration Network (AMIRNet) that can effectively capture and utilize accurate degradation representation for image restoration. AMIRNet learns a degradation representation for unknown degraded images by progressively constructing a tree structure through clustering, without any prior knowledge of degradation information. This tree-structured representation explicitly reflects the consistency and discrepancy of various distortions, providing a specific clue for image restoration. To further enhance the performance of the image restoration network and overcome domain gaps caused by unknown distortions, we design a feature transform block (FTB) that aligns domains and refines features with the guidance of the degradation representation. We conduct extensive experiments on multiple distorted datasets, demonstrating the effectiveness of our method and its advantages over state-of-the-art restoration methods both qualitatively and quantitatively. Yu Zhu 0004, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001 |
ACM Multimedia | 5 |
| 2023 | Toward Re-Identifying Any AnimalabstractThe current state of re-identification (ReID) models poses limitations to their applicability in the open world, as they are primarily designed and trained for specific categories like person or vehicle. In light of the importance of ReID technology for tracking wildlife populations and migration patterns, we propose a new task called ``Re-identify Any Animal in the Wild'' (ReID-AW). This task aims to develop a ReID model capable of handling any unseen wildlife category it encounters. To address this challenge, we have created a comprehensive dataset called Wildlife-71, which includes ReID data from 71 different wildlife categories. This dataset is the first of its kind to encompass multiple object categories in the realm of ReID. Furthermore, we have developed a universal re-identification model named UniReID specifically for the ReID-AW task. To enhance the model's adaptability to the target category, we employ a dynamic prompting mechanism using category-specific visual prompts. These prompts are generated based on knowledge gained from a set of pre-selected images within the target category. Additionally, we leverage explicit semantic knowledge derived from the large-scale pre-trained language model, GPT-4. This allows UniReID to focus on regions that are particularly useful for distinguishing individuals within the target category. Extensive experiments have demonstrated the remarkable generalization capability of our UniReID model. It showcases promising performance in handling arbitrary wildlife categories, offering significant advancements in the field of ReID for wildlife conservation and research purposes. Bingliang Jiao, Lingqiao Liu, Liying Gao, Ruiqi Wu 0001, Guosheng Lin, Peng Wang 0015, Yanning Zhang 0001 |
NeurIPS | 7 |
| 2023 | An Internal-External Constrained Distillation Framework for Continual Semantic Segmentation
Qingsen Yan, Shengqiang Liu, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001 |
PRCV (3) | 6 |
| 2023 | DP-Authentication: A novel deep learning based drone pilot authentication schemeabstractUnmanned Aerial Vehicles (UAVs), also known as drones, have recently been proposed as flying base stations for providing reliable service to IoT devices. However, due to the lack of effective authentication schemes, UAVs are often hijacked by adversaries, which raises a high potential for sensitive information leakage. Therefore, designing a real-time authentication scheme is essential to enhance UAV safety. Up to the present, several works exist about pilot authentication by classifying radio-control signals. As propagating through the open environment, radio-control signals can be sniffed, analyzed, and simulated, posing significant threats to UAV security. For this reason, we propose a novel deep learning-based drone pilot authentication scheme, DP-Authentication, to protect UAVs from malicious radio-manipulated attacks. Specifically, we collect UAV flight data from the onboard PX4 flight stack and feed them into the authentication scheme to validate pilot legal status dynamically. As verified by comprehensive experiments, the proposed authentication scheme can authenticate pilots with an accuracy of 95.24% and detect malicious hijacking with an accuracy of 96.82%. Thanks to the low system overhead, it holds great promise for deployment on the UAV side to monitor pilot legal status in real-time. Liyao Han, Yijie Xun, Jiajia Liu 0001, Abderrahim Benslimane, Yanning Zhang 0001 |
Ad Hoc Networks | 5 |
| 2023 | Searching sharing relationship for instance segmentation decoder
Yuling Xi, Ning Wang 0020, Shaohua Wan 0001, Xiaoming Wang 0010, Peng Wang 0015, Yanning Zhang 0001 |
Appl. Intell. | 6 |
| 2023 | Hyperspectral anomaly detection via weighted-sparsity-regularized tensor linear representationabstractAbstract Anomaly detection aims at locating the spectral different objects of a specific scene without any prior information, and has gained increasing attention. By decomposing the input hyperspectral image (HSI) into a background tensor and an anomaly tensor, the tensor approximation is an efficient tool for detecting the anomalies. Low rankness is usually utilized as the regularizer during the background reconstruction process. Different from most existing hyperspectral anomaly detection methods which compute the truncated nuclear norm of the third folding of the original HSI, a novel weighted‐sparsity‐regularized tensor linear representation (WsrTLR) method is proposed for hyperspectral anomaly detection in this paper. Tensor linear representation is utilized to formulate the background HSI by a three‐dimensional (3D) representation base and the corresponding 3D representation coefficient. Low rankness is applied to constrict the representation coefficient, an operation which avoids destroying the multi‐way structure and losing information during the matrixing process, and ensures a satisfactory detection accuracy. Meanwhile, by incorporating the weighted‐sparsity‐regularized tensor linear representation to reconstruct the background tensor, the anomalies can be easily detected by eliminating the background tensor from the original scene. In addition, to avoid negative influence caused by the redundant bands and noisy bands in the representation process, informative bands have been first selected via an optimal neighborhood reconstruction strategy. Experimental results and data analysis on four real hyperspectral datasets, which contain anomalies with different sizes, have demonstrated the effectiveness of the proposed method. Jinqiu Sun, Yong Xia 0001, Yanning Zhang 0001 |
IET Image Process. | 4 |
| 2023 | A Dynamic Feature Interaction Framework for Multi-task Visual Perception
Yuling Xi, Hao Chen 0041, Ning Wang 0020, Peng Wang 0015, Yanning Zhang 0001, Chunhua Shen, Yifan Liu 0001 |
Int. J. Comput. Vis. | 5 |
| 2023 | Efficient spatiotemporal context modeling for action recognition
Congqi Cao, Yue Lu 0008, Yifan Zhang 0001, Dongmei Jiang, Yanning Zhang 0001 |
Neurocomputing | 5 |
| 2023 | Context-aware style learning and content recovery networks for neural style transfer
Lianwei Wu, Pusheng Liu, Yuheng Yuan, Yanning Zhang 0001 |
Inf. Process. Manag. | 5 |
| 2023 | Explore unsupervised exposure correction via illumination component divided guidance
Wei Sun 0036, Linyang Tian, Qianzhou Wang, Ruijia Cui, Xiaobao Yang 0001, Yanning Zhang 0001 |
Knowl. Based Syst. | 7 |
| 2023 | Efficient Maximum-Likelihood Estimation of Equivalent Number of Looks for PolSAR ImageabstractThe complex Wishart distribution is a widely used statistical model for multilook PolSAR image data, of which the equivalent number of looks (ENL) is a critical parameter. Over the past decades, various estimators have been developed to estimate the ENL of complex Wishart distribution, of which the maximum likelihood (ML) estimator is important since it is asymptotically unbiased and has small variance. However, this estimator is very time-consuming since it has no analytical solution and is usually solved numerically. To address this problem, this letter proposes an efficient ML estimator of ENL by deriving an approximate closed-form solution. Moreover, to estimate the ENL map of a PolSAR image, we also develop an efficient way to compute the local sample statistics parallelly. The experimental results on two PolSAR images show that our method yields highly approximate ENL values as the traditional ML estimator while is much more efficient. It costs less than 0.8 seconds on a general laptop to estimate the ENL map of a PolSAR image with 900×1024 pixels. Xianxiang Qin, Yanning Zhang 0001, Ying Li 0017 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Visible and Infrared Object Tracking via Convolution-Transformer Network With Joint Multimodal Feature LearningabstractThe existing Transformer-based RGBT tracker mainly focus on the enhancement of features extracted by Convolutional Neural Network (CNN). The potential of Transformer in representation learning remains under-explored. In this paper, we propose a Convolution-Transformer network with joint multimodal feature learning, in which both representation learning and feature fusion leverage Transformer. Specifically, we use the multi-branch Convolution-Transformer feature extraction network to process the extraction task of local modality-independent features and global modality-shared features respectively. Several simplified Transformer encoder layers form the Transformer backbone network, which is more suitable for the real-time object tracking. Besides, we found that inter-modality correlation is an important factor for modality interactions and mutual exploitation. Therefore, we propose a Joint Multimodal Feature Learning (JMFL) module, which uses cross-attention to capture the dependencies of cross-modal and enhance multimodal fusion by bidirectional guidance of multimodal information. The proposed method is fully experimented on two large benchmark datasets and compared with some current well-performing methods. The experimental results show that the proposed method performs well in terms of tracking accuracy and speed. Jiazhu Qiu, Rui Yao 0006, Yong Zhou 0003, Peng Wang 0015, Yanning Zhang 0001, Hancheng Zhu |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | Learning depth via leveraging semantics: Self-supervised monocular depth estimation with both implicit and explicit semantic guidance
Rui Li 0013, Danna Xue, Shaolin Su, Xiantuo He, Qing Mao, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001 |
Pattern Recognit. | 8 |
| 2023 | Enhancing 3D-2D Representations for Convolution Occupancy Networks
Qing Mao, Rui Li 0013, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001 |
Pattern Recognit. | 5 |
| 2023 | Object detection based on cortex hierarchical activation in border sensitive mechanism and classification-GIou joint representation
Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001 |
Pattern Recognit. | 6 |
| 2023 | From Distortion Manifold to Perceptual Quality: a Data Efficient Blind Image Quality Assessment Approach
Shaolin Su, Qingsen Yan, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001 |
Pattern Recognit. | 5 |
| 2023 | Multi-stage image denoising with the wavelet transform
Chunwei Tian, Menghua Zheng, Wangmeng Zuo, Bob Zhang 0001, Yanning Zhang 0001, David Zhang 0001 |
Pattern Recognit. | 5 |
| 2023 | FP-DARTS: Fast parallel differentiable neural architecture search for image classification
Wenna Wang, Xiuwei Zhang 0001, Hengfei Cui, Hanlin Yin, Yanning Zhang 0001 |
Pattern Recognit. | 5 |
| 2023 | 3D Medical image segmentation using parallel transformers
Qingsen Yan, Shengqiang Liu, Songhua Xu, Caixia Dong, Zongfang Li, Qinfeng Shi, Yanning Zhang 0001, Duwei Dai |
Pattern Recognit. | 7 |
| 2023 | AugFCOS: Augmented fully convolutional one-stage object detection network
Xiuwei Zhang 0001, Yinghui Xing, Wenna Wang, Hanlin Yin, Yanning Zhang 0001 |
Pattern Recognit. | 6 |
| 2023 | Meta-hallucinating prototype for few-shot learning promotion
Lei Zhang 0054, Fei Zhou 0008, Wei Wei 0008, Yanning Zhang 0001 |
Pattern Recognit. | 4 |
| 2023 | FastICENet: A real-time and accurate semantic segmentation model for aerial remote sensing river ice image
Xiuwei Zhang 0001, Lingyan Ran, Yinghui Xing, Wenna Wang, Zeze Lan, Hanlin Yin, Houjun He, Qixing Liu, Baosen Zhang, Yanning Zhang 0001 |
Signal Process. | 11 |
| 2023 | Addressing Information Inequality for Text-Based Person Search via Pedestrian-Centric Visual Denoising and Bias-Aware AlignmentsabstractText-based person search is an important task in video surveillance, which aims to retrieve the corresponding pedestrian images with a given description. In this fine-grained retrieval task, accurate cross-modal information matching is an essential yet challenging problem. However, existing methods usually ignore the information inequality between modalities, which could introduce great difficulties to cross-modal matching. Specifically, in this task, the images inevitably contain some pedestrian-irrelevant noise like background and occlusion, and the descriptions could be biased to partial pedestrian content in images. With that in mind, in this paper, we propose a Text-Guided Denoising and Alignment (TGDA) model to alleviate the information inequality and realize effective cross-modal matching. In TGDA, we first design a prototype-based denoising module, which integrates pedestrian knowledge from textual features into a prototype vector and uses it as guidance to filter out pedestrian-irrelevant noise from visual features. Thereafter, a bias-aware alignment module is introduced, which guides our model to focus on the description-biased pedestrian content in cross-modal features consistently. Through extensive experiments, the effectiveness of both modules has been validated. Besides, our TGDA achieves state-of-the-art performance on various related benchmarks. Liying Gao, Kai Niu 0002, Bingliang Jiao, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Learnable Locality-Sensitive Hashing for Video Anomaly DetectionabstractVideo anomaly detection (VAD) mainly refers to identifying anomalous events that have not occurred in the training set where only normal samples are available. Existing works usually formulate VAD as a reconstruction or prediction problem. However, the adaptability and scalability of these methods are limited. In this paper, we propose a novel distance-based VAD method to take advantage of all the available normal data efficiently and flexibly. In our method, the smaller the distance between a testing sample and normal samples, the higher the probability that the testing sample is normal. Specifically, we propose to use locality-sensitive hashing (LSH) to map the samples whose similarity exceeds a certain threshold into the same bucket in advance. To utilize multiple hashes and further alleviate the computation and memory usage, we propose to use the hash codes rather than the features as the representations of the samples. In this manner, the complexity of near neighbor search is cut down significantly. To make the samples that are semantically similar get closer and those not similar get further apart, we propose a novel learnable version of LSH that embeds LSH into a neural network and optimizes the hash functions with contrastive learning strategy. The proposed method is robust to data imbalance and can handle the large intra-class variations in normal data flexibly. Besides, it has a good ability of scalability. Extensive experiments demonstrate the superiority of our method, which achieves new state-of-the-art results on VAD benchmarks. Yue Lu 0008, Congqi Cao, Yifan Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Learning to Class-Adaptively Manipulate Embeddings for Few-Shot LearningabstractIn few-shot learning (FSL), meta-learning approach (MLA) mainly focuses on learning transferable knowledge from plenty of auxiliary FSL tasks to facilitate fast generalization to a new task. For a given FSL task, due to the inter-class distribution discrepancy, each class necessitates a specific embedding (i.e., a mapping function) to map samples into an ideal semantic space where samples from this class can be well separately from other classes. Moreover, these embeddings may vary with different tasks. Hence, one crucial knowledge for MLA is how to separately construct optimal embeddings for each class based on a few training samples given in a FSL task. However, most existing MLAs rarely consider this and thus show limited generalization capacity. To mitigate this problem, instead of directly construct class-adaptive embeddings, we present a new MLA that aims at learning to class-adaptively manipulate the features of samples for accurate classification in a new FSL task. In a specific, for a new FSL task, the proposed MLA first learns to generate some class-specific weights based on training samples via exploiting the inter-class distribution discrepancy between this class and the others. Then, the generated weights are utilized to compute the Hadamard product of features produced by a task-agnostic embedding module. By doing this, the proposed MLA can dynamically enhance or depress some specific semantic dimensions of sample features depending on the distribution of each class for accurate classification, and thus equals to constructing class-adaptive embeddings for each class but in a simpler way which can appropriately avoid over-fitting and is scalable to cases with extensive classes. To show its efficacy, we test the proposed MLA on four benchmark FSL datasets under various settings and report superior performance over existing state-of-the-arts. Fei Zhou 0008, Wei Wei 0008, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Semi-Supervised Neural Architecture Search for Hyperspectral Imagery Classification Method With Dynamic Feature ClusteringabstractHyperspectral image(HSI) contains rich spatial and spectral information, which makes HSI classification task the research focus of HSI analysis within remote sensing community. Though deep learning based HSI classification methods obtain good performance in recent years, how to learn network structure better suitable for a given HSI instead of utilizing a manually designed one for HSI classification is still a challenging problem, especially providing only small amount of labeled samples. To address this problem, we propose the first semi-supervised HSI classification network constructed via the neural architecture search. Specifically, we propose a two-head semi-supervised HSI classification framework utilizing both labeled and unlabeled data, which consists of a shared feature extraction module, a classifier module for labeled samples together with a clustering module for unlabeled samples. To boost the performance of the constructed two-head network, we propose to utilize deep features instead of the original pixels for HSI clustering to generate pseudo labels for the unlabeled data. Within the conducted semi-supervised network, we specifically design a method to automatically search for the shared feature extraction module better suitable for the given HSI data, which leads to better HSI classification results. Experimental results on three HSI datasets demonstrate the effectiveness of the proposed method, providing only limited number of labeled training samples. Wei Wei 0008, Shuyi Zhao, Songzheng Xu, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Dynamic Super-Pixel Normalization for Robust Hyperspectral Image ClassificationabstractDeep neural networks (DNNs) have underpinned most of recent progress of hyperspectral image (HSI) classification. One premise of their success lies in the high image quality without noise corruption. However, due to the limitation of the imaging sensor and imaging conditions, HSIs captured in practice inevitably suffer from random noise, which will degrade the generalization performance and robustness of most existing DNN-based methods. In this study, we propose a dynamic super-pixel normalization (DSN) based DNN for HSI classification, which can adaptively relieve the negative effect caused by various types of noise corruption and improve the generalization performance. To achieve this goal, we propose a DSN module, for a given super-pixel which normalizes the inner pixel features using parameters dynamically generated based on themselves. By doing this, such a module enables adaptively restoring the similarity among pixels within the super-pixel corrupted by random noise through aligning their feature distribution, thus enhancing the generalization performance on noisy HSI. Moreover, it can be directly plugged into any other existing DNN architectures. To appropriately train the proposed DNN model, we further present a semi-supervised learning framework, which integrates the cross entropy loss and Kullback-Leibler (KL) divergence loss on labeled samples with the information entropy loss on the unlabeled samples for joint learning to well sidestep over-fitting. Experiments on three benchmark HSI classification datasets demonstrate the advantages of the proposed method over several state-of-the-art competitors in handling HSIs under different types of noise corruption. Cong Wang 0013, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Lightweighted Hyperspectral Image Classification Network by Progressive Bi-QuantizationabstractConvolutional neural network (CNN) has shown its powerful ability for hyperspectral image (HSI) classification, which however, is difficult to deploy on resource-limited or low-latency platforms due to its parameter and computation redundancy. Though binary neural network (BNN) has attracted attention for its extreme compressing and speeding up ability by binarizing both weights and activations, it has rarely been explored for HSI classification. In this study, we elaborately design a BNN with good performance for HSI classification task. Specifically, an adaptive gradient scale module is proposed to flexibly modify the gradient during training stage to better optimize the BNN and does not add any extra computation for inference. Furthermore, a curriculum learning-based progressive binarization strategy is utilized to improve the performance. Compared with the existing BNN works, our method can increase the HSI classification accuracy by a large margin while maintaining the compressing ratio. Abundant experiments on three datasets demonstrate the effectiveness of the proposed method. Wei Wei 0008, Chongxing Song, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Pansharpening via Frequency-Aware Fusion Network With Explicit Similarity ConstraintsabstractThe process of fusing a high spatial resolution (HR) panchromatic (PAN) image and a low spatial resolution (LR) multispectral (MS) image to obtain an HRMS image is known as pansharpening. With the development of convolutional neural networks, the performance of pansharpening methods has been improved, however, the blurry effects and the spectral distortion still exist in their fusion results due to the insufficiency in details learning and the frequency mismatch between MS and PAN. Therefore, the improvement of spatial details at the premise of reducing spectral distortion is still a challenge. In this paper, we propose a frequency-aware fusion network (FAFNet) together with a novel high-frequency feature similarity loss to address above mentioned problems. FAFNet is mainly composed of two kinds of blocks, where the frequency aware blocks aim to extract features in the frequency domain with the help of discrete wavelet transform (DWT) layers, and the frequency fusion blocks reconstruct and transform the features from frequency domain to spatial domain with the assistance of inverse DWT (IDWT) layers. Finally, the fusion results are obtained through a convolutional block. In order to learn the correspondence, we also propose a high-frequency feature similarity loss to constrain the HF features derived from PAN and MS branches, so that HF features of PAN can reasonably be used to supplement that of MS. Experimental results on three datasets at both reduced- and full-resolution demonstrate the superiority of the proposed method compared with several state-of-the-art pansharpening models. The codes are available at https://github.com/YinghuiXing/FAFNet. Yinghui Xing, Yan Zhang 0127, Houjun He, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Progressive Modality-Alignment for Unsupervised Heterogeneous Change DetectionabstractChange detection based on heterogeneous images is of great importance in some applications, such as disaster monitoring and damage assessment. However, due to the huge modality discrepancy in heterogeneous images, it is difficult to accurately detect the changed regions. In this paper, we analyze the interference of modality-alignment and changed areas to each other, and propose a progressive modality-alignment based unsupervised change detection model for heterogeneous images. Specifically, the modality alignment is achieved in an iterative manner, which can improve the detection accuracy progressively. To reduce the influence of modality discrepancy and the changed regions to each other, a pseudo-label self-learning strategy is designed, where the pseudo-labels learned by the model itself are used to act as a guidance of change detection, and they are in turn refined by the proposed progressive model. Experimental results on different real heterogeneous images verify the effectiveness and robustness of proposed method. Yinghui Xing, Lingyan Ran, Xiuwei Zhang 0001, Hanlin Yin, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Improving Inconspicuous Attributes Modeling for Person Search by LanguageabstractPerson search by language aims to retrieve the interested pedestrian images based on natural language sentences. Although great efforts have been made to address the cross-modal heterogeneity, most of the current solutions suffer from only capturing salient attributes while ignoring inconspicuous ones, being weak in distinguishing very similar pedestrians. In this work, we propose the Adaptive Salient Attribute Mask Network (ASAMN) to adaptively mask the salient attributes for cross-modal alignments, and therefore induce the model to simultaneously focus on inconspicuous attributes. Specifically, we consider the uni-modal and cross-modal relations for masking salient attributes in the Uni-modal Salient Attribute Mask (USAM) and Cross-modal Salient Attribute Mask (CSAM) modules, respectively. Then the Attribute Modeling Balance (AMB) module is presented to randomly select a proportion of masked features for cross-modal alignments, ensuring the balance of modeling capacity of both salient attributes and inconspicuous ones. Extensive experiments and analyses have been carried out to validate the effectiveness and generalization capacity of our proposed ASAMN method, and we have obtained the state-of-the-art retrieval performance on the widely-used CUHK-PEDES and ICFG-PEDES benchmarks. Kai Niu 0002, Linjiang Huang, Liang Wang 0001, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2023 | Rethinking and Improving Feature Pyramids for One-Stage Referring Expression ComprehensionabstractReferring Expression Comprehension (REC) is an important task in the vision-and-language community, since it is an essential step for many cross-modal tasks such as VQA, image retrieval and image caption. To obtain a better trade-off between speed and accuracy, existing researches usually follow a one-stage paradigm, where this task can be considered as a language-conditioned object detection task. Meanwhile, previous one-stage REC frameworks provide many different research perspectives, such as the strategies of fusion, the stage of fusion and the design of detection head. Surprisingly, these works mostly ignore the value of integrating multi-level features and even only apply single-scale features to locate the target. In this paper, we focus on rethinking and improving feature pyramids for one-stage REC. By experimental validations, we first prove that although multi-scale fusion is an effective approach for improving performance, the mature neck structures from object detection (e.g., FPN, BFN and HRFPN) have a limited impact on this task. Further, we visualize the outputs of FPN and find the underlying reason is that these coarse-grained FPN fusion strategies suffer from semantic ambiguity problem. Based on the above insights, we propose a new Language-Guided FPN (LG-FPN) method, which can dynamically allocate and select the fine-grained information by stacking language-gate and union-gate. A large number of contrastive and ablative experiments show that our LG-FPN is an effective and reliable module that can adapt to different visual backbones, fusion strategies and detection heads. Finally, our method achieves state-of-the-art performance on four referring expression datasets. Wei Suo, Mengyang Sun, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | SharpFormer: Learning Local Feature Preserving Global Representations for Image DeblurringabstractThe goal of dynamic scene deblurring is to remove the motion blur presented in a given image. To recover the details from the severe blurs, conventional convolutional neural networks (CNNs) based methods typically increase the number of convolution layers, kernel-size, or different scale images to enlarge the receptive field. However, these methods neglect the non-uniform nature of blurs, and cannot extract varied local and global information. Unlike the CNNs-based methods, we propose a Transformer-based model for image deblurring, named SharpFormer, that directly learns long-range dependencies via a novel Transformer module to overcome large blur variations. Transformer is good at learning global information but is poor at capturing local information. To overcome this issue, we design a novel Locality preserving Transformer (LTransformer) block to integrate sufficient local information into global features. In addition, to effectively apply LTransformer to the medium-resolution features, a hybrid block is introduced to capture intermediate mixed features. Furthermore, we use a dynamic convolution (DyConv) block, which aggregates multiple parallel convolution kernels to handle the non-uniform blur of inputs. We leverage a powerful two-stage attentive framework composed of the above blocks to learn the global, hybrid, and local features effectively. Extensive experiments on the GoPro and REDS datasets show that the proposed SharpFormer performs favourably against the state-of-the-art methods in blurred image restoration. Qingsen Yan, Dong Gong, Zhen Zhang 0008, Yanning Zhang 0001, Qinfeng Shi |
IEEE Trans. Image Process. | 5 |
| 2023 | An Improved Combination of Faster R-CNN and U-Net Network for Accurate Multi-Modality Whole Heart SegmentationabstractDetailed information of substructures of the whole heart is usually vital in the diagnosis of cardiovascular diseases and in 3D modeling of the heart. Deep convolutional neural networks have been demonstrated to achieve state-of-the-art performance in 3D cardiac structures segmentation. However, when dealing with high-resolution 3D data, current methods employing tiling strategies usually degrade segmentation performances due to GPU memory constraints. This work develops a two-stage multi-modality whole heart segmentation strategy, which adopts an improved Combination of Faster R-CNN and 3D U-Net (CFUN+). More specifically, the bounding box of the heart is first detected by Faster R-CNN, and then the original Computed Tomography (CT) and Magnetic Resonance Imaging (MRI) images of the heart aligned with the bounding box are input into 3D U-Net for segmentation. The proposed CFUN+ method redefines the bounding box loss function by replacing the previous Intersection over Union (IoU) loss with Complete Intersection over Union (CIoU) loss. Meanwhile, the integration of the edge loss makes the segmentation results more accurate, and also improves the convergence speed. The proposed method achieves an average Dice score of 91.1% on the Multi-Modality Whole Heart Segmentation (MM-WHS) 2017 challenge CT dataset, which is 5.2% higher than the baseline CFUN model, and achieves state-of-the-art segmentation results. In addition, the segmentation speed of a single heart has been dramatically improved from a few minutes to less than 6 seconds. Hengfei Cui, Yifan Wang 0033, Yan Li 0129, Di Xu 0010, Lei Jiang 0015, Yong Xia 0001, Yanning Zhang 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2023 | Self-Supervised Monocular Depth Estimation With Frequency-Based Recurrent RefinementabstractSelf-supervised monocular depth estimation has succeeded in learning scene geometry from only image pairs or sequences. However, it is still highly ill-posed for self-supervised depth estimation to generate high-quality depth maps with both global high accuracy and local fine details. To address this issue, we propose a novel frequency-based recurrent refinement scheme to improve the self-supervised depth estimation. Since the global and local depth representation can be correlated to high/low frequency coefficients in the frequency domain, we propose a frequency-based recurrent depth coefficient refinement (RDCR) scheme, which progressively refines both low frequency and high frequency depth coefficients with an RNN-based architecture in a multi-level manner. During the recurrent process, the depth coefficients generated from the previous time step are used as the input to generate the current depth coefficients, yielding progressively optimized depth estimations. Meanwhile, considering that the depth details often appear in areas with high image frequency, we further improve depth details during the RDCR process by leveraging the image-based high frequency components. Specifically, in each RDCR module, we enhance the high frequency depth representations by selecting and feeding the informative image-based high frequency features with a learned feature weighting mask. Extensive experiments show that the proposed method achieves globally accurate estimation with fine local details, outperforming other self-supervised methods in both quantitative and qualitative comparisons. Rui Li 0013, Danna Xue, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001 |
IEEE Trans. Multim. | 6 |
| 2023 | A Proposal-Free One-Stage Framework for Referring Expression Comprehension and Generation via Dense Cross-AttentionabstractReferring Expression Comprehension (REC) and Generation (REG) have become one of the most important tasks in visual reasoning, since it is an essential step for many vision-and-language tasks such as visual question answering or visual dialogue. However, it has not been widely used in many downstream tasks, mainly for the following reasons: 1) mainstream two-stage methods rely on additional annotations or off-the-shelf detectors to generate proposals. It would heavily degrade the generalization ability of models and lead to inevitable error accumulation. 2) Although one-stage strategies for REC have been proposed, these methods have to depend on lots of hyper-parameters (such as anchors) to generate bounding box. In this paper, we present a proposal-free one-stage (PFOS) framework that can directly regress the region-of-interest from the image or generate unambiguous descriptions in an end-to-end manner. Instead of using the dominant two-stage fashion, we take the dense-grid of images as input for a cross-attention transformer that learns multi-modal correspondences. The final bounding box or sentence is directly predicted from the image without the anchor selection or the computation of visual difference. Furthermore, we expand the traditional two-stage listener-speaker framework to jointly train by a one-stage learning paradigm. Our model achieves state-of-the-art performance on both accuracy and speed for comprehension and competitive results for generation. Mengyang Sun, Wei Suo, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | VOID: 3D object recognition based on voxelization in invariant distance space
Jiaqi Yang 0002, Shichao Fan, Siwen Quan, Yanning Zhang 0001 |
Vis. Comput. | 6 |
| 2022 | Exploring and Evaluating Image Restoration Potential in Dynamic ScenesabstractIn dynamic scenes, images often suffer from dynamic blur due to superposition of motions or low signal-noise ratio resulted from quick shutter speed when avoiding motions. Recovering sharp and clean results from the captured images heavily depends on the ability of restoration methods and the quality of the input. Although existing research on image restoration focuses on developing models for obtaining better restored results, fewer have studied to evaluate how and which input image leads to superior restored quality. In this paper, to better study an image's potential value that can be explored for restoration, we propose a novel concept, referring to image restoration potential (IRP). Specifically, We first establish a dynamic scene imaging dataset containing composite distortions and applied image restoration processes to validate the rationality of the existence to IRP. Based on this dataset, we investigate several properties of IRP and propose a novel deep model to accurately predict IRP values. By gradually distilling and selective fusing the degradation features, the proposed model shows its superiority in IRP prediction. Thanks to the proposed model, we are then able to validate how various image restoration related applications are benefited from IRP prediction. We show the potential usages of IRP as a filtering principle to select valuable frames, an auxiliary guidance to improve restoration models, and also an indicator to optimize camera settings for capturing better images under dynamic scenarios. Shaolin Su, Yu Zhu 0004, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001 |
CVPR | 6 |
| 2022 | Dynamically Transformed Instance Normalization Network for Generalizable Person Re-Identification
Bingliang Jiao, Lingqiao Liu, Liying Gao, Guosheng Lin, Lu Yang 0016, Shizhou Zhang, Peng Wang 0015, Yanning Zhang 0001 |
ECCV (14) | 8 |
| 2022 | A Simple and Robust Correlation Filtering Method for Text-Based Person Search
Wei Suo, Mengyang Sun, Kai Niu 0002, Yiqi Gao, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001 |
ECCV (35) | 6 |
| 2022 | Real-World Image Super-Resolution Via Kernel Augmentation And Stochastic VariationabstractDeep learning (DL) based single image super-resolution (SISR) algorithms have now achieved highly satisfactory evaluation and visualization results on synthetic datasets. However, in some practical applications, especially when restoring some real-world low-resolution (LR) photos, the limitation and unicity of the most commonly used bicubic down-sampling kernel often lead to significant performance degradation of models trained under ideal conditions. Thus, we first propose a kernel augmentation (KA) strategy based on generative adversarial networks (GANs) to improve the generalization ability and robustness of current SISR models. Then, we intend to reconstruct the stochastic variation (SV) features that are widely present in natural images to obtain a more realistic feature representation. In the end, extensive experiments demonstrate the feasibility and effectiveness of our approach in dealing with real-world SISR problems. Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001 |
ICIP | 4 |
| 2022 | Non-Local Proposal Dynamic Enhancement Learning for Few-Shot Object Detection in Remote Sensing ImagesabstractDeep neural networks have underpinned much of recent progress in few-shot object detection (FSOD) in remote sensing images. The key lies in accurately inferring the object categories and bounding boxes depending on the feature of each proposal region. However, due to lack of sufficient labeled samples for training model well-fitting, the feature of each proposal fails to be discriminative and informative enough for accurate inference, thus limiting the generalization capacity. To mitigate this problem, we propose a non-local proposal dynamic enhancement learning (NPDEL) methods for FSOD in remote sensing images. In contrast to directly utilizing the proposal features extracted from the backbone, we propose to enhance them before inference using a non-local dynamic enhancement module which first carries out a non-local graph convolution on all proposal features and then dynamically fuses the convolved results with the original features for enhancement. By doing this, the enhanced proposal features can adaptively aggregate the related semantic information from the whole image, thus improving their discriminability as well as the generalization capacity in FSOD. Experiments results on different FSOD tasks demonstrate the efficacy of the proposed method. Haoyu Wang 0016, Lei Zhang 0054, Wei Wei 0008, Chen Ding 0002, Yanning Zhang 0001 |
IGARSS | 5 |
| 2022 | Wavefusion: Wavelet Assistant Fusion Model for Pan-SharpeningabstractPan-sharpening refers to obtain a high-resolution multispectral (HRMS) image by fusing a panchromatic (PAN) image and a low-resolution multispectral (LRMS) image. Recently, convolutional neural networks (CNNs) have achieved great success in pan-sharpening. However, the down-sampling operations in commonly used CNN-based models lead to information loss, and the corresponding up-sampling operations usually introduce some undesirable artifacts, resulting in suboptimal fusion results. In this paper, we propose a simple but effective wavelet assistant fusion model (WaveFusion) to address aforementioned issue. The proposed model consists of three parts, namely a wavelet feature extraction (WFE) part, a wavelet feature fusion (WFF) part and a reconstruction part. With the assistance of the wavelet transform and also a simple alignment operation, WaveFusion obtains the best fusion result compared with some state-of-the-art methods, especially for the fusion at the full resolution. Yinghui Xing, Yan Zhang 0127, Yanning Zhang 0001 |
IGARSS | 3 |
| 2022 | Dynamic Long-Short Range Structure Learning for Low-Illumination Remote Sensing Imagery HDR ReconstructionabstractA promising way for low-illumination (LI) remote sensing images high-dynamic range (HDR) reconstruction is to model the mapping function from the input LI images to the corresponding high-quality counterpart using deep convolution neural networks. Due to various image contents, the key for achieving pleasing performance lies on comprehensively exploit the image-specific long-rang (e.g., non-local similarity, low-rank) and short-range (e.g., local similarity, texture etc.) structures in the LI images using appropriate network architecture. However, most existing methods can only exploit either short-range or long-range structures that are contentagnostic shared across all images, thus limiting their generalization capacity. To tackle this problem, we propose a dynamic long-short range structure learning framework for LR remote sensing images HDR reconstruction. In contrast to existing methods, we introduce a novel two-branch network architecture including a pixel-aware dynamic module that can adaptively exploit the pixel-aware short-range structure surrounding each pixel depending on its feature representation, and a long-range transformer module that dynamically exploit the long-range correlation between image patchesin the deep feature space. Then, the learned long-short range structures are integrated and cast into pixel-wise scaling factors of an illumination enhance module to restore the LI image. It empowers us to effectively exploit the image-specific long-short range structures of each input IL images for accurate HDR reconstruction. Experimental results on remote sensing images with different levels of IL demonstrate the effectiveness of the proposed method. Lei Zhang 0054, Wei Wei 0008, Chen Ding 0002, Yanning Zhang 0001 |
IGARSS | 5 |
| 2022 | Pluggable Weakly-Supervised Cross-View Learning for Accurate Vehicle Re-IdentificationabstractLearning cross-view consistent feature representation is the key for accurate vehicle Re-identification (ReID), since the visual appearance of vehicles changes significantly under different viewpoints. To this end, many existing approaches resort to the supervised cross-view learning using extensive extra viewpoints annotations, which however, is difficult to deploy in real applications due to the expensive labelling cost and the continous viewpoint variation that makes it hard to define discrete viewpoint labels. In this study, we present a pluggable Weakly-supervised Cross-View Learning (WCVL) module for vehicle ReID. Through hallucinating the cross-view samples as the hardest positive counterparts with small luminance difference and large local feature variance, we can learn the consistent feature representation via minimizing the cross-view feature distance based on vehicle IDs only without using any viewpoint annotation. More importantly, the proposed method can be seamlessly plugged into most existing vehicle ReID baselines for cross-view learning without re-training the baselines. To demonstrate its efficacy, we plug the proposed method into a bunch of off-the-shelf baselines and obtain significant performance improvement on four public benchmark datasets, i.e., VeRi-776, VehicleID, VRIC and VRAI. Lu Yang 0016, Hongbang Liu, Lingqiao Liu, Jinghao Zhou, Lei Zhang 0054, Peng Wang 0015, Yanning Zhang 0001 |
ICMR | 7 |
| 2022 | Cross-modal Co-occurrence Attributes Alignments for Person Search by LanguageabstractPerson search by language refers to retrieving the interested pedestrian images based on a free-form natural language description, which has important applications in smart video surveillance. Although great efforts have been made to align images with sentences, the challenge of reporting bias, i.e., attributes are only partially matched across modalities, still incurs large noise and influences the accurate retrieval seriously. To address this challenge, we propose a novel cross-modal matching method named Cross-modal Co-occurrence Attributes Alignments (C2A2), which can better deal with noise and obtain significant improvements in retrieval performance for person search by language. First, we construct visual and textual attribute dictionaries relying on matrix decomposition, and carry out cross-modal alignments using denoising reconstruction features to address the noise from pedestrian-unrelated elements. Second, we re-gather pixels of image and words of sentence under the guidance of learned attribute dictionaries, to adaptively constitute more discriminative co-occurrence attributes in both modalities. And the re-gathered co-occurrence attributes are carefully captured by imposing explicit cross-modal one-to-one alignments which consider relations across modalities, better alleviating the noise from non-correspondence attributes. The whole C_2A_2 method can be trained end-to-end without any pre-processing, i.e., requiring negligible additional computation overheads. It significantly outperforms the existing solutions, and finally achieves the new state-of-the-art retrieval performance on two large-scale benchmarks, CUHK-PEDES and RSTPReid datasets. Kai Niu 0002, Linjiang Huang, Yan Huang 0008, Peng Wang 0015, Liang Wang 0001, Yanning Zhang 0001 |
ACM Multimedia | 6 |
| 2022 | SlimSeg: Slimmable Semantic Segmentation with Boundary SupervisionabstractAccurate semantic segmentation models typically require significant computational resources, inhibiting their use in practical applications. Recent works rely on well-crafted lightweight models to achieve fast inference. However, these models cannot flexibly adapt to varying accuracy and efficiency requirements. In this paper, we propose a simple but effective slimmable semantic segmentation (SlimSeg) method, which can be executed at different capacities during inference depending on the desired accuracy-efficiency tradeoff. More specifically, we employ parametrized channel slimming by stepwise downward knowledge distillation during training. Motivated by the observation that the differences between segmentation results of each submodel are mainly near the semantic borders, we introduce an additional boundary guided semantic segmentation loss to further improve the performance of each submodel. We show that our proposed SlimSeg with various mainstream networks can produce flexible models that provide dynamic adjustment of computational cost and better performance than independent models. Extensive experiments on semantic segmentation benchmarks, Cityscapes and CamVid, demonstrate the generalization ability of our framework. Danna Xue, Fei Yang 0004, Luis Herranz, Jinqiu Sun, Yu Zhu 0004, Yanning Zhang 0001 |
ACM Multimedia | 7 |
| 2022 | Semantic-Augmented Local Decision Aggregation Network for Action Recognition
Congqi Cao, Jiakang Li, Qinyi Lv, Runping Xi, Yanning Zhang 0001 |
PRCV (3) | 5 |
| 2022 | Beyond Vision: A Semantic Reasoning Enhanced Model for Gesture Recognition with Improved Spatiotemporal Capacity
Congqi Cao, Yanning Zhang 0001 |
PRCV (3) | 3 |
| 2022 | Dual spin-image: A bi-directional spin-image variant using multi-scale radii for 3D local shape description
Daryl L. Bibissi, Jiaqi Yang 0002, Siwen Quan, Yanning Zhang 0001 |
Comput. Graph. | 4 |
| 2022 | Accurate localization of moving objects in dynamic environment for small unmanned aerial vehicle platform using global averagingabstractAbstract Small unmanned aerial vehicles (UAVs) have developed rapidly and are widely used for disaster relief, traffic monitoring and military surveillance. To perform these tasks better, it is necessary to improve the environmental perception ability of UAVs in a dynamic environment, including their static and dynamic perception ability. Specifically, both three‐dimensional reconstruction for a static scene and localization for moving objects are required. Simultaneous Localization And Mapping technology has made great progress in static scene structure reconstruction and UAV self‐motion estimation. However, accurate real‐time localization of moving objects is still challenging. In this article, a global averaging based localization method is proposed to locate moving objects for a small UAV platform. Inspired by global structure from motion, this idea is applied to the localization of moving objects. To solve moving object localization, the relative motion estimation and global position optimisation methods are proposed. The proposed method was tested in various scenarios with a several trajectories. The extensive experimental results demonstrate the robustness and effectiveness of the proposed method. Xiuchuan Xie, Tao Yang 0006, Yanning Zhang 0001, Bang Liang, Linfeng Liu 0002 |
IET Comput. Vis. | 3 |
| 2022 | Dual-Attention-Guided Network for Ghost-Free High Dynamic Range Imaging
Qingsen Yan, Dong Gong, Qinfeng Shi, Anton van den Hengel, Chunhua Shen, Ian D. Reid 0001, Yanning Zhang 0001 |
Int. J. Comput. Vis. | 7 |
| 2022 | One-shot Video Graph Generation for Explainable Action Reasoning
Tao Zhuo, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001, Mohan Kankanhalli |
Neurocomputing | 6 |
| 2022 | Video summarization with a dual-path attentive network
Guoqiang Liang 0001, Yanbing Lv, Shucheng Li, Xiahong Wang, Yanning Zhang 0001 |
Neurocomputing | 5 |
| 2022 | Deep U-Net architecture with curriculum learning for myocardial pathology segmentation in multi-sequence cardiac magnetic resonance images
Hengfei Cui, Lei Jiang 0015, Chang Yuwen, Yong Xia 0001, Yanning Zhang 0001 |
Knowl. Based Syst. | 5 |
| 2022 | Hyperspectral and Multispectral Image Fusion via Variational Tensor Subspace DecompositionabstractThe fusion of hyperspectral image (HSI) and multispectral image (MSI) refers to enhance the spatial resolution of HSI with the help of a corresponding MSI that has a high spatial resolution to finally obtain an HSI with high resolution in both spatial and spectral domains. In this letter, we propose a variational tensor subspace decomposition-based fusion method to fully explore the differences and correlations among three modes of the HSI tensor. Experimental results on two HSI datasets show that the proposed method can achieve superior performance compared with existing state-of-the-art fusion methods with high computational efficiency. Yinghui Xing, Yan Zhang 0127, Shuyuan Yang 0001, Yanning Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | SSA-Net: Spatial Scale Attention Network for Image-Based Geo-LocalizationabstractImage-based geo-localization is estimating the location of a query image by matching it to a large amount of images in geo-tagged database. This matching task is very challenging due to the vast differences in visual appearance or modality of image pairs on different platforms, for example, one image from the RGB camera, the other from the light detection and ranging (LiDAR) sensor. The spatial layout of the scene can provide important clues and significantly reduce matching ambiguity. Therefore, we propose a novel deep network that embeds spatial configuration of the scenes into feature representation. Specifically, we design a spatial-scale attention (SSA) module to highlight the salience correspondence layout features at different scales. The encoded features not only represent the emergence of certain objects, but also reflect the relative locations of the objects. By this way, we learn more discriminative deep feature representations, leading to a higher recall. The experimental results on two standard cross-view benchmark datasets (CVUSA and CVACT) and a cross-modal dataset (GRAL) demonstrate that our method performs better than the state-of-the-art methods. Remarkably, the recall rate@top-1 improves from 27.6% to 40.5% on the GRAL dataset. Xiuwei Zhang 0001, Xiangchuang Meng, Hanlin Yin, Yuanzeng Yue, Yinghui Xing, Yanning Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2022 | DifUnet++: A Satellite Images Change Detection Network Based on Unet++ and Differential PyramidabstractChange detection (CD) is one of the most important topics in the field of remote sensing. In this letter, we propose an effective satellite images CD network named DifUnet++. As the presentation of explicit difference is more conducive to extract change features, we design a differential pyramid of two input images as the input of Unet++. Considering the scale diversity of changed regions in remote sensing images, a multiply side-outs fusion strategy is adopted to predict the detection results of different scales. Furthermore, a learning upsampling method is utilized to refine the details of CD. The proposed architecture is evaluated on two public satellite image CD data sets. The experimental results show that our method performs much better than state-of-the-art methods. Xiuwei Zhang 0001, Yuanzeng Yue, Wenxiang Gao, Shuai Yun, Qian Su, Hanlin Yin, Yanning Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2022 | Improving data augmentation for low resource speech-to-text translation with diverse paraphrasing
Chenggang Mi 0001, Lei Xie 0001, Yanning Zhang 0001 |
Neural Networks | 3 |
| 2022 | Video summarization with a convolutional attentive adversarial network
Guoqiang Liang 0001, Yanbing Lv, Shucheng Li, Shizhou Zhang, Yanning Zhang 0001 |
Pattern Recognit. | 5 |
| 2022 | Multi-scale attention-based pseudo-3D convolution neural network for Alzheimer's disease diagnosis using structural MRI
Zhao Pei, Zhiyang Wan, Yanning Zhang 0001, Miao Wang 0008, Chengcai Leng, Yee-Hong Yang |
Pattern Recognit. | 3 |
| 2022 | Video super-resolution via mixed spatial-temporal convolution and selective fusion
Wei Sun 0036, Dong Gong, Qinfeng Shi, Anton van den Hengel, Yanning Zhang 0001 |
Pattern Recognit. | 5 |
| 2022 | High dynamic range imaging via gradient-aware context aggregation network
Qingsen Yan, Dong Gong, Qinfeng Shi, Anton van den Hengel, Jinqiu Sun, Yu Zhu 0004, Yanning Zhang 0001 |
Pattern Recognit. | 7 |
| 2022 | Center Prediction Loss for Re-identification
Lu Yang 0016, Yunlong Wang 0008, Lingqiao Liu, Peng Wang 0015, Yanning Zhang 0001 |
Pattern Recognit. | 5 |
| 2022 | Adaptive Graph Convolutional Networks for Weakly Supervised Anomaly Detection in VideosabstractFor weakly supervised anomaly detection, most existing work is limited to the problem of inadequate video representation due to the inability of modeling long-term contextual information. To solve this, we propose a novel weakly supervised adaptive graph convolutional network (WAGCN) to model the complex contextual relationship among video segments. By which, we fully consider the influence of other video segments on the current one when generating the anomaly probability score for each segment. Firstly, we combine the temporal consistency as well as feature similarity of video segments to construct a global graph, which makes full use of the association information among spatial-temporal features of anomalous events in videos. Secondly, we propose a graph learning layer in order to break the limitation of setting topology manually, which can extract graph adjacency matrix based on data adaptively and effectively. Extensive experiments on two public datasets (i.e., UCF-Crime dataset and ShanghaiTech dataset) demonstrate the effectiveness of our approach which achieves state-of-the-art performance. Congqi Cao, Xin Zhang 0168, Shizhou Zhang, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Signal Process. Lett. | 5 |
| 2022 | MS2Net: Multi-Scale and Multi-Stage Feature Fusion for Blurred Image Super-ResolutionabstractAt present, most mainstream algorithms for single image super-resolution (SISR) assume the image degradation process as an ideal degradation process (e.g. bicubic downscaling), which violates the actual degeneration conditions. In real-world image capturing, objects often move in a dynamic environment, and camera shake also often occurs, which results in serious blurs. Our work focuses on the task of image super-resolution with heavy motion blur, for which we adopt a network with two branches: one branch for image deblurring and the other one for super-resolution. Since the features obtained by the deblurring are rich in details, we apply their features as supplementary information to the super-resolution branch. Based on the adopted dual-branch framework, our major technical novelties lie in two novel modules: Multi-Scale Feature Fusion (MSFF1) module which fuses features of different scale from the deblurring branch to get local and global information, and Multi-Stage Feature Fusion (MSFF2) module which further filters useful information with attention. We evaluate the proposed method under various blur scenarios on the benchmark datasets, demonstrating competitive performance against existing methods. Axi Niu, Yu Zhu 0004, Chaoning Zhang, Jinqiu Sun, In-So Kweon, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2022 | Toward Efficient and Robust Metrics for RANSAC Hypotheses and 3D Rigid RegistrationabstractThis paper focuses on developing efficient and robust evaluation metrics for RANSAC hypotheses to achieve accurate 3D rigid registration. Estimating six-degree-of-freedom (6-DoF) pose from feature correspondences remains a popular approach to 3D rigid registration, where random sample consensus (RANSAC) is a well-known solution to this problem. However, existing metrics for RANSAC hypotheses are either time-consuming or sensitive to common nuisances, parameter variations, and different application scenarios, resulting in performance deterioration with respect to overall registration accuracy and speed. We alleviate this problem by first analyzing the contributions of inliers and outliers and then proposing several efficient and robust metrics with different designing motivations for RANSAC hypotheses. Comparative experiments on four standard datasets with different nuisances and application scenarios verify that our considered metrics can significantly improve the registration performance and are more robust than several state-of-the-art competitors, making them good gifts to practical applications. This work also draws an interesting conclusion, i.e., not all inliers are equal while all outliers should be equal, which may shed new light on this research problem. Jiaqi Yang 0002, Siwen Quan, Qian Zhang 0046, Yanning Zhang 0001, Zhiguo Cao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Toward Effective Hyperspectral Image Classification Using Dual-Level Deep Spatial Manifold RepresentationabstractHyperspectral image (HSI) contains an abundant spatial structure that can be embedded into feature extraction (FE) or classifier (CL) components for pixelwise classification enhancement. Although some existing works have exploited some simple spatial structures (e.g., local similarity) to enhance either the FE or CL component, few of them consider the latent manifold structure and how to simultaneously embed the manifold structure into both components seamlessly. Thus, their performance is still limited, especially in cases with limited or noisy training samples. To solve both problems with one stone, we present a novel dual-level deep spatial manifold representation (SMR) network for HSI classification, which consists of two kinds of blocks: an SMR-based FE block and an SMR-based CL block. In both blocks, graph convolution is utilized to adaptively model the latent manifold structure lying in each local spatial area. The difference is that the former block condenses the SMR in deep feature space to form the representation for each center pixel, while the later block leverages the SMR to propagate the label information of other pixels within the local area to the center one. To train the network well, we impose an unsupervised information loss on unlabeled samples and a supervised cross-entropy loss on the labeled samples for joint learning, which empowers the network to utilize sufficient samples for SMR learning. Extensive experiments on two benchmark HSI data set demonstrate the efficacy of the proposed method in terms of pixelwise classification, especially in the cases with limited or noisy training samples. Cong Wang 0013, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Unsupervised Recurrent Hyperspectral Imagery Super-Resolution Using Pixel-Aware RefinementabstractUnsupervised fusion-based hyperspectral imagery (HSI) super-resolution (SR) is an essential task of HSI processing, which aims to reconstruct a high-resolution (HR) HSI using only an observed low-resolution HSI and a conventional HR image. Although a large number of unsupervised HSI SR methods have been proposed, the heuristic handcrafted image priors adopted by the majority of these methods restrict their capacity to capture specific characteristics of the HSI, as well as their ability to generalize to noisy observation images. In this study, we investigate a fusion-based HSI SR framework with the deep image prior, in which the deep neural network (rather than a heuristic handcrafted image prior) is exploited to capture plenty of image statistics. Within this framework, we further propose an unsupervised recurrence-based HSI SR method using pixel-aware refinement, which utilizes the intermediate reconstruction results to self-supervise unsupervised learning. Due to containing the information of the image-specific characteristic, the proposed method achieves better performance, in terms of both accuracy and robustness to noise, compared with the existing methods. Extensive experiments on four HSI data sets demonstrate the effectiveness of the proposed method. Wei Wei 0008, Jiangtao Nie, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Boosting Hyperspectral Image Classification With Unsupervised Feature LearningabstractThe deep learning-based method has shown promising competence in image classification. Its success can be attributed to the ability to learn discriminative feature representation given plenty of labeled data. However, in real-hyperspectral image (HSI) classification applications, since pixel labeling is difficult and costly, the labels we can obtain within an HSI are always limited and noisy (i.e., inaccurate), which consequently causes overfitting of the deep learning-based method. To address this problem, we propose a novel unified deep learning network to employ both labeled and unlabeled data for training, with which the unsupervised structure knowledge, e.g., intracluster similarity and intercluster dissimilarity, inherently contained in those unlabeled data can be exploited to boost the conventional supervised classification. Specifically, we first explore the unsupervised structure knowledge in unlabeled data via a clustering method and formulate a supervised clustering task on those data with the obtained cluster labels. Then, we propose a multitask network to jointly address both the conventional classification task and the formulated supervised clustering task. With a shared feature extraction module and a high-level feature fusion module, the unsupervised structure knowledge contained in unlabeled data can be effectively introduced into the classification task, which is beneficial to learn a more discriminative feature representation and, thus, well mitigates the overfitting problem and improves the classification results. Experimental results on three data sets demonstrate the proposed method can effectively label the unlabeled data within an HSI, especially when the training labels are limited and noisy. Wei Wei 0008, Songzheng Xu, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Correspondence Selection With Loose-Tight Geometric Voting for 3-D Point Cloud RegistrationabstractThis article presents a simple yet effective method for 3-D correspondence selection and point cloud registration. It first models the initial correspondence set as a graph with nodes representing correspondences and edges connecting geometrically compatible nodes. Such graphs offer either loose or tight geometric constraints for judging the correctness of correspondence, e.g., edges, loops, and cliques. Then, we render these constraints dynamic voters to judge the correctness of a node. More specifically, we develop a loose–tight geometric voting (LT-GV) method that employs both loose and tight geometric constraints in the graph to score 3-D feature correspondences. The motivation behind this is to strike a balanced performance in terms of precision and recall because loose and tight constraints are complementary to each other. Under the dynamic voting scheme with both loose and tight voters, consistent correspondences can be retrieved based on the voting score. Both feature-matching and 3-D point cloud registration experiments on datasets with different modalities, challenges, application scenarios, and comparisons with state-of-the-art methods (including deep learned methods) verify that our LT-GV is effective for correspondence selection, robust to a number of nuisances, and able to dramatically boost 3-D point cloud registration performance. Jiaqi Yang 0002, Siwen Quan, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | SAC-COT: Sample Consensus by Sampling Compatibility Triangles in Graphs for 3-D Point Cloud RegistrationabstractSix-degree-of-freedom (6-DOF) pose estimation from feature correspondences remains a popular and robust approach for 3-D registration. However, heavy outliers that existed in the initial correspondence set pose a great challenge to this problem. This article presents a simple yet effective estimator called SAmple Consensus by sampling COmpatibility Triangles in graphs (SAC-COT) for robust 6-DOF pose estimation and 3-D registration. The key novelty is a guided three-point sampling approach. It is based on a novel correspondence sample representation, i.e., COmpatibility Triangle (COT). We first model the correspondence set as a graph with nodes connecting compatible correspondences. Then, by ranking and sampling COTs formed by ternary loops, we show that correct hypotheses can be generated in early iteration stage. Finally, the hypothesis generated by the COT yielding to the maximum consensus is the output of SAC-COT. Extensive experiments on six data sets and extensive comparisons with the state-of-the-art estimators confirm that: 1) SAC-COT can achieve accurate registrations with a few iterations and 2) SAC-COT outperforms all competitors and is ultrarobust when confronted with Gaussian noise, data decimation, holes, clutter, partial overlap, varying scales of input correspondences, and data modality variation. Jiaqi Yang 0002, Siwen Quan, Zhaoshuai Qi, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | ADHR-CDNet: Attentive Differential High-Resolution Change Detection Network for Remote Sensing ImagesabstractWith the development of deep learning, change detection technology has gained great progress. However, how to effectively extract multi-scale substantive changed features and accurately detect small changed objects as well as the accurate details is still a challenge. To solve the problem, we propose Attentived Differential High-Resolution Change Detection Network (ADHR-CDNet) for remote sensing images. In ADHR-CDNet, a novel high-resolution backbone with a Differential Pyramid Module (DPM) is proposed to extract multi-level and multi-scale substantive changed features. The backbone structure with four interconnected sub-network branches of different resolution is helpful to extract multi-level and multi-scale features. DPM is capable of distinguishing between substantive changes and pseudo changes induced by illumination, shadow, seasonal variation, and so on. Then, a novel Multi-Scale Spatial feature Attention Module (MSSAM) is presented to effectively fuse the spatial detail information of different scale features produced by our backbone to generate finer prediction. We conduct quantitative and qualitative experiments on three public change detection datasets: the Lebedev, the LEVIR-CD, and the WHU Building dataset. The proposed ADHR-CDNet reaches F1-score of 97.2% (improved 3.1%) on the Lebedev dataset, 91.4% (improved 1.6%) on the LEVIR-CD dataset, and 90.9% (improved 1.2%) on the WHU Building dataset. The experimental results demonstrate that our method performs much better than the state-of-the-art methods. The visualization comparison results show that our method can effectively detect small changed objects and significantly improve the details of detected changed objects. Our code is available at https://github.com/w-here/ASGO-113lab/tree/main/ADHR-CDNet. Xiuwei Zhang 0001, Mu Tian, Yinghui Xing, Yuanzeng Yue, Hanlin Yin, Runliang Xia, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2022 | Learning to Compare Relation: Semantic Alignment for Few-Shot LearningabstractFew-shot learning is a fundamental and challenging problem since it requires recognizing novel categories from only a few examples. The objects for recognition have multiple variants and can locate anywhere in images. Directly comparing query images with example images can not handle content misalignment. The representation and metric for comparison are critical but challenging to learn due to the scarcity and wide variation of the samples in few-shot learning. In this paper, we present a novel semantic alignment model to compare relations, which is robust to content misalignment. We propose to add two key ingredients to existing few-shot learning frameworks for better feature and metric learning ability. First, we introduce a semantic alignment loss to align the relation statistics of the features from samples that belong to the same category. And second, local and global mutual information maximization is introduced, allowing for representations that contain locally-consistent and intra-class shared information across structural locations in an image. Furthermore, we introduce a principled approach to weigh multiple loss functions by considering the homoscedastic uncertainty of each stream. We conduct extensive experiments on several few-shot learning datasets. Experimental results show that the proposed method is capable of comparing relations with semantic alignment strategies, and achieves state-of-the-art performance. Congqi Cao, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Learning Spectral Cues for Multispectral and Panchromatic Image FusionabstractRecently, deep learning based multispectral (MS) and panchromatic (PAN) image fusion methods have been proposed, which extracted features automatically and hierarchically by a series of non-linear transformations to model the complicated imaging discrepancy. But they always pay more attention to the extraction and compensation of spatial details and use the mean squared error or mean absolute error as a loss function, regardless of the preservation of spectral information contained in multispectral images. For the sake of the improvements in both spatial and spectral resolution, this paper presents a novel fusion model that takes the spectral preservation into consideration, and learns the spectral cues from the process of generating a spectrally refined multispectral image, which is constrained by a spectral loss between the generated image and the reference image. Then these spectral cues are used to modulate the PAN features to obtain final fusion result. Experimental results on reduced-resolution and full-resolution datasets demonstrate that the proposed method can obtain a better fusion result in terms of visual inspection and evaluation indices when compared with current state-of-the-art methods. Yinghui Xing, Shuyuan Yang 0001, Yan Zhang 0127, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Identity-Aware Facial Expression Recognition Via Deep Metric Learning Based on Synthesized ImagesabstractPerson-dependent facial expression recognition has received considerable research attention in recent years. Unfortunately, different identities can adversely influence recognition accuracy, and the recognition task becomes challenging. Other adverse factors, including limited training data and improper measures of facial expressions, can further contribute to the above dilemma. To solve these problems, a novel identity-aware method is proposed in this study. Furthermore, this study also represents the first attempt to fulfill the challenging person-dependent facial expression recognition task based on deep metric learning and facial image synthesis techniques. Technically, a StarGAN is incorporated to synthesize facial images depicting different but complete basic emotions for each identity to augment the training data. Then, a deep-convolutional-neural-network-based network is employed to automatically extract latent features from both real facial images and all synthesized facial images. Next, a Mahalanobis metric network trained based on extracted latent features outputs a learned metric that measures facial expression differences between images, and the recognition task can thus be realized. Extensive experiments based on several well-known publicly available datasets are carried out in this study for performance evaluations. Person-dependent datasets, including CK+, Oulu (all 6 subdatasets), MMI, ISAFE, ISED, etc., are all incorporated. After comparing the new method with several popular or state-of-the-art facial expression recognition methods, its superiority in person-dependent facial expression recognition can be proposed from a statistical point of view. Wei Huang 0013, Peng Zhang 0005, Yufei Zha, Yuming Fang 0001, Yanning Zhang 0001 |
IEEE Trans. Multim. | 6 |
| 2021 | KonIQ++: Boosting No-Reference Image Quality Assessment in the Wild by Jointly Predicting Image Quality and Defects
Shaolin Su, Vlad Hosu, Hanhe Lin, Yanning Zhang 0001, Dietmar Saupe |
BMVC | 4 |
| 2021 | A Comprehensive CT Dataset for Liver Computer Assisted Diagnosis
Qingsen Yan, Bo Wang 0011, Dong Gong, Dingwen Zhang, Yang Yang 0009, Zheng You, Yanning Zhang 0001, Qinfeng Shi |
BMVC | 7 |
| 2021 | Dual Attention Guided R2 U-Net Architecture for Right Ventricle Segmentation in MRI Images
Lei Jiang 0015, Hengfei Cui, Chang Yuwen, Yanning Zhang 0001 |
ICIG (2) | 4 |
| 2021 | Meta Transfer Learning for Few-Shot Hyperspectral Image ClassificationabstractWe propose a novel meta-learning approach for few-shot hyperspectral image (HSI) classification, which learns to distil transferable prior knowledge from a base dataset with sufficient labeled samples and generalize the knowledge to an unseen dataset with extremely limited labeled samples for performance improvement. Specifically, we first construct a backbone classification model using an embedding module and a linear classifier. Then, we sample extensive synthetic few-shot tasks from the base dataset, each of which consists of a support set with limited labeled samples and a query set with some unlabeled test samples. Given these tasks, we propose to optimize the embedding module using an episode learning scheme where for each task we train the linear classier based on an initialized embedding module using the support set and ultimately optimize the embedding module based on the test error on the query set until the test error on all tasks is minimized. By doing this, the resultant embedding module is able to appropriately generalize to an unseen few-shot classification task and lead to good performance with the linear classifier. Experiments on two standard classification benchmarks under different few-shot settings demonstrate the efficacy of the proposed method. Fei Zhou 0008, Lei Zhang 0054, Wei Wei 0008, Zongwen Bai, Yanning Zhang 0001 |
IGARSS | 5 |
| 2021 | Local-enhanced Interaction for Temporal Moment LocalizationabstractTemporal moment localization via language aims to localize a video span in an untrimmed video which best matches the given natural language query. In most previous works, they try to match the whole query feature with multiple moment proposals, or match a global video embedding with phrase or word level query features. However, these coarse interaction models will become insufficient when the query-video contains more complex relationship. To address this issue, we propose a multi-branches interaction model for temporal moment localization. Specifically, the query sentence and video are encoded into multiple feature embeddings over several semantic sub-spaces. Then, each phrase embedding filters on a video feature to generate an attention sequence, which is used to re-weight the video features. Moreover, a dynamic pointer decoder is developed to iteratively regress the temporal boundary, which can prevent our model from falling into a local optimum. To validate the proposed method, we have conducted extensive experiments on two popular benchmark datasets Charade-STA and TACoS. The experimental performance surpasses other state-of-the-arts methods, which demonstrates the effectiveness of our proposed model. Guoqiang Liang 0001, Shiyu Ji, Yanning Zhang 0001 |
ICMR | 3 |
| 2021 | Unsupervised Cross-Modal Distillation for Thermal Infrared TrackingabstractThe target representation learned by convolutional neural networks plays an important role in Thermal Infrared (TIR) tracking. Currently, most of the top-performing TIR trackers are still employing representations learned by the model trained on the RGB data. However, this representation does not take into account the information in the TIR modality itself, limiting the performance of TIR tracking. Jingxian Sun 0003, Lichao Zhang 0001, Yufei Zha, Abel Gonzalez-Garcia, Peng Zhang 0005, Wei Huang 0013, Yanning Zhang 0001 |
ACM Multimedia | 7 |
| 2021 | 3D Correspondence Grouping with Compatibility Features
Jiaqi Yang 0002, Zhiguo Cao 0001, Yanning Zhang 0001 |
PRCV (2) | 5 |
| 2021 | Few-shot action recognition with implicit temporal alignment and pair similarity optimization
Congqi Cao, Qinyi Lv, Peng Wang 0015, Yanning Zhang 0001 |
Comput. Vis. Image Underst. | 5 |
| 2021 | Multiple object tracking based on multi-task learning with strip attentionabstractAbstract Multiple object tracking (MOT) framework based on bifurcate strategy was usually challenged by data association of different model path, which work for object localisation and appearance embedding independently. By incorporating the re‐identification (re‐ID) as appearance embedding model, more recent studies on task combination of a single network have made a great progress in tracking performance. Unfortunately, the contributive improvement from re‐ID model is hard to balance the accuracy and efficiency for the whole framework. For more effective enhancement of the overall tracking performance, a real‐time detection needs to be taken into consideration with other auxiliary means for MOT modelling. Therefore, in this study, a one‐shot multiple object tracking is proposed based on multi‐task learning to obtain satisfactory performance in both speed and robustness. With updated re‐training strategy for the backbone model of detection, a D2LA network is proposed to achieve more characteristic fine‐grained feature extraction in branching task of pedestrian recognition. Additionally, a strip attention module is also introduced to further strengthen the feature discriminative capability of the tracking framework in occlusion. Experiments on the 2DMOT15, MOT16, MOT17, and MOT20 benchmark data sets have shown a superior performance in comparison to other state‐of‐the‐art tracking approaches. Yaoye Song, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Tao You, Yanning Zhang 0001 |
IET Image Process. | 6 |
| 2021 | NAS-FCOS: Efficient Search for Object Detection Architectures
Ning Wang 0020, Yang Gao 0001, Hao Chen 0041, Peng Wang 0015, Zhi Tian, Chunhua Shen, Yanning Zhang 0001 |
Int. J. Comput. Vis. | 7 |
| 2021 | Adversarial learning with collaborative attention for facial makeup removal
Xueling Chen, Yu Zhu 0004, Yanning Zhang 0001 |
Neurocomputing | 3 |
| 2021 | A light-weight, efficient, and general cross-modal image fusion network
Aiqing Fang, Jiaqi Yang 0002, Beibei Qin, Yanning Zhang 0001 |
Neurocomputing | 5 |
| 2021 | Deep-learning-based reading eye-movement analysis for aiding biometric recognition
Xiaoming Wang 0010, Yanning Zhang 0001 |
Neurocomputing | 3 |
| 2021 | Towards accurate HDR imaging with learning generator constraints
Qingsen Yan, Bo Wang 0011, Lei Zhang 0054, Zheng You, Qinfeng Shi, Yanning Zhang 0001 |
Neurocomputing | 7 |
| 2021 | Anomaly Detection of Hyperspectral Image via Tensor CompletionabstractIn this letter, a novel method of anomaly detection for the hyperspectral image (HSI) is proposed. This method originates from two ideas. First, compared with the anomalies, the spectral curves of some (not all) backgrounds are usually easy to be accurately found. Second, the spectral curves of the missing pixels in the background can be recovered via the tensor completion technology. In this way, anomalies can be picked out via discriminating the background tensor and the original HSI. Specifically, some background pixels with low response in the detection map are first detected. Then, the selected background pixels with their spectral curves are utilized to reconstruct a three-order tensor whose elements are missing at some extent. Tensor completion technology is applied to this tensor and achieves a complete tensor, which depicts the background of the scene. Finally, the reconstructed tensor originated from the background pixels is discriminated from the original HSI to pick out the anomalies. The experimental data and performance analysis have demonstrated the effectiveness of the proposed method. Yong Xia 0001, Yanning Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | Triple attention learning for classification of 14 thoracic diseases using chest radiography
Hongyu Wang 0011, Shanshan Wang 0002, Zibo Qin, Yanning Zhang 0001, Ruijiang Li, Yong Xia 0001 |
Medical Image Anal. | 4 |
| 2021 | River ice monitoring and change detection with multi-spectral and SAR images: application over yellow river
Xiuwei Zhang 0001, Yuanzeng Yue, Fei Li 0011, Xiuzhong Yuan, Minhao Fan, Yanning Zhang 0001 |
Multim. Tools Appl. | 7 |
| 2021 | A Performance Evaluation of Correspondence Grouping Methods for 3D Rigid Data MatchingabstractSeeking consistent point-to-point correspondences between 3D rigid data (point clouds, meshes, or depth maps) is a fundamental problem in 3D computer vision. While a number of correspondence selection methods have been proposed in recent years, their advantages and shortcomings remain unclear regarding different applications and perturbations. To fill this gap, this paper gives a comprehensive evaluation of nine state-of-the-art 3D correspondence grouping methods. A good correspondence grouping algorithm is expected to retrieve as many as inliers from initial feature matches, giving a rise in both precision and recall as well as facilitating accurate transformation estimation. Toward this rule, we deploy experiments on three benchmarks with different application contexts, including shape retrieval, 3D object recognition, and point cloud registration. We also investigate various perturbations such as noise, point density variation, clutter, occlusion, partial overlap, different scales of initial correspondences, and different combinations of keypoint detectors and descriptors. The rich variety of application scenarios and nuisances result in different spatial distributions and inlier ratios of initial feature correspondences, thus enabling a thorough evaluation. Based on the outcomes, we give a summary of the traits, merits, and demerits of evaluated approaches and indicate some potential future research directions. Jiaqi Yang 0002, Ke Xian, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Non-linear and selective fusion of cross-modal images
Aiqing Fang, Jiaqi Yang 0002, Yanning Zhang 0001 |
Pattern Recognit. | 4 |
| 2021 | All-in-focus synthetic aperture imaging using generative adversarial network-based semantic inpainting
Zhao Pei, Yanning Zhang 0001, Miao Ma, Yee-Hong Yang |
Pattern Recognit. | 3 |
| 2021 | Non-uniform motion deblurring with blurry component divided guidance
Wei Sun 0036, Qingsen Yan, Axi Niu, Rui Li 0013, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001 |
Pattern Recognit. | 8 |
| 2021 | Improving visible-thermal ReID with structural common space embedding and part models
Lingyan Ran, Yujun Hong, Shizhou Zhang, Yanning Zhang 0001 |
Pattern Recognit. Lett. | 5 |
| 2021 | COVID-19 Chest CT Image Segmentation Network by Multi-Scale Fusion and Enhancement OperationsabstractA novel coronavirus disease 2019 (COVID-19) was detected and has spread rapidly across various countries around the world since the end of the year 2019. Computed Tomography (CT) images have been used as a crucial alternative to the time-consuming RT-PCR test. However, pure manual segmentation of CT images faces a serious challenge with the increase of suspected cases, resulting in urgent requirements for accurate and automatic segmentation of COVID-19 infections. Unfortunately, since the imaging characteristics of the COVID-19 infection are diverse and similar to the backgrounds, existing medical image segmentation methods cannot achieve satisfactory performance. In this article, we try to establish a new deep convolutional neural network tailored for segmenting the chest CT images with COVID-19 infections. We first maintain a large and new chest CT image dataset consisting of 165,667 annotated chest CT images from 861 patients with confirmed COVID-19. Inspired by the observation that the boundary of the infected lung can be enhanced by adjusting the global intensity, in the proposed deep CNN, we introduce a feature variation block which adaptively adjusts the global properties of the features for segmenting COVID-19 infection. The proposed FV block can enhance the capability of feature representation effectively and adaptively for diverse cases. We fuse features at different scales by proposing Progressive Atrous Spatial Pyramid Pooling to handle the sophisticated infection areas with diverse appearance and shapes. The proposed method achieves state-of-the-art performance. Dice similarity coefficients are 0.987 and 0.726 for lung and COVID-19 segmentation, respectively. We conducted experiments on the data collected in China and Germany and show that the proposed deep CNN can produce impressive performance effectively. The proposed network enhances the segmentation ability of the COVID-19 infection, makes the connection with other techniques and contributes to the development of remedying COVID-19 infection. Qingsen Yan, Bo Wang 0011, Dong Gong, Chuan Luo 0003, Jianhu Shen, Jingyang Ai, Qinfeng Shi, Yanning Zhang 0001, Liang Zhang 0010, Zheng You |
IEEE Trans. Big Data | 9 |
| 2021 | Movement Aware CoMP Handover in Heterogeneous Ultra-Dense NetworksabstractThe densification of base station (BS) deployments is driving the evolution of network structures towards heterogeneous ultra-dense networks (UDN), making coordinated multipoint (CoMP) a viable and promising transmission solution. However, the BS cooperation regions formed by applying CoMP in the UDN are small and irregular, which causes frequent handover for mobile users. Different from most existing work that focus on the trigger time of handover, we explore how to choose the appropriate BS cooperation set to reduce handover rate. In this paper, we consider movement aware CoMP handover (MACH). By estimating cell dwell time, a user would be intelligently assigned to macro cell or small cell according to its movement trend. To enhance reliability, we further proposed improved MACH (iMACH) to achieve a trade-off between BSs with long dwell time and the current best performed BS for multipoint cooperation while user moving. Using stochastic geometry method, expressions of coverage probability, handover probability and throughput that characterize performance of the proposed schemes are derived. The numerical results indicate that the theoretical analyses fit the simulation results well and the proposed schemes surpass the existing schemes in terms of the aforementioned metrics, and more intelligent and suitable for ultra-dense scenarios. Wen Sun 0004, Lu Wang 0050, Jiajia Liu 0001, Nei Kato, Yanning Zhang 0001 |
IEEE Trans. Commun. | 5 |
| 2021 | Club Ideas and Exertions: Aggregating Local Predictions for Action RecognitionabstractRecognizing the actions performed in a video is challenging for an intelligent system since there are wide variations and enormous information in the video. Attention mechanism pays attention to key target areas, ignores irrelevant information and extracts more discriminant features. In recent years, attention mechanism has been introduced into video recognition. Although a rich literature has been spawned, most of the research on attention aims to aggregate local features by attention. Instead of feature aggregation, we propose to aggregate decisions based on local spatio-temporal attention regions for action recognition, which is inspired by ensemble learning. The proposed decision fusion module is easy to interpret and architecture-independent. In this article, the regions around the body joints are regarded as the key regions. We use the corresponding regions of the body joints in the 3-D feature maps as the basic local features for local classification. Finally, all the local classification results are combined to make a global decision. Furthermore, when training the network, we can selectively add supervision to the local and global decisions. We experimentally show that the proposed mechanism can improve the recognition performance on multiple datasets which demonstrates its effectiveness. Congqi Cao, Jiakang Li, Runping Xi, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Where to Look and How to Describe: Fashion Image Retrieval With an Attentional Heterogeneous Bilinear NetworkabstractFashion products typically feature in compositions of a variety of styles at different clothing parts. In order to distinguish images of different fashion products, we need to extract both appearance (i.e., “how to describe”) and localization (i.e., “where to look”) information, and their interactions. To this end, we propose a biologically inspired framework for image-based fashion product retrieval, which mimics the hypothesized two-stream visual processing system of human brain. The proposed attentional heterogeneous bilinear network (AHBN) consists of two branches: a deep CNN branch to extract fine-grained appearance attributes and a fully convolutional branch to extract landmark localization information. A joint channel-wise attention mechanism is further applied to the extracted heterogeneous features to focus on important channels, followed by a compact bilinear pooling layer to model the interaction of the two streams. Our proposed framework achieves satisfactory performance on three image-based fashion product retrieval benchmarks. Haibo Su, Peng Wang 0015, Lingqiao Liu, Hui Li 0031, Zhen Li 0068, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Learning to Zoom-In via Learning to Zoom-Out: Real-World Super-Resolution by Generating and Adapting DegradationabstractMost learning-based super-resolution (SR) methods aim to recover high-resolution (HR) image from a given low-resolution (LR) image via learning on LR-HR image pairs. The SR methods learned on synthetic data do not perform well in real-world, due to the domain gap between the artificially synthesized and real LR images. Some efforts are thus taken to capture real-world image pairs. However, the captured LR-HR image pairs usually suffer from unavoidable misalignment, which hampers the performance of end- to-end learning. Here, focusing on the real-world SR, we ask a different question: since misalignment is unavoidable, can we propose a method that does not need LR-HR image pairing and alignment at all and utilizes real images as they are? Hence we propose a framework to learn SR from an arbitrary set of unpaired LR and HR images and see how far a step can go in such a realistic and "unsupervised" setting. To do so, we firstly train a degradation generation network to generate realistic LR images and, more importantly, to capture their distribution (i.e., learning to zoom out). Instead of assuming the domain gap has been eliminated, we minimize the discrepancy between the generated data and real data while learning a degradation adaptive SR network (i.e., learning to zoom in). The proposed unpaired method achieves state-of- the-art SR results on real-world images, even in the datasets that favour the paired-learning methods more. Wei Sun 0036, Dong Gong, Qinfeng Shi, Anton van den Hengel, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | Embarrassingly Simple Binarization for Deep Single Imagery Super-Resolution NetworksabstractDeep convolutional neural networks (DCCNs) have shown pleasing performance in single image super-resolution (SISR). To deploy them onto real devices with limited storage and computational resources, a promising solution is to binarize the network, i.e., quantize each float-point weight and activation into 1 bit. However, existing works on binarizing DCNNs still suffer from severe performance degradation in SISR. To mitigate this problem, we argue that the performance degradation mainly comes from no appropriate constraint on the network weights, which causes it difficult to sensitively reverse the binarization results of these weights using the backpropagated gradient during training and thus limits the flexibility of network in respect of fitting extensive training samples. Inspired by this, we present an embarrassingly simple but effective binarization scheme for SISR, which can obviously relieve the performance degeneration resulted from network binarization and is applicable to different DCNN architectures. Specifically, we force each weight to follow a compact uniform prior, with which the weight will be given a very small absolute value close to zero and its binarization result can be straightforwardly reversed even by a small backpropagated gradient. By doing this, the flexibility and the generalization performance of the binarized network can be improved. Moreover, such a prior performs much better when introducing real identity shortcuts into the network. In addition, to avoid falling into bad local minima during training, we employ a pixel-wise curriculum learning strategy to learn the constrained weights in an easy-to-hard manner. Experiments on four SISR benchmark datasets demonstrate the effectiveness of the proposed binarization method in terms of binarizing different SISR network architectures, e.g., it even achieves performance comparable to the baseline with 5 quantization bits. Lei Zhang 0054, Zhiqiang Lang, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Attend to the Difference: Cross-Modality Person Re-Identification via Contrastive CorrelationabstractThe problem of cross-modality person re-identification has been receiving increasing attention recently, due to its practical significance. Motivated by the fact that human usually attend to the difference when they compare two similar objects, we propose a dual-path cross-modality feature learning framework which preserves intrinsic spatial structures and attends to the difference of input cross-modality image pairs. Our framework is composed by two main components: a Dual-path Spatial-structure-preserving Common Space Network (DSCSN) and a Contrastive Correlation Network (CCN). The former embeds cross-modality images into a common 3D tensor space without losing spatial structures, while the latter extracts contrastive features by dynamically comparing input image pairs. Note that the representations generated for the input RGB and Infrared images are mutually dependant to each other. We conduct extensive experiments on two public available RGB-IR ReID datasets, SYSU-MM01 and RegDB, and our proposed method outperforms state-of-the-art algorithms by a large margin with both full and simplified evaluation modes. Shizhou Zhang, Peng Wang 0015, Guoqiang Liang 0001, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 6 |
| 2021 | Attention-Guided Deep Neural Network With Multi-Scale Feature Fusion for Liver Vessel SegmentationabstractLiver vessel segmentation is fast becoming a key instrument in the diagnosis and surgical planning of liver diseases. In clinical practice, liver vessels are normally manual annotated by clinicians on each slice of CT images, which is extremely laborious. Several deep learning methods exist for liver vessel segmentation, however, promoting the performance of segmentation remains a major challenge due to the large variations and complex structure of liver vessels. Previous methods mainly using existing UNet architecture, but not all features of the encoder are useful for segmentation and some even cause interferences. To overcome this problem, we propose a novel deep neural network for liver vessel segmentation, called LVSNet, which employs special designs to obtain the accurate structure of the liver vessel. Specifically, we design Attention-Guided Concatenation (AGC) module to adaptively select the useful context features from low-level features guided by high-level features. The proposed AGC module focuses on capturing rich complemented information to obtain more details. In addition, we introduce an innovative multi-scale fusion block by constructing hierarchical residual-like connections within one single residual block, which is of great importance for effectively linking the local blood vessel fragments together. Furthermore, we construct a new dataset containing 40 thin thickness cases (0.625 mm) which consist of CT volumes and annotated vessels. To evaluate the effectiveness of the method with minor vessels, we also propose an automatic stratification method to split major and minor liver vessels. Extensive experimental results demonstrate that the proposed LVSNet outperforms previous methods on liver vessel segmentation datasets. Additionally, we conduct a series of ablation studies that comprehensively support the superiority of the underlying concepts. Qingsen Yan, Bo Wang 0011, Wei Zhang 0098, Chuan Luo 0003, Wei Xu 0005, Zhengqing Xu, Yanning Zhang 0001, Qinfeng Shi, Liang Zhang 0010, Zheng You |
IEEE J. Biomed. Health Informatics | 7 |
| 2021 | A Robust Attentional Framework for License Plate Recognition in the WildabstractRecognizing car license plates in natural scene images is an important yet still challenging task in realistic applications. Many existing approaches perform well for license plates collected under constrained conditions,e.g., shooting in frontal and horizontal view-angles and under good lighting conditions. However, their performance drops significantly in an unconstrained environment that features rotation, distortion, occlusion, blurring, shading or extreme dark or bright conditions. In this work, we propose a robust framework for license plate recognition in the wild. It is composed of a tailored CycleGAN model for license plate image generation and an elaborate designed image-to-sequence network for plate recognition. On one hand, the CycleGAN based plate generation engine alleviates the exhausting human annotation work. Massive amount of training data can be obtained with a more balanced character distribution and various shooting conditions, which helps to boost the recognition accuracy to a large extent. On the other hand, the 2D attentional based license plate recognizer with an Xception-based CNN encoder is capable of recognizing license plates with different patterns under various scenarios accurately and robustly. Without using any heuristics rule or post-processing, our method achieves the state-of-the-art performance on four public datasets, which demonstrates the generality and robustness of our framework. Moreover, we released a new license plate dataset, named “CLPD”, with 1200 images from all 31 provinces in mainland China. The dataset can be available from:https://github.com/wangpengnorman/CLPD_dataset. Linjiang Zhang, Peng Wang 0015, Hui Li 0031, Zhen Li 0068, Chunhua Shen, Yanning Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2021 | Person Re-Identification in Aerial ImageryabstractNowadays, with the rapid development of consumer Unmanned Aerial Vehicles (UAVs), visual surveillance by utilizing the UAV platform has been very attractive. Most of the research works for UAV captured visual data are mainly focused on the tasks of object detection and tracking. However, limited attention has been paid to the task of person Re-identification (ReID) which has been widely studied in ordinary surveillance cameras with fixed emplacements. In this paper, to facilitate the research of person ReID in aerial imagery, we collect a large scale airborne person ReID dataset named as Person ReID in Aerial Imagery (PRAI-1581), which consists of 39,461 images of 1581 person identities. The images of the dataset are shot by two DJI consumer UAVs flying at an altitude ranging from 20 to 60 meters above the ground, which covers most of the real UAV surveillance scenarios. In addition, we propose to utilize subspace pooling of convolution feature maps to represent the input person images. Our method can learn a discriminative and compact feature representation for ReID in aerial imagery and can be trained in an end-to-end fashion efficiently. We conduct extensive experiments on the proposed dataset and the experimental results demonstrate that re-identifying persons in aerial imagery is a challenging problem, where our method performs favorably against state of the arts. Shizhou Zhang, Xing Wei 0001, Peng Wang 0015, Bingliang Jiao, Yanning Zhang 0001 |
IEEE Trans. Multim. | 7 |
| 2021 | Deep Blind Hyperspectral Image Super-ResolutionabstractThe production of a high spatial resolution (HR) hyperspectral image (HSI) through the fusion of a low spatial resolution (LR) HSI with an HR multispectral image (MSI) has underpinned much of the recent progress in HSI super-resolution. The premise of these signs of progress is that both the degeneration from the HR HSI to LR HSI in the spatial domain and the degeneration from the HR HSI to HR MSI in the spectral domain are assumed to be known in advance. However, such a premise is difficult to achieve in practice. To address this problem, we propose to incorporate degeneration estimation into HSI super-resolution and present an unsupervised deep framework for "blind" HSIs super-resolution where the degenerations in both domains are unknown. In this framework, we model the latent HR HSI and the unknown degenerations with deep network structures to regularize them instead of using handcrafted (or shallow) priors. Specifically, we generate the latent HR HSI with an image-specific generator network and structure the degenerations in spatial and spectral domains through a convolution layer and a fully connected layer, respectively. By doing this, the proposed framework can be formulated as an end-to-end deep network learning problem, which is purely supervised by those two input images (i.e., LR HSI and HR MSI) and can be effectively solved by the backpropagation algorithm. Experiments on both natural scene and remote sensing HSI data sets show the superior performance of the proposed method in coping with unknown degeneration either in the spatial domain, spectral domain, or even both of them. Lei Zhang 0054, Jiangtao Nie, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | Pixel-Aware Deep Function-Mixture Network for Spectral Super-ResolutionabstractSpectral super-resolution (SSR) aims at generating a hyperspectral image (HSI) from a given RGB image. Recently, a promising direction is to learn a complicated mapping function from the RGB image to the HSI counterpart using a deep convolutional neural network. This essentially involves mapping the RGB context within a size-specific receptive field centered at each pixel to its spectrum in the HSI. The focus thereon is to appropriately determine the receptive field size and establish the mapping function from RGB context to the corresponding spectrum. Due to their differences in category or spatial position, pixels in HSIs often require different-sized receptive fields and distinct mapping functions. However, few efforts have been invested to explicitly exploit this prior.To address this problem, we propose a pixel-aware deep function-mixture network for SSR, which is composed of a new class of modules, termed function-mixture (FM) blocks. Each FM block is equipped with some basis functions, i.e., parallel subnets of different-sized receptive fields. Besides, it incorporates an extra subnet as a mixing function to generate pixel-wise weights, and then linearly mixes the outputs of all basis functions with those generated weights. This enables us to pixel-wisely determine the receptive field size and the mapping function. Moreover, we stack several such FM blocks to further increase the flexibility of the network in learning the pixel-wise mapping. To encourage feature reuse, intermediate features generated by the FM blocks are fused in late stage, which proves to be effective for boosting the SSR performance. Experimental results on three benchmark HSI datasets demonstrate the superiority of the proposed method. Lei Zhang 0054, Zhiqiang Lang, Peng Wang 0023, Wei Wei 0008, Shengcai Liao, Ling Shao 0001, Yanning Zhang 0001 |
AAAI | 7 |
| 2020 | Unsupervised Adaptation Learning for Hyperspectral Imagery Super-ResolutionabstractThe key for fusion based hyperspectral image (HSI) super-resolution (SR) is to infer the posteriori of a latent HSI using appropriate image prior and likelihood that depends on degeneration. However, in practice the priors of high-dimensional HSIs can be extremely complicated and the degeneration is often unknown. Consequently most existing approaches that assume a shallow hand-crafted image prior and a pre-defined degeneration, fail to well generalize in real applications. To tackle this problem, we present an unsupervised adaptation learning (UAL) framework. Instead of directly modelling the complicated image prior, we propose to first implicitly learn a general image prior using deep networks and then adapt it to a specific HSI. Following this idea, we develop a two-stage SR network that leverages two consecutive modules: a fusion module and an adaptation module, to recover the latent HSI in a coarse-to-fine scheme. The fusion module is pretrained in a supervised manner on synthetic data to capture a spatial-spectral prior that is general across most HSIs. To adapt the learned general prior to the specific HSI under unknown degeneration, we introduce a simple degeneration network to assist learning both the adaptation module and the degeneration in an unsupervised way. In this way, the resultant image-specific prior and the estimated degeneration can benefit the inference of a more accurate posteriori, thereby increasing generalization capacity. To verify the efficacy of UAL, we extensively evaluate it on four benchmark datasets and report strong results that surpass existing approaches. Lei Zhang 0054, Jiangtao Nie, Wei Wei 0008, Yanning Zhang 0001, Shengcai Liao, Ling Shao 0001 |
CVPR | 4 |
| 2020 | Blindly Assess Image Quality in the Wild Guided by a Self-Adaptive Hyper NetworkabstractBlind image quality assessment (BIQA) for authentically distorted images has always been a challenging problem, since images captured in the wild include varies contents and diverse types of distortions. The vast majority of prior BIQA methods focus on how to predict synthetic image quality, but fail when applied to real-world distorted images. To deal with the challenge, we propose a self-adaptive hyper network architecture to blind assess image quality in the wild. We separate the IQA procedure into three stages including content understanding, perception rule learning and quality predicting. After extracting image semantics, perception rule is established adaptively by a hyper network, and then adopted by a quality prediction network. In our model, image quality can be estimated in a self-adaptive manner, thus generalizes well on diverse images captured in the wild. Experimental results verify that our approach not only outperforms the state-of-the-art methods on challenging authentic image databases but also achieves competing performances on synthetic image databases, though it is not explicitly designed for the synthetic task. Shaolin Su, Qingsen Yan, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001 |
CVPR | 7 |
| 2020 | NAS-FCOS: Fast Neural Architecture Search for Object DetectionabstractThe success of deep neural networks relies on significant architecture engineering. Recently neural architecture search (NAS) has emerged as a promise to greatly reduce manual effort in network design by automatically searching for optimal architectures, although typically such algorithms need an excessive amount of computational resources, e.g., a few thousand GPU-days. To date, on challenging vision tasks such as object detection, NAS, especially fast versions of NAS, is less studied. Here we propose to search for the decoder structure of object detectors with search efficiency being taken into consideration. To be more specific, we aim to efficiently search for the feature pyramid network (FPN) as well as the prediction head of a simple anchor-free object detector, namely FCOS, using a tailored reinforcement learning paradigm. With carefully designed search space, search algorithms and strategies for evaluating network quality, we are able to efficiently search a top-performing detection architecture within 4 days using 8 V100 GPUs. The discovered architecture surpasses state-of-the-art object detection models (such as Faster R-CNN, RetinaNet and FCOS) by 1.5 to 3.5 points in AP on the COCO dataset, with comparable computation complexity and memory footprint, demonstrating the efficacy of the proposed NAS for object detection. Ning Wang 0020, Yang Gao 0001, Hao Chen 0041, Peng Wang 0015, Zhi Tian, Chunhua Shen, Yanning Zhang 0001 |
CVPR | 7 |
| 2020 | Distributed Q-Learning-Assisted Grant-Free NORA for Massive Machine-Type CommunicationsabstractLarge-scale connectivity support is a critical challenge in the massive machine-type communications scenario. Grant-free random access (RA) is a promising solution because it can reduce severe signaling overhead in contention-based RA procedure. However, there will still be collisions due to the random selection of spectrum resources by the devices. Therefore, we propose a distributed Q-learning-assisted grant-free RA scheme to alleviate the collisions between devices. Considering the characteristic of the machine-type communications devices with bursty traffic, the random packet arrival model is adopted in this paper. In order to cope with the difficulties brought by the random transmission of devices to Q-learning, an action reward based on the active probabilities of devices is designed. In addition, we introduce the power domain nor-orthogonal multiple access to further enhance the number of accessible devices. Numerical results demonstrate the advantages of the proposed scheme from the devices' successful access probability. Zhenjiang Shi, Wei Gao 0047, Jiajia Liu 0001, Nei Kato, Yanning Zhang 0001 |
GLOBECOM | 5 |
| 2020 | Unsupervised Deep Hyperspectral Super-Resolution With Unregistered ImagesabstractFusion based hyperspectral image (HSI) super-resolution has long been the research focus of hyperspectral image processing since it can generate a high-resolution (HR) HSI in both spatial and spectral domains. However, the success of the existing fusion based HSI super-resolution methods depends on the premise that the images utilized for fusion (i.e. the input low-spatial-resolution HSI and the low-spectral-resolution multispectral image) are exactly registered. Although such a premise is too idealistic to comply with in real cases, few efforts have considered this problem. To fill this gap, we propose to incorporate image registration into HSI super-resolution for joint unsupervised learning in this study. Specifically, a spatial transformer network (STN) is introduced to learn the parameters of the affine transformation between the input two images. In order to avoid over-fitting, we constrain the STN with a novel constraint during learning. By doing this, both the STN and super-resolution network can be cast into a weighted joint learning model without any supervision from the latent HR HSI. Experimental results demonstrate the effectiveness of the proposed method in coping with unregistered input images. Jiangtao Nie, Lei Zhang 0054, Wei Wei 0008, Chen Ding 0002, Yanning Zhang 0001 |
ICME | 5 |
| 2020 | Attention-Based Network For Low-Light Image EnhancementabstractThe captured images under low-light conditions often suffer insufficient brightness and notorious noise. Hence, low-light image enhancement is a key challenging task in computer vision. A variety of methods have been proposed for this task, but these methods often failed in an extreme low-light environment and amplified the underlying noise in the input image. To address such a difficult problem, this paper presents a novel attention-based neural network to generate high-quality enhanced low-light images from the raw sensor data. Specifically, we first employ attention strategy (i.e. spatial attention and channel attention modules) to suppress undesired chromatic aberration and noise. The spatial attention module focuses on denoising by taking advantage of the non-local correlation in the image. The channel attention module guides the network to refine redundant colour features. Furthermore, we propose a new pooling layer, called inverted shuffle layer, which adaptively selects useful information from previous features. Extensive experiments demonstrate the superiority of the proposed network in terms of suppressing the chromatic aberration and noise artifacts in enhancement, especially when the low-light image has severe noise. Qingsen Yan, Yu Zhu 0004, Xianjun Li, Jinqiu Sun, Yanning Zhang 0001 |
ICME | 6 |
| 2020 | Deep Self-Supervised Learning for Few-Shot Hyperspectral Image ClassificationabstractDespite the success of deep learning based methods for hyperspectral imagery (HSI) classification, they demand amounts of labeled samples for training whereas the labeled samples in lots of applications are always insufficient due to the expensive manual annotation cost. To address this problem, we propose a two-branch deep learning based method for few-shot HSI classification, where two branches separately accomplish HSI classification in a cube-wise level and a cube-pair level. With a shared feature extractor sub-network, the self-supervised knowledge contained in the cube-pair branch provides an effective way to regularize the original few-shot HSI classification branch (i.e., cube-wise branch) with limited labeled samples, which thus improves the performance of HSI classification. The superiority of the proposed method on few-shot HSI classification is demonstrated experimentally on two HSI benchmark datasets. Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
IGARSS | 4 |
| 2020 | Enhancing Self-supervised Monocular Depth Estimation via Incorporating Robust ConstraintsabstractSelf-supervised depth estimation has shown great prospects in inferring 3D structures using purely unannotated images. However, its performance usually drops when trained on the images with changing brightness and moving objects. In this paper, we address this issue by enhancing the robustness of the self-supervised paradigm using a set of image-based and geometry-based constraints. Our contributions are threefold, 1) we propose a gradient-based robust photometric loss which restrains the false supervisory signals caused by brightness changes, 2) we propose to filter out the unreliable areas that violate the rigid assumption by a novel combined selective mask, which is computed on the forward pass of the network by leveraging the inter-loss consistency and the loss-gradient consistency, and 3) we constrain the motion estimation network to generate across-frame consistent motions via proposing a triplet-based cycle consistency constraint. Extensive experiments conducted on KITTI, Cityscape and Make3D datasets demonstrate the superiority of our method, that the proposed method can effectively handle complex scenes with changing brightness and object motions. Both qualitative and quantitative results show that the proposed method outperforms the state-of-the-art methods. Rui Li 0013, Xiantuo He, Yu Zhu 0004, Xianjun Li, Jinqiu Sun, Yanning Zhang 0001 |
ACM Multimedia | 6 |
| 2020 | Emotion recognition from spatiotemporal EEG representations with hybrid convolutional recurrent neural networks via wearable multi-channel headset
Jingxia Chen, Dongmei Jiang, Yanning Zhang 0001 |
Comput. Commun. | 3 |
| 2020 | Ghost Removal via Channel Attention in Exposure Fusion
Qingsen Yan, Bo Wang 0011, Xianjun Li, Qinfeng Shi, Zheng You, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001 |
Comput. Vis. Image Underst. | 10 |
| 2020 | Adaptive Importance Learning for Improving Lightweight Image Super-Resolution Network
Lei Zhang 0054, Peng Wang 0023, Chunhua Shen, Lingqiao Liu, Wei Wei 0008, Yanning Zhang 0001, Anton van den Hengel |
Int. J. Comput. Vis. | 6 |
| 2020 | Blur kernel estimation of noisy-blurred image via dynamic structure prior
Xueling Chen, Yu Zhu 0004, Wei Liu 0044, Jinqiu Sun, Yanning Zhang 0001 |
Neurocomputing | 5 |
| 2020 | Cross-modal image fusion guided by subjective visual attention
Aiqing Fang, Yanning Zhang 0001 |
Neurocomputing | 3 |
| 2020 | Autonomous deep learning: A genetic DCNN designer for image classification
Benteng Ma, Yong Xia 0001, Yanning Zhang 0001 |
Neurocomputing | 4 |
| 2020 | Video super-resolution via dense non-local spatial-temporal convolutional network
Wei Sun 0036, Jinqiu Sun, Yu Zhu 0004, Yanning Zhang 0001 |
Neurocomputing | 4 |
| 2020 | Attention-guided dual spatial-temporal non-local network for video super-resolution
Wei Sun 0036, Yanning Zhang 0001 |
Neurocomputing | 2 |
| 2020 | A holistic representation guided attention network for scene text recognition
Lu Yang 0016, Peng Wang 0015, Hui Li 0031, Zhen Li 0068, Yanning Zhang 0001 |
Neurocomputing | 5 |
| 2020 | Topology Poisoning Attack in SDN-Enabled Vehicular Edge NetworkabstractThe development of the Internet of Vehicles (IoV) has made people's lives and travels safer, more efficient, and more comfortable. The combination of edge computing and IoV can provide processing and storage capabilities close to vehicles, thus becoming a potential paradigm. At this time, the software-defined networking (SDN) architecture is extremely necessary to realize centralized control and convenient management for complex and dynamic vehicular edge networks. However, as the brain of the SDN architecture, little attention has been paid to the security of the SDN controller. Once the controller is threatened, severe global chaos may happen. Therefore, in this article, we study the attack against the SDN controller, which is the topology poisoning attack. We successfully implement this attack in four mainstream controllers and analyze its impact from multiple levels. To the best of our knowledge, we are the first to study this attack in the vehicular edge network. In addition, in view of the counter-attacks of the existing defence mechanisms, we propose an attack-tolerance scheme based on deep reinforcement learning (DRL) to enhance the vehicular edge network with a certain degree of self-recovery. Jiadai Wang, Yawen Tan, Jiajia Liu 0001, Yanning Zhang 0001 |
IEEE Internet Things J. | 4 |
| 2020 | Smart and Resilient EV Charging in SDN-Enhanced Vehicular Edge Computing NetworksabstractSmart grid delivers power with two-way flows of electricity and information with the support of information and communication technologies. Electric vehicles (EVs) with rechargeable batteries can be powered by external sources of electricity from the grid, and thus charging scheduling that guides low-battery EVs to charging services is significant for service quality improvement of EV drivers. The revolution of communications and data analytics driven by massive data in smart grid brings many challenges as well as chances for EV charging scheduling, and how to schedule EV charging in a smart and resilient way has inevitably become a crucial problem. Toward this end, we in this paper leverage the techniques of software defined networking and vehicular edge computing to investigate a joint problem of fast charging station selection and EV route planning. Our objective is to minimize the total overhead from users' perspective, including time and charging fares in the whole process, considering charging availability and electricity price fluctuation. A deep reinforcement learning (DRL) based solution is proposed to determine an optimal charging scheduling policy for low-battery EVs. Besides, in response to dynamic EV charging, we further develop a resilient EV charging strategy based on incremental update, with EV drivers' user experience being well considered. Extensive simulations demonstrate that our proposed DRL-based solution obtains near-optimal EV charging overhead with good adaptivity, and the solution with incremental update achieves much higher computation efficiency than conventional game-theoretical method in dynamic EV charging. Jiajia Liu 0001, Hongzhi Guo 0005, Jingyu Xiong, Nei Kato, Jie Zhang 0052, Yanning Zhang 0001 |
IEEE J. Sel. Areas Commun. | 6 |
| 2020 | Hyperspectral Image Classification With Transfer Learning and Markov Random FieldsabstractThis letter provides a brand new way of feature extraction, which can be applied in the supervised classification of hyperspectral image. The convolutional neural network (CNN) has been proven to be an effective method of image classification. However, due to its long training time, it requires a large amount of the labeled data to achieve the expected outcome. To decrease the training time and reduce the dependence on large labeled data set, we propose using the method of transfer learning by taking the advantage of Bayesian framework to integrate with spectrum and spatial information, making use of the Markov property of images to distinguish and separate the ones with class tags, and employing the CNN trained by band samples randomly selected from the data sets. The method of classification mentioned in our letter makes use of the real hyperspectral data sets to perform the experimental evaluation. The result demonstrates that our method is superior to the previous methods. Yue Zhang 0036, Yi Li 0034, Yanning Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2020 | Hyperspectral Image Classification With Data Augmentation and Classifier FusionabstractRecently, deep convolutional neural network (DCNN)-based methods have achieved much success in hyperspectral image (HSI) classification, when sufficient labeled samples are provided during training. However, due to the expensive cost of labeling in HSIs, only limited labeled samples can be given in practice, which often causes these methods to be overfitting. To address this problem, we present a new HSI classification method in this study, which is constructed in the following two steps. First, we establish a data mixture model to augment the labeled training set quadratically and train a DCNN-based classifier on it. Then, through randomly sampling the coefficient in the data mixture model, we obtain several independent classifiers and fuse them with a voting strategy to produce the final classification results. Since both data augmentation and classifier fusion are effective to deal with limited samples, the proposed method shows superior performance in the classification of HSIs, which can be demonstrated by the experimental results on two benchmark HSI data sets. Cong Wang 0013, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2020 | Dim small target detection based on convolutinal neural network in star image
Danna Xue, Jinqiu Sun, Yaoqi Hu, Yushu Zheng, Yu Zhu 0004, Yanning Zhang 0001 |
Multim. Tools Appl. | 6 |
| 2020 | NFN+: A novel network followed network for retinal vessel segmentation
Yicheng Wu 0001, Yong Xia 0001, Yang Song 0001, Yanning Zhang 0001, Tom Weidong Cai |
Neural Networks | 4 |
| 2020 | Robust Visual Tracking based on Adversarial Unlabeled Instance Generation with Label Smoothing Loss Regularization
Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Garth Douglas Cooper, Yanning Zhang 0001 |
Pattern Recognit. | 6 |
| 2020 | Visual question answering model based on visual relationship detection
Yuling Xi, Yanning Zhang 0001, Songtao Ding, Shaohua Wan 0001 |
Signal Process. Image Commun. | 2 |
| 2020 | Loanword Identification in Low-Resource Languages with Minimal SupervisionabstractBilingual resources play a very important role in many natural language processing tasks, especially the tasks in cross-lingual scenarios. However, it is expensive and time consuming to build such resources. Lexical borrowing happens in almost every language. This inspires us to detect these loanwords effectively, and to use the “loanword (in receipt language)”-“donor word (in donor language)” to extend the bilingual resource for NLP tasks in low-resource languages. In this article, we propose a novel method to identify loanwords in Uyghur. The most important advantage of this method is that the model only relies on large amount of monolingual corpora and only a small scale of annotated data. Our loanword identification model includes two parts: loanword candidate generation and loanword prediction. In the first part, we use two large-scale monolingual corpora and a small bilingual dictionary to train a cross-lingual embedding model. Since semantic unrelated words often cannot be treated as loanword pairs, a loanword candidate list will be generated according to this model and a word list in Uyghur. In the second part, we predict from the preceding candidates based on a log-linear model that integrates several features such as pronunciation similarity, part-of-speech tags, and hybrid language modeling. To evaluate the effectiveness of our proposed method, we conduct two types of experiments: loanword identification and OOV translation. Experimental results showed that (1) our proposed method achieved significant F1 improvements compared to other models in all four loanword identification tasks in Uyghur, and (2) after extending the existing translation models with loanword identification results, OOV rates in several language pairs reduced significantly and the translation performance improved. Chenggang Mi 0001, Lei Xie 0001, Yanning Zhang 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2020 | Towards Effective Deep Embedding for Zero-Shot LearningabstractZero-shot learning (ZSL) can be formulated as a cross-domain matching problem: after being projected into a joint embedding space, a visual sample will match against all candidate class-level semantic descriptions and be assigned to the nearest class. In this process, the embedding space underpins the success of such matching and is crucial for ZSL. In this paper, we conduct an in-depth study on the construction of embedding space for ZSL and posit that an ideal embedding space should satisfy two criteria: intra-class compactness and inter-class separability. While the former encourages the embeddings of visual samples of one class to distribute tightly close to the semantic description embedding of this class, the latter requires embeddings from different classes to be well separated from each other. Towards this goal, we present a simple but effective two-branch network to simultaneously map semantic descriptions and visual samples into a joint space, on which visual embeddings are forced to regress to their class-level semantic embeddings and the embeddings crossing classes are required to be distinguishable by a trainable classifier. Furthermore, we extend our method to a transductive setting to better handle the model bias problem in ZSL (i.e., samples from unseen classes tend to be categorized into seen classes) with minimal extra supervision. Specifically, we propose a pseudo labeling strategy to progressively incorporate the testing samples into the training process and thus balance the model between seen and unseen classes. Experimental results on five standard ZSL datasets show the superior performance of the proposed method and its transductive extension. Lei Zhang 0054, Peng Wang 0023, Lingqiao Liu, Chunhua Shen, Wei Wei 0008, Yanning Zhang 0001, Anton van den Hengel |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | Automobile Driver Fingerprinting: A New Machine Learning Based Authentication SchemeabstractAdvanced technologies are constantly emerging in automobile industry, which not only provides drivers with a comfortable driving experience, but also enhances the safety of passengers. However, there are still some security issues need to be solved in automobiles, such as automobile driver fingerprinting. At present, identification technologies, such as fingerprint recognition and iris recognition, cannot monitor the driver's identity in real-time manner. Therefore, it is of great significance to design a real-time automobile driver fingerprinting scheme to ensure the safety of people's properties and even lives. Different from previous work concerning automobile driver fingerprinting, in this article, we conduct a comprehensive study on behavioral characteristics of drivers in two vehicles, namely Luxgen U5 SUV and Buick Regal. We exploit the actual data of the controller area network to construct a driver identity comparison library by extracting and processing the feature data. Then, we construct a combined model based on convolutional neural network and support vector domain description to achieve efficient automobile driver fingerprinting. Extensive experimental results show that the proposed driver fingerprinting scheme can dynamically match the driver's identity in real time without affecting the normal driving. Yijie Xun, Jiajia Liu 0001, Nei Kato, Yongqiang Fang, Yanning Zhang 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2020 | Deep HDR Imaging via A Non-Local NetworkabstractOne of the most challenging problems in reconstructing a high dynamic range (HDR) image from multiple low dynamic range (LDR) inputs is the ghosting artifacts caused by the object motion across different inputs. When the object motion is slight, most existing methods can well suppress the ghosting artifacts through aligning LDR inputs based on optical flow or detecting anomalies among them. However, they often fail to produce satisfactory results in practice, since the real object motion can be very large. In this study, we present a novel deep framework, termed NHDRRnet, which adopts an alternative direction and attempts to remove ghosting artifacts by exploiting the non-local correlation in inputs. In NHDRRnet, we first adopt an Unet architecture to fuse all inputs and map the fusion results into a low-dimensional deep feature space. Then, we feed the resultant features into a novel global non-local module which reconstructs each pixel by weighted averaging all the other pixels using the weights determined by their correspondences. By doing this, the proposed NHDRRnet is able to adaptively select the useful information (e.g., which are not corrupted by large motions or adverse lighting conditions) in the whole deep feature space to accurately reconstruct each pixel. In addition, we also incorporate a triple-pass residual module to capture more powerful local features, which proves to be effective in further boosting the performance. Extensive experiments on three benchmark datasets demonstrate the superiority of the proposed NDHRnet in terms of suppressing the ghosting artifacts in HDR reconstruction, especially when the objects have large motions. Qingsen Yan, Lei Zhang 0054, Yu Liu 0029, Yu Zhu 0004, Jinqiu Sun, Qinfeng Shi, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 7 |
| 2020 | Evaluating Local Geometric Feature Representations for 3D Rigid Data MatchingabstractLocal geometric descriptors act as an essential component for 3D rigid data matching. A rotational invariant local geometric descriptor usually consists of two components: local reference frame (LRF) and feature representation. However, existing evaluation efforts have mainly been paid on the LRF or the overall descriptor and the quantitative comparison of feature representations remains unexplored. This paper fills the gap by comprehensively evaluating nine state-of-the-art local geometric feature representations. In particular, our evaluation assesses feature representations based on ground-truth LRFs such that the ranking of tested methods is more convincing as compared with existing studies. The experiments are deployed on six standard datasets with various application scenarios (shape retrieval, point cloud registration, and object recognition) and data modalities (LiDAR, Kinect, and Space Time) as well as perturbations including Gaussian noise, shot noise, data decimation, clutter, occlusion, and limited overlap. The evaluated terms cover the major concerns for a feature representation, e.g., distinctiveness, robustness, compactness, and efficiency. The outcomes present interesting findings that may shed new light on this community and provide complementary perspectives to existing evaluations on the topic of local geometric feature description. A summary of evaluated methods regarding their peculiarities is finally presented to guide real-world applications and new descriptor crafting. Jiaqi Yang 0002, Siwen Quan, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | 3D APA-Net: 3D Adversarial Pyramid Anisotropic Convolutional Network for Prostate Segmentation in MR ImagesabstractAccurate and reliable segmentation of the prostate gland using magnetic resonance (MR) imaging has critical importance for the diagnosis and treatment of prostate diseases, especially prostate cancer. Although many automated segmentation approaches, including those based on deep learning have been proposed, the segmentation performance still has room for improvement due to the large variability in image appearance, imaging interference, and anisotropic spatial resolution. In this paper, we propose the 3D adversarial pyramid anisotropic convolutional deep neural network (3D APA-Net) for prostate segmentation in MR images. This model is composed of a generator (i.e., 3D PA-Net) that performs image segmentation and a discriminator (i.e., a six-layer convolutional neural network) that differentiates between a segmentation result and its corresponding ground truth. The 3D PA-Net has an encoder-decoder architecture, which consists of a 3D ResNet encoder, an anisotropic convolutional decoder, and multi-level pyramid convolutional skip connections. The anisotropic convolutional blocks can exploit the 3D context information of the MR images with anisotropic resolution, the pyramid convolutional blocks address both voxel classification and gland localization issues, and the adversarial training regularizes 3D PA-Net and thus enables it to generate spatially consistent and continuous segmentation results. We evaluated the proposed 3D APA-Net against several state-of-the-art deep learning-based segmentation approaches on two public databases and the hybrid of the two. Our results suggest that the proposed model outperforms the compared approaches on three databases and could be used in a routine clinical workflow. Haozhe Jia, Yong Xia 0001, Yang Song 0001, Donghao Zhang 0004, Heng Huang 0001, Yanning Zhang 0001, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 6 |
| 2020 | Ensemble Tracking Based on Diverse Collaborative Framework With Multi-Cue Dynamic FusionabstractTracking with deep neural networks has been verified to arrive at a new level accuracy in many challenging scenarios, but the tracking robustness has been still challenged by model singularity and self-learning loop mechanism. As a promising solution for the limitations, to ensemble diverse tracking strategies into a highly-interactive framework has shown a potential effectiveness in recent studies. In this work, a collaborative tracking framework is proposed by exploiting both discriminative correlation filters and deep classifiers into an ensembling framework. With a multi-cue dynamic fusion scheme performed on all the ensembled members’ outputs, a robust long-term tracking can be achieved by calculating the optimal robustness scores based on a dynamic weighted sum of multi-cue metrics. Meanwhile, the obtained reliable and diverse training samples are also utilized to adaptively update the tracker in each branch with heuristic frequency, which is able to alleviate the training samples’ contamination and model corruption. Experiments on the OTB-2015, Temple color 128, UAV123, VOT2016, and VOT2018 benchmark datasets have shown superior performance in comparison to other state-of-the-art tracking approaches. Peng Zhang 0005, Tao Zhuo, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001 |
IEEE Trans. Multim. | 6 |
| 2020 | Learning Deep Gradient Descent Optimization for Image DeconvolutionabstractAs an integral component of blind image deblurring, non-blind deconvolution removes image blur with a given blur kernel, which is essential but difficult due to the ill-posed nature of the inverse problem. The predominant approach is based on optimization subject to regularization functions that are either manually designed or learned from examples. Existing learning-based methods have shown superior restoration quality but are not practical enough due to their restricted and static model design. They solely focus on learning a prior and require to know the noise level for deconvolution. We address the gap between the optimization- and learning-based approaches by learning a universal gradient descent optimizer. We propose a recurrent gradient descent network (RGDN) by systematically incorporating deep neural networks into a fully parameterized gradient descent scheme. A hyperparameter-free update unit shared across steps is used to generate the updates from the current estimates based on a convolutional neural network. By training on diverse examples, the RGDN learns an implicit image prior and a universal update rule through recursive supervision. The learned optimizer can be repeatedly used to improve the quality of diverse degenerated observations. The proposed method possesses strong interpretability and high generalization. Extensive experiments on synthetic benchmarks and challenging real-world images demonstrate that the proposed deep optimization method is effective and robust to produce favorable results as well as practical for real-world image deblurring applications. Dong Gong, Zhen Zhang 0008, Qinfeng Shi, Anton van den Hengel, Chunhua Shen, Yanning Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2020 | Accurate Tensor Completion via Adaptive Low-Rank RepresentationabstractLow-rank representation-based approaches that assume low-rank tensors and exploit their low-rank structure with appropriate prior models have underpinned much of the recent progress in tensor completion. However, real tensor data only approximately comply with the low-rank requirement in most cases, viz., the tensor consists of low-rank (e.g., principle part) as well as non-low-rank (e.g., details) structures, which limit the completion accuracy of these approaches. To address this problem, we propose an adaptive low-rank representation model for tensor completion that represents low-rank and non-low-rank structures of a latent tensor separately in a Bayesian framework. Specifically, we reformulate the CANDECOMP/PARAFAC (CP) tensor rank and develop a sparsity-induced prior for the low-rank structure that can be used to determine tensor rank automatically. Then, the non-low-rank structure is modeled using a mixture of Gaussians prior that is shown to be sufficiently flexible and powerful to inform the completion process for a variety of real tensor data. With these two priors, we develop a Bayesian minimum mean-squared error estimate framework for inference. The developed framework can capture the important distinctions between low-rank and non-low-rank structures, thereby enabling more accurate model, and ultimately, completion. For various applications, compared with the state-of-the-art methods, the proposed model yields more accurate completion results. Lei Zhang 0054, Wei Wei 0008, Qinfeng Shi, Chunhua Shen, Anton van den Hengel, Yanning Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2019 | Attention-Guided Network for Ghost-Free High Dynamic Range ImagingabstractGhosting artifacts caused by moving objects or misalignments is a key challenge in high dynamic range (HDR) imaging for dynamic scenes. Previous methods first register the input low dynamic range (LDR) images using optical flow before merging them, which are error-prone and cause ghosts in results. A very recent work tries to bypass optical flows via a deep network with skip-connections, however, which still suffers from ghosting artifacts for severe movement. To avoid the ghosting from the source, we propose a novel attention-guided end-to-end deep neural network (AHDRNet) to produce high-quality ghost-free HDR images. Unlike previous methods directly stacking the LDR images or features for merging, we use attention modules to guide the merging according to the reference image. The attention modules automatically suppress undesired components caused by misalignments and saturation and enhance desirable fine details in the non-reference images. In addition to the attention model, we use dilated residual dense block (DRDB) to make full use of the hierarchical features and increase the receptive field for hallucinating the missing details. The proposed AHDRNet is a non-flow-based method, which can also avoid the artifacts generated by optical-flow estimation error. Experiments on different datasets show that the proposed AHDRNet can achieve state-of-the-art quantitative and qualitative results. Qingsen Yan, Dong Gong, Qinfeng Shi, Anton van den Hengel, Chunhua Shen, Ian D. Reid 0001, Yanning Zhang 0001 |
CVPR | 7 |
| 2019 | MACH: Movement Aware CoMP Handover in Heterogeneous Ultra-Dense NetworksabstractThe densification of small cells, ultimately towards ultra-dense networks (UDN), makes coordinated multipoint (CoMP) a feasible transmission solution for mobile users. However, CoMP may increase the handover rate, as users move across small and irregular BS cooperation regions. In this paper, we consider movement aware CoMP handover (MACH) in heterogeneous UDNs. Unlike most prior works, which focus on the handover trigger time, we explore the appropriate selection of BS cooperation set to reduce handover rate. By estimating cell dwell time, a user would be intelligently assigned to macro cell or small cell according to its movement trend. Moreover, we achieve a balance between BSs with long dwell time and the current best performed BS for multipoint cooperation while user moving. The performance of the proposed MACH is analyzed in terms of coverage probability and handover probability using stochastic geometry. Through extensive simulations, we show that the analytical results fit well with simulations, and the proposed MACH outperforms the existing works in both handover probability and coverage probability. Wen Sun 0004, Lu Wang 0050, Jiajia Liu 0001, Nei Kato, Yanning Zhang 0001 |
GLOBECOM | 5 |
| 2019 | Collaborative Computation Offloading at UAV-Enhanced EdgeabstractIn conventional terrestrial cellular networks, mobile devices at the cell edge often suffer from poor channel conditions, and thus unmanned aerial vehicles (UAVs) are introduced in recent years to improve the reliability of communication links. However, with the rapid development of Internet of Things (IoT) technology, the emerging IoT applications have blooming demands for high computation capacity from the resource-constrained IoT mobile devices (IMDs), motivated by which, mobile edge computing has been envisioned as an appealing solution to the resource bottleneck problem of IMDs. In order to cope with poor communication performance and high computation demands of cell-edge IMDs, we in this paper leverage UAV-aided edge computing to collaboratively assist computation offloading, taking account of the limited battery life of both IMDs and the UAV. We investigate a joint optimization problem of collaborative computation offloading, bandwidth portion, bit allocation, and UAV trajectory design, aiming to minimize the weighted energy consumption of IMDs and the UAV. Extensive numerical results validate the necessity of introducing UAV-aided edge computing to cellular networks, and the advantages of our proposed scheme on energy savings. Jingyu Xiong, Hongzhi Guo 0005, Jiajia Liu 0001, Nei Kato, Yanning Zhang 0001 |
GLOBECOM | 5 |
| 2019 | Vehicle Re-Identification in Aerial Imagery: Dataset and ApproachabstractIn this work, we construct a large-scale dataset for vehicle re-identification (ReID), which contains 137k images of 13k vehicle instances captured by UAV-mounted cameras. To our knowledge, it is the largest UAV-based vehicle ReID dataset. To increase intra-class variation, each vehicle is captured by at least two UAVs at different locations, with diverse view-angles and flight-altitudes. We manually label a variety of vehicle attributes, including vehicle type, color, skylight, bumper, spare tire and luggage rack. Furthermore, for each vehicle image, the annotator is also required to mark the discriminative parts that helps them to distinguish this particular vehicle from others. Besides the dataset, we also design a specific vehicle ReID algorithm to make full use of the rich annotation information. It is capable of explicitly detecting discriminative parts for each specific vehicle and significantly outperforming the evaluated baselines and state-of-the-art vehicle ReID approaches. Peng Wang 0015, Bingliang Jiao, Lu Yang 0016, Shizhou Zhang, Wei Wei 0008, Yanning Zhang 0001 |
ICCV | 7 |
| 2019 | Robust and Accurate Hybrid Structure-From-MotiabstractIn this paper, we propose a hybrid Structure-from-Motion scheme which combines the strength of both global and local incremental SfM methods to get a drift-free and accurate estimation with lower time consumption. More specifically, we propose to construct a robust maximum leaf spanning tree (RMLST) from the initial scene graph and further expand it to a robust graph (RG) to grasp the global picture of camera distribution and scene structure. Then the views in the robust graph are solved in global manner as an initial estimation. After that, the remaining views are estimated with the proposed community-based local incremental approach to guarantee local accuracy and scalability. Bundle adjustment is conducted to optimize the estimation. Experiments show that our method is robust and free from the scene drift as global SfM, and shows much better efficiency than incremental approaches. Besides, our algorithm achieves higher accuracy compared with the state-of-the-art methods. Rui Li 0013, Dong Gong, Jinqiu Sun, Yu Zhu 0004, Ziwei Wei, Yanning Zhang 0001 |
ICIP | 6 |
| 2019 | Deep Spectral Super-Resolution with Noisy InputabstractLearning based methods, e.g., sparse coding or deep convolutional neural networks (DCNNs) have underpinned much of recent progress in increasing the spectral resolution of an RGB image for hyperspectral image (HSI) super-resolution. However, these methods suffer severe performance loss, when the test RGB image distributed differently from the training set, e.g., being corrupted with random noise. To mitigate this problem, we propose an unsupervised deep spectral super-resolution method, which employs a DCNN to generate the latent HSI from an input RGB and encourages it to fit the input RGB image through down-sampling in spectral domain as well as a sparse gradient prior in spatial domain. Due to the powerful capacity of DCNN in capturing the low-level image statistics, the proposed method is able to automatically accommodate the noise corruption in the input RGB image. Experimental results shows the superior performance of the proposed method. Zhiqiang Lang, Lei Zhang 0054, Wei Wei 0008, Jiangtao Nie, Chunna Tian, Yanning Zhang 0001 |
IGARSS | 6 |
| 2019 | Unsupervised deep domain adaptation for hyperspectral image classificationabstractDeep neural networks have been proven to be a promising way for hyperspectral image (HSI) classification. Their success depends on a premise that source domain (i.e., training) and target domain (i.e., test) samples are identically distributed. However, due to various imaging environments, in practice obvious distribution discrepancy often exists between these two domains, which can dramatically reduce the capacity of the classifier trained in source domain generalizing to target domain. To mitigate this problem, we present a novel deep unsupervised domain adaptation framework for HSI classification, which can simultaneously align the distributions of two domains and learn a classifier in source domain. Firstly, we employ two auto-encoder networks to separately project the samples from two domains into two low-dimensional feature spaces. Then, a multi-level maximum mean discrepancy (MMD) loss is imposed on the feature space to reduce the distribution discrepancy between two domains. Given the resultant features, a classification subnet is further learned to classify the labeled samples in source domain. Since the classifier is trained based on the domain-invariant features, it can well generalize to the target domain. Experimental results on one benchmark cross-domain HSI datasets prove the superior performance of the proposed method. Wei Li 0219, Wei Wei 0008, Lei Zhang 0054, Cong Wang 0013, Yanning Zhang 0001 |
IGARSS | 5 |
| 2019 | Robust Deep Hyperspectral Imagery Super-ResolutionabstractFusing a low spatial resolution (LR) hyperspectral image (HSI) with a high spatial resolution (HR) multi-spectral image (MSI) is an effective way for HSI super-resolution. When the input LR HSI and the HR MSI are clean, most of existing fusion based methods can produce pleasing results. However, the input HSI and MSI are often corrupted with random noise in practice, which can greatly degrade the performance of these methods. To address this problem, we present a robust deep HSI super-resolution method in this study. In contrast to leveraging a heuristic shallow sparsity or low-rank prior in previous methods, we propose to employ a deep convolution neural network as the prior of the latent HR HSI. With such a prior, the fusion based HSI super-resolution can be formulated as an end-to-end deep learning problem, which can be effectively solved with the back-propagation algorithm. Due to the deep structure, the proposed image prior is able to capture more powerful statistics of the latent HR HSI, and thus can still produce pleasing results with noisy input images. Experimental results on two benchmark datasets demonstrate the effectiveness of the proposed method. Jiangtao Nie, Lei Zhang 0054, Cong Wang 0013, Wei Wei 0008, Yanning Zhang 0001 |
IGARSS | 5 |
| 2019 | Improving Hyperspectral Image Classification with Unsupervised Knowledge LearningabstractRecently, deep convolutional neural networks(DCNNs) based methods have shown pleasing performance in hyperspectral image(HSI) classification. However, due to extensive coefficients resulted by the deep structure, these methods are prone to be overfitting during training, especially when the labeled samples are limited. To address this problem, we propose to learn the unsupervised knowledge from both unlabeled and labeled samples to regularize the conventional supervised learning. Following this idea, we present a two-branch network, in which two branches are separately utilized to perform the clustering and classification based on a shared feature extraction module. Thanks to the shared structure, the crucial unsupervised information (e.g., intra-cluster similarity & inter-cluster dissimilarity, etc.) can be injected into the supervised learning procedure, and thus leads to improved generalization capacity. Experiments on two widely used HSI datasets show the superior performance of the proposed method. Wei Wei 0008, Lei Zhang 0054, Yanning Zhang 0001 |
IGARSS | 4 |
| 2019 | Person Re-identification with Neural Architecture Search
Shizhou Zhang, Xing Wei 0001, Peng Wang 0015, Yanning Zhang 0001 |
PRCV (1) | 5 |
| 2019 | Multi-Scale Dense Networks for Deep High Dynamic Range ImagingabstractGenerating a high dynamic range (HDR) image from a set of sequential exposures is a challenging task for dynamic scenes. The most common approaches are aligning the input images to a reference image before merging them into an HDR image, but artifacts often appear in cases of large scene motion. The state-of-the-art method using deep learning can solve this problem effectively. In this paper, we propose a novel deep convolutional neural network to generate HDR, which attempts to produce more vivid images. The key idea of our method is using the coarse-to-fine scheme to gradually reconstruct the HDR image with the multi-scale architecture and residual network. By learning the relative changes of inputs and ground truth, our method can produce not only artificial free image but also restore missing information. Furthermore, we compare to existing methods for HDR reconstruction, and show high-quality results from a set of low dynamic range (LDR) images. We evaluate the results in qualitative and quantitative experiments, our method consistently produces excellent results than existing state-of-the-art approaches in challenging scenes. Qingsen Yan, Dong Gong, Qinfeng Shi, Jinqiu Sun, Ian D. Reid 0001, Yanning Zhang 0001 |
WACV | 7 |
| 2019 | ARSAC: Efficient model estimation via adaptively ranked sample consensus
Rui Li 0013, Jinqiu Sun, Dong Gong, Yu Zhu 0004, Haisen Li, Yanning Zhang 0001 |
Neurocomputing | 6 |
| 2019 | Complementary coded aperture set for compressive high-resolution imaging
Wei Sun 0036, Jinqiu Sun, Yu Zhu 0004, Yaoqi Hu, Chen Ding 0002, Haisen Li, Yanning Zhang 0001 |
Neurocomputing | 7 |
| 2019 | EMS-Net: Ensemble of Multiscale Convolutional Neural Networks for Classification of Breast Cancer Histology Images
Zhanbo Yang, Lingyan Ran, Shizhou Zhang, Yong Xia 0001, Yanning Zhang 0001 |
Neurocomputing | 5 |
| 2019 | Accurate imagery recovery using a multi-observation patch model
Lei Zhang 0054, Wei Wei 0008, Qinfeng Shi, Chunhua Shen, Anton van den Hengel, Yanning Zhang 0001 |
Inf. Sci. | 6 |
| 2019 | A Coarse-to-Fine Optimization for Hyperspectral Band SelectionabstractHyperspectral band selection is a feature selection method that selects a most representative set of bands to achieve a good performance in several tasks such as classification and anomaly detection. It reduces the burden of storage, transmission, and computation. In this letter, a two-stage band selection algorithm is introduced. It selects bands and refines the result using a linear reconstruction error criterion. Then a coarse-to-fine band selection (CFBS) strategy is applied to the two-stage band selection in order to achieve a better result. CFBS selects bands group by group. Each group is selected based on bands that are not well represented by the previous groups, trying to minimize the linear reconstruction error. Experiments show that the proposed method has a significant advancement compared with other competitors. Jianzhe Lin, Yanning Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2019 | Robust Hyperspectral Image Domain Adaptation With Noisy LabelsabstractIn hyperspectral image (HSI) classification, domain adaptation (DA) methods have been proved effective to address unsatisfactory classification results caused by the distribution difference between training (i.e., source domain) and testing (i.e., target domain) pixels. However, these methods rely on accurate labels in source domain, and seldom consider the performance drop resulted by noisy label, which often happens since labeling pixel in HSI is a challenging task. To improve the robustness of DA method to label noise, we propose a new unsupervised HSI DA method, which is constructed from both feature-level and classifier-level. First, a linear transformation function is learned in feature-level to align the source (domain) subspace with the target (domain) subspace. Then, a robust low-rank representation based classifier is developed to well cope with the features obtained from the aligned subspace. Since both subspace alignment and the classifier are immune to noisy labels, the proposed method obtains good classification results when confronting with noisy labels in source domain. Experimental results on two DA benchmarks demonstrate the effectiveness of the proposed method. Wei Wei 0008, Wei Li 0219, Lei Zhang 0054, Cong Wang 0013, Peng Zhang 0005, Yanning Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2019 | Robust artifact-free high dynamic range imaging of dynamic scenes
Qingsen Yan, Yu Zhu 0004, Yanning Zhang 0001 |
Multim. Tools Appl. | 3 |
| 2019 | Fast-Convergent Fully Connected Deep Learning Model Using Constrained Nodes Input
Chen Ding 0002, Ying Li 0017, Lei Zhang 0054, Lu Yang 0016, Wei Wei 0008, Yong Xia 0001, Yanning Zhang 0001 |
Neural Process. Lett. | 8 |
| 2019 | Human trajectory prediction in crowded scene using social-affinity Long Short-Term Memory
Zhao Pei, Xiaoning Qi, Yanning Zhang 0001, Miao Ma, Yee-Hong Yang |
Pattern Recognit. | 3 |
| 2019 | M3Net: A multi-model, multi-size, and multi-view deep neural network for brain magnetic resonance image segmentation
Yong Xia 0001, Yanning Zhang 0001 |
Pattern Recognit. | 3 |
| 2019 | Enhancing image visuality by multi-exposure fusion
Qingsen Yan, Yu Zhu 0004, Jinqiu Sun, Lei Zhang 0054, Yanning Zhang 0001 |
Pattern Recognit. Lett. | 6 |
| 2019 | Skeleton-Based Action Recognition With Gated Convolutional Neural NetworksabstractFor skeleton-based action recognition, most of the existing works used recurrent neural networks. Using convolutional neural networks (CNNs) is another attractive solution considering their advantages in parallelization, effectiveness in feature learning, and model base sufficiency. Besides these, skeleton data are low-dimensional features. It is natural to arrange a sequence of skeleton features chronologically into an image, which retains the original information. Therefore, we solve the sequence learning problem as an image classification task using CNNs. For better learning ability, we build a classification network with stacked residual blocks and having a special design called linear skip gated connection which can benefit information propagation across multiple residual blocks. When arranging the coordinates of body joints in one frame into a skeleton feature, we systematically investigate the performance of part-based, chain-based, and traversal-based orders. Furthermore, a fully convolutional permutation network is designed to learn an optimized order for data rearrangement. Without any bells and whistles, our proposed model achieves state-of-the-art performance on two challenging benchmark datasets, outperforming existing methods significantly. Congqi Cao, Cuiling Lan, Yifan Zhang 0001, Wenjun Zeng 0001, Hanqing Lu, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2019 | Semantics-Aware Visual Object TrackingabstractIn this paper, we propose a semantics-aware visual object tracking method, which introduces semantics into the tracking procedure and extends the model of an object with explicit semantics prior to enhancing the robustness of three key aspects of the tracking framework, i.e., appearance model, search scheme, and scale adaptation. We first present a semantic object proposal generation method for video sequences to generate high-quality category-oriented object proposals. Then, a hybrid semantics-aware tracking algorithm with semantic compatibility is proposed. This algorithm takes full advantages of globally sparse semantic object proposal prediction and locally dense prediction with a template model and semantic distractor-aware color appearance model. Furthermore, we propose to exploit semantics to localize object accurately via an energy minimization framework-based scale adaptation method, which jointly integrates dense location prior, instance-specific color, and category-specific semantic information. Extensive experiments are conducted on two widely used benchmarks, and the results demonstrate that our method achieves the state-of-the-art performance. Rui Yao 0006, Guosheng Lin, Chunhua Shen, Yanning Zhang 0001, Qinfeng Shi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Normalized Non-Negative Sparse Encoder for Fast Image RepresentationabstractImage representation based on sparse coding generalizes the bag of words model. Although it reduces the reconstruction error for local features to achieve the state-of-the-art image classification performance, the large computational cost hinders the application of sparse coding-based image features. In this paper, we propose approximating a sparse code using the output of a simple neural network. The resulting parameter learning model for the neural network automatically incorporates non-negative and shift-invariant constraints, leading to an efficient normalized non-negative sparse coding (N3SC) sparse encoder. Without the use of the traditional iterative process to solve the sparse coding objective, the sparse encoder directly “converts” each local feature into a sparse code. We also introduce a method for training the encoder based on the auto-encoder method. In addition, we formally propose the corresponding sparse coding scheme called N3SC, which enforces both the non-negative constraint and the shift-invariant constraint in addition to the traditional sparse coding criteria. As demonstrated by several experiments, the obtained N3SC encoder requires only 3%-10% of the processing time for image feature extraction compared with the standard sparse coding scheme. At the same time, the features extracted using the exact solutions of the N3SC coding scheme and the N3SC encoder offer superior image classification accuracy compared to the accuracy of many existing sparse coding-based representations. Shizhou Zhang, Jinjun Wang, Weiwei Shi 0003, Yihong Gong, Yong Xia 0001, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2019 | Unsupervised Domain Adaptation Using Robust Class-Wise MatchingabstractUnsupervised domain adaptation (DA) enables a classifier trained on data from one domain to be applied to data from another without labels. Given that the key to transferring a classifier across domains is to mitigate the data distribution mismatch for each class, most previous works completely or partially focus on global distribution matching across domains. The global data space, however, can be complicated, which makes modeling the global distribution difficult. To mitigate this problem, we present a novel unsupervised DA framework where the DA problem is addressed by proposing a robust class-wise matching strategy. Specifically, through minimizing a maximum mean discrepancy-based class-wise fisher discriminant across domains, this framework jointly optimizes two modules: a transferable feature learning module that reduces the distribution discrepancy between the same classes as well as increasing the distribution discrepancy between different classes across domains by a linear projection, and a robust classifier that exploits both the supervised information in source domain and the unsupervised low-rank property of target domain. In experiments on three DA benchmark data sets, the proposed framework shows the state-of-the-art performance. Lei Zhang 0054, Peng Wang 0023, Wei Wei 0008, Hao Lu 0003, Chunhua Shen, Anton van den Hengel, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2019 | Superpixel-Based Fast Fuzzy C-Means Clustering for Color Image SegmentationabstractA great number of improved fuzzy c-means (FCM) clustering algorithms have been widely used for grayscale and color image segmentation. However, most of them are time-consuming and unable to provide desired segmentation results for color images due to two reasons. The first one is that the incorporation of local spatial information often causes a high computational complexity due to the repeated distance computation between clustering centers and pixels within a local neighboring window. The other one is that a regular neighboring window usually breaks up the real local spatial structure of images and thus leads to a poor segmentation. In this work, we propose a superpixel-based fast FCM clustering algorithm that is significantly faster and more robust than stateof-the-art clustering algorithms for color image segmentation. To obtain better local spatial neighborhoods, we first define a multiscale morphological gradient reconstruction operation to obtain a superpixel image with accurate contour. In contrast to traditional neighboring window of fixed size and shape, the superpixel image provides better adaptive and irregular local spatial neighborhoods that are helpful for improving color image segmentation. Second, based on the obtained superpixel image, the original color image is simplified efficiently and its histogram is computed easily by counting the number of pixels in each region of the superpixel image. Finally, we implement FCM with histogram parameter on the superpixel image to obtain the final segmentation result. Experiments performed on synthetic images and real images demonstrate that the proposed algorithm provides better segmentation results and takes less time than state-of-the-art clustering algorithms for color image segmentation. Tao Lei 0003, Xiaohong Jia 0002, Yanning Zhang 0001, Shigang Liu, Hongying Meng, Asoke K. Nandi |
IEEE Trans. Fuzzy Syst. | 3 |
| 2019 | Intracluster Structured Low-Rank Matrix Analysis Method for Hyperspectral DenoisingabstractHyperspectral images (HSIs) denoising aims at eliminating the noise generated during the acquisition and transmission of HSIs. Since denoising is an ill-posed problem, utilizing proper knowledge of HSIs as regularization is essential for a good denoiser. Many HSI denoising methods have been proposed to leverage various prior knowledge, e.g., total variation, sparsity, and so on. Among those knowledge, a low-rank property has been shown to be effective for HSI denoising since it has the ability to deal with the missing values. However, most existing low-rank methods seldom consider mining the useful structures inside the low-rank matrix for a better denoising result. In addition, the rank number needs to be assigned manually. To address these problems, we propose an intracluster structured low-rank matrix analysis method for HSI denoising. First, we divide the original HSI into some clusters by taking advantages of both local similarity and nonlocal similarity structures, with which the resulted clusters are simpler and show more obvious low-rank property. Second, with singular value decomposition on the low-rank matrix in each cluster, the structured sparsity is modeled among the singular values to capture the structure of the low-rank matrix. Finally, an efficient optimization method is proposed to learn the structured sparsity adaptively from the data, as well as to inversely estimate the latent clean HSI from the noisy counterpart. The proposed method can not only obtain better denoising results compared with the-state-of-the-art methods but also automatically determine the rank number. Extensive experimental results demonstrate the effectiveness of the proposed method. Wei Wei 0008, Lei Zhang 0054, Yining Jiao, Chunna Tian, Cong Wang 0013, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2019 | Learning Discriminative Compact Representation for Hyperspectral Imagery ClassificationabstractAbundant spectral information of hyperspectral images (HSIs) has shown an obvious advantage in improving the performance of classification in the remote sensing domain. However, this is paid by the expensive consumption on the computation, transmission, as well as storage of HSIs. To address this problem, we propose to learn the discriminative compact representation for HSIs classification, which not only greatly reduces the data redundancy in the image but also preserves the discriminative information required for pixelwise classification in HSIs. To this end, we present a multi-task deep learning framework, which integrates HSIs autoencoding and classification into a two-branch deep neural network for jointly learning. In the network, we employ an encoder block to learn the compact representation of the input HSI via compression in the spectral domain. Being fed with the compact representation, the autoencoding branch then employs a decoder block to reconstruct the input HSI, while the classification branch utilizes a classifier block to predict the label for each pixel. Through end-to-end joint learning, the compact representation is not only informative enough to accurately reconstruct the original HSI but also discriminative enough to appropriately label each pixel with the trained classier. Sufficient experimental results on four HSIs classification data sets demonstrate the effectiveness of the proposed framework. Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | MPTV: Matching Pursuit-Based Total Variation Minimization for Image DeconvolutionabstractTotal variation (TV) regularization has proven effective for a range of computer vision tasks through its preferential weighting of sharp image edges. Existing TV-based methods, however, often suffer from the over-smoothing issue and solution bias caused by the homogeneous penalization. In this paper, we consider addressing these issues by applying inhomogeneous regularization on different image components. We formulate the inhomogeneous TV minimization problem as a convex quadratic constrained linear programming problem. Relying on this new model, we propose a matching pursuit-based total variation minimization method (MPTV), specifically for image deconvolution. The proposed MPTV method is essentially a cutting-plane method that iteratively activates a subset of nonzero image gradients and then solves a subproblem focusing on those activated gradients only. Compared with existing methods, the MPTV is less sensitive to the choice of the trade-off parameter between data fitting and regularization. Moreover, the inhomogeneity of MPTV alleviates the over-smoothing and ringing artifacts and improves the robustness to errors in blur kernel. Extensive experiments on different tasks demonstrate the superiority of the proposed method over the current state of the art. Dong Gong, Mingkui Tan, Qinfeng Shi, Anton van den Hengel, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | Two-Stream Convolutional Networks for Blind Image Quality AssessmentabstractTraditional image quality assessment (IQA) methods do not perform robustly due to the shallow hand-designed features. It has been demonstrated that deep neural network can learn more effective features than ever. In this paper, we describe a new deep neural network to predict the image quality accurately without relying on the reference image. To learn more effective feature representations for non-reference IQA, we propose a two-stream convolution network that includes two subcomponents for image and gradient image. The motivation for this design is using a two-stream scheme to capture different-level information of inputs and easing the difficulty of extracting features from one steam. The gradient stream focuses on extracting structure features in details, and the image stream pays more attention to the information in intensity. In addition, to consider the locally non-uniform distribution of distortion in images, we add a region-based fully convolutional layer for using the information around the center of the input image patch. The final score of the overall image is calculated by averaging of the patch scores. The proposed network performs in an end-to-end manner in both the training and testing phases. The experimental results on a series of benchmark datasets, e.g., LIVE, CISQ, IVC, TID2013, and Waterloo Exploration Database, show that the proposed algorithm outperforms the state-of-the-art methods, which verifies the effectiveness of our network architecture. Qingsen Yan, Dong Gong, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | A randomization and trial supply management system for adaptive clinical studies of TCM and its scientific research application in recurrent tuberculosis
Tiancai Wen, Baoyan Liu, Liyun He, Xiaoying Lv, Xin Wang 0121, Yanning Zhang 0001 |
BIBM | 6 |
| 2018 | Large-scale 3D Point Cloud Classification Based On Feature Description Matrix By CNNabstractLarge-scale 3D Point cloud classification is a basic topic for various applications. Traditional geometries features are usually independent of each other and difficult to adapt to a fixed classification model. With the rise of the neural network, deep learning is considered in 3D point cloud application. 3D points are difficult to feed the neural network directly based on deep learning, as they cannot be arranged in a fixed order as image pixels. In this paper, we combine traditional feature-based methods with the Convolutional neural network(CNN) to finish the classification task. The core idea is to construct a feasible structure called Feature Description Matrix(FDM) which encapsulates the local feature of the point to feed CNN for training and testing. By extracting geometry features and designed Feature Description Vectors(FDV) for FDM, a simple mechanism for point cloud classification is given, and experiments validate the effectiveness of our method, with higher classification accuracy compared to state-of-art works. Lei Wang 0089, Weiliang Meng, Runping Xi, Yanning Zhang 0001, Ling Lu, Xiaopeng Zhang 0001 |
CASA | 4 |
| 2018 | Holoscopic 3D Micro-Gesture Recognition Based on Fast Preprocessing and Deep Learning TechniquesabstractIt is a challenge to recognize holoscopic 3D (H3D) micro-gesture based on general vision techniques because images captured by H3D imaging system are unclear, i.e., the captured images include a large number of blurred grids. Many feature extraction methods can not be directly used for H3D images because the edge information of the grids will be captured. In this paper, we propose a fast and robust preprocessing method for H3D image reconstruction. The reconstructed images are clear and can be used directly for feature extraction or feature learning. Two contributions are presented in this paper. Firstly, we propose a bi-directional morphological filter used for enhancing the grids in an H3D image. Secondly, we propose a fast clustering algorithm with spatial information to extract grids from the H3D image. Because bi-directional morphological filter is able to incorporate local spatial information to the objective function of the fast clustering algorithm, the grids in H3D images are removed completely. Moreover, because the fast clustering algorithm perform clustering on gray levels of H3D images, a small computational cost is required. The proposed method is used to reconstruct H3D images to obtain multiple images with low resolution captured for 3D gesture recognition. Experiments show that the proposed preprocessing method is not only able to obtain better images that are clear and suitable for feature extraction or feature learning, but also is able to improve recognition accuracy in the micro-gesture recognition based on H3D imaging systems. Tao Lei 0003, Xiaohong Jia 0002, Yanning Zhang 0001, Xuhui Su, Shigang Liu |
FG | 4 |
| 2018 | Analysis of Disease Comorbidity Patterns in a Large-Scale China Population
Mengfei Guo, Tiancai Wen, Baoyan Liu, Jin Zhang 0044, Runshun Zhang, Yanning Zhang 0001, Xuezhong Zhou |
ICIC (2) | 8 |
| 2018 | A Subpixel Spatial-Spectral Feature Mining for Hyperspectral Image ClassificationabstractThis paper presents a subpixel spatial-spectral feature mining approach for hyperspectral image classification. First, a regional clustering-based spatial preprocessing (RCSPP) strategy is introduced to identify the endmember signatures from the original image. Then, a partial unmixing model of mixture tuned matched filtering (MTMF) is adopted to estimate the abundance maps. Finally, the morphological component analysis (MCA) is adopted to decompose the abundance map into different spatial morphological components, and the smoothness components are chosen for classification. The experimental results reveal that the obtained subpixel spatial-spectral feature can lead to very good classification accuracies. Xiang Xu 0002, Jun Li 0009, Yanning Zhang 0001, Shutao Li 0001 |
IGARSS | 3 |
| 2018 | Urban Impervious Surface Extraction Based on the Integration of Remote Sensing Images and Social Media DataabstractThis paper presents an inspiring approach for accurate estimation of impervious surfaces, which exploits the strength of two kind of heterogeneous features, i.e., physical features derived from satellite images and social features derived from social media datasets, respectively. On the one hand, we use a morphological attribute profiles guided spectral mixture analysis model to achieve estimates of physical features. On the other hand, we mine the social features from textual information of social media datasets. Then, a multivariable linear regression model is conducted to obtain the impervious surfaces. Experiment results, conducted with multi-spectral images collected by LANDSAT-8 and social media datasets scraped from Sina Weibo of Guangzhou city, suggest that our approach could lead to reliable and good estimation of the imperviousness. Wei Wei 0008, Jun Li 0009, Yanning Zhang 0001 |
IGARSS | 4 |
| 2018 | Multiscale Network Followed Network Model for Retinal Vessel Segmentation
Yicheng Wu 0001, Yong Xia 0001, Yang Song 0001, Yanning Zhang 0001, Tom Weidong Cai |
MICCAI (2) | 4 |
| 2018 | An Improved Camouflage Target Detection Using Hyperspectral Image Based on Block-Diagonal and Low-Rank Representation
Fei Li 0011, Xiuwei Zhang 0001, Lei Zhang 0054, Yanning Zhang 0001, Dongmei Jiang, Genping Zhao |
PRCV (4) | 4 |
| 2018 | Deep Classification and Segmentation Model for Vessel Extraction in Retinal Images
Yicheng Wu 0001, Yong Xia 0001, Yanning Zhang 0001 |
PRCV (2) | 3 |
| 2018 | Blind Image Quality Assessment via Deep Recursive Convolutional Network with Skip Connection
Qingsen Yan, Jinqiu Sun, Shaolin Su, Yu Zhu 0004, Haisen Li, Yanning Zhang 0001 |
PRCV (2) | 6 |
| 2018 | Accurate Spectral Super-Resolution from Single RGB Image Using Multi-scale CNN
Yiqi Yan, Lei Zhang 0054, Jun Li 0009, Wei Wei 0008, Yanning Zhang 0001 |
PRCV (2) | 5 |
| 2018 | TMFUF: a triple matrix factorization-based unified framework for predicting comprehensive drug-drug interactions of new drugsabstractBACKGROUND: A significant number of adverse drug reactions is caused by unexpected Drug-drug interactions (DDIs). The identification of DDIs becomes crucial before the co-prescription of multiple drugs is made. Such a task in clinics or in drug discovery usually requires high costs and numerous limitations, while computational approaches are able to predict potential DDIs effectively by utilizing diverse drug attributes (e.g. side effects). Nevertheless, they're incapable when required to predict enhancive and degressive DDIs, which change increasingly and decreasingly the pharmacological behavior of interacting drugs respectively. The pharmacological change of DDIs is one of the most important factors when making a multi-drug prescription. RESULTS: In this work, we design a Triple Matrix Factorization-based Unified Framework (TMFUF) to address the above issue. By leveraging a group of side effect entries of drugs, TMFUF achieves the inspiring result (AUC = 0.842 and AUPR = 0.526) in the case of conventional DDI prediction under the traditional screening task. In the comparison with two state-of-the-art approaches, TMFUF demonstrates it superiority by ~ 7% and ~ 20% improvement in terms of AUC and AUPR respectively. More importantly, TMFUF shows its ability in the comprehensive DDI prediction under different screening tasks. Finally, a utilization TMFUF reveals the significant pairs of side effects, which contribute to form enhancive and degressive DDIs, for further clinical validation. CONCLUSIONS: The proposed TMFUF is first capable to predict both conventional binary DDIs and comprehensive DDIs such that it captures the pharmacological changes caused by DDIs. Furthermore, it provides a unified solution of DDI prediction for two screening scenarios, which involves newly given drugs having no prior interaction. Another advantage is its ability to indicate how significantly the pairs of drug features contribute to form DDIs. Jianyu Shi, Yanning Zhang 0001, Siu-Ming Yiu |
BMC Bioinform. | 5 |
| 2018 | BMCMDA: a novel model for predicting human microbe-disease associations via binary matrix completionabstractBACKGROUND: Human Microbiome Project reveals the significant mutualistic influence between human body and microbes living in it. Such an influence lead to an interesting phenomenon that many noninfectious diseases are closely associated with diverse microbes. However, the identification of microbe-noninfectious disease associations (MDAs) is still a challenging task, because of both the high cost and the limitation of microbe cultivation. Thus, there is a need to develop fast approaches to screen potential MDAs. The growing number of validated MDAs enables us to meet the demand in a new insight. Computational approaches, especially machine learning, are promising to predict MDA candidates rapidly among a large number of microbe-disease pairs with the advantage of no limitation on microbe cultivation. Nevertheless, a few computational efforts at predicting MDAs are made so far. RESULTS: In this paper, grouping a set of MDAs into a binary MDA matrix, we propose a novel predictive approach (BMCMDA) based on Binary Matrix Completion to predict potential MDAs. The proposed BMCMDA assumes that the incomplete observed MDA matrix is the summation of a latent parameterizing matrix and a noising matrix. It also assumes that the independently occurring subscripts of observed entries in the MDA matrix follows a binomial model. Adopting a standard mean-zero Gaussian distribution for the nosing matrix, we model the relationship between the parameterizing matrix and the MDA matrix under the observed microbe-disease pairs as a probit regression. With the recovered parameterizing matrix, BMCMDA deduces how likely a microbe would be associated with a particular disease. In the experiment under leave-one-out cross-validation, it exhibits the inspiring performance (AUC = 0.906, AUPR =0.526) and demonstrates its superiority by ~ 7% and ~ 5% improvements in terms of AUC and AUPR respectively in the comparison with the pioneering approach KATZHMDA. CONCLUSIONS: Our BMCMDA provides an effective approach for predicting MDAs and can be also extended to other similar predicting tasks of binary relationship (e.g. protein-protein interaction, drug-target interaction). Jianyu Shi, Yanning Zhang 0001, Jiang-Bo Cao, Siu-Ming Yiu |
BMC Bioinform. | 3 |
| 2018 | Cluster Sparsity Field: An Internal Hyperspectral Imagery Prior for Reconstruction
Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Chunhua Shen, Anton van den Hengel, Qinfeng Shi |
Int. J. Comput. Vis. | 3 |
| 2018 | Blind image deblurring by promoting group sparsity
Dong Gong, Rui Li 0013, Yu Zhu 0004, Haisen Li, Jinqiu Sun, Yanning Zhang 0001 |
Neurocomputing | 6 |
| 2018 | Pedestrian search in surveillance videos by learning discriminative deep features
Shizhou Zhang, De Cheng, Yihong Gong, Dahu Shi, Xi Qiu, Yong Xia 0001, Yanning Zhang 0001 |
Neurocomputing | 7 |
| 2018 | NODULe: Combining constrained multi-scale LoG filters with densely dilated 3D deep convolutional neural network for pulmonary nodule detection
Yong Xia 0001, Haoyue Zeng, Yanning Zhang 0001 |
Neurocomputing | 4 |
| 2018 | Salient object detection in hyperspectral imagery using multi-scale spectral-spatial gradient
Lei Zhang 0054, Yanning Zhang 0001, Hangqi Yan, Wei Wei 0008 |
Neurocomputing | 2 |
| 2018 | Adaptive Unsymmetrical Trim-Based Morphological Filter for High-Density Impulse Noise Removal
Tao Lei 0003, Yanning Zhang 0001, Yi Wang 0069, Shigang Liu |
Multim. Tools Appl. | 2 |
| 2018 | Visible and infrared image registration based on region features and edginess
Yanjia Chen, Xiuwei Zhang 0001, Yanning Zhang 0001, Stephen J. Maybank, Zhipeng Fu |
Mach. Vis. Appl. | 3 |
| 2018 | Validation of right coronary artery lumen area from cardiac computed tomography against intravascular ultrasound
Hengfei Cui, Yong Xia 0001, Yanning Zhang 0001, Liang Zhong 0001 |
Mach. Vis. Appl. | 3 |
| 2018 | Going deeper with two-stream ConvNets for action recognition in video surveillance
Peng Zhang 0005, Tao Zhuo, Wei Huang 0013, Yanning Zhang 0001 |
Pattern Recognit. Lett. | 5 |
| 2018 | Significantly Fast and Robust Fuzzy C-Means Clustering Algorithm Based on Morphological Reconstruction and Membership FilteringabstractAs fuzzy c-means clustering (FCM) algorithm is sensitive to noise, local spatial information is often introduced to an objective function to improve the robustness of the FCM algorithm for image segmentation. However, the introduction of local spatial information often leads to a high computational complexity, arising out of an iterative calculation of the distance between pixels within local spatial neighbors and clustering centers. To address this issue, an improved FCM algorithm based on morphological reconstruction and membership filtering (FRFCM) that is significantly faster and more robust than FCM is proposed in this paper. First, the local spatial information of images is incorporated into FRFCM by introducing morphological reconstruction operation to guarantee noise-immunity and image detail-preservation. Second, the modification of membership partition, based on the distance between pixels within local spatial neighbors and clustering centers, is replaced by local membership filtering that depends only on the spatial neighbors of membership partition. Compared with state-of-the-art algorithms, the proposed FRFCM algorithm is simpler and significantly faster, since it is unnecessary to compute the distance between pixels within local spatial neighbors and clustering centers. In addition, it is efficient for noisy image segmentation because membership filtering are able to improve membership partition matrix efficiently. Experiments performed on synthetic and real-world images demonstrate that the proposed algorithm not only achieves better results, but also requires less time than the state-of-the-art algorithms for image segmentation. Tao Lei 0003, Xiaohong Jia 0002, Yanning Zhang 0001, Lifeng He, Hongying Meng, Asoke K. Nandi |
IEEE Trans. Fuzzy Syst. | 3 |
| 2018 | Exploiting Structured Sparsity for Hyperspectral Anomaly DetectionabstractSparse representation-based background modeling facilitates much recent progress in hyperspectral anomaly detection (AD). The sparse representation of background often exhibits underlying structure, which is crucial to distinguish between background and anomaly. However, how to exploit such underlying structure is still challenging. To address this problem, we present a novel hyperspectral AD method, which can exploit the structured sparsity in modeling the background more accurately. With the plausible background area detected by a local RX detector, a robust background spectrum dictionary is learned in a principal component analysis way. A reweighted Laplace prior-based structured sparse representation model is then employed to reconstruct the spectrum of each pixel. With considering the structured sparsity in representation, the background pixels can be reconstructed more accurately than the anomaly ones, which thus can be detected based on the reconstruction error. To further improve the detection performance, an intracluster reconstruction model is developed to exploit the spatial similarity among the background pixels in the same cluster. The anomaly pixels can then be detected based on the cost of intracluster reconstruction error. By linearly combining these two detection results, improvement is obviously achieved on detection accuracy. Experimental results on both simulated and real-world data sets demonstrate that the proposed method outperforms several state-of-the-art hyperspectral AD methods. Fei Li 0011, Xiuwei Zhang 0001, Lei Zhang 0054, Dongmei Jiang, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2018 | Exploiting Clustering Manifold Structure for Hyperspectral Imagery Super-ResolutionabstractFusing a low-resolution hyperspectral image (HSI) with a high-resolution (HR) conventional image into an HR HSI has become a prevalent HSIs super-resolution scheme. However, in most previous works, little attention has been paid on exploiting the underlying manifold structure in the spatial domain of the latent HR HSI. In this paper, we advance a provable prior knowledge that the clustering manifold structure of the latent HSI can be well preserved in the spatial domain of the input conventional image. Inspired by this, we first conduct clustering in the spatial domain of the input conventional image and adopt the intra-cluster self-expressiveness model to implicitly depict the clustering manifold structure, which enables learning the complicated manifold structure via solving a constrained ridge regression model without knowing the exact form of the manifold. Then, we incorporate the learned structure into a variational super-resolution framework to regularize the latent HSI. The resulted framework can be effectively optimized by a standard alternating direction method of multipliers. Since the learned structure can well depict the underlying spatial manifold of the latent HSI, the proposed method shows the state-of-the-art super-resolution performance on two benchmark data sets. Lei Zhang 0054, Wei Wei 0008, Chengcheng Bai, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2017 | MPGL: An Efficient Matching Pursuit Method for Generalized LASSOabstractUnlike traditional LASSO enforcing sparsity on the variables, Generalized LASSO (GL) enforces sparsity on a linear transformation of the variables, gaining flexibility and success in many applications. However, many existing GL algorithms do not scale up to high-dimensional problems, and/or only work well for a specific choice of the transformation. We propose an efficient Matching Pursuit Generalized LASSO (MPGL) method, which overcomes these issues, and is guaranteed to converge to a global optimum. We formulate the GL problem as a convex quadratic constrained linear programming (QCLP) problem and tailor-make a cutting plane method. More specifically, our MPGL iteratively activates a subset of nonzero elements of the transformed variables, and solves a subproblem involving only the activated elements thus gaining significant speed-up. Moreover, MPGL is less sensitive to the choice of the trade-off hyper-parameter between data fitting and regularization, and mitigates the long-standing hyper-parameter tuning issue in many existing methods. Experiments demonstrate the superior efficiency and accuracy of the proposed method over the state-of-the-arts in both classification and image processing tasks. Dong Gong, Mingkui Tan, Yanning Zhang 0001, Anton van den Hengel, Qinfeng Shi |
AAAI | 3 |