Jing Zhang 0052

dblp:05/3499-52 · DBLP profile ↗
← Back
77ranked-venue papers
12as first author
68since 2021 · last 2026
0000-0002-8516-0913ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 55 · 8 first-author · 47 since 2021Artificial intelligence and machine learning · 48 · 10 first-author · 42 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021
YearPublicationVenuePosition
2026 Uncertainty-Aware 3D Edge Reconstruction with Difference of Gaussians
abstract
3D edge reconstruction from posed multi-view images remains a critical yet under-explored task. While 3D Gaussian Splatting (3DGS)-based methods have recently achieved promising performance, they face two main challenges. First, edges exhibit clear discontinuities from the background, but the intrinsic smoothness of Gaussian kernels makes it challenging to model such discontinuities. Second, due to the absence of multi-view edge annotations, models are trained with pseudo labels instead. These pseudo labels extracted by pre-trained$2 D$edge detectors often exhibit cross-view inconsistencies, leading to degraded performance. To address these issues, we propose a novel uncertainty-aware 3D edge reconstruction using Difference of Gaussians (DoG) as kernels, called EdgeDoG. First, we incorporate DoG kernels to model edge discontinuities explicitly. Second, we design a dual-uncertainty strategy: primitive-level uncertainty is estimated via multi-view Fisher information to eliminate noisy 3D primitives, while pixel-level uncertainty is computed from gradients of rendered depth maps to reweight the training loss, thereby compensating for inconsistent 2D pseudo labels with robust 3D geometric cues. Extensive experiments on diverse datasets demonstrate that our method achieves superior performance compared to previous approaches.
Caixia Zhou, Haibin Ling, Jing Zhang 0052
3DV5
2026 Learning Spatial Decay for Vision Transformers
abstract
Vision Transformers (ViTs) have revolutionized computer vision, yet their self-attention mechanism lacks explicit spatial inductive biases, leading to suboptimal performance on spatially-structured tasks. Existing approaches introduce data-independent spatial decay based on fixed distance metrics, applying uniform attention weighting regardless of image content and limiting adaptability to diverse visual scenarios. Inspired by recent advances in large language models where content-aware gating mechanisms (e.g., GLA, HGRN2, FOX) significantly outperform static alternatives, we present the first successful adaptation of data-dependent spatial decay to 2D vision transformers. We introduce Spatial Decay Transformer (SDT), featuring a novel Context-Aware Gating (CAG) mechanism that generates dynamic, data-dependent decay for patch interactions. Our approach learns to modulate spatial attention based on both content relevance and spatial proximity. We address the fundamental challenge of 1D-to-2D adaptation through a unified spatial-content fusion framework that integrates manhattan distance-based spatial priors with learned content representations. Extensive experiments on ImageNet-1K classification and generation tasks demonstrate consistent improvements over strong baselines. Our work establishes data-dependent spatial decay as a new paradigm for enhancing spatial attention in vision transformers.
Yuxin Mao, Zhen Qin 0003, Jinxing Zhou, Bin Fan 0002, Jing Zhang 0052, Yiran Zhong, Yuchao Dai
AAAI5
2026 SteerMusic: Enhanced Musical Consistency for Zero-shot Text-Guided and Personalized Music Editing
abstract
Music editing is an important step in music production, which has broad applications, including game development and film production. Most existing zero-shot text-guided editing methods rely on pretrained diffusion models by involving forward-backward diffusion processes. However, these methods often struggle to preserve the musical content. Additionally, text instructions alone usually fail to accurately describe the desired music. In this paper, we propose two music editing methods that improve the consistency between the original and edited music by leveraging score distillation. The first method, SteerMusic, is a coarse-grained zero-shot editing approach using delta denoising score. The second method, SteerMusic+, enables fine-grained personalized music editing by manipulating a concept token that represents a user-defined musical style. SteerMusic+ allows for the editing of music into user-defined musical styles that cannot be achieved by the text instructions alone. Experimental results show that our methods outperform existing approaches in preserving both music content consistency and editing fidelity. User studies further validate that our methods achieve superior music editing quality.
Xinlei Niu, Kin Wai Cheuk, Jing Zhang 0052, Naoki Murata, Chieh-Hsin Lai, Michele Mancusi, Woosung Choi, Giorgio Fabbro, Wei-Hsiang Liao 0001, Charles P. Martin, Yuki Mitsufuji
AAAI3
2026 Unlocking Vision-Language Models for Video Anomaly Detection via Fine-Grained Prompting
abstract
Prompting has emerged as a practical way to adapt frozen vision-language models (VLMs) for video anomaly detection (VAD). Yet, existing prompts are often overly abstract, overlooking the fine-grained human–object interactions or action semantics that define complex anomalies in surveillance videos. We propose ASK-Hint, a structured prompting framework that leverages action-centric knowledge to elicit more accurate and interpretable reasoning from frozen VLMs. Our approach organizes prompts into semantically coherent groups (e.g. violence, property crimes, public safety) and formulates fine-grained guiding questions that align model predictions with discriminative visual cues. Extensive experiments on UCF-Crime and XD-Violence show that ASK-Hint consistently improves AUC over prior baselines, achieving state-of-the-art performance compared to both fine-tuned and training-free methods. Beyond accuracy, our framework provides interpretable reasoning traces towards anomaly and demonstrates strong generalization across datasets and VLM backbones. These results highlight the critical role of prompt granularity and establish ASK-Hint as a new training-free and generalizable solution for explainable video anomaly detection.
Shu Zou, Lukas Wesemann, Fabian Waschkowski, Zhaoyuan Yang, Jing Zhang 0052
WACV6
2026 A Generative Victim Model for Segmentation
Aixuan Li, Jing Zhang 0052, Zhexiong Wan, Yiran Zhong, Yuchao Dai
Int. J. Comput. Vis.2
2026 Exploiting Continuity for Unsupervised Single Depth Map Super-Resolution
abstract
Depth map super-resolution (DSR) aims at reconstructing high-resolution depth maps from low-resolution input. Existing DSR methods rely on using the high-resolution RGB images as guidance and are typically trained in a supervised manner. However, obtaining well-aligned RGB-Depth pairs is challenging, and these approaches also suffer from texture over-transfer issues. To address these limitations, we propose the Implicit Depth Fitting Network (IDFN), a zero-shot, unsupervised framework that relies solely on a single low-resolution depth map. We formulate DSR as learning a resolution-independent continuous depth field. To accurately capture complex scene geometries, our framework combines Fourier feature encoding with periodic activation functions, effectively balancing smooth surface reconstruction with sharp edge preservation. Furthermore, we introduce a Dithering Strategy to model the sensor degradation process. This strategy enables the network to learn the underlying continuous signal from discrete area-integrated observations, effectively suppressing aliasing artifacts. Extensive experiments show that IDFN not only achieves performance comparable to state-of-the-art unsupervised approaches, but also yields results competitive with leading supervised approaches, demonstrating the effectiveness of our method.
Ruobing Jian, Jing Zhang 0052, Yuchao Dai
IEEE Signal Process. Lett.2
2026 Structure-Preserving Frequency-Regularized Text-Guided Optimal Transport for Unpaired Rain Streaks and Raindrops Removal
abstract
The removal of rain streaks and raindrops is crucial for enhancing the image visibility and mitigating the weather degradations. However, most existing approaches rely on the paired rainy and clean images, which are challenging to obtain in real-world scenarios. To this end, we propose a novel structure-preserving frequency-regularized text-guided optimal transport (SFTOT) framework, which formulates the unpaired rain streaks and raindrops removal as an optimal transport problem. Specifically, we introduce a structure-preserving transport cost, incorporating the structural similarity constraint to minimize the duality gap between the primal and dual formulations, while preserving the structural details of reconstructed images. Furthermore, by embedding the inherent frequency sparsity of rain streaks and raindrops into the transport cost, we derive a frequency-regularized optimal transport objective, ensuring consistency in frequency distributions between the generated and clean images. Additionally, we employ a pre-trained one-step stable diffusion model as the restoration network, which is fine-tuned using the low-rank adaptation (LoRA) adapters and zero convolutional layers, while integrating the domain-specific text prompts for both degraded and clean images to guide the generation process. Extensive experiments demonstrate that our method surpasses the existing well-performing unpaired learning approaches, achieving notable improvements in both the fidelity and photo-realism.
Yuanbo Wen 0002, Tao Gao 0001, Qianxi Zhang, Jing Zhang 0052, Ting Chen 0003, Lidong Liu
IEEE Trans. Multim.5
2025 Patch-level Sounding Object Tracking for Audio-Visual Question Answering
abstract
Answering questions related to audio-visual scenes, i.e., the AVQA task, is becoming increasingly popular. A critical challenge is accurately identifying and tracking sounding objects related to the question along the timeline. In this paper, we present a new Patch-level Sounding Object Tracking (PSOT) method. It begins with a Motion-driven Key Patch Tracking (M-KPT) module, which relies on visual motion information to identify salient visual patches with significant movements that are more likely to relate to sounding objects and questions. We measure the patch-wise motion intensity map between neighboring video frames and utilize it to construct and guide a motion-driven graph network. Meanwhile, we design a Sound-driven KPT (S-KPT) module to explicitly track sounding patches. This module also involves a graph network, with the adjacency matrix regularized by the audio-visual correspondence map. The M-KPT and S-KPT modules are performed in parallel for each temporal segment, allowing balanced tracking of salient and sounding objects. Based on the tracked patches, we further propose a Question-driven KPT (Q-KPT) module to retain patches highly relevant to the question, ensuring the model focuses on the most informative clues. The audio-visual-question features are updated during the processing of these modules, which are then aggregated for final answer prediction. Extensive experiments on standard datasets demonstrate the effectiveness of our method, achieving competitive performance even compared to recent large-scale pretraining-based approaches.
Zhangbin Li, Jinxing Zhou, Jing Zhang 0052, Shengeng Tang, Kun Li 0008, Dan Guo 0001
AAAI3
2025 Multi-axis Prompt and Multi-dimension Fusion Network for All-in-one Weather-degraded Image Restoration
abstract
Existing approaches aiming to remove adverse weather degradations compromise the image quality and incur the long processing time. To this end, we introduce a multi-axis prompt and multi-dimension fusion network (MPMF-Net). Specifically, we develop a multi-axis prompts learning block (MPLB), which learns the prompts along three separate axis planes, requiring fewer parameters and achieving superior performance. Moreover, we present a multi-dimension feature interaction block (MFIB), which optimizes intra-scale feature fusion by segregating features along height, width and channel dimensions. This strategy enables more accurate mutual attention and adaptive weight determination. Additionally, we propose the coarse-scale degradation-free implicit neural representations (CDINR) to normalize the degradation levels of different weather conditions. Extensive experiments demonstrate the significant improvements of our model over the recent well-performing approaches in both reconstruction fidelity and inference time.
Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Ting Chen 0003
AAAI3
2025 Brain-Inspired Spiking Neural Networks for Energy-Efficient Object Detection
abstract
Brain-inspired spiking neural networks (SNNs) have the capability of energy-efficient processing of temporal information. However, leveraging the rich dynamic characteristics of SNNs and prior works in artificial neural networks (ANNs) to construct an effective object detection model for visual tasks remains an open question for further exploration. To develop a directly-trained , low energy consumption and high-performance multi-scale SNN model, we propose a novel interpretable object detection framework Multi-scale Spiking Detector (MSD). Initially, we propose a spiking convolutional neuron as a core component of the Optic Nerve Nucleus Block (ONNB), designed to significantly enhance the deep feature extraction capabilities of SNNs. ONNB enables direct training with improved energy efficiency, demonstrating superior performance compared to state-of-the-art ANN-to-SNN conversion and SNN techniques. In addition, we propose a Multi-scale Spiking Detection Framework to emulate the biological response and comprehension of stimuli from different objects. Wherein, spiking multi-scale fusion and the spiking detector are employed to integrate features across different depths and to detect response outcomes, respectively. Our method outperforms state-of-the-art ANN detectors, with only 7.8 M parameters and 6.43 mJ energy consumption. MSD obtains the mean average precision (mAP) of 62.0% and 66.3% on COCO and Gen1 datasets, respectively.
Tao Gao 0001, Yisheng An, Ting Chen 0003, Jing Zhang 0052, Yuanbo Wen 0002, Mengkun Liu, Qianxi Zhang
CVPR5
2025 Identifying and Mitigating Position Bias of Multi-image Vision-Language Models
abstract
The evolution of Large Vision-Language Models (LVLMs) has progressed from single to multi-image reasoning. Despite this advancement, our findings indicate that LVLMs struggle to robustly utilize information across multiple images, with predictions significantly affected by the alteration of image positions. To further explore this issue, we introduce Position-wise Question Answering (PQA), a meticulously designed task to quantify reasoning capabilities at each position. Our analysis reveals a pronounced position bias in LVLMs: open-source models excel in reasoning with images positioned later but underperform with those in the middle or at the beginning, while proprietary models show improved comprehension for images at the beginning and end but struggle with those in the middle. Motivated by this, we propose SoFt Attention (SoFA), a simple, training-free approach that mitigates this bias by employing linear interpolation between inter-image causal attention and bidirectional counterparts. Experimental results demonstrate that SoFA reduces position bias and enhances the reasoning performance of existing LVLMs. The code will be available https://github.com/xytian1008/sofa.
Shu Zou, Zhaoyuan Yang, Jing Zhang 0052
CVPR4
2025 Probability Density Geodesics in Image Diffusion Latent Space
abstract
Diffusion models indirectly estimate the probability density over a data space, which can be used to study its structure. In this work, we show that geodesics can be computed in diffusion latent space, where the norm induced by the spatially-varying inner product is inversely proportional to the probability density. In this formulation, a path that traverses a high density (that is, probable) region of image latent space is shorter than the equivalent path through a low density region. We present algorithms for solving the associated initial and boundary value problems and show how to compute the probability density along the path and the geodesic distance between two points. Using these techniques, we analyze how closely video clips approximate geodesics in a pre-trained image diffusion space. Finally, we demonstrate how these techniques can be applied to training-free image sequence interpolation and extrapolation, given a pre-trained image diffusion model.
Qingtao Yu, Zhaoyuan Yang, Peter H. Tu, Jing Zhang 0052, Hongdong Li, Richard I. Hartley, Dylan Campbell
CVPR5
2025 Black Sheep in the Herd: Playing with Spuriously Correlated Attributes for Vision-Language Recognition
abstract
Few-shot adaptation for Vision-Language Models (VLMs) presents a dilemma: balancing in-distribution accuracy with out-of-distribution generalization. Recent research has utilized low-level concepts such as visual attributes to enhance generalization. However, this study reveals that VLMs overly rely on a small subset of attributes on decision-making, which co-occur with the category but are not inherently part of it, termed spuriously correlated attributes. This biased nature of VLMs results in poor generalization. To address this, 1) we first propose Spurious Attribute Probing (SAP), identifying and filtering out these problematic attributes to significantly enhance the generalization of existing attribute-based methods; 2) We introduce Spurious Attribute Shielding (SAS), a plug-and-play module that mitigates the influence of these attributes on prediction, seamlessly integrating into various Parameter-Efficient Fine-Tuning (PEFT) methods. In experiments, SAP and SAS significantly enhance accuracy on distribution shifts across 11 datasets and 3 generalization tasks without compromising downstream performance, establishing a new state-of-the-art benchmark.
Shu Zou, Zhaoyuan Yang, Mengqi He, Jing Zhang 0052
ICLR5
2025 Audio-Visual Segmentation with Semantics
Jinxing Zhou, Xuyang Shen, Weixuan Sun, Jing Zhang 0052, Stanley T. Birchfield, Dan Guo 0001, Lingpeng Kong, Meng Wang 0001, Yiran Zhong
Int. J. Comput. Vis.6
2025 MSNet: Multi-Scale Network for Object Detection in Remote Sensing Images
Tao Gao 0001, Shilin Xia, Mengkun Liu, Jing Zhang 0052, Ting Chen 0003
Pattern Recognit.4
2025 Generative Transformer for Accurate and Reliable Salient Object Detection
abstract
We explore the impact of transformers on accurate and reliable salient object detection. For accuracy, we integrate the transformer with a deterministic model and delineate its advantages in structural modeling. Regarding reliability, we address the transformer’s tendency to produce overly confident, incorrect predictions. To gauge reliability implicitly, we introduce a latent variable model within the transformer framework, termed the inferential generative adversarial network (iGAN). The stochastic nature of the latent variable facilitates the estimation of predictive uncertainty, which serves as an auxiliary measure of the model’s prediction reliability. Different from the conventional GAN, which defines the distribution of the latent variable as fixed standard normal distribution$\mathcal {N}(0,\mathbf {I})$. The proposed iGAN infers the latent variable by gradient-based Markov Chain Monte Carlo (MCMC), namely Langevin dynamics, leading to an input-dependent latent variable model. We apply our proposed iGAN to fully supervised salient object detection, explaining that iGAN within the transformer framework leads to both accurate and reliable salient object detection. The source code and experimental results are publicly available via our project page:https://npucvr.github.io/TransformerSOD.
Yuxin Mao, Jing Zhang 0052, Zhexiong Wan, Aixuan Li, Yunqiu Lv, Yuchao Dai
IEEE Trans. Circuits Syst. Video Technol.2
2025 Contrastive Conditional Latent Diffusion for Audio-Visual Segmentation
abstract
Audio-visual Segmentation (AVS) is conceptualized as a conditional generation task, where audio is considered as the conditional variable for segmenting the sound producer(s). In this case, audio should be extensively explored to maximize its contribution for the final segmentation task. We propose a contrastive conditional latent diffusion model for audio-visual segmentation (AVS) to thoroughly investigate the impact of audio, where the correlation between audio and the final segmentation map is modeled to guarantee the strong correlation between them. To achieve semantic-correlated representation learning, our framework incorporates a latent diffusion model. The diffusion model learns the conditional generation process of the ground-truth segmentation map, resulting in ground-truth aware inference during the denoising process at the test stage. As our model is conditional, it is vital to ensure that the conditional variable contributes to the model output. We thus extensively model the contribution of the audio signal by minimizing the density ratio between the conditional probability of the multimodal data, e.g. conditioned on the audio-visual data, and that of the unimodal data, e.g. conditioned on the audio data only. In this way, our latent diffusion model via density ratio optimization explicitly maximizes the contribution of audio for AVS, which can then be achieved with contrastive learning as a constraint, where the diffusion part serves as the main objective to achieve maximum likelihood estimation, and the density ratio optimization part imposes the constraint. By adopting this latent diffusion model via contrastive learning, we effectively enhance the contribution of audio for AVS. The effectiveness of our solution is validated through experimental results on the benchmark dataset. Code and results are online via our project page: https://github.com/OpenNLPLab/DiffusionAVS.
Yuxin Mao, Jing Zhang 0052, Mochu Xiang, Yunqiu Lv, Dong Li 0033, Yiran Zhong, Yuchao Dai
IEEE Trans. Image Process.2
2025 All-in-One Weather-Degraded Image Restoration Via Adaptive Degradation-Aware Self-Prompting Model
abstract
Existing approaches for all-in-one weather-degraded image restoration suffer from inefficiencies in leveraging degradation-aware priors, resulting in sub-optimal performance in adapting to different weather conditions. To this end, we develop an adaptive degradation-aware self-prompting model (ADSM) for all-in-one weather-degraded image restoration. Specifically, our model employs the contrastive language-image pre-training model (CLIP) to facilitate the training of our proposed latent prompt generators (LPGs), which represent three types of latent prompts to characterize the degradation type, degradation property and image caption. Moreover, we integrate the acquired degradation-aware prompts into the time embedding of diffusion model to improve degradation perception. Meanwhile, we employ the latent caption prompt to guide the reverse sampling process using the cross-attention mechanism, thereby guiding the accurate image reconstruction. Furthermore, to accelerate the reverse sampling procedure of diffusion model and address the limitations of frequency perception, we introduce a wavelet-oriented noise estimating network (WNE-Net). Extensive experiments conducted on eight publicly available datasets demonstrate the effectiveness of our proposed approach in both task-specific and all-in-one applications.
Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Kaihao Zhang, Ting Chen 0003
IEEE Trans. Multim.4
2024 LaViP: Language-Grounded Visual Prompting
abstract
We introduce a language-grounded visual prompting method to adapt the visual encoder of vision-language models for downstream tasks. By capitalizing on language integration, we devise a parameter-efficient strategy to adjust the input of the visual encoder, eliminating the need to modify or add to the model's parameters. Due to this design choice, our algorithm can operate even in black-box scenarios, showcasing adaptability in situations where access to the model's parameters is constrained. We will empirically demonstrate that, compared to prior art, grounding visual prompts with language enhances both the accuracy and speed of adaptation. Moreover, our algorithm excels in base-to-novel class generalization, overcoming limitations of visual prompting and exhibiting the capacity to generalize beyond seen classes. We thoroughly assess and evaluate our method across a variety of image recognition datasets, such as EuroSAT, UCF101, DTD, and CLEVR, spanning different learning situations, including few-shot adaptation, base-to-novel class generalization, and transfer learning.
Nilakshan Kunananthaseelan, Jing Zhang 0052, Mehrtash Harandi
AAAI2
2024 Adversarial Purification with the Manifold Hypothesis
abstract
In this work, we formulate a novel framework for adversarial robustness using the manifold hypothesis. This framework provides sufficient conditions for defending against adversarial examples. We develop an adversarial purification method with this framework. Our method combines manifold learning with variational inference to provide adversarial robustness without the need for expensive adversarial training. Experimentally, our approach can provide adversarial robustness even if attackers are aware of the existence of the defense. In addition, our method can also serve as a test-time defense mechanism for variational autoencoders.
Zhaoyuan Yang, Jing Zhang 0052, Richard I. Hartley, Peter H. Tu
AAAI3
2024 ArGue: Attribute-Guided Prompt Tuning for Vision-Language Models
abstract
Although soft prompt tuning is effective in efficiently adapting Vision-Language (V&L) models for downstream tasks, it shows limitations in dealing with distribution shifts. We address this issue with Attribute-Guided Prompt Tuning (ArGue), making three key contributions. 1) In contrast to the conventional approach of directly appending soft prompts preceding class names, we align the model with primitive visual attributes generated by Large language Models (LLMs). We posit that a model's ability to express high confidence in these attributes signifies its ca-pacity to discern the correct class rationales. 2) We intro-duce attribute sampling to eliminate disadvantageous at-tributes, thus only semantically meaningful attributes are preserved. 3) We propose negative prompting, explicitly enumerating class-agnostic attributes to activate spurious correlations and encourage the model to generate highly orthogonal probability distributions in relation to these neg-ative features. In experiments, our method significantly out-performs current state-of-the-art prompt tuning methods on both novel class prediction and out-of-distribution general-ization tasks. The code is available https://github.com/Liam-Tian/ArGue.
Shu Zou, Zhaoyuan Yang, Jing Zhang 0052
CVPR4
2024 SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation
abstract
We present SoundLoCD, a novel text-to-sound generation framework, which incorporates a LoRA-based conditional discrete contrastive latent diffusion model. Unlike recent large-scale sound generation models, our model can be efficiently trained under limited computational resources. The integration of a contrastive learning strategy further enhances the connection between text conditions and the generated outputs, resulting in coherent and high-fidelity performance. Our experiments demonstrate that SoundLoCD outperforms the baseline with greatly reduced computational resources. A comprehensive ablation study further validates the contribution of each component within SoundLoCD1.
Xinlei Niu, Jing Zhang 0052, Christian Walder, Charles P. Martin
ICASSP2
2024 Multi-Dimension Queried and Interacting Network for Stereo Image Deraining
abstract
Eliminating the rain degradation in stereo images poses a formidable challenge, which necessitates the efficient exploitation of mutual information present between the dual views. To this end, we devise MQINet, which employs multi-dimension queries and interactions for stereo image deraining. More specifically, our approach incorporates a context-aware dimension-wise queried block (CDQB). This module leverages dimension-wise queries that are independent of the input features and employs global context-aware attention (GCA) to capture essential features while avoiding the entanglement of redundant or irrelevant information. Meanwhile, we introduce an intra-view physics-aware attention (IPA) based on the inverse physical model of rainy images. IPA extracts shallow features that are sensitive to the physics of rain degradation, facilitating the reduction of rain-related artifacts during the early learning period. Furthermore, we integrate a cross-view multi-dimension interacting attention mechanism (CMIA) to foster comprehensive feature interaction between the two views across multiple dimensions. Extensive experimental evaluations demonstrate the superiority of our model over EPRRNet and StereoIRR, achieving respective improvements of 4.18 dB and 0.45 dB in PSNR. Code and models are available at https://github.com/chdwyb/MQINet.
Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Ting Chen 0003
ICASSP4
2024 Encoder-Minimal and Decoder-Minimal Framework for Remote Sensing Image Dehazing
abstract
Haze obscures remote sensing images, hindering valuable information extraction. To this end, we propose RSHazeNet, an encoder-minimal and decoder-minimal framework for efficient remote sensing image dehazing. Specifically, regarding the process of merging features within the same level, we develop an innovative module called intra-level transposed fusion module (ITFM). This module employs adaptive transposed self-attention to capture comprehensive context-aware information, facilitating the robust context-aware feature fusion. Meanwhile, we present a cross-level multi-view interaction module (CMIM) to enable effective interactions between features from various levels, mitigating the loss of information due to the repeated sampling operations. In addition, we propose a multi-view progressive extraction block (MPEB) that partitions the features into four distinct components and employs convolution with varying kernel sizes, groups, and dilation factors to facilitate view-progressive feature learning. Extensive experiments demonstrate the superiority of our proposed RSHazeNet. We release the source code and all pre-trained models at https://github.com/chdwyb/RSHazeNet.
Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Ting Chen 0003
ICASSP4
2024 IMPUS: Image Morphing with Perceptually-Uniform Sampling Using Diffusion Models
abstract
We present a diffusion-based image morphing approach with perceptually-uniform sampling (IMPUS) that produces smooth, direct and realistic interpolations given an image pair. The embeddings of two images may lie on distinct conditioned distributions of a latent diffusion model, especially when they have significant semantic difference. To bridge this gap, we interpolate in the locally linear and continuous text embedding space and Gaussian latent space. We first optimize the endpoint text embeddings and then map the images to the latent space using a probability flow ODE. Unlike existing work that takes an indirect morphing path, we show that the model adaptation yields a direct path and suppresses ghosting artifacts in the interpolated images. To achieve this, we propose a heuristic bottleneck constraint based on a novel relative perceptual path diversity score that automatically controls the bottleneck size and balances the diversity along the path with its directness. We also propose a perceptually-uniform sampling technique that enables visually smooth changes between the interpolated images. Extensive experiments validate that our IMPUS can achieve smooth, direct, and realistic image morphing and is adaptable to several other generative tasks.
Zhaoyuan Yang, Jing Zhang 0052, Dylan Campbell, Peter H. Tu, Richard I. Hartley
ICLR5
2024 Latent Optimal Paths by Gumbel Propagation for Variational Bayesian Dynamic Programming
abstract
We propose the stochastic optimal path which solves the classical optimal path problem by a probability-softening solution. This unified approach transforms a wide range of DP problems into directed acyclic graphs in which all paths follow a Gibbs distribution. We show the equivalence of the Gibbs distribution to a message-passing algorithm by the properties of the Gumbel distribution and give all the ingredients required for variational Bayesian inference of a latent path, namely Bayesian dynamic programming (BDP). We demonstrate the usage of BDP in the latent space of variational autoencoders (VAEs) and propose the BDP-VAE which captures structured sparse optimal paths as latent variables. This enables end-to-end training for generative tasks in which models rely on unobserved structural information. At last, we validate the behavior of our approach and showcase its applicability in two real-world applications: text-to-speech and singing voice synthesis. Our implementation code is available at https://github.com/XinleiNIU/LatentOptimalPathsBayesianDP.
Xinlei Niu, Christian Walder, Jing Zhang 0052, Charles P. Martin
ICML3
2024 HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
abstract
We introduce HybridVC, a voice conversion (VC) framework built upon a pre-trained conditional variational autoencoder (CVAE) that combines the strengths of a latent model with contrastive learning. HybridVC supports text and audio prompts, enabling more flexible voice style conversion. HybridVC models a latent distribution conditioned on speaker embeddings acquired by a pretrained speaker encoder and optimises style text embeddings to align with the speaker style information through contrastive learning in parallel. Therefore, HybridVC can be efficiently trained under limited computational resources. Our experiments demonstrate HybridVC's superior training efficiency and its capability for advanced multimodal voice style conversion. This underscores its potential for widespread applications such as user-defined personalised voice in various social media platforms. A comprehensive ablation study further validates the effectiveness of our method.
Xinlei Niu, Jing Zhang 0052, Charles P. Martin
INTERSPEECH2
2024 TAVGBench: Benchmarking Text to Audible-Video Generation
abstract
The Text to Audible-Video Generation (TAVG) task involves generating videos with accompanying audio based on text descriptions. Achieving this requires skillful alignment of both audio and video elements. To support research in this field, we have developed a comprehensive Text to Audible-Video Generation Benchmark (TAVGBench), which contains over 1.7 million clips with a total duration of 11.8 thousand hours. We propose an automatic annotation pipeline to ensure each audible video has detailed descriptions for both its audio and video contents. We also introduce the Audio-Visual Harmoni score (AVHScore) to provide a quantitative measure of the alignment between the generated audio and video modalities. Additionally, we present a baseline model for TAVG called TAVDiffusion, which uses a two-stream latent diffusion model to provide a fundamental starting point for further research in this area. We achieve the alignment of audio and video by employing cross-attention and contrastive learning. Through extensive experiments and evaluations on TAVGBench, we demonstrate the effectiveness of our proposed model under both conventional metrics and our proposed metrics. The dataset and code can be found on this page https://npucvr.github.io/TAVGBench/ and on github https://github.com/OpenNLPLab/TAVGBench.
Yuxin Mao, Xuyang Shen, Jing Zhang 0052, Zhen Qin 0003, Jinxing Zhou, Mochu Xiang, Yiran Zhong, Yuchao Dai
ACM Multimedia3
2024 DreamSteerer: Enhancing Source Image Conditioned Editability using Personalized Diffusion Models
abstract
Recent text-to-image (T2I) personalization methods have shown great premise in teaching a diffusion model user-specified concepts given a few images for reusing the acquired concepts in a novel context. With massive efforts being dedicated to personalized generation, a promising extension is personalized editing, namely to edit an image using personalized concepts, which can provide more precise guidance signal than traditional textual guidance. To address this, one straightforward solution is to incorporate a personalized diffusion model with a text-driven editing framework. However, such solution often shows unsatisfactory editability on the source image. To address this, we propose DreamSteerer, a plug-in method for augmenting existing T2I personalization methods. Specifically, we enhance the source image conditioned editability of a personalized diffusion model via a novel Editability Driven Score Distillation (EDSD) objective. Moreover, we identify a mode trapping issue with EDSD, and propose a mode shifting regularization with spatial feature guided sampling to avoid such issue. We further employ two key modifications on the Delta Denoising Score framework that enable high-fidelity local editing with personalized concepts. Extensive experiments validate that DreamSteerer can significantly improve the editability of several T2I personalization baselines while being computationally efficient.
Zhaoyuan Yang, Jing Zhang 0052
NeurIPS3
2024 A novel dual-stage progressive enhancement network for single image deraining
Tao Gao 0001, Yuanbo Wen 0002, Jing Zhang 0052, Ting Chen 0003
Eng. Appl. Artif. Intell.3
2024 Synergizing triple attention with depth quality for RGB-D salient object detection
Peipei Song, Peiyan Zhong, Jing Zhang 0052, Piotr Koniusz, Feng Duan 0006, Nick Barnes
Neurocomputing4
2024 A Self-Supplementary and Revised Network for Remote Sensing Object Detection
abstract
Object detection is an essential and crucial task in interpretation of optical remote sensing images (RSIs). However, its performance is usually limited due to the complex background and multiscale characteristics of targets. To overcome these limitations, a self-supplementary and revised anchor-free detector is proposed. First, to reduce the computational cost of detection, a partial bottleneck (PBottleneck) structure is designed to efficiently extract multiscale feature information in a lightweight manner. Second, pure spatial feature pyramid network (PSFPN) attaches importance to description of distance and suppresses environmental disturbance by a devised multidirectional distance attention (MDDA) mechanism. In addition, pure fusion strategy (PFS) is created to boost information with no occlusion between various features. Third, toward the multiscale objects issue, self-learning supplementary and revised module (SSRM) is explored to generate more abundant and balanced expression by adaptively incorporating the supplementary and corrected information from adjacent features. Finally, comprehensive experiments are conducted on several publicly available datasets, demonstrating effectiveness of our proposed detector, leading to a new benchmark.
Tao Gao 0001, Zixiang Liu, Guiping Wu, Yuanbo Wen 0002, Lidong Liu, Ting Chen 0003, Jing Zhang 0052
IEEE Geosci. Remote. Sens. Lett.8
2024 From heavy rain removal to detail restoration: A faster and better network
Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Kaihao Zhang, Ting Chen 0003
Pattern Recognit.3
2024 Frequency-Oriented Efficient Transformer for All-in-One Weather-Degraded Image Restoration
abstract
Adverse weather conditions, such as rain, raindrop, snow and haze, consistently degrade images in an unpredictable manner, thereby rendering existing task-specific and task-aligned methods inadequate in addressing this formidable problem. To this end, we investigate the application of Transformer in image restoration and introduce an efficient frequency-oriented method called AIRFormer, which is designed to restore weather-degraded images comprehensively and holistically. Specifically, we identify that the initial self-attention mechanism exhibits distinctive properties akin to a low-pass filter. Therefore, we construct a frequency-guided Transformer encoder by incorporating wavelet-based prior information to guide the extraction of image features. Additionally, considering the non-specific frequency characteristics of self-attention in the later stages, we develop a frequency-refined Transformer decoder that incorporates learnable task-specific queries across spatial dimensions, channel dimensions, and wavelet domains. To facilitate the training of our proposed method, we curate a comprehensive benchmark dataset named AIR40K that, encompasses a wide range of challenging scenarios. Extensive experimental evaluations demonstrate the superiority of our AIRFormer over both task-aligned and all-in-one methods across 15 publicly available datasets. Notably, AIRFormer achieves the best trade-off between the inference time and quality of reconstructed image, comparing with existing methods such as TransWeather and Restormer. The source code, dataset and pre-trained models will be available at https://github.com/chdwyb/AIRFormer.
Tao Gao 0001, Yuanbo Wen 0002, Kaihao Zhang, Jing Zhang 0052, Ting Chen 0003, Lidong Liu, Wenhan Luo
IEEE Trans. Circuits Syst. Video Technol.4
2024 Mutual Information Regularization for Weakly-Supervised RGB-D Salient Object Detection
abstract
In this paper, we present a weakly-supervised RGB-D salient object detection model via scribble supervision. Specifically, as a multimodal learning task, we focus on effective multimodal representation learning via inter-modal mutual information regularization. In particular, following the principle of disentangled representation learning, we introduce a mutual information upper bound with a mutual information minimization regularizer to encourage the disentangled representation of each modality for salient object detection. Based on our multimodal representation learning framework, we introduce an asymmetric feature extractor for our multimodal data, which is proven more effective than the conventional symmetric backbone setting. We also introduce multimodal variational auto-encoder as stochastic prediction refinement techniques, which takes pseudo labels from the first training stage as supervision and generates refined prediction. Experimental results on benchmark RGB-D salient object detection datasets verify both effectiveness of our explicit multimodal disentangled representation learning method and the stochastic prediction refinement strategy, achieving comparable performance with the state-of-the-art fully supervised models. Our code and data are available at:https://npucvr.github.io/MIRV/.
Aixuan Li, Yuxin Mao, Jing Zhang 0052, Yuchao Dai
IEEE Trans. Circuits Syst. Video Technol.3
2024 Measuring and Modeling Uncertainty Degree for Monocular Depth Estimation
abstract
Effectively measuring and modeling the reliability of a trained model is essential to the real-world deployment of monocular depth estimation (MDE) models. However, the intrinsic ill-posedness and ordinal-sensitive nature of MDE pose major challenges to the estimation of uncertainty degree of the trained models. On the one hand, utilizing current uncertainty modeling methods may increase memory consumption and usually take more time. On the other hand, measuring the uncertainty based on model accuracy can also be problematic, where uncertainty reliability and prediction accuracy are not well decoupled. In this paper, we propose to model the uncertainty of MDE models from the perspective of the inherent probability distributions originating from the depth probability volume and its extensions, and to assess it more fairly with more comprehensive metrics. By simply introducing additional training regularization terms, our model, with surprisingly simple formations and without requiring extra modules or multiple inferences, can provide uncertainty estimations with state-of-the-art reliability, and can be further improved when combined with ensemble or sampling methods. A series of experiments demonstrate the effectiveness of our methods. Code and results are available at https://github.com/npucvr/MDEUncertainty.
Mochu Xiang, Jing Zhang 0052, Nick Barnes, Yuchao Dai
IEEE Trans. Circuits Syst. Video Technol.2
2024 Weakly-Supervised Contrastive Learning for Unsupervised Object Discovery
abstract
Unsupervised object discovery (UOD) refers to the task of discriminating the whole region of objects from the background within a scene without relying on labeled datasets, which benefits the task of bounding-box-level localization and pixel-level segmentation. This task is promising due to its ability to discover objects in a generic manner. We roughly categorize existing techniques into two main directions, namely the generative solutions based on image resynthesis, and the clustering methods based on self-supervised models. We have observed that the former heavily relies on the quality of image reconstruction, while the latter shows limitations in effectively modeling semantic correlations. To directly target at object discovery, we focus on the latter approach and propose a novel solution by incorporating weakly-supervised contrastive learning (WCL) to enhance semantic information exploration. We design a semantic-guided self-supervised learning model to extract high-level semantic features from images, which is achieved by fine-tuning the feature encoder of a self-supervised model, namely DINO, via WCL. Subsequently, we introduce Principal Component Analysis (PCA) to localize object regions. The principal projection direction, corresponding to the maximal eigenvalue, serves as an indicator of the object region(s). Extensive experiments on benchmark unsupervised object discovery datasets demonstrate the effectiveness of our proposed solution. The source code and experimental results are publicly available via our project page at https://github.com/npucvr/WSCUOD.git.
Yunqiu Lv, Jing Zhang 0052, Nick Barnes, Yuchao Dai
IEEE Trans. Image Process.2
2023 Modeling the Distributional Uncertainty for Salient Object Detection Models
abstract
Most of the existing salient object detection (SOD) models focus on improving the overall model performance, without explicitly explaining the discrepancy between the training and testing distributions. In this paper, we investigate a particular type of epistemic uncertainty, namely distributional uncertainty, for salient object detection. Specifically, for the first time, we explore the existing class-aware distribution gap exploration techniques, i.e. long-tail learning, single-model uncertainty modeling and test-time strategies, and adapt them to model the distributional uncertainty for our class-agnostic task. We define test sample that is dissimilar to the training dataset as being “out-of-distribution” (OOD) samples. Different from the conventional OOD definition, where OOD samples are those not belonging to the closed-world training categories, OOD samples for SOD are those break the basic priors of saliency, i.e. center prior, color contrast prior, compactness prior and etc., indicating OOD as being “continuous” instead of being discrete for our task. We've carried out extensive experimental results to verify effectiveness of existing distribution gap modeling techniques for SOD, and conclude that both train-time single-model uncertainty estimation techniques and weight-regularization solutions that preventing model activation from drifting too much are promising directions for modeling distributional uncertainty for SOD.
Jing Zhang 0052, Mochu Xiang, Yuchao Dai
CVPR2
2023 P2C: Self-Supervised Point Cloud Completion from Single Partial Clouds
abstract
Point cloud completion aims to recover the complete shape based on a partial observation. Existing methods require either complete point clouds or multiple partial observations of the same object for learning. In contrast to previous approaches, we present Partial2Complete (P2C), the first self-supervised framework that completes point cloud objects using training samples consisting of only a single incomplete point cloud per object. Specifically, our framework groups incomplete point clouds into local patches as input and predicts masked patches by learning prior information from different partial objects. We also propose Region-Aware Chamfer Distance to regularize shape mismatch without limiting completion capability, and devise the Normal Consistency Constraint to incorporate a local planarity assumption, encouraging the recovered shape surface to be continuous and complete. In this way, P2C no longer needs multiple observations or complete point clouds as ground truth. Instead, structural cues are learned from a category-specific dataset to complete partial point clouds of objects. We demonstrate the effectiveness of our approach on both synthetic ShapeNet data and real-world ScanNet data, showing that P2C produces comparable results to methods trained with complete shapes, and outperforms methods learned with multiple partial observations. Code is available at https://github.com/CuiRuikai/Partial2Complete.
Ruikai Cui, Shi Qiu 0001, Saeed Anwar, Jiawei Liu 0005, Chaoyue Xing, Jing Zhang 0052, Nick Barnes
ICCV6
2023 Model Calibration in Dense Classification with Adaptive Label Perturbation
abstract
For safety-related applications, it is crucial to produce trustworthy deep neural networks whose prediction is associated with confidence that can represent the likelihood of correctness for subsequent decision-making. Existing dense binary classification models are prone to being over-confident. To improve model calibration, we propose Adaptive Stochastic Label Perturbation (ASLP) which learns a unique label perturbation level for each training image. ASLP employs our proposed Self-Calibrating Binary Cross Entropy (SC-BCE) loss, which unifies label perturbation processes including stochastic approaches (like DisturbLabel), and label smoothing, to correct calibration while maintaining classification rates. ASLP follows Maximum Entropy Inference of classic statistical mechanics to maximise prediction entropy with respect to missing information. It performs this while: (1) preserving classification accuracy on known data as a conservative solution, or (2) specifically improves model calibration degree by minimising the gap between the prediction accuracy and expected confidence of the target training label. Extensive results demonstrate that ASLP can significantly improve calibration degrees of dense binary classification models on both in-distribution and out-of-distribution data. The code is available on https://github.com/Carlisle-Liu/ASLP.
Jiawei Liu 0005, Changkun Ye, Shan Wang 0010, Ruikai Cui, Jing Zhang 0052, Kaihao Zhang, Nick Barnes
ICCV5
2023 Multimodal Variational Auto-encoder based Audio-Visual Segmentation
abstract
We propose an Explicit Conditional Multimodal Variational Auto-Encoder (ECMVAE) for audio-visual segmentation (AVS), aiming to segment sound sources in the video sequence. Existing AVS methods focus on implicit feature fusion strategies, where models are trained to fit the discrete samples in the dataset. With a limited and less diverse dataset, the resulting performance is usually unsatisfactory. In contrast, we address this problem from an effective representation learning perspective, aiming to model the contribution of each modality explicitly. Specifically, we find that audio contains critical category information of the sound producers, and visual data provides candidate sound producer(s). Their shared information corresponds to the target sound producer(s) shown in the visual data. In this case, cross-modal shared representation learning is especially important for AVS. To achieve this, our ECMVAE factorizes the representations of each modality with a modality-shared representation and a modality-specific representation. An orthogonality constraint is applied between the shared and specific representations to maintain the exclusive attribute of the factorized latent code. Further, a mutual information maximization regularizer is introduced to achieve extensive exploration of each modality. Quantitative and qualitative evaluations on the AVSBench demonstrate the effectiveness of our approach, leading to a new state-of-the-art for AVS, with a 3.84 mIOU performance leap on the challenging MS3 subset for multiple sound source segmentation.
Yuxin Mao, Jing Zhang 0052, Mochu Xiang, Yiran Zhong, Yuchao Dai
ICCV2
2023 RPEFlow: Multimodal Fusion of RGB-PointCloud-Event for Joint Optical Flow and Scene Flow Estimation
abstract
Recently, the RGB images and point clouds fusion methods have been proposed to jointly estimate 2D optical flow and 3D scene flow. However, as both conventional RGB cameras and LiDAR sensors adopt a frame-based data acquisition mechanism, their performance is limited by the fixed low sampling rates, especially in highly-dynamic scenes. By contrast, the event camera can asynchronously capture the intensity changes with a very high temporal resolution, providing complementary dynamic information of the observed scenes. In this paper, we incorporate RGB images, Point clouds and Events for joint optical flow and scene flow estimation with our proposed multi-stage multimodal fusion model, RPEFlow. First, we present an attention fusion module with a cross-attention mechanism to implicitly explore the internal cross-modal correlation for 2D and 3D branches, respectively. Second, we introduce a mutual information regularization term to explicitly model the complementary information of three modalities for effective multimodal feature learning. We also contribute a new synthetic dataset to advocate further research. Experiments on both synthetic and real datasets show that our model outperforms the existing state-of-theart by a wide margin. Code and dataset is available at https://npucvr.github.io/RPEFlow.
Zhexiong Wan, Yuxin Mao, Jing Zhang 0052, Yuchao Dai
ICCV3
2023 Progressive dilation dense residual fusion network for single-image deraining
abstract
Abstract Rain removal is very important for many applications in computer vision, and it is a challenging problem due to its ill‐posed nature, especially for single‐image deraining. In order to remove rain streaks more thoroughly, as well as to retain more details, a progressive dilation dense residual fusion network is proposed. The entire network is designed in a cascade manner with multiple fusion blocks. The fusion block consists of a dilation dense residual block (DDRB) and a dense residual feature fusion block (DRFFB), where DDRB is created for feature extraction and DRFFB is mainly designed for feature fusion operation. Meanwhile, detail compensation memory mechanism (DCMM) is leveraged between each of two cascade modules to retain more background details. Compared with previous state‐of‐the‐art methods, extensive experiments show that the proposed method can achieve better results, in terms of rain streaks removal and background details preservation. Furthermore, the authors’ network also shows its superiority for image noise removal.
Xiaolin Kong, Tao Gao 0001, Ting Chen 0003, Jing Zhang 0052
IET Image Process.4
2023 Salient Objects in Clutter
abstract
In this paper, we identify and address a serious design bias of existing salient object detection (SOD) datasets, which unrealistically assume that each image should contain at least one clear and uncluttered salient object. This design bias has led to a saturation in performance for state-of-the-art SOD models when evaluated on existing datasets. However, these models are still far from satisfactory when applied to real-world scenes. Based on our analyses, we propose a new high-quality dataset and update the previous saliency benchmark. Specifically, our dataset, called Salient Objects in Clutter (SOC), includes images with both salient and non-salient objects from several common object categories. In addition to object category annotations, each salient image is accompanied by attributes that reflect common challenges in common scenes, which can help provide deeper insight into the SOD problem. Further, with a given saliency encoder, e.g., the backbone network, existing saliency models are designed to achieve mapping from the training image set to the training ground-truth set. We therefore argue that improving the dataset can yield higher performance gains than focusing only on the decoder design. With this in mind, we investigate several dataset-enhancement strategies, including label smoothing to implicitly emphasize salient boundaries, random image augmentation to adapt saliency models to various scenarios, and self-supervised learning as a regularization strategy to learn from small datasets. Our extensive results demonstrate the effectiveness of these tricks. We also provide a comprehensive benchmark for SOD, which can be found in our repository: https://github.com/DengPingFan/SODBenchmark.
Deng-Ping Fan, Jing Zhang 0052, Ming-Ming Cheng, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 An Energy-Based Prior for Generative Saliency
abstract
We propose a novel generative saliency prediction framework that adopts an informative energy-based model as a prior distribution. The energy-based prior model is defined on the latent space of a saliency generator network that generates the saliency map based on a continuous latent variables and an observed image. Both the parameters of saliency generator and the energy-based prior are jointly trained via Markov chain Monte Carlo-based maximum likelihood estimation, in which the sampling from the intractable posterior and prior distributions of the latent variables are performed by Langevin dynamics. With the generative saliency model, we can obtain a pixel-wise uncertainty map from an image, indicating model confidence in the saliency prediction. Different from existing generative models, which define the prior distribution of the latent variables as a simple isotropic Gaussian distribution, our model uses an energy-based informative prior which can be more expressive in capturing the latent space of the data. With the informative energy-based prior, we extend the Gaussian distribution assumption of generative models to achieve a more representative distribution of the latent space, leading to more reliable uncertainty estimation. We apply the proposed frameworks to both RGB and RGB-D salient object detection tasks with both transformer and convolutional neural network backbones. We further propose an adversarial learning algorithm and a variational inference algorithm as alternatives to train the proposed generative framework. Experimental results show that our generative saliency model with an energy-based prior can achieve not only accurate saliency predictions but also reliable uncertainty maps that are consistent with human perception.
Jing Zhang 0052, Jianwen Xie, Nick Barnes, Ping Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Toward Deeper Understanding of Camouflaged Object Detection
abstract
Preys in the wild evolve to be camouflaged to avoid being recognized by predators. In this way, camouflage acts as a key defence mechanism across species that is critical to survival. To detect and segment the whole scope of a camouflaged object, camouflaged object detection (COD) is introduced as a binary segmentation task, with the binary ground truth camouflage map indicating the exact regions of the camouflaged objects. In this paper, we revisit this task and argue that the binary segmentation setting fails to fully understand the concept of camouflage. We find that explicitly modeling the conspicuousness of camouflaged objects against their particular backgrounds can not only lead to a better understanding about camouflage, but also provide guidance to designing more sophisticated camouflage techniques. Furthermore, we observe that it is some specific parts of camouflaged objects that make them detectable by predators. With the above understanding about camouflaged objects, we present the first triple-task learning framework to simultaneouslylocalize, segment, and rankcamouflaged objects, indicating the conspicuousness level of camouflage. As no corresponding datasets exist for either the localization model or the ranking model, we generate localization maps with an eye tracker, which are then processed according to the instance level labels to generate our ranking-based training and testing dataset. We also contribute the largest COD testing set to comprehensively analyse performance of the COD models. Experimental results show that our triple-task learning framework achieves new state-of-the-art, leading to a more explainable COD network. Our code, data, and results are available at:https://github.com/JingZhang617/COD-Rank-Localize-and-Segment.
Yunqiu Lv, Jing Zhang 0052, Yuchao Dai, Aixuan Li, Nick Barnes, Deng-Ping Fan
IEEE Trans. Circuits Syst. Video Technol.2
2023 A Task-Balanced Multiscale Adaptive Fusion Network for Object Detection in Remote Sensing Images
abstract
Object detection is essential in the interpretation of remote sensing images. However, the blurred background and objects with vast variances are identified as the two main challenges of the task. We propose a novel detector adapted to complicated background and multi-scale objects, namely, task-balanced multi-scale adaptive fusion network (TMAFNet), targeting directly on the above two challenges. Firstly, a depth separable global context module (DSGC) is constructed to understand contextual relations among pixels from a global perspective, which is extraordinarily necessary to distinguish objects from the environment. Most importantly, DSGC reduces the computational cost by decoupling the acquisition of global information into single-channel global interaction and multi-channel single-point interaction. Secondly, in order to eliminate disturbance and enhance representation ability of objects, hidden recursive feature pyramid network (HRFPN) is explored, which encodes the information of difference before and after using the multi-scale fusion. HRFPN is proven to enhance the target features by reducing the background noise. Thirdly, a semi-coupling task-balanced head (SCTB) is presented to guarantee the consistency of detection. We have conducted comprehensive experiments on several publicly available datasets, and the results illustrate that our modules improve adaptability and robustness of the network, leading to a new state-of-the-art.
Tao Gao 0001, Zixiang Liu, Jing Zhang 0052, Guiping Wu, Ting Chen 0003
IEEE Trans. Geosci. Remote. Sens.3
2023 Global to Local: A Scale-Aware Network for Remote Sensing Object Detection
abstract
With the wide application of remote sensing images (RSIs) in military and civil fields, remote sensing object detection (RSOD) has gradually become a hot research direction. However, we observe two main challenges for remote sensing object detection, namely the complicated background and the small objects issues. Given the different appearances of generic objects and remote sensing objects, the detection algorithms designed for the former usually cannot perform well for the latter. We propose a novel global to local scale-aware detection network (GLSANet) for remote sensing object detection, aiming to solve the above mentioned two challenges. Firstly, we design a global semantic information interaction module (GSIIM) to excavate and reinforce the high-level semantic information in the deep feature map, which alleviates the obstacles of complex background on foreground objects. Secondly, we optimize the feature pyramid network to improve the performance of multiscale object detection in RSIs. Finally, a local attention pyramid (LAP) is introduced to highlight the feature representation of small objects gradually while suppressing the background and noise in the shallower feature maps. Extensive experiments on three public datasets demonstrate that the proposed method achieves superior performance compared with the state-of-the-art detectors, especially on small object detection dataset. Specifically, our algorithm reaches 94.57% mAP on NWPU VHR-10 dataset, 95.93% mAP on RSOD dataset and 77.9% mAP on DIOR dataset, respectively.
Tao Gao 0001, Qianqian Niu, Jing Zhang 0052, Ting Chen 0003, Shaohui Mei, Ahmad Jubair
IEEE Trans. Geosci. Remote. Sens.3
2023 Encoder-Free Multiaxis Physics-Aware Fusion Network for Remote Sensing Image Dehazing
abstract
Current methods for remote sensing image dehazing confront noteworthy computational intricacies and yield suboptimal dehazed outputs, thereby circumscribing their pragmatic applicability. To this end, we propose EMPF-Net, a novel encoder-free multi-axis physics-aware fusion network that exhibits both light-weighted characteristics and computational efficiency. In our pipeline, we contend that conventional u-shaped networks allocate substantial computational resources to encode haze-degraded features, which play a subordinate role in the reconstruction process. Consequently, our encoder stages solely incorporate down-sampling operations. To improve the representation efficiency and enhance the generalization capabilities, we devise a multi-axis partial queried learning block (MPQLB) that primarily concentrates on learning dimension-wise queries, instead of relying solely on strictly-correlated content of the input features. Furthermore, we augment the reconstruction procedure by incorporating ground truth supervision into each stage via a supervised cross-scale transposed attention module (SCTAM). It calculates attention maps under the guidance of clean images, thereby suppressing less informative features to propagate to the subsequent level. In addition, to address the challenge of ineffective intral-level feature fusion, which result in insufficient elimination of haze-degraded information and negatively impact the quality of reconstructed images, we introduce a physics-aware intra-level fusion module (PIFM). This module harnesses a physical inversion model to facilitate the intra-level feature interaction and alleviate the interference of dehazing-irrelevant information. Our proposed EMPF-Net is evaluated on 12 publicly available datasets, and the experimental results substantiate our superiority in terms of both metrical scores and visual quality, despite being equipped with a modest parameter count of 300 K. Our approach is readily accessible at https://github.com/chdwyb/EMPF-Net.
Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Ting Chen 0003
IEEE Trans. Geosci. Remote. Sens.3
2023 Predictive Uncertainty Estimation for Camouflaged Object Detection
abstract
Uncertainty is inherent in machine learning methods, especially those for camouflaged object detection aiming to finely segment the objects concealed in background. The strong enquote center bias of the training dataset leads to models of poor generalization ability as the models learn to find camouflaged objects around image center, which we define as enquote model bias. Further, due to the similar appearance of camouflaged object and its surroundings, it is difficult to label the accurate scope of the camouflaged object, especially along object boundaries, which we term as enquote data bias. To effectively model the two types of biases, we resort to uncertainty estimation and introduce predictive uncertainty estimation technique, which is the sum of model uncertainty and data uncertainty, to estimate the two types of biases simultaneously. Specifically, we present a predictive uncertainty estimation network (PUENet) that consists of a Bayesian conditional variational auto-encoder (BCVAE) to achieve predictive uncertainty estimation, and a predictive uncertainty approximation (PUA) module to avoid the expensive sampling process at test-time. Experimental results show that our PUENet achieves both highly accurate prediction, and reliable uncertainty estimation representing the biases within both model parameters and the datasets.
Yi Zhang 0076, Jing Zhang 0052, Wassim Hamidouche, Olivier Déforges
IEEE Trans. Image Process.2
2022 Transmission-Guided Bayesian Generative Model for Smoke Segmentation
abstract
Smoke segmentation is essential to precisely localize wildfire so that it can be extinguished in an early phase. Although deep neural networks have achieved promising results on image segmentation tasks, they are prone to be overconfident for smoke segmentation due to its non-rigid shape and transparent appearance. This is caused by both knowledge level uncertainty due to limited training data for accurate smoke segmentation and labeling level uncertainty representing the difficulty in labeling ground-truth. To effectively model the two types of uncertainty, we introduce a Bayesian generative model to simultaneously estimate the posterior distribution of model parameters and its predictions. Further, smoke images suffer from low contrast and ambiguity, inspired by physics-based image dehazing methods, we design a transmission-guided local coherence loss to guide the network to learn pair-wise relationships based on pixel distance and the transmission feature. To promote the development of this field, we also contribute a high-quality smoke segmentation dataset, SMOKE5K, consisting of 1,400 real and 4,000 synthetic images with pixel-wise annotation. Experimental results on benchmark testing datasets illustrate that our model achieves both accurate predictions and reliable uncertainty maps representing model ignorance about its prediction. Our code and dataset are publicly available at: https://github.com/redlessme/Transmission-BVM.
Siyuan Yan, Jing Zhang 0052, Nick Barnes
AAAI2
2022 Energy-Based Generative Cooperative Saliency Prediction
abstract
Conventional saliency prediction models typically learn a deterministic mapping from an image to its saliency map, and thus fail to explain the subjective nature of human attention. In this paper, to model the uncertainty of visual saliency, we study the saliency prediction problem from the perspective of generative models by learning a conditional probability distribution over the saliency map given an input image, and treating the saliency prediction as a sampling process from the learned distribution. Specifically, we propose a generative cooperative saliency prediction framework, where a conditional latent variable model~(LVM) and a conditional energy-based model~(EBM) are jointly trained to predict salient objects in a cooperative manner. The LVM serves as a fast but coarse predictor to efficiently produce an initial saliency map, which is then refined by the iterative Langevin revision of the EBM that serves as a slow but fine predictor. Such a coarse-to-fine cooperative saliency prediction strategy offers the best of both worlds. Moreover, we propose a ``cooperative learning while recovering" strategy and apply it to weakly supervised saliency prediction, where saliency annotations of training images are partially observed. Lastly, we find that the learned energy function in the EBM can serve as a refinement module that can refine the results of other pre-trained saliency prediction models. Experimental results show that our model can produce a set of diverse and plausible saliency maps of an image, and obtain state-of-the-art performance in both fully supervised and weakly supervised saliency prediction tasks.
Jing Zhang 0052, Jianwen Xie, Zilong Zheng, Nick Barnes
AAAI1
2022 A General Divergence Modeling Strategy for Salient Object Detection
Jing Zhang 0052, Yuchao Dai
ACCV (7)2
2022 Energy-Based Residual Latent Transport for Unsupervised Point Cloud Completion
Ruikai Cui, Shi Qiu 0001, Saeed Anwar, Jing Zhang 0052, Nick Barnes
BMVC4
2022 Audio-Visual Segmentation
Jinxing Zhou, Weixuan Sun, Jing Zhang 0052, Stanley T. Birchfield, Dan Guo 0001, Lingpeng Kong, Meng Wang 0001, Yiran Zhong
ECCV (37)5
2022 Multi-Modal Transformer for RGB-D Salient Object Detection
abstract
The main focus of existing RGB-D salient object detection models is achieving effective multi-modal fusion. Due to the limited receptive field of conventional convolutional neural networks (CNNs), CNN-based multi-modal fusion strategies fail to extensively model the correlation between the two modalities (appearance information from the RGB image and geometric information from the depth data). Given the success of transformer networks for long-range dependency modeling, we investigate multi-modal transformer networks for RGB-D salient object detection. Specifically, a transformer-based multi-modal fusion module is presented to effectively fuse appearance features and geometric features. Experimental results on six challenging benchmark RGB-D salient object detection datasets demonstrate the effectiveness of our approach.
Peipei Song, Jing Zhang 0052, Piotr Koniusz, Nick Barnes
ICIP2
2022 Modeling Aleatoric Uncertainty for Camouflaged Object Detection
abstract
Aleatoric uncertainty captures noise within the observations. For camouflaged object detection, due to similar appearance of the camouflaged foreground and the back-ground, it’s difficult to obtain highly accurate annotations, especially annotations around object boundaries. We argue that training directly with the "noisy" camouflage map may lead to a model of poor generalization ability. In this paper, we introduce an explicitly aleatoric uncertainty estimation technique to represent predictive uncertainty due to noisy labeling. Specifically, we present a confidence-aware camouflaged object detection (COD) framework using dynamic supervision to produce both an accurate camouflage map and a reliable "aleatoric uncertainty". Different from existing techniques that produce deterministic prediction following the point estimation pipeline, our framework formalises aleatoric uncertainty as probability distribution over model output and the input image. We claim that, once trained, our confidence estimation network can evaluate the pixel-wise accuracy of the prediction without relying on the ground truth camouflage map. Extensive results illustrate the superior performance of the proposed model in explaining the camouflage prediction. Our codes are available at https://github.com/Carlisle-Liu/OCENet
Jiawei Liu 0005, Jing Zhang 0052, Nick Barnes
WACV2
2022 Inferring the Class Conditional Response Map for Weakly Supervised Semantic Segmentation
abstract
Image-level weakly supervised semantic segmentation (WSSS) relies on class activation maps (CAMs) for pseudo labels generation. As CAMs only highlight the most discriminative regions of objects, the generated pseudo labels are usually unsatisfactory to serve directly as supervision. To solve this, most existing approaches follow a multi-training pipeline to refine CAMs for better pseudo-labels, which includes: 1) re-training the classification model to generate CAMs; 2) post-processing CAMs to obtain pseudo labels; and 3) training a semantic segmentation model with the obtained pseudo labels. However, this multi-training pipeline requires complicated adjustment and additional time. To address this, we propose a class-conditional inference strategy and an activation aware mask refinement loss function to generate better pseudo labels without retraining the classifier. The class conditional inference-time approach is presented to separately and iteratively reveal the classification network’s hidden object activation to generate more complete response maps. Further, our activation aware mask refinement loss function introduces a novel way to exploit saliency maps during segmentation training and refine the foreground object masks without suppressing background objects. Our method achieves superior WSSS results without requiring re-training of the classifier. https://github.com/weixuansun/InferCam
Weixuan Sun, Jing Zhang 0052, Nick Barnes
WACV2
2022 Uncertainty Inspired RGB-D Saliency Detection
abstract
We propose the first stochastic framework to employ uncertainty for RGB-D saliency detection by learning from the data labeling process. Existing RGB-D saliency detection models treat this task as a point estimation problem by predicting a single saliency map following a deterministic learning pipeline. We argue that, however, the deterministic solution is relatively ill-posed. Inspired by the saliency data labeling process, we propose a generative architecture to achieve probabilistic RGB-D saliency detection which utilizes a latent variable to model the labeling variations. Our framework includes two main models: 1) a generator model, which maps the input image and latent variable to stochastic saliency prediction, and 2) an inference model, which gradually updates the latent variable by sampling it from the true or approximate posterior distribution. The generator model is an encoder-decoder saliency network. To infer the latent variable, we introduce two different solutions: i) a Conditional Variational Auto-encoder with an extra encoder to approximate the posterior distribution of the latent variable; and ii) an Alternating Back-Propagation technique, which directly samples the latent variable from the true posterior distribution. Qualitative and quantitative results on six challenging RGB-D benchmark datasets show our approach's superior performance in learning the distribution of saliency maps. The source code is publicly available via our project page: https://github.com/JingZhang617/UCNet.
Jing Zhang 0052, Deng-Ping Fan, Yuchao Dai, Saeed Anwar, Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Nick Barnes
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Semi-supervised Active Salient Object Detection
Yunqiu Lv, Bowen Liu 0012, Jing Zhang 0052, Yuchao Dai, Aixuan Li, Tong Zhang 0023
Pattern Recognit.3
2022 Weakly Supervised RGB-D Salient Object Detection With Prediction Consistency Training and Active Scribble Boosting
abstract
RGB-D salient object detection (SOD) has attracted increasingly more attention as it shows more robust results in complex scenes compared with RGB SOD. However, state-of-the-art RGB-D SOD approaches heavily rely on a large amount of pixel-wise annotated data for training. Such densely labeled annotations are often labor-intensive and costly. To reduce the annotation burden, we investigate RGB-D SOD from a weakly supervised perspective. More specifically, we use annotator-friendly scribble annotations as supervision signals for model training. Since scribble annotations are much sparser compared to ground-truth masks, some critical object structure information might be neglected. To preserve such structure information, we explicitly exploit the complementary edge information from two modalities (i.e., RGB and depth). Specifically, we leverage the dual-modal edge guidance and introduce a new network architecture with a dual-edge detection module and a modality-aware feature fusion module. In order to use the useful information of unlabeled pixels, we introduce a prediction consistency training scheme by comparing the predictions of two networks optimized by different strategies. Moreover, we develop an active scribble boosting strategy to provide extra supervision signals with negligible annotation cost, leading to significant SOD performance improvement. Extensive experiments on seven benchmarks validate the superiority of our proposed method. Remarkably, the proposed method with scribble annotations achieves competitive performance in comparison to fully supervised state-of-the-art methods.
Yunqiu Xu, Xin Yu 0002, Jing Zhang 0052, Linchao Zhu, Dadong Wang
IEEE Trans. Image Process.3
2021 Uncertainty-Aware Joint Salient Object and Camouflaged Object Detection
abstract
Visual salient object detection (SOD) aims at finding the salient object(s) that attract human attention, while camouflaged object detection (COD) on the contrary intends to discover the camouflaged object(s) that hidden in the surrounding. In this paper, we propose a paradigm of lever-aging the contradictory information to enhance the detection ability of both salient object detection and camouflaged object detection. We start by exploiting the easy positive samples in the COD dataset to serve as hard positive samples in the SOD task to improve the robustness of the SOD model. Then, we introduce a "similarity measure" module to explicitly model the contradicting attributes of these two tasks. Furthermore, considering the uncertainty of labeling in both tasks’ datasets, we propose an adversarial learning network to achieve both higher order similarity measure and network confidence estimation. Experimental results on benchmark datasets demonstrate that our solution leads to state-of-the-art (SOTA) performance for both tasks1.
Aixuan Li, Jing Zhang 0052, Yunqiu Lv, Bowen Liu 0012, Tong Zhang 0023, Yuchao Dai
CVPR2
2021 Simultaneously Localize, Segment and Rank the Camouflaged Objects
abstract
Camouflage is a key defence mechanism across species that is critical to survival. Common strategies for camouflage include background matching, imitating the color and pattern of the environment, and disruptive coloration, disguising body outlines [37]. Camouflaged object detection (COD) aims to segment camouflaged objects hiding in their surroundings. Existing COD models are built upon binary ground truth to segment the camouflaged objects without illustrating the level of camouflage. In this paper, we revisit this task and argue that explicitly modeling the conspicuousness of camouflaged objects against their particular backgrounds can not only lead to a better understanding about camouflage and evolution of animals, but also provide guidance to design more sophisticated camouflage techniques. Furthermore, we observe that it is some specific parts of the camouflaged objects that make them detectable by predators. With the above understanding about camouflaged objects, we present the first ranking based COD network (Rank-Net) to simultaneously localize, segment and rank camouflaged objects. The localization model is proposed to find the discriminative regions that make the camouflaged object obvious. The segmentation model segments the full scope of the camouflaged objects. Further, the ranking model infers the detectability of different camouflaged objects. Moreover, we contribute a large COD testing set to evaluate the generalization ability of COD models. Experimental results show that our model achieves new state-of-the-art, leading to a more interpretable COD network1.
Yunqiu Lv, Jing Zhang 0052, Yuchao Dai, Aixuan Li, Bowen Liu 0012, Nick Barnes, Deng-Ping Fan
CVPR2
2021 Weakly Supervised Video Salient Object Detection
abstract
Significant performance improvement has been achieved for fully-supervised video salient object detection with the pixel-wise labeled training datasets, which are time-consuming and expensive to obtain. To relieve the burden of data annotation, we present the first weakly super-vised video salient object detection model based on relabeled “fixation guided scribble annotations”. Specifically, an "Appearance-motion fusion module" and bidirectional ConvLSTM based framework are proposed to achieve effective multi-modal learning and long-term temporal context modeling based on our new weak annotations. Further, we design a novel foreground-background similarity loss to further explore the labeling similarity across frames. A weak annotation boosting strategy is also introduced to boost our model performance with a new pseudo-label generation technique. Extensive experimental results on six benchmark video saliency detection datasets illustrate the effectiveness of our solution1.
Wangbo Zhao, Jing Zhang 0052, Long Li 0008, Nick Barnes, Nian Liu 0002, Junwei Han 0001
CVPR2
2021 RGB-D Saliency Detection via Cascaded Mutual Information Minimization
abstract
Existing RGB-D saliency detection models do not explicitly encourage RGB and depth to achieve effective multi-modal learning. In this paper, we introduce a novel multistage cascaded learning framework via mutual information minimization to explicitly model the multi-modal information between RGB image and depth data. Specifically, we first map the feature of each mode to a lower dimensional feature vector, and adopt mutual information minimization as a regularizer to reduce the redundancy between appearance features from RGB and geometric features from depth. We then perform multi-stage cascaded learning to impose the mutual information minimization constraint at every stage of the network. Extensive experiments on benchmark RGB-D saliency datasets illustrate the effectiveness of our framework. Further, to prosper the development of this field, we contribute the largest (7× larger than NJU2K) COME15K dataset, which contains 15,625 image pairs with high quality polygon-/scribble-/object-/instance-/rank-level annotations. Based on these rich labels, we additionally construct four new benchmarks with strong baselines and observe some interesting phenomena, which can motivate future model design. Source code and dataset are available at https://github.com/JingZhang617/cascaded_rgbd_sod.
Jing Zhang 0052, Deng-Ping Fan, Yuchao Dai, Xin Yu 0002, Yiran Zhong, Nick Barnes, Ling Shao 0001
ICCV1
2021 Learning structure-aware semantic segmentation with image-level supervision
abstract
Compared with expensive pixel-wise annotations, image-level labels make it possible to learn semantic segmentation in a weakly-supervised manner. Within this pipeline, the class activation map (CAM) is obtained and further processed to serve as a pseudo label to train the semantic segmentation model in a fully-supervised manner. In this paper, we argue that the lost structure information in CAM limits its application in downstream semantic segmentation, leading to deteriorated predictions. Furthermore, the inconsistent class activation scores inside the same object contradicts the common sense that each region of the same object should belong to the same semantic category. To produce sharp prediction with structure information, we introduce an auxiliary semantic boundary detection module, which penalizes the deteriorated predictions. Furthermore, we adopt smoothness loss to encourage prediction inside the object to be consistent. Experimental results on the PASCAL-VOC dataset illustrate the effectiveness of the proposed solution.
Jiawei Liu 0005, Jing Zhang 0052, Yicong Hong, Nick Barnes
IJCNN2
2021 Learning Generative Vision Transformer with Energy-Based Latent Space for Saliency Prediction
abstract
Vision transformer networks have shown superiority in many computer vision tasks. In this paper, we take a step further by proposing a novel generative vision transformer with latent variables following an informative energy-based prior for salient object detection. Both the vision transformer network and the energy-based prior model are jointly trained via Markov chain Monte Carlo-based maximum likelihood estimation, in which the sampling from the intractable posterior and prior distributions of the latent variables are performed by Langevin dynamics. Further, with the generative vision transformer, we can easily obtain a pixel-wise uncertainty map from an image, which indicates the model confidence in predicting saliency from the image. Different from the existing generative models which define the prior distribution of the latent variables as a simple isotropic Gaussian distribution, our model uses an energy-based informative prior which can be more expressive to capture the latent space of the data. We apply the proposed framework to both RGB and RGB-D salient object detection tasks. Extensive experimental results show that our framework can achieve not only accurate saliency predictions but also meaningful uncertainty maps that are consistent with the human perception.
Jing Zhang 0052, Jianwen Xie, Nick Barnes, Ping Li 0001
NeurIPS1
2021 Learning Saliency From Single Noisy Labelling: A Robust Model Fitting Perspective
abstract
The advances made in predicting visual saliency using deep neural networks come at the expense of collecting large-scale annotated data. However, pixel-wise annotation is labor-intensive and overwhelming. In this paper, we propose to learn saliency prediction from a single noisy labelling, which is easy to obtain (e.g., from imperfect human annotation or from unsupervised saliency prediction methods). With this goal, we address a natural question: Can we learn saliency prediction while identifying clean labels in a unified framework? To answer this question, we call on the theory of robust model fitting and formulate deep saliency prediction from a single noisy labelling as robust network learning and exploit model consistency across iterations to identify inliers and outliers (i.e., noisy labels). Extensive experiments on different benchmark datasets demonstrate the superiority of our proposed framework, which can learn comparable saliency prediction with state-of-the-art fully supervised saliency methods. Furthermore, we show that simply by treating ground truth annotations as noisy labelling, our framework achieves tangible improvements over state-of-the-art methods.
Jing Zhang 0052, Yuchao Dai, Tong Zhang 0023, Mehrtash Harandi, Nick Barnes, Richard I. Hartley
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 3D Guided Weakly Supervised Semantic Segmentation
Weixuan Sun, Jing Zhang 0052, Nick Barnes
ACCV (1)2
2020 UC-Net: Uncertainty Inspired RGB-D Saliency Detection via Conditional Variational Autoencoders
abstract
In this paper, we propose the first framework (UCNet) to employ uncertainty for RGB-D saliency detection by learning from the data labeling process. Existing RGB-D saliency detection methods treat the saliency detection task as a point estimation problem, and produce a single saliency map following a deterministic learning pipeline. Inspired by the saliency data labeling process, we propose probabilistic RGB-D saliency detection network via conditional variational autoencoders to model human annotation uncertainty and generate multiple saliency maps for each input image by sampling in the latent space. With the proposed saliency consensus process, we are able to generate an accurate saliency map based on these multiple predictions. Quantitative and qualitative evaluations on six challenging benchmark datasets against 18 competing algorithms demonstrate the effectiveness of our approach in learning the distribution of saliency maps, leading to a new state-of-the-art in RGB-D saliency detection.
Jing Zhang 0052, Deng-Ping Fan, Yuchao Dai, Saeed Anwar, Fatemehsadat Saleh, Tong Zhang 0023, Nick Barnes
CVPR1
2020 Weakly-Supervised Salient Object Detection via Scribble Annotations
abstract
Compared with laborious pixel-wise dense labeling, it is much easier to label data by scribbles, which only costs 1~2 seconds to label one image. However, using scribble labels to learn salient object detection has not been explored. In this paper, we propose a weakly-supervised salient object detection model to learn saliency from such annotations. In doing so, we first relabel an existing large-scale salient object detection dataset with scribbles, namely S-DUTS dataset. Since object structure and detail information is not identified by scribbles, directly training with scribble labels will lead to saliency maps of poor boundary localization. To mitigate this problem, we propose an auxiliary edge detection task to localize object edges explicitly, and a gated structure-aware loss to place constraints on the scope of structure to be recovered. Moreover, we design a scribble boosting scheme to iteratively consolidate our scribble annotations, which are then employed as supervision to learn high-quality saliency maps. As existing saliency evaluation metrics neglect to measure structure alignment of the predictions, the saliency map ranking may not comply with human perception. We present a new metric, termed saliency structure measure, as a complementary metric to evaluate sharpness of the prediction. Extensive experiments on six benchmark datasets demonstrate that our method not only outperforms existing weakly-supervised/unsupervised methods, but also is on par with several fully-supervised state-of-the-art models (Our code and data is publicly available at: https://github.com/JingZhang617/Scribble_Saliency).
Jing Zhang 0052, Xin Yu 0002, Aixuan Li, Peipei Song, Bowen Liu 0012, Yuchao Dai
CVPR1
2020 Learning Noise-Aware Encoder-Decoder from Noisy Labels by Alternating Back-Propagation for Saliency Detection
Jing Zhang 0052, Jianwen Xie, Nick Barnes
ECCV (17)1
2019 Few-Shot Learning via Saliency-Guided Hallucination of Samples
abstract
Learning new concepts from a few of samples is a standard challenge in computer vision. The main directions to improve the learning ability of few-shot training models include (i) a robust similarity learning and (ii) generating or hallucinating additional data from the limited existing samples. In this paper, we follow the latter direction and present a novel data hallucination model. Currently, most datapoint generators contain a specialized network (i.e., GAN) tasked with hallucinating new datapoints, thus requiring large numbers of annotated data for their training in the first place. In this paper, we propose a novel less-costly hallucination method for few-shot learning which utilizes saliency maps. To this end, we employ a saliency network to obtain the foregrounds and backgrounds of available image samples and feed the resulting maps into a two-stream network to hallucinate datapoints directly in the feature space from viable foreground-background combinations. To the best of our knowledge, we are the first to leverage saliency maps for such a task and we demonstrate their usefulness in hallucinating additional datapoints for few-shot learning. Our proposed network achieves the state of the art on publicly available datasets.
Jing Zhang 0052, Piotr Koniusz
CVPR2
2018 Deep Unsupervised Saliency Detection: A Multiple Noisy Labeling Perspective
abstract
The success of current deep saliency detection methods heavily depends on the availability of large-scale supervision in the form of per-pixel labeling. Such supervision, while labor-intensive and not always possible, tends to hinder the generalization ability of the learned models. By contrast, traditional handcrafted features based unsupervised saliency detection methods, even though have been surpassed by the deep supervised methods, are generally dataset-independent and could be applied in the wild. This raises a natural question that "Is it possible to learn saliency maps without using labeled data while improving the generalization ability?". To this end, we present a novel perspective to unsupervised saliency detection through learning from multiple noisy labeling generated by "weak" and "noisy" unsupervised handcrafted saliency methods. Our end-to-end deep learning framework for unsupervised saliency detection consists of a latent saliency prediction module and a noise modeling module that work collaboratively and are optimized jointly. Explicit noise modeling enables us to deal with noisy saliency maps in a probabilistic way. Extensive experimental results on various benchmarking datasets show that our model not only outperforms all the unsupervised saliency methods with a large margin but also achieves comparable performance with the recent state-of-the-art supervised deep saliency methods.
Jing Zhang 0052, Tong Zhang 0023, Yuchao Dai, Mehrtash Harandi, Richard I. Hartley
CVPR1
2017 Integrated deep and shallow networks for salient object detection
abstract
Deep convolutional neural network (CNN) based salient object detection methods have achieved state-of-the-art performance and outperform those unsupervised methods with a wide margin. In this paper, we propose to integrate deep and unsupervised saliency for salient object detection under a unified framework. Specifically, our method takes results of unsupervised saliency (Robust Background Detection, RBD) and normalized color images as inputs, and directly learns an end-to-end mapping between inputs and the corresponding saliency maps. The color images are fed into a Fully Convolutional Neural Networks (FCNN) adapted from semantic segmentation to exploit high-level semantic cues for salient object detection. Then the results from deep FCNN and RBD are concatenated to feed into a shallow network to map the concatenated feature maps to saliency maps. Finally, to obtain a spatially consistent saliency map with sharp object boundaries, we fuse superpixel level saliency map at multi-scale. Extensive experimental results on 8 benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art approaches with a margin.
Jing Zhang 0052, Bo Li 0090, Yuchao Dai, Fatih Porikli, Mingyi He
ICIP1
2017 Deep Salient Object Detection by Integrating Multi-level Cues
abstract
A key problem in salient object detection is how to effectively exploit the multi-level saliency cues in a unified and data-driven manner. In this paper, building upon the recent success of deep neural networks, we propose a fully convolutional neural network based approach empowered with multi-level fusion to salient object detection. By integrating saliency cues at different levels through fully convolutional neural networks and multi-level fusion, our approach could effectively exploit both learned semantic cues and higher-order region statistics for edge-accurate salient object detection. First, we fine-tune a fully convolutional neural network for semantic segmentation to adapt it to salient object detection to learn a suitable yet coarse perpixel saliency prediction map. This map is often smeared across salient object boundaries since the local receptive fields in the convolutional network apply naturally on both sides of such boundaries. Second, to enhance the resolution of the learned saliency prediction and to incorporate higher-order cues that are omitted by the neural network, we propose a multi-level fusion approach where super-pixel level coherency in saliency is exploited. Our extensive experimental results on various benchmark datasets demonstrate that the proposed method outperforms the state-of the-art approaches.
Jing Zhang 0052, Yuchao Dai, Fatih Porikli
WACV1
2016 Hyperspectral image classification based on deep stacking network
abstract
Hyperspectral image (HIS) classification is a hot topic in remote sensing community and most of the existing methods extract the features of original Hyperspectral data using shallow layer networks such as neural network (NN) and support vector machine (SVM). As deep learning recently achieves great success in machine learning and pattern recognition area for its ability in deep feature extraction and representations, two deep networks i.e. deep convolutional network (DCN) and deep belief network (DBN) have been used for hyperspectral image classification and better results have been achieved. Differing from those deep networks for HSI classification, in this paper, we propose a new method for hyperspectral image classification based on deep stacking network (DSN), which owns advantages to other deep models for its simplicity when processing in batch-mode learning - not requiring stochastic gradient descent that other DNNs require. The feature extraction is gradually obtained by employing nonlinear activation function on the hidden layer nodes of each module, which is different from those DSNs that usually use linear weights between the hidden layer and the output layer. Experimental results on AVIRIS hyperspectral images show that the proposed method achieves improved classification performance when compared with that via SVM and NN methods.
Mingyi He, Yifan Zhang 0006, Jing Zhang 0052
IGARSS4