Bihan Wen

dblp:158/9840 · DBLP profile ↗
← Back
150ranked-venue papers
11as first author
119since 2021 · last 2026
0000-0002-6874-6453ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 98 · 10 first-author · 71 since 2021Artificial intelligence and machine learning · 60 · 2 first-author · 53 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 12 since 2021Security and privacy · 7 · 6 since 2021Systems, architecture and hardware · 5 · 5 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Conformal Prediction for Multi-Source Detection on a Network
abstract
Detecting the origin of information or infection spread in networks is a fundamental challenge with applications in misinformation tracking, epidemiology, and beyond. We study the multi-source detection problem: given snapshot observations of node infection status on a graph, estimate the set of source nodes that initiated the propagation. Existing methods either lack statistical guarantees or are limited to specific diffusion models and assumptions. We propose a novel conformal prediction framework that provides statistically valid recall guarantees for source set detection, independent of the underlying diffusion process or data distribution. Our approach introduces principled score functions to quantify the alignment between predicted probabilities and true sources, and leverages a calibration set to construct prediction sets with user-specified recall and coverage levels. The method is applicable to both single- and multi-source scenarios, supports general network diffusion dynamics, and is computationally efficient for large graphs. Empirical results demonstrate that our method achieves rigorous coverage with competitive accuracy, outperforming existing baselines in both reliability and scalability.
Xingchao Jian, Purui Zhang 0001, Lan Tian, Wenfei Liang 0001, Wee-Peng Tay, Bihan Wen, Felix Krahmer
AAAI7
2026 SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision
abstract
Large Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucination issues still limit their real-world applicability. Although previous mitigation methods effectively reduce hallucinations in photographic images, they largely overlook the potential risks posed by stylized images, which play crucial roles in critical scenarios such as game scene understanding, art education, and medical analysis. In this work, we first construct a dataset comprising photographic images and their corresponding stylized versions with carefully annotated caption labels. We then conduct head-to-head comparisons on both discriminative and generative tasks by benchmarking 13 advanced LVLMs on the collected datasets. Our findings reveal that stylized images tend to induce significantly more hallucinations than their photographic counterparts. To address this issue, we propose Style-Aware Visual Early Revision (SAVER), a novel mechanism that dynamically adjusts LVLMs' final outputs based on the token-level visual attention patterns, leveraging early-layer feedback to mitigate hallucinations caused by stylized images. Extensive experiments demonstrate that SAVER achieves state-of-the-art performance in hallucination mitigation across various models, datasets, and tasks.
Zhaoxu Li, Chenqi Kong, Yi Yu 0011, Qiangqiang Wu, Xinghao Jiang, Ngai-Man Cheung, Bihan Wen, Alex Chichung Kot, Xudong Jiang 0001
AAAI7
2026 IL-DiffTSF: Invertible Latent Diffusion for Probabilistic Time Series Forecasting
abstract
Internet of Things (IoT) devices generate large volumes of time series data that are often volatile and complex, making probabilistic time series forecasting (TSF) essential for modeling the distribution of future outcomes. Recently, diffusion-based TSF methods have gained attention for their ability to learn complex distributions. However, they typically apply the diffusion process directly in the time domain, which may struggle to capture complex temporal dependencies, thus limiting the full potential of the diffusion process. Besides, they obtain probabilistic forecasts by sampling multiple plausible outcomes from the learned distribution, which is time-consuming and less effective. To solve these problems, we propose Invertible Latent Diffusion for probabilistic Time Series Forecasting (IL-DiffTSF), a novel approach based on a latent diffusion model. Specifically, we design an invertible latent projection between time series and latent space, where a conditional diffusion process is applied. This design ensures bidirectional consistency and minimal information loss, enabling more accurate TSF. Moreover, instead of sampling-based probabilistic forecasting, IL-DiffTSF represents uncertainties by directly learning a mapping from latent representations to prediction errors, achieving faster and more reliable uncertainty estimates. Experiments on univariate and multivariate benchmarks validate the efficiency and effectiveness of IL-DiffTSF. The code for this project is available at https://github.com/vanerkz/IL-DiffTSF.
Van Kwan Zhi Koh, Songnan Lin, Zhiping Lin 0001, Bihan Wen
IEEE Internet Things J.5
2026 WBCAtt+: Fine-grained pixel-level morphological annotations for white blood cell images
Satoshi Tsutsui, Winnie Pang, Shuting He, Bihan Wen
Medical Image Anal.4
2026 Dynamic-Aware video distillation: Adaptive temporal partitioning based on video semantics for edge device
Yinjie Zhao, Heng Zhao 0004, Yew-Soon Ong, Bihan Wen, Joey Tianyi Zhou
Neural Networks4
2026 A regularized deep self-expression feature augmentation network for few-shot unconstrained palmprint recognition
Kunlei Jing, Hebo Ma, Chen Zhang 0013, Zhiyuan Zha, Bihan Wen
Pattern Recognit.5
2026 Single-Image Reflection Removal via Iterative Prompt Learning of Reflection Level
abstract
Single-image reflection removal (SIRR) aims to restore the latent background layer from a reflection-contaminated image. Despite the promising progress achieved by deep learning-based methods, the roles of negative training samples and descriptive prompts for the reflection severity are underexplored in most existing deep SIRR approaches, limiting their reflection removal performance and generalization capability. In this work, we introduce a novel training framework that synergistically leverages learnable prompts and image data to optimize the restoration network. To this end, we define reflection levels corresponding to varying degrees of reflection interference on the background content and learn reflection-level prompts to supervise the SIRR process. We propose an Iterative Reflection Level Reduction (IRLR) framework composed of a Restoration Network Training Module (RNTM) and a Reflection Level Learning Module (RLLM). Specifically, RNTM predicts the background layer under the guidance of prompts learned by RLLM, while RLLM in turn refines these prompts using outputs from RNTM. The two modules are trained iteratively to progressively reduce the reflection levels of estimated background layers. To initialize the prompts, we construct a dedicated reflection-level dataset for pretraining. For adaptively supervising RNTM, we design a new reflection-level-aware strategy to address the challenge of directly aligning the output background with the minimal reflection level. Comprehensive experimental results demonstrate that the proposed method significantly outperforms state-of-the-art methods on average performance across several released datasets, improving PSNR by 0.82 dB and SSIM by 0.0120, respectively. The source code and dataset are available at https://github.com/NamecantbeNULL/IRLR_SIRR.
Binbin Song, Jiantao Zhou 0001, Shuning Xu, Xina Liu, Haiwei Wu, Xiaopeng Fan 0001, Bihan Wen
IEEE Trans. Image Process.7
2026 MSF-Mamba: Motion-Aware State Fusion Mamba for Efficient Micro-Gesture Recognition
abstract
Micro-gesture recognition (MGR) targets the identification of subtle and fine-grained human motions and requires accurate modeling of both long-range and local spatiotemporal dependencies. While convolutional neural networks (CNNs) are effective at capturing local patterns, they struggle with long-range dependencies due to their limited receptive fields. Transformer-based models address this limitation through self-attention mechanisms but suffer from high computational costs. Recently, Mamba has shown promise as an efficient model, leveraging state space models (SSMs) to enable linear-time processing. However, directly applying the vanilla Mamba to MGR may not be optimal. This is because Mamba processes inputs as 1D sequences, with state updates relying solely on the previous state, and thus lacks the ability to model local spatiotemporal dependencies. In addition, previous methods lack a design of motion-awareness, which is crucial in MGR. To overcome these limitations, we propose motion-aware state fusion mamba (MSF-Mamba), which enhances Mamba with local spatiotemporal modeling by fusing local contextual neighboring states. Our design introduces a motion-aware state fusion module based on central frame difference (CFD). Furthermore, a multiscale version named MSF-Mamba$^{+}$has been proposed. Specifically, MSF-Mamba$^{+}$supports multiscale motion-aware state fusion, as well as an adaptive scale weighting module that dynamically weighs the fused states across different scales. These enhancements explicitly address the limitations of vanilla Mamba by enabling motion-aware local spatiotemporal modeling, allowing MSF-Mamba and MSF-Mamba$^{+}$to effectively capture subtle motion cues for MGR. Experiments on two public MGR datasets (i.e., SMG and iMiGUE) demonstrate that even the lightweight version, namely, MSF-Mamba, achieves state-of-the-art performance, outperforming existing CNN-, Transformer-, and SSM-based models while maintaining high efficiency. For example, MSF-Mamba improves Top-1 accuracy by +2.2% and +1.5% over VideoMamba on SMG and iMiGUE, respectively. MSF-Mamba$^{+}$outperforms VideoMamba on SMG and iMiGUE, achieving Top-1 accuracy improvements of 2.9% and 3.0%, respectively. The code will be accessible onhttps://github.com/Leedeng/MSF-Mamba
Deng Li 0002, Bohao Xing, Rong Gao 0005, Bihan Wen, Heikki Kälviäinen, Xin Liu 0012
IEEE Trans. Multim.5
2025 ZoRI: Towards Discriminative Zero-Shot Remote Sensing Instance Segmentation
abstract
Instance segmentation algorithms in remote sensing are typically based on conventional methods, limiting their application to seen scenarios and closed-set predictions. In this work, we propose a novel task called zero-shot remote sensing instance segmentation, aimed at identifying aerial objects that are absent from training data. Challenges arise when classifying aerial categories with high inter-class similarity and intra-class variance. Besides, the domain gap between vision-language models’ pretraining datasets and remote sensing datasets hinders the zero-shot capabilities of the pretrained model when it is directly applied to remote sensing images. To address these challenges, we propose a Zero-Shot Remote Sensing Instance Segmentation framework, dubbed ZoRI. Our approach features a discrimination-enhanced classifier that uses refined textual embeddings to increase the awareness of class disparities. Instead of direct fine-tuning, we propose a knowledge-maintained adaptation strategy that decouples semantic-related information to preserve vision-language alignment while adjusting features to capture remote sensing domain-specific visual cues. Additionally, we introduce a prior-injected prediction with cache bank of aerial visual prototypes to supplement the semantic richness of text embeddings and seamlessly integrate aerial representations, adapting to the remote sensing domain. We establish new experimental protocols and benchmarks, and extensive experiments demonstrate that ZoRI achieves the state-of-art performance on the zero-shot remote sensing instance segmentation task.
Shuting He, Bihan Wen
AAAI3
2025 SoftShadow: Leveraging Soft Masks for Penumbra-Aware Shadow Removal
abstract
Recent advancements in deep learning have yielded promising results for the image shadow removal task. However, most existing methods rely on binary pre-generated shadow masks. The binary nature of such masks could potentially lead to artifacts near the boundary between shadow and non-shadow areas. In view of this, inspired by the physical model of shadow formation, we introduce novel soft shadow masks specifically designed for shadow removal. To achieve such soft masks, we propose a SoftShadow framework by leveraging the prior knowledge of pretrained SAM and integrating physical constraints. Specifically, we jointly tune the SAM and the subsequent shadow removal network using penumbra formation constraint loss, mask reconstruction loss, and shadow removal loss. This framework enables accurate predictions of penumbra (partially shaded) and umbra (fully shaded) areas while simultaneously facilitating end-to-end shadow removal. Through extensive experiments on popular datasets, we found that our Soft-Shadow framework, which generates soft masks, can better restore boundary artifacts, achieve state-of-the-art performance, and demonstrate superior generalizability.
Xinrui Wang 0004, Lanqing Guo, Siyu Huang, Bihan Wen
CVPR5
2025 Reconciling Stochastic and Deterministic Strategies for Zero-shot Image Restoration using Diffusion Model in Dual
abstract
Plug-and-play (PnP) methods offer an iterative strategy for solving image restoration (IR) problems in a zero-shot manner, using a learned discriminative denoiser as the implicit prior. More recently, a sampling-based variant of this approach, which utilizes a pre-trained generative diffusion model, has gained great popularity for solving IR problems through stochastic sampling. The IR results using PnP with a pre-trained diffusion model demonstrate distinct advantages compared to those using discriminative denoisers, i.e.,improved perceptual quality while sacrificing the data fidelity. The unsatisfactory results are due to the lack of integration of these strategies in the IR tasks. In this work, we propose a novel zero-shot IR scheme, dubbed Reconciling Diffusion Model in Dual (RDMD), which leverages only a single pre-trained diffusion model to construct two complementary regularizers. Specifically, the diffusion model in RDMD will iteratively perform deterministic denoising and stochastic sampling, aiming to achieve highfidelity image restoration with appealing perceptual quality. RDMD also allows users to customize the distortion-perception tradeoff with a single hyperparameter, enhancing the adaptability of the restoration process in different practical scenarios. Extensive experiments on several IR tasks demonstrate that our proposed method could achieve superior results compared to existing approaches on both the FFHQ and ImageNet datasets. Code is available at https://github.com/chongwang1024/rdmd.
Chong Wang 0011, Lanqing Guo, Zixuan Fu, Siyuan Yang 0001, Hao Cheng 0016, Alex Chichung Kot, Bihan Wen
CVPR7
2025 Discovering Hidden Visual Concepts Beyond Linguistic Input in Infant Learning
abstract
Infants develop complex visual understanding rapidly, even preceding of the acquisition of linguistic skills. As computer vision seeks to replicate the human vision system, understanding infant visual development may offer valuable insights. In this paper, we present an interdisciplinary study exploring this question: can a computational model that imitates the infant learning process develop broader visual concepts that extend beyond the vocabulary it has heard, similar to how infants naturally learn? To investigate this, we analyze a recently published model in Science by Vong et al., which is trained on longitudinal, egocentric images of a single child paired with transcribed parental speech. We perform neuron labeling to identify visual concept neurons hidden in the model’s internal representations. We then demonstrate that these neurons can recognize objects beyond the model’s original vocabulary. Furthermore, we compare the differences in representation between infant models and those in modern computer vision models, such as CLIP and ImageNet pre-trained model. Ultimately, our work bridges cognitive science and computer vision by analyzing the internal representations of a computational model trained on an infant visual and linguistic inputs. Our code is available at https://github.com/Kexueyi/discover_infant_vis.
Xueyi Ke, Satoshi Tsutsui, Yayun Zhang, Bihan Wen
CVPR4
2025 Hyperspectral Image Reconstruction with Unseen Material Detection
abstract
Reconstruction of hyperspectral images (HSIs) from their RGB measurements is an ill-posed inverse problem. The key to successful reconstruction relies on establishing an effective HSI prior, for which deep learning techniques have achieved impressive performance. However, the scarcity of large-scale HSI datasets poses a significant challenge, limiting the practical application of deep HSI reconstruction methods and often leading to incorrect results when dealing with unseen substances or materials. To tackle this challenge, we propose a deep RGB-to-HSI reconstruction model based on the sparse prior of hyperspectral signals. The network can effectively correlate HSI and RGB features via shared sparse codes, representing the weights of spectral-unique materials. Besides, testing images with unseen materials can be detected by measuring their sparse modeling errors. Experimental results demonstrate that the proposed method achieves promising results on RGB-to-HSI reconstruction. Further, the sparse modeling error evidently demonstrates its efficacy as an indicator for unseen materials.
Songnan Lin, Bihan Wen
ICASSP3
2025 Distribution Alignment Informed Thresholding for Semi-Supervised Curvilinear Structure Segmentation
abstract
Curvilinear structure segmentation using deep neural networks is often limited by the high cost of annotation. Semi-supervised learning (SSL) helps mitigate this dependency on extensive annotated data. State-of-the-art SSL approaches generate pseudo-labels for unlabeled data, which are then used for further model training. These methods primarily focus on calibrating thresholds to binarize the predictions. In this work, we assume that when labeled and unlabeled data are similar, the foreground-to-background ratio should be consistent between them. To leverage this assumption, we calibrate the threshold by minimizing the distribution gap between labeled ground truth and pseudo-labels on unlabeled data. Our proposed threshold calibration can be integrated with existing SSL methods. We evaluate its effectiveness on four datasets, demonstrating that our method outperforms current state-of-the-art SSL techniques, especially in scenarios with very low labeled data.
Yuhao Mo, Bihan Wen, Xulei Yang, Ce Zhu, Xun Xu 0002
ICASSP3
2025 Low-Rank Transformer Adaptation for Arbitrary Style Transfer
abstract
Arbitrary style transfer aims to apply artistic characteristics from a style reference to an image while preserving the image’s original content. Although many methods have achieved remarkable results in style transfer, they typically rely on largescale datasets for training, which increases both the cost and complexity of data collection. To address this issue, we propose a Low-rank Transformer Adaptation method for style transfer which leverages the efficiency of low-rank adaptation to reduce the model’s complexity without compromising performance. It not only accelerates the training process but also delivers high-quality image generation with a small-scale dataset. Additionally, we introduce an edge detection loss to enhance the preservation of content outlines further, ensuring that the fine details of the image are maintained during the style transfer process Experimental results demonstrate that our method achieves competitive performance even with significantly less data, and it exhibits superiority in both visual quality and evaluation metrics.
Meichen Liu, Bihan Wen
ICASSP3
2025 SCORE: Scene Context Matters in Open-Vocabulary Remote Sensing Instance Segmentation
Shuting He, Huaiyuan Qin, Bihan Wen
ICCV4
2025 GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding
abstract
Sequential grounding in 3D point clouds (SG3D) refers to locating sequences of objects by following text instructions for a daily activity with detailed steps. Current 3D visual grounding (3DVG) methods treat text instructions with multiple steps as a whole, without extracting useful temporal information from each step. However, the instructions in SG3D often contain pronouns such as "it", "here" and "the same" to make language expressions concise. This requires grounding methods to understand the context and retrieve relevant information from previous steps to correctly locate object sequences. Due to the lack of an effective module for collecting related historical information, state-of-the-art 3DVG methods face significant challenges in adapting to the SG3D task. To fill this gap, we propose GroundFlow -- a plug-in module for temporal reasoning on 3D point cloud sequential grounding. Firstly, we demonstrate that integrating GroundFlow improves the task accuracy of 3DVG baseline methods by a large margin (+7.5\% and +10.2\%) in the SG3D benchmark, even outperforming a 3D large language model pre-trained on various datasets. Furthermore, we selectively extract both short-term and long-term step information based on its relevance to the current instruction, enabling GroundFlow to take a comprehensive view of historical information and maintain its temporal understanding advantage as step counts increase. Overall, our work introduces temporal reasoning capabilities to existing 3DVG models and achieves state-of-the-art performance in the SG3D benchmark across five datasets.
Shuting He, Cheston Tan, Bihan Wen
ICCV4
2025 Training-Free Text-Guided Image Editing with Visual Autoregressive Model
abstract
Text-guided image editing is an essential task that enables users to modify images through natural language descriptions. Recent advances in diffusion models and rectified flows have significantly improved editing quality, primarily relying on inversion techniques to extract structured noise from input images. However, inaccuracies in inversion can propagate errors, leading to unintended modifications and compromising fidelity. Moreover, even with perfect inversion, the entanglement between textual prompts and image features often results in global changes when only local edits are intended. To address these challenges, we propose a novel text-guided image editing framework based on VAR (Visual AutoRegressive modeling), which eliminates the need for explicit inversion while ensuring precise and controlled modifications. Our method introduces a caching mechanism that stores token indices and probability distributions from the original image, capturing the relationship between the source prompt and the image. Using this cache, we design an adaptive fine-grained masking strategy that dynamically identifies and constrains modifications to relevant regions, preventing unintended changes. A token reassembling approach further refines the editing process, enhancing diversity, fidelity, and control. Our framework operates in a training-free manner and achieves high-fidelity editing with faster inference speeds, processing a 1K resolution image in as fast as 1.2 seconds. Extensive experiments demonstrate that our method achieves performance comparable to, or even surpassing, existing diffusion- and rectified flow-based approaches in both quantitative metrics and visual quality. The code will be released.
Yufei Wang 0006, Lanqing Guo, Jiaxing Huang 0001, Pichao Wang, Bihan Wen
ICCV6
2025 M-SpecGene: Generalized Foundation Model for RGBT Multispectral Vision
Kailai Zhou, Fuqiang Yang, Shixian Wang, Bihan Wen, Chongde Zi, Linsen Chen, Qiu Shen, Xun Cao
ICCV4
2025 Harnessing Forecast Uncertainty in Deep Learning for Time Series Anomaly Detection with Posterior Distribution Scoring
abstract
Time series anomaly detection tools (TSAD) are widely applicable across industries, such as monitoring time series data of water pipeline pressure, network traffic activities, and hardware telemetry. The primary objective is to identify anomalous segments and alert users to potential issues before any consequences. A closely related tool is time series forecasting, and some practitioners leverage it for anomaly detection. As the main objective of a forecasting model is to minimize errors, it tends to over-fit the time series, making it challenging to distinguish whether the forecasting errors occur due to model limitations or true anomalous segments. This paper introduces a method called the posterior anomaly scoring criterion, which uses deep learning time series forecasting models to estimate forecast uncertainties for TSAD. We propose replacing the forecasting model’s output layer to estimate forecast distributions and compute the probability of the posterior distribution to attain anomaly scores. These scores are processed through an automated threshold criterion to classify the anomalous segments. The experiments have demonstrated our model performs the best in four out of five datasets across seven benchmark models.
Van Kwan Zhi Koh, Ehsan Shafiee, Zhiping Lin 0001, Bihan Wen
ISCAS5
2025 TWavefussion: Wavelet-based Diffusion with Transformer for Multivariate Time Series Anomaly Detection
abstract
Multivariate Time Series (MTS) anomaly detection is challenging in distinguishing anomalous data from normal data in high-dimensional, complex distributions. Even some deep learning methods still have difficulties capturing intricate MTS patterns. Recent advancements in Diffusion Models (DM) for sample generation have inspired us to explore their potential in MTS anomaly detection. In this paper, TWavefussion, an unsupervised diffusion model for MTS anomaly detection combining wavelet-based diffusion model and transformer autoencoder, is proposed. The wavelet-based diffusion model captures fine-grained local features in the high-frequency components of latent features and helps fuse both local and global MTS features better. Comparative experiments show TWavefussion achieves leading performance on three of four datasets.
Hongjun Sheng, Xinggan Peng, Van Kwan Zhi Koh, Bihan Wen, Zhiping Lin 0001
ISCAS4
2025 DEEMO: De-identity Multimodal Emotion Recognition and Reasoning
abstract
Emotion understanding is a critical yet challenging task. Most existing approaches rely heavily on identity-sensitive information, such as facial expressions and speech, which raises concerns about personal privacy. To address this, we introduce the De-identity Multimodal Emotion Recognition and Reasoning ( DEEMO ), a novel task designed to enable emotion understanding using de-identified video and audio inputs. The DEEMO dataset consists of two subsets: DEEMO-NFBL , which includes rich annotations of Non-Facial Body Language (NFBL), and DEEMO-MER , an instruction dataset for Multimodal Emotion Recognition and Reasoning using identity-free cues. This design supports emotion understanding without compromising identity privacy. In addition, we propose DEEMO-LLaMA, a Multimodal Large Language Model (MLLM) that integrates de-identified audio, video, and textual information to enhance both emotion recognition and reasoning. Extensive experiments show that DEEMO-LLaMA achieves state-of-the-art performance on both tasks, outperforming existing MLLMs by a significant margin, achieving 74.49% accuracy and 74.45% F1-score in de-identity emotion recognition, and 6.20 clue overlap and 7.66 label overlap in de-identity emotion reasoning. Our work contributes to ethical AI by advancing privacy-preserving emotion understanding and promoting responsible affective computing. The dataset and codes will be available at https://github.com/Leedeng/DEEMO.
Deng Li 0002, Bohao Xing, Xin Liu 0012, Baiqiang Xia, Bihan Wen, Heikki Kälviäinen
ACM Multimedia5
2025 KaRF: Weakly-Supervised Kolmogorov-Arnold Networks-based Radiance Fields for Local Color Editing
abstract
Recent advancements have suggested that neural radiance fields (NeRFs) show great potential in color editing within the 3D domain. However, most existing NeRF-based editing methods continue to face significant challenges in local region editing, which usually lead to imprecise local object boundaries, difficulties in maintaining multi-view consistency, and over-reliance on annotated data. To address these limitations, in this paper, we propose a novel weakly-supervised method called KaRF for local color editing, which facilitates high-fidelity and realistic appearance edits in arbitrary regions of 3D scenes. At the core of the proposed KaRF approach is a unified two-stage Kolmogorov-Arnold Networks (KANs)-based radiance fields framework, comprising a segmentation stage followed by a local recoloring stage. This architecture seamlessly integrates geometric priors from NeRF to achieve weakly-supervised learning, leading to superior performance. More specifically, we propose a residual adaptive gating KAN structure, which integrates KAN with residual connections, adaptive parameters, and gating mechanisms to effectively enhance segmentation accuracy and refine specific editing effects. Additionally, we propose a palette-adaptive reconstruction loss, which can enhance the accuracy of additive mixing results. Extensive experiments demonstrate that the proposed KaRF algorithm significantly outperforms many state-of-the-art methods both qualitatively and quantitatively. Our code and more results are available at: https://github.com/PaiDii/KARF.git.
Wudi Chen, Zhiyuan Zha, Shigang Wang 0003, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Zipei Fan, Ce Zhu
NeurIPS4
2025 Sparc3D: Sparse Representation and Construction for High-Resolution 3D Shapes Modeling
abstract
High-fidelity 3D object synthesis remains significantly more challenging than 2D image generation due to the unstructured nature of mesh data and the cubic complexity of dense volumetric grids. Existing two-stage pipelines—compressing meshes with a VAE (using either 2D or 3D supervision), followed by latent diffusion sampling—often suffer from severe detail loss caused by inefficient representations and modality mismatches introduced in VAE. We introduce Sparc3D, a unified framework that combines a sparse deformable marching cubes representation Sparcubes with a novel encoder Sparconv-VAE. Sparcubes converts raw meshes into high-resolution ($1024^3$) surfaces with arbitrary topology by scattering signed distance and deformation fields onto a sparse cube, allowing differentiable optimization. Sparconv-VAE is the first modality-consistent variational autoencoder built entirely upon sparse convolutional networks, enabling efficient and near-lossless 3D reconstruction suitable for high-resolution generative modeling through latent diffusion. Sparc3D achieves state-of-the-art reconstruction fidelity on challenging inputs, including open surfaces, disconnected components, and intricate geometry. It preserves fine-grained shape details, reduces training and inference cost, and integrates naturally with latent diffusion models for scalable, high-resolution 3D generation.
Yufei Wang 0006, Heliang Zheng, Yihao Luo, Bihan Wen
NeurIPS5
2025 Compressed Event Sensing (CES) Volumes for Event Cameras
Songnan Lin, Jing Chen 0018, Bihan Wen
Int. J. Comput. Vis.4
2025 Revisiting One-Stage Deep Uncalibrated Photometric Stereo via Fourier Embedding
abstract
This paper introduces a one-stage deep uncalibrated photometric stereo (UPS) network, namely Fourier Uncalibrated Photometric Stereo Network (FUPS-Net), for non-Lambertian objects under unknown light directions. It departs from traditional two-stage methods that first explicitly learn lighting information and then estimate surface normals. Two-stage methods were deployed because the interplay of lighting with shading cues presents challenges for directly estimating surface normals without explicit lighting information. However, these two-stage networks are disjointed and separately trained so that the error in explicit light calibration will propagate to the second stage and cannot be eliminated. In contrast, the proposed FUPS-Net utilizes an embedded Fourier transform network to implicitly learn lighting features by decomposing inputs, rather than employing a disjointed light estimation network. Our approach is motivated from observations in the Fourier domain of photometric stereo images: lighting information is mainly encoded in amplitudes, while geometry information is mainly associated with phases. Leveraging this property, our method "decomposes" geometry and lighting in the Fourier domain as guidance, via the proposed Fourier Embedding Extraction (FEE) block and Fourier Embedding Aggregation (FEA) block, which generate lighting and geometry features for the FUPS-Net to implicitly resolve the geometry-lighting ambiguity. Furthermore, we propose a Frequency-Spatial Weighted (FSW) block that assigns weights to combine features extracted from the frequency domain and those from the spatial domain for enhancing surface reconstructions. FUPS-Net overcomes the limitations of two-stage UPS methods, offering better training stability, a concise end-to-end structure, and avoiding accumulated errors in disjointed networks. Experimental results on synthetic and real datasets demonstrate the superior performance of our approach, and its simpler training setup, potentially paving the way for a new strategy in deep learning-based UPS methods.
Yakun Ju, Boxin Shi, Bihan Wen, Kin-Man Lam 0001, Xudong Jiang 0001, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 On the Adversarial Vulnerabilities of Transfer Learning in Remote Sensing
abstract
Transfer learning with pretrained models from general computer vision tasks is widely adopted in remote sensing, offering reduced training costs and improved performance across applications like scene classification and semantic segmentation. However, this practice exposes downstream tasks to significant vulnerabilities, as adversaries can exploit publicly available pre-trained models to craft attacks that compromise model integrity. This paper introduces Adversarial Neuron Manipulation (ANM), a novel attack strategy with two variants—ANM-S (Single Neuron) and ANM-M (Multiple Neurons)—that generates highly transferable perturbations by selectively targeting fragile neurons in pretrained models. Unlike existing adversarial attacks, ANM requires no domain-specific knowledge, enhancing its applicability and efficiency across diverse settings. Evaluated on various benchmark datasets, ANM significantly degrades model performance. ANM-M, in particular, reduces accuracy by up to 84.24% in scene classification, 42.47% in semantic segmentation, and 75.33% in object detection (mAP). These results reveal critical weaknesses in deep learning models, especially their susceptibility to transferable attacks across CNN and Transformer architectures. By exposing the security risks inherent in transfer learning for remote sensing, ANM underscores the urgent need for robust defense mechanisms to safeguard safety-critical applications. Our findings advocate for enhanced model resilience and motivate future research into countering sophisticated adversarial threats.
Xingjian Tian, Yonghao Xu, Bihan Wen
IEEE Trans. Geosci. Remote. Sens.4
2025 Uncertainty-Aware With Adaptive Geometric Correction for Multimodal Land-Cover Classification
abstract
Land cover classification (LCC) is a fundamental task in remote sensing and geographic information science. Multi-modal fusion has shown great potential for enhancing LCC performance, for example, by combining optical and synthetic aperture radar (SAR) imagery to leverage their complementary strengths. However, two key challenges hinder effective fusion:1) local geometric mismatches caused by distinct imaging geometries, and2) inconsistent reliability (the ability of a modality to deliver accurate and stable information) in LCC arising from different modalities and their acquisition conditions. To address these issues, we propose Uncertainty-Aware Fusion with Adaptive Geometric Correction (UAG), which comprises three main components. First, the Adaptive Geometric Correction Module (AGCM) applies learnable pixel shifts to establish bidirectional local correlations between multiscale optical and SAR features, thereby mitigating spatial inconsistencies. Second, the Adaptive Uncertainty-Aware Dynamic Fusion Module (ADFM) employs evidential deep learning to model uncertainty, defined as the extent of reliability deficiency, for each modality using the Dirichlet distribution and subjective logic, enabling confidence-aware feature weighting. Third, a lightweight multiscale decoder integrates hierarchical features through a hybrid MLP-convolutional architecture, improving both segmentation efficiency and accuracy. We evaluate UAG on WHU-OPT-SAR and DFC23 datasets, where experimental results demonstrate substantial improvements over state-of-the-art methods. The code will be released at https://github.com/cccwbin/UAGNet.
Xu Wang 0015, Yi Xiao 0003, Wenxin Huang, Bihan Wen, Xian Zhong
IEEE Trans. Geosci. Remote. Sens.5
2025 Triply Laplacian Scale Mixture Modeling for Seismic Data Noise Suppression
Sirui Pan, Zhiyuan Zha, Shigang Wang 0003, Yue Li 0003, Zipei Fan, Bihan Wen, Ce Zhu
IEEE Trans. Geosci. Remote. Sens.8
2025 Texture-Consistent 3D Scene Style Transfer via Transformer-Guided Neural Radiance Fields
abstract
Recent advancements have suggested that neural radiance fields (NeRFs) show great potential in 3D style transfer. However, most existing NeRF-based style transfer methods still face considerable challenges in generating stylized images that simultaneously preserve clear scene textures and maintain strong cross-view consistency. To address these limitations, in this paper, we propose a novel transformer-guided approach for 3D scene style transfer. Specifically, we first design a transformer-based style transfer network to capture long-range dependencies and generate 2D stylized images with initial consistency, which serve as supervision for the 3D stylized generation. To enable fine-grained control over style, we propose a latent style vector as a conditional feature and design a style network that projects this style information into the 3D space. We further develop a merge network that integrates style features with scene geometry to render 3D stylized images that are both visually coherent and stylistically consistent. In addition, we propose a texture consistency loss to preserve scene structure and enhance texture fidelity across views. Extensive quantitative and qualitative experimental results demonstrate that our proposed approach outperforms many state-of-the-art methods in terms of visual perception, image quality and multi-view consistency. Our code and more results are available at: https://github.com/PaiDii/TGTC-Style.git.
Wudi Chen, Zhiyuan Zha, Shigang Wang 0003, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu
IEEE Trans. Image Process.5
2025 Digital Staining With Knowledge Distillation: A Unified Framework for Unpaired and Paired-but-Misaligned Data
abstract
Staining is essential in cell imaging and medical diagnostics but poses significant challenges, including high cost, time consumption, labor intensity, and irreversible tissue alterations. Recent advances in deep learning have enabled digital staining through supervised model training. However, collecting large-scale, perfectly aligned pairs of stained and unstained images remains difficult. In this work, we propose a novel unsupervised deep learning framework for digital cell staining that reduces the need for extensive paired data using knowledge distillation. We explore two training schemes: (1) unpaired and (2) paired-but-misaligned settings. For the unpaired case, we introduce a two-stage pipeline, comprising light enhancement followed by colorization, as a teacher model. Subsequently, we obtain a student staining generator through knowledge distillation with hybrid non-reference losses. To leverage the pixel-wise information between adjacent sections, we further extend to the paired-but-misaligned setting, adding the Learning to Align module to utilize pixel-level information. Experiment results on our dataset demonstrate that our proposed unsupervised deep staining method can generate stained images with more accurate positions and shapes of the cell targets in both settings. Compared with competing methods, our method achieves improved results both qualitatively and quantitatively (e.g., NIQE and PSNR). We applied our digital staining method to the White Blood Cell (WBC) dataset, investigating its potential for medical applications.
Ziwang Xu, Lanqing Guo, Satoshi Tsutsui, Alex Chichung Kot, Bihan Wen
IEEE Trans. Medical Imaging6
2025 Learned Focused Plenoptic Image Compression With Local-Global Correlation Learning
abstract
The dense light field sampling of focused plenoptic images (FPIs) yields substantial amounts of redundant data, necessitating efficient compression in practical applications. However, the presence of discontinuous structures and long-distance properties in FPIs poses a challenge. In this paper, we propose a novel end-to-end approach for learned focused plenoptic image compression (LFPIC). Specifically, we introduce a local-global correlation learning strategy to build the nonlinear transforms. This strategy can effectively handle the discontinuous structures and leverage long-distance correlations in FPI for high compression efficiency. Additionally, we propose a spatial-wise context model tailored for LFPIC to help emphasize the most related symbols during coding and further enhance the rate-distortion performance. Experimental results demonstrate the effectiveness of our proposed method, achieving a 22.16% BD-rate reduction (measured in PSNR) on the public dataset compared to the recent state-of-the-art LFPIC method. This improvement holds significant promise for benefiting the applications of focused plenoptic cameras.
Gaosheng Liu, Huanjing Yue, Bihan Wen, Jing-Yu Yang 0002
IEEE Trans. Multim.3
2025 Heterogeneous Prototype Learning From Contaminated Faces Across Domains via Disentangling Latent Factors
abstract
This article studies an emerging practical problem called heterogeneous prototype learning (HPL). Unlike the conventional heterogeneous face synthesis (HFS) problem that focuses on precisely translating a face image from a source domain to another target one without removing facial variations, HPL aims at learning the variation-free prototype of an image in the target domain while preserving the identity characteristics. HPL is a compounded problem involving two cross-coupled subproblems, that is, domain transfer and prototype learning (PL), thus making most of the existing HFS methods that simply transfer the domain style of images unsuitable for HPL. To tackle HPL, we advocate disentangling the prototype and domain factors in their respective latent feature spaces and then replacing the source domain with the target one for generating a new heterogeneous prototype. In doing so, the two subproblems in HPL can be solved jointly in a unified manner. Based on this, we propose a disentangled HPL framework, dubbed DisHPL, which is composed of one encoder-decoder generator and two discriminators. The generator and discriminators play adversarial games such that the generator embeds contaminated images into a prototype feature space only capturing identity information and a domain-specific feature space, while generating realistic-looking heterogeneous prototypes. Experiments on various heterogeneous datasets with diverse variations validate the superiority of DisHPL.
Binghui Wang, Mang Ye, Yiu-Ming Cheung, Yintao Zhou, Wei Huang 0013, Bihan Wen
IEEE Trans. Neural Networks Learn. Syst.7
2025 DenseKD: Dense Knowledge Distillation by Exploiting Region and Sample Importance
abstract
Knowledge distillation (KD) can compress deep neural networks (DNNs) by transferring the knowledge of the redundant teacher model to the resource-friendly student model, where cross-layer KD (CKD) conducts KD between each stage of students and the multiple stages of teachers. However, previous CKD schemes select the coarse-grained stagewise features of teachers to teach students, leading to improper channel alignment. Also, most of these methods conduct uniform distillation for all the knowledge, limiting students to focus more on important knowledge. To address these problems, we propose a dense KD (DenseKD) in this article, dubbed as DenseKD. First, to achieve more accurate feature alignment in CKD, we construct the learnable dense architecture to make each channel of student flexibly capture more diverse channelwise features from teacher. Moreover, we introduce region importance to investigate the region's guiding potential, it distinguishes the influence of different regions by the variation of representations of teacher models. In addition, to make students pay more attention to useful samples in KD, we calculate sample importance by the loss of teacher models. Consistent improvements over state-of-the-art approaches are observed in experiments on multiple vision tasks. For example, in the classification task, DenseKD achieves 72.30% accuracy of ResNet-20 on CIFAR-100, which is higher than the results of previous CKD methods. In addition, in the object detection task, DenseKD gains 2.84% mean average precision (mAP) improvements of Faster R-CNN with ResNet-18 against vanilla KD.
Haonan Zhang 0002, Longjun Liu, Yi Zhang 0140, Fei Hui, Bihan Wen
IEEE Trans. Neural Networks Learn. Syst.6
2024 Progressive Divide-and-Conquer via Subsampling Decomposition for Accelerated MRI
abstract
Deep unfolding networks (DUN) have emerged as a pop-ular iterative framework for accelerated magnetic reso-nance imaging (MRI) reconstruction. However, conventional DUN aims to reconstruct all the missing information within the entire null space in each iteration. Thus it could be challenging when dealing with highly ill-posed degradation, often resulting in subpar reconstruction. In this work, we propose a Progressive Divide-And-Conquer (PDAC) strategy, aiming to break down the subsampling process in the actual severe degradation and thus per-form reconstruction sequentially. Starting from decomposing the original maximum-a-posteriori problem of accel-erated MRI, we present a rigorous derivation of the pro-posed PDAC framework, which could be further unfolded into an end-to-end trainable network. Each PDAC iter-ation specifically targets a distinct segment of moderate degradation, based on the decomposition. Furthermore, as part of the PDAC iteration, such decomposition is adaptively learned as an auxiliary task through a degradation predictor which provides an estimation of the decomposed sampling mask. Following this prediction, the sampling mask is further integrated via a severity conditioning mod-ule to ensure awareness of the degradation severity at each stage. Extensive experiments demonstrate that our pro-posed method achieves superior performance on the pub-licly available fastMRI and Stanford2D FSE datasets in both multi-coil and single-coil settings. Code is available at https://github.com/ChongWang1024/PDAC.
Chong Wang 0011, Lanqing Guo, Yufei Wang 0006, Hao Cheng 0016, Yi Yu 0011, Bihan Wen
CVPR6
2024 SinSR: Diffusion-Based Image Super-Resolution in a Single Step
abstract
While super-resolution (SR) methods based on diffusion models exhibit promising results, their practical application is hindered by the substantial number of required inference steps. Recent methods utilize the degraded images in the initial state, thereby shortening the Markov chain. Nevertheless, these solutions either rely on a precise formulation of the degradation process or still necessitate a relatively lengthy generation path (e.g., 15 iterations). To enhance inference speed, we propose a simple yet effective method for achieving single-step SR generation, named SinSR. Specifically, we first derive a deterministic sampling process from the most recent state-of-the-art (SOTA) method for accelerating diffusion-based SR. This allows the mapping between the input random noise and the generated high-resolution image to be obtained in a reduced and acceptable number of inference steps during training. We show that this deterministic mapping can be distilled into a student model that performs SR within only one inference step. Additionally, we propose a novel consistency-preserving loss to simultaneously leverage the ground-truth image during the distillation process, ensuring that the performance of the student model is not solely bound by the feature manifold of the teacher model, resulting in further performance improvement. Extensive experiments conducted on synthetic and real-world datasets demonstrate that the proposed method can achieve comparable or even superior performance compared to both previous SOTA methods and the teacher model, in just one sampling step, resulting in a remarkable up to × 10 speedup for inference. Our code will be released at https://github.com/wyf0912/SinSR/.
Yufei Wang 0006, Wenhan Yang, Yaohui Wang 0001, Lanqing Guo, Lap-Pui Chau, Ziwei Liu 0002, Yu Qiao 0001, Alex Chichung Kot, Bihan Wen
CVPR10
2024 CaKDP: Category-Aware Knowledge Distillation and Pruning Framework for Lightweight 3D Object Detection
abstract
Knowledge distillation (KD) possesses immense potential to accelerate the deep neural networks (DNNs) for LiDAR-based 3D detection. However, in most of prevailing approaches, the suboptimal teacher models and insufficient student architecture investigations limit the performance gains. To address these issues, we propose a simple yet effective Category-aware Knowledge Distillation and Pruning (CaKDP) framework for compressing 3D detectors. Firstly, CaKDP transfers the knowledge of two-stage detector to one-stage student one, mitigating the impact of inadequate teacher models. To bridge the gap between the heterogeneous detectors, we investigate their differences, and then introduce the student-motivated category-aware KD to align the category prediction between distillation pairs. Secondly, we propose a category-aware pruning scheme to obtain the customizable architecture of compact student model. The method calculates the category prediction gap before and after removing each filter to evaluate the importance of filters, and retains the important filters. Finally, to further improve the student performance, a modified IOU-aware refinement module with negligible computations is leveraged to remove the redundant false positive predictions. Experiments demonstrate that CaKDP achieves the compact detector with high performance. For example, on WOD, CaKDP accelerates CenterPoint by half while boosting L2 mAPH by 1.61%. The code is available at https://github.com/zhnxjtu/CaKDP.
Haonan Zhang 0002, Longjun Liu, Bihan Wen
CVPR6
2024 STSP: Spatial-Temporal Subspace Projection for Video Class-Incremental Learning
Hao Cheng 0016, Siyuan Yang 0001, Chong Wang 0011, Joey Tianyi Zhou, Alex Chichung Kot, Bihan Wen
ECCV (28)6
2024 Temporal As a Plugin: Unsupervised Video Denoising with Pre-trained Image Denoisers
Zixuan Fu, Lanqing Guo, Chong Wang 0011, Yufei Wang 0006, Bihan Wen
ECCV (56)6
2024 Make a Cheap Scaling: A Self-Cascade Diffusion Model for Higher-Resolution Adaptation
Lanqing Guo, Yingqing He, Haoxin Chen, Menghan Xia, Xiaodong Cun, Yufei Wang 0006, Siyu Huang, Yong Zhang 0034, Xintao Wang 0002, Qifeng Chen 0001, Ying Shan, Bihan Wen
ECCV (36)12
2024 🤖 SegPoint: Segment Any Point Cloud via Large Language Model
Shuting He, Henghui Ding, Xudong Jiang 0001, Bihan Wen
ECCV (22)4
2024 Joint RGB-Spectral Decomposition Model Guided Image Enhancement in Mobile Photography
Kailai Zhou, Lijing Cai, Yibo Wang 0004, Bihan Wen, Qiu Shen, Xun Cao
ECCV (13)5
2024 Video-Text Prompting for Weakly Supervised Spatio-Temporal Video Grounding
abstract
Weakly-supervised Spatio-Temporal Video Grounding(STVG) aims to localize target object tube given a text query, without densely annotated training data.Existing methods extract each candidate tube feature independently by cropping objects from video frame feature, discarding all contextual information such as position change and inter-entity relationship.In this paper, we propose Video-Text Prompting(VTP) to construct candidate feature.Instead of cropping tube region from feature map, we draw visual markers(e.g.red circle) over objects tubes as video prompts; corresponding text prompt(e.g. in red circle) is also inserted after the subject word of query text to highlight its presence.Nevertheless, each candidate feature may look similar without cropping.To address this, we further propose Contrastive VTP(CVTP) by introducing negative contrastive samples whose candidate object is erased instead of being highlighted; by comparing the difference between VTP candidate and the contrastive sample, the gap of matching score between correct candidate and the rest is enlarged.Extensive experiments and ablations are conducted on several STVG datasets and our results surpass existing weakly-supervised methods by a great margin, demonstrating the effectiveness of our proposed methods.
Heng Zhao 0004, Yinjie Zhao, Bihan Wen, Yew-Soon Ong, Joey Tianyi Zhou
EMNLP3
2024 Benchmarking Adversarial Robustness of Image Shadow Removal with Shadow-Adaptive Attacks
abstract
Shadow removal is a task aimed at erasing regional shadows present in images and reinstating visually pleasing natural scenes with consistent illumination. While recent deep learning techniques have demonstrated impressive performance in image shadow removal, their robustness against adversarial attacks remains largely unexplored. Furthermore, many existing attack frameworks typically allocate a uniform budget for perturbations across the entire input image, which may not be suitable for attacking shadow images. This is primarily due to the unique characteristic of spatially varying illumination within shadow images. In this paper, we propose a novel approach, called shadow-adaptive adversarial attack. Different from standard adversarial attacks, our attack budget is adjusted based on the pixel intensity in different regions of shadow images. Consequently, the optimized adversarial noise in the shadowed regions becomes visually less perceptible while permitting a greater tolerance for perturbations in non-shadow regions. The proposed shadow-adaptive attacks naturally align with the varying illumination distribution in shadow images, resulting in perturbations that are less conspicuous. Building on this, we conduct a comprehensive empirical evaluation of existing shadow removal methods, subjecting them to various levels of attack on publicly available datasets.
Chong Wang 0011, Yi Yu 0011, Lanqing Guo, Bihan Wen
ICASSP4
2024 Compress Clean Signal from Noisy Raw Image: A Self-Supervised Approach
abstract
Raw images offer unique advantages in many low-level visual tasks due to their unprocessed nature. However, this unprocessed state accentuates noise, making raw images challenging to compress effectively. Current compression methods often overlook the ubiquitous noise in raw space, leading to increased bitrates and reduced quality. In this paper, we propose a novel raw image compression scheme that selectively compresses the noise-free component of the input, while discarding its real noise using a self-supervised approach. By excluding noise from the bitstream, both the coding efficiency and reconstruction quality are significantly enhanced. We curate an full-day dataset of raw images with calibrated noise parameters and reference images to evaluate the performance of models under a wide range of input signal-noise ratios. Experimental results demonstrate that our method surpasses existing compression techniques, achieving a more advantageous rate-distortion balance with improvements ranging from +2 to +10dB and yielding a bit saving of 2 to 50 times. The code will be released upon paper acceptance.
Yufei Wang 0006, Alex Chichung Kot, Bihan Wen
ICML4
2024 Learning-Based Human Detection via Radar for Dynamic and Cluttered Indoor Environments
abstract
Radar-based human detection draws significant attention in response to growing safety concerns driven by advances in factory automation and smart home technologies. However, much of this research typically operates in controlled environments characterized by minimal clutter and noise, which limits their effectiveness in real-world scenarios such as urban areas, factories, and smart homes. In this study, we address this limitation by collecting a real-world radar dataset in dynamic and cluttered indoor environments, spanning five distinctive environments. To simulate non-human targets, we introduce a moving trolley. Subsequently, we propose a system including simple Radar Signal Processing steps and a learning-based model using unsupervised domain adaptation to enhance its adaptability to unseen environments. Through a series of comprehensive experiments employing popular learning-based methods on our dataset, we demonstrate the model’s efficacy in mitigating environmental interference and successfully adapting to previously unseen environments.
Songnan Lin, Hao Cheng 0016, Weixian Liu, Bihan Wen
ISCAS5
2024 Fusing EO and LiDAR for SAR Image Translation with Multi-Modal Generative Adversarial Networks
abstract
To generate high-quality Synthetic Aperture Radar (SAR) images, the translation of Electro-optical (EO) to SAR images has been studied in recent years. However, these methods cannot make use of other modal information, such as Light Detection and Ranging (LiDAR), to enhance their performance, due to the absence of well-registered multi-modal remote sensing data. Therefore, in this paper, we first construct an organized high-resolution real-world benchmark dataset, which comprises 479 registered multi-modal images. We then explore how to adapt the state-of-the-art method, named the query-selected attention model (QS-Attn), to make it more suitable for integrating EO and LiDAR images into real SAR image translation. Extensive experiments have been conducted to evaluate our proposed method and compare it with state-of-the-art methods. The experimental results validate the superiority and effectiveness of our approach.
Yuanyuan Qing, Zhiping Lin 0001, Bihan Wen
ISCAS4
2024 Spectral Convergence of Simplicial Complex Signals
abstract
Topological signal processing (TSP) utilizes simplicial complexes to model structures with higher order than vertices and edges. In this paper, we study the transferability of TSP via a generalized higher-order version of graphon, known as complexon. We recall the notion of a complexon as the limit of a simplicial complex sequence [1]. Inspired by the graphon shift operator and message-passing neural network, we construct a marginal complexon and complexon shift operator (CSO) according to components of all possible dimensions from the complexon. We investigate the CSO's eigenvalues and eigenvectors and relate them to a new family of weighted adjacency matrices. We prove that when a simplicial complex signal sequence converges to a complexon signal, the eigenvalues, eigenspaces, and Fourier transform of the corresponding CSOs converge to that of the limit complexon signal. This conclusion is further verified by two numerical experiments. These results hint at learning transferability on large simplicial complexes or simplicial complex sequences, which generalize the graphon signal processing framework.
Purui Zhang 0001, Xingchao Jian, Wee-Peng Tay, Bihan Wen
ISIT5
2024 Integrating Clinical Knowledge into Concept Bottleneck Models
Winnie Pang, Xueyi Ke, Satoshi Tsutsui, Bihan Wen
MICCAI (4)4
2024 Dual-head Genre-instance Transformer Network for Arbitrary Style Transfer
abstract
Arbitrary style transfer aims to render artistic features from a style reference onto an image while retaining its original content. Previous methods either focus on learning the holistic style from a specific artist or extracting instance features from a single artwork. However, they often fail to apply style elements uniformly across the entire image and lack adaptation to the style of different artworks. To solve these issues, our key insight is that the art genre has better generality and adaptability than the overall features of the artist. To this end, we propose a Dual-head Genre-instance Transformer (DGiT) framework to simultaneously capture the genre and instance features for arbitrary style transfer. To the best of our knowledge, this is the first work to integrate the genre features and instance features to generate a high-quality stylized image. Moreover, we design two contrastive losses to enhance the capability of the network to capture two style features. Our approach ensures the uniform distribution of the overall style across the stylized image while enhancing the details of textures and strokes in local regions. Qualitative and quantitative evaluations demonstrate that our approach exhibits superior visual quality and efficiency.
Meichen Liu, Shuting He, Songnan Lin, Bihan Wen
ACM Multimedia4
2024 Evolving Storytelling: Benchmarks and Methods for New Character Customization with Diffusion Models
abstract
Diffusion-based models for story visualization have shown promise in generating content-coherent images for storytelling tasks. However, how to effectively integrate new characters into existing narratives while maintaining character consistency remains an open problem, particularly with limited data. Two major limitations hinder the progress: (1) the absence of a suitable benchmark due to potential character leakage and inconsistent text labeling, and (2) the challenge of distinguishing between new and old characters, leading to ambiguous results. To address these challenges, we introduce the NewEpisode benchmark, comprising refined datasets designed to evaluate generative models' adaptability in generating new stories with fresh characters using just a single example story. The refined dataset involves refined text prompts and eliminates character leakage. Additionally, to mitigate the character confusion of generated results, we propose EpicEvo, a method that customizes a diffusion-based visual story generation model with a single story featuring the new characters seamlessly integrating them into established character dynamics. EpicEvo introduces a novel adversarial character alignment module to align the generated images progressively in the diffusive process, with exemplar images of new characters, while applying knowledge distillation to prevent forgetting of characters and background details. Our evaluation quantitatively demonstrates that EpicEvo outperforms existing baselines on the NewEpisode benchmark, and qualitative studies confirm its superior customization of visual story generation in diffusion models. In summary, EpicEvo provides an effective way to incorporate new characters using only one example story, unlocking new possibilities for applications such as serialized cartoons.
Yufei Wang 0006, Satoshi Tsutsui, Weisi Lin, Bihan Wen, Alex Chichung Kot
ACM Multimedia5
2024 From Chaos to Clarity: 3DGS in the Dark
abstract
Novel view synthesis from raw images provides superior high dynamic range (HDR) information compared to reconstructions from low dynamic range RGB images. However, the inherent noise in unprocessed raw images compromises the accuracy of 3D scene representation. Our study reveals that 3D Gaussian Splatting (3DGS) is particularly susceptible to this noise, leading to numerous elongated Gaussian shapes that overfit the noise, thereby significantly degrading reconstruction quality and reducing inference speed, especially in scenarios with limited views. To address these issues, we introduce a novel self-supervised learning framework designed to reconstruct HDR 3DGS from a limited number of noisy raw images. This framework enhances 3DGS by integrating a noise extractor and employing a noise-robust reconstruction loss that leverages a noise distribution prior. Experimental results show that our method outperforms LDR/HDR 3DGS and previous state-of-the-art (SOTA) self-supervised and supervised pre-trained models in both reconstruction quality and inference speed on the RawNeRF dataset across a broad range of training views. We will release the code upon paper acceptance.
Yufei Wang 0006, Alex Chichung Kot, Bihan Wen
NeurIPS4
2024 ContextGS : Compact 3D Gaussian Splatting with Anchor Level Context Model
abstract
Recently, 3D Gaussian Splatting (3DGS) has become a promising framework for novel view synthesis, offering fast rendering speeds and high fidelity. However, the large number of Gaussians and their associated attributes require effective compression techniques. Existing methods primarily compress neural Gaussians individually and independently, i.e., coding all the neural Gaussians at the same time, with little design for their interactions and spatial dependence. Inspired by the effectiveness of the context model in image compression, we propose the first autoregressive model at the anchor level for 3DGS compression in this work. We divide anchors into different levels and the anchors that are not coded yet can be predicted based on the already coded ones in all the coarser levels, leading to more accurate modeling and higher coding efficiency. To further improve the efficiency of entropy coding, e.g., to code the coarsest level with no already coded anchors, we propose to introduce a low-dimensional quantized feature as the hyperprior for each anchor, which can be effectively compressed. Our work pioneers the context model in the anchor level for 3DGS representation, yielding an impressive size reduction of over 100 times compared to vanilla 3DGS and 15 times compared to the most recent state-of-the-art work Scaffold-GS, while achieving comparable or even higher rendering quality.
Yufei Wang 0006, Lanqing Guo, Wenhan Yang, Alex Chichung Kot, Bihan Wen
NeurIPS6
2024 A joint learning method with consistency-aware for low-resolution facial expression recognition
Yuanlun Xie, Wenhong Tian, Ruini Xue, Zhiyuan Zha, Bihan Wen
Expert Syst. Appl.6
2024 Beyond Learned Metadata-Based Raw Image Reconstruction
Yufei Wang 0006, Yi Yu 0011, Wenhan Yang, Lanqing Guo, Lap-Pui Chau, Alex Chichung Kot, Bihan Wen
Int. J. Comput. Vis.7
2024 Intrinsic-style distribution matching for arbitrary style transfer
Meichen Liu, Songnan Lin, Hengmin Zhang, Zhiyuan Zha, Bihan Wen
Knowl. Based Syst.5
2024 Structured residual sparsity for video compressive sensing reconstruction
Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiachao Zhang, Jiantao Zhou 0001, Ce Zhu
Signal Process.2
2024 Toward Robust Image Denoising via Flow-Based Joint Image and Noise Model
abstract
One of the fundamental challenges in image restoration is denoising, where the objective is to estimate the clean image from its noisy measurements. Existing denoising approaches generally focus on exploiting effective natural image priors to remove the noise. However, the utilization and analysis of the noise model are often ignored, although the noise model can provide complementary information to the denoising algorithms. As a result, they are very sensitive to different noise distributions. To tackle this issue and hence towards a robust image denoiser in practice, in this paper, we propose a novel Flow-based joint Image and NOise model (FINO) that distinctly decouples the image and noise in the latent space and losslessly reconstructs them via a series of invertible transformations. We further present a variable swapping strategy to align structural information in images and a noise correlation matrix to constrain the noise based on spatially minimized correlation information. Experimental results demonstrate FINO’s capacity to remove both synthetic additive white Gaussian noise (AWGN) and real noise. Furthermore, the generalization of FINO to the removal of spatially variant noise and noise with inaccurate estimation surpasses that of the popular and state-of-the-art methods by large margins.
Lanqing Guo, Siyu Huang, Haosen Liu 0001, Bihan Wen
IEEE Trans. Circuits Syst. Video Technol.4
2024 Cross-Image Disentanglement for Low-Light Enhancement in Real World
abstract
Images captured in the low-light condition suffer from low visibility and various imaging artifacts, e.g., real noise. Existing supervised algorithms for low-light image enhancement require a large set of pixel-aligned training image pairs, which are hard to prepare in practice. Though some recent unsupervised methods can alleviate such data challenges, many real world artifacts inevitably get falsely amplified in the enhanced results due to the lack of corresponding supervision. In this paper, instead of using perfectly aligned images for training, we creatively employ the misaligned real world images as the guidance, which are considerably easier to collect. Specifically, we propose a Cross-Image Disentanglement Network (CIDN) with weakly supervised learning, to separately extract cross-image brightness and image-specific content features from low/normal-light images. Based on that, CIDN can simultaneously correct the brightness and suppress image artifacts in the feature domain, which largely increases the robustness of the pixel shifts between training pairs. By considering real world corruptions, we propose a new training dataset with misaligned and noisy image pairs and its corresponding evaluation dataset. Experimental results show that our model achieves state-of-the-art performances on both the newly proposed dataset and other popular low-light datasets. The code implementation is publicly available at:https://github.com/GuoLanqing/CIDN.
Lanqing Guo, Renjie Wan, Wenhan Yang, Alex Chichung Kot, Bihan Wen
IEEE Trans. Circuits Syst. Video Technol.5
2024 Salient Object Detection Toward Single-Pixel Imaging
abstract
Replacing CCD and CMOS image sensors in conventional cameras with digital micromirror devices (DMD), single-pixel cameras low-costly shot images by capturing compressed measurements and computation. However, the compressed measurements lack explicit spatial information, causing difficulties for high-level tasks such as salient object detection (SOD) that are usually designed to have visual inputs. To address the issue, we propose a single-pixel imaging-based SOD network called SPISODNet that enables predicting saliency maps directly from compressed measurements with high accuracy. Specifically, we first design an underlying feature inversion module (UFIM) to capture the underlying scene information, and then develop a context-aware flow (CAF) consisting of a feature focus module (FFM), three bidirectional attention modules (BAMs), and a spatial information-induced attention module (SIAM) to acquire and polish saliency predictions. Extensive experiments demonstrate that our method achieves superior performance for single-pixel imaging-based SOD.
HuiHui Yue, Jichang Guo, Xiangjun Yin, Yi Zhang 0107, Bihan Wen, Chongyi Li
IEEE Trans. Circuits Syst. Video Technol.5
2024 Accelerated PALM for Nonconvex Low-Rank Matrix Recovery With Theoretical Analysis
abstract
Low-rank matrix recovery is a major challenge in machine learning and computer vision, particularly for large-scale data matrices, as popular methods involving nuclear norm and singular value decomposition (SVD) are associated with high computational costs and biased estimators. To overcome this challenge, we propose a novel approach to learning low-rank matrices based on the matrix volume and a nonconvex logarithmic function. The matrix volume is the product of all the nonzero singular values of a matrix and has unique geometric properties and connections with other convex and nonconvex functions. We establish a generalized nonconvex regularization problem using the penalty function strategy and introduce an accelerated proximal alternating linearized minimization (AccPALM) algorithm with double acceleration, which combines Nesterov’s acceleration and power strategy. The algorithm reduces computational costs and has provable convergence results under the Kurdyka-Łojasiewicz (KŁ) inequality with mild conditions. Our approach shows superior accuracy, efficiency, and convergence behavior compared to other low-rank matrix learning methods on robust matrix completion (RMC) and low-rank representation (LRR) tasks. We analyze the impact of algorithm parameters on convergence and performance and present visually appealing results to further demonstrate the effectiveness of our approach. The proposed methodology represents a promising advance in the field of low-rank matrix recovery, and its effectiveness has been validated via extensive numerical experiments. The source code for the proposed algorithms is accessible at https://github.com/ZhangHengMin/AccPALMcodes.
Hengmin Zhang, Bihan Wen, Zhiyuan Zha, Bob Zhang 0001, Yang Tang 0001, Guo Yu 0001, Wenli Du
IEEE Trans. Circuits Syst. Video Technol.2
2024 Multiple Complementary Priors for Multispectral Image Compressive Sensing Reconstruction
abstract
Compressive sensing (CS) techniques using a few compressed measurements have drawn considerable interest in reconstructing multispectral imagery (MSI). Nonlocal-based tensor methods have been widely used for MSI-CS reconstruction, which employ the nonlocal self-similarity (NSS) property of MSI to obtain satisfactory results. However, such methods only consider the internal priors of MSI while ignoring important external image information, for example deep-driven priors learned from a corpus of natural image datasets. Meanwhile, they usually suffer from annoying ringing artifacts due to the aggregation of overlapping patches. In this article, we propose a novel approach for highly effective MSI-CS reconstruction using multiple complementary priors (MCPs). The proposed MCP jointly exploits nonlocal low-rank and deep image priors under a hybrid plug-and-play framework, which contains multiple pairs of complementary priors, namely, internal and external, shallow and deep, and NSS and local spatial priors. To make the optimization tractable, a well-known alternating direction method of multiplier (ADMM) algorithm based on the alternating minimization framework is developed to solve the proposed MCP-based MSI-CS reconstruction problem. Extensive experimental results demonstrate that the proposed MCP algorithm outperforms many state-of-the-art CS techniques in MSI reconstruction. The source code of the proposed MCP-based MSI-CS reconstruction algorithm is available at: https://github.com/zhazhiyuan/MCP_MSI_CS_Demo.git.
Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiachao Zhang, Jiantao Zhou 0001, Xudong Jiang 0001, Ce Zhu
IEEE Trans. Cybern.2
2024 Stealthy Adversarial Examples for Semantic Segmentation in Remote Sensing
abstract
Deep learning methods have been proven effective in remote sensing image analysis and interpretation, where semantic segmentation plays a vital role. These deep segmentation methods are susceptible to adversarial attacks, while most of the existing attack methods tend to manipulate the image globally, leading to noticeable perturbations and chaotic segmentation. In this work, we propose a novel Stealthy Attack for Semantic Segmentation (SASS), which can largely increase the effectiveness and stealthiness from the existing attack methods on remote sensing images. SASS manipulates specific victim classes or objects of interest while preserving the original segmentation results for other classes or objects. In practice, as different inference mechanisms, overlapped inference, can be applied in segmentation, the efficacy of SASS may be degraded. To this end, we further introduce the Masked Stealthy Attack for Semantic Segmentation (MSASS), which generates augmented adversarial perturbations that only affect victim areas. We evaluate the effectiveness of SASS and MSASS using four state-of-the-art semantic segmentation models on the Vaihingen and Zurich Summer datasets. Extensive experiments demonstrate that our SASS and MSASS methods achieve superior attack performances on victim areas while maintaining high accuracies of other areas (drop less than 2%). The detection success rates of adversarial examples for segmentation, as characterized by Xiaoet al. [1], significantly drop from 97.78% for the untargeted PGD attack to 28.71% for our MSASS method on Zurich Summer dataset. Our work contributes to the field of adversarial attacks in semantic segmentation for remote sensing images by improving stealthiness, flexibility, and robustness. We anticipate that our findings will inspire the development of defense methods to enhance the security and reliability of semantic segmentation models against our stealthy attack.
Yonghao Xu, Bihan Wen
IEEE Trans. Geosci. Remote. Sens.4
2024 Detection of Adversarial Attacks via Disentangling Natural Images and Perturbations
abstract
The vulnerability of deep neural networks against adversarial attacks,i.e., imperceptible adversarial perturbations can easily give rise to wrong predictions, poses a huge threat to the security of their real-world deployments. In this paper, a novel Adversarial Detection method via Disentangling Natural images and Perturbations (ADDNP) is proposed. Compared to natural images that can typically be modeled by lower-dimensional subspaces or manifolds, the distributions of adversarial perturbations are much more complex,e.g., one normal example’s adversarial counterparts generated by different attack strategies can be significantly distinct. The proposed ADDNP exploits such distinct properties for the detection of adversarial attacks amongst normal examples. Specifically, we use a dual-branch disentangling framework to encode natural images and perturbations of inputs separately, followed by joint reconstruction. During inference, the reconstruction discrepancy (RD) measured in the learned latent feature space is used as an indicator of adversarial perturbations. The proposed ADDNP algorithm is evaluated on three popular datasets,i.e., CIFAR-10, CIFAR-100, andminiImageNet with increasing data complexity, across multiple popular attack strategies. Compared to the existing and state-of-the-art detection methods, ADDNP has demonstrated promising performance on adversarial detection, with significant improvements on more challenging datasets.
Yuanyuan Qing, Zhuotao Liu, Pierre Moulin, Bihan Wen
IEEE Trans. Inf. Forensics Secur.5
2024 Efficient Image Classification via Structured Low-Rank Matrix Factorization Regression
abstract
In real-world applications involving sparse coding and low-rank matrix recovery problems, linear regression methods usually struggle to effectively capture the structured correlations present in data matrices. This limitation arises from representation approaches that treat images as vectors and handle testing samples individually, overlooking these correlations. To address these challenges, we propose a novel approach that leverages the low-rank property to capture the global and intrinsic structure of residual and coefficient matrices, departing from the assumption of independent and identically distributed (I.I.D) data. Our method introduces nonconvex and nonsmooth low-rank matrix regression models guided by the extended matrix variate power exponential distribution (M.P.E.D). By incorporating factorization strategies into the regression coefficient matrix and utilizing the Schatten-$p$norm with three distinct values of$p$, we enhance computational efficiency. Our formulation enables efficient subproblem solving through the introduction of auxiliary variables and the use of singular value threshold operators. We achieve closed-form solutions using the proposed multi-variable alternating direction method of multipliers (ADMM). Theoretical analysis establishes the local convergence properties and computational complexity of our optimization algorithm. Furthermore, we conduct numerical experiments on various image datasets, including face, object, and digital, to demonstrate the superior performance and computational efficiency of our methods compared to several related regression approaches. The source codes for our method are available athttps://github.com/ZhangHengMin/TIFS_SLRMFR.
Hengmin Zhang, Jian Yang 0003, Jianjun Qian, Guangwei Gao, Xiangyuan Lan, Zhiyuan Zha, Bihan Wen
IEEE Trans. Inf. Forensics Secur.7
2024 SOUL-Net: A Sparse and Low-Rank Unrolling Network for Spectral CT Image Reconstruction
abstract
Spectral computed tomography (CT) is an emerging technology, that generates a multienergy attenuation map for the interior of an object and extends the traditional image volume into a 4-D form. Compared with traditional CT based on energy-integrating detectors, spectral CT can make full use of spectral information, resulting in high resolution and providing accurate material quantification. Numerous model-based iterative reconstruction methods have been proposed for spectral CT reconstruction. However, these methods usually suffer from difficulties such as laborious parameter selection and expensive computational costs. In addition, due to the image similarity of different energy bins, spectral CT usually implies a strong low-rank prior, which has been widely adopted in current iterative reconstruction models. Singular value thresholding (SVT) is an effective algorithm to solve the low-rank constrained model. However, the SVT method requires a manual selection of thresholds, which may lead to suboptimal results. To relieve these problems, in this article, we propose a sparse and low-rank unrolling network (SOUL-Net) for spectral CT image reconstruction, that learns the parameters and thresholds in a data-driven manner. Furthermore, a Taylor expansion-based neural network backpropagation method is introduced to improve the numerical stability. The qualitative and quantitative results demonstrate that the proposed method outperforms several representative state-of-the-art algorithms in terms of detail preservation and artifact reduction.
Xiang Chen 0015, Wenjun Xia, Ziyuan Yang 0001, Hu Chen 0002, Yan Liu 0052, Jiliu Zhou, Yang Chen 0008, Bihan Wen, Yi Zhang 0018
IEEE Trans. Neural Networks Learn. Syst.9
2024 Disentangled Feature Representation for Few-Shot Image Classification
abstract
Learning the generalizable feature representation is critical to few-shot image classification. While recent works exploited task-specific feature embedding using meta-tasks for few-shot learning, they are limited in many challenging tasks as being distracted by the excursive features such as the background, domain, and style of the image samples. In this work, we propose a novel disentangled feature representation (DFR) framework, dubbed DFR, for few-shot learning applications. DFR can adaptively decouple the discriminative features that are modeled by the classification branch, from the class-irrelevant component of the variation branch. In general, most of the popular deep few-shot learning methods can be plugged in as the classification branch, thus DFR can boost their performance on various few-shot tasks. Furthermore, we propose a novel FS-DomainNet dataset based on DomainNet, for benchmarking the few-shot domain generalization (DG) tasks. We conducted extensive experiments to evaluate the proposed DFR on general, fine-grained, and cross-domain few-shot classification, as well as few-shot DG, using the corresponding four benchmarks, i.e., mini-ImageNet, tiered-ImageNet, Caltech-UCSD Birds 200-2011 (CUB), and the proposed FS-DomainNet. Thanks to the effective feature disentangling, the DFR-based few-shot classifiers achieved state-of-the-art results on all datasets.
Hao Cheng 0016, Yufei Wang 0006, Haoliang Li, Alex Chichung Kot, Bihan Wen
IEEE Trans. Neural Networks Learn. Syst.5
2024 Temporal Output Discrepancy for Loss Estimation-Based Active Learning
abstract
While deep learning succeeds in a wide range of tasks, it highly depends on the massive collection of annotated data which is expensive and time-consuming. To lower the cost of data annotation, active learning has been proposed to interactively query an oracle to annotate a small proportion of informative samples in an unlabeled dataset. Inspired by the fact that the samples with higher loss are usually more informative to the model than the samples with lower loss, in this article we present a novel deep active learning approach that queries the oracle for data annotation when the unlabeled sample is believed to incorporate high loss. The core of our approach is a measurement temporal output discrepancy (TOD) that estimates the sample loss by evaluating the discrepancy of outputs given by models at different optimization steps. Our theoretical investigation shows that TOD lower-bounds the accumulated sample loss thus it can be used to select informative unlabeled samples. On basis of TOD, we further develop an effective unlabeled data sampling strategy as well as an unsupervised learning criterion for active learning. Due to the simplicity of TOD, our methods are efficient, flexible, and task-agnostic. Extensive experimental results demonstrate that our approach achieves superior performances than the state-of-the-art active learning methods on image classification and semantic segmentation tasks. In addition, we show that TOD can be utilized to select the best model of potentially the highest testing accuracy from a pool of candidate models.
Siyu Huang, Tianyang Wang 0004, Haoyi Xiong, Bihan Wen, Jun Huan, Dejing Dou
IEEE Trans. Neural Networks Learn. Syst.4
2023 ShadowFormer: Global Context Helps Shadow Removal
abstract
Recent deep learning methods have achieved promising results in image shadow removal. However, most of the existing approaches focus on working locally within shadow and non-shadow regions, resulting in severe artifacts around the shadow boundaries as well as inconsistent illumination between shadow and non-shadow regions. It is still challenging for the deep shadow removal model to exploit the global contextual correlation between shadow and non-shadow regions. In this work, we first propose a Retinex-based shadow model, from which we derive a novel transformer-based network, dubbed ShandowFormer, to exploit non-shadow regions to help shadow region restoration. A multi-scale channel attention framework is employed to hierarchically capture the global information. Based on that, we propose a Shadow-Interaction Module (SIM) with Shadow-Interaction Attention (SIA) in the bottleneck stage to effectively model the context correlation between shadow and non-shadow regions. We conduct extensive experiments on three popular public datasets, including ISTD, ISTD+, and SRD, to evaluate the proposed method. Our method achieves state-of-the-art performance by using up to 150X fewer model parameters.
Lanqing Guo, Siyu Huang, Ding Liu 0001, Hao Cheng 0016, Bihan Wen
AAAI5
2023 sRGB Real Noise Synthesizing with Neighboring Correlation-Aware Noise Model
abstract
Modeling and synthesizing real noise in the standard RGB (sRGB) domain is challenging due to the complicated noise distribution. While most of the deep noise generators proposed to synthesize sRGB real noise using an end-to-end trained model, the lack of explicit noise modeling degrades the quality of their synthesized noise. In this work, we propose to model the real noise as not only dependent on the underlying clean image pixel intensity, but also highly correlated to its neighboring noise realization within the local region. Correspondingly, we propose a novel noise synthesizing framework by explicitly learning its neighboring correlation on top of the signal dependency. With the proposed noise model, our framework greatly bridges the distribution gap between synthetic noise and real noise. We show that our generated “real” sRGB noisy images can be used for training supervised deep denoisers, thus to improve their real denoising results with a large margin, comparing to the popular classic denoisers or the deep denoisers that are trained on other sRGB noise generators. The code will be available at https://github.com/xuan611/sRGB-Real-Noise-Synthesizing.
Zixuan Fu, Lanqing Guo, Bihan Wen
CVPR3
2023 ShadowDiffusion: When Degradation Prior Meets Diffusion Model for Shadow Removal
abstract
Recent deep learning methods have achieved promising results in image shadow removal. However, their restored images still suffer from unsatisfactory boundary artifacts, due to the lack of degradation prior embedding and the deficiency in modeling capacity. Our work addresses these issues by proposing a unified diffusion framework that integrates both the image and degradation priors for highly effective shadow removal. In detail, we first propose a shadow degradation model, which inspires us to build a novel unrolling diffusion model, dubbed ShandowDiffusion. It remarkably improves the model's capacity in shadow removal via progressively refining the desired output with both degradation prior and diffusive generative prior, which by nature can serve as a new strong baseline for image restoration. Furthermore, ShadowDiffusion progressively refines the estimated shadow mask as an auxiliary task of the diffusion generator, which leads to more accurate and robust shadow-free image generation. We conduct extensive experiments on three popular public datasets, including ISTD, ISTD+, and SRD, to validate our method's effectiveness. Compared to the state-of-the-art methods, our model achieves a significant improvement in terms of PSNR, increasing from 31.69dB to 34. 73dB over SRD dataset.11https://github.com/GuoLanqing/ShadowDiffusion
Lanqing Guo, Chong Wang 0011, Wenhan Yang, Siyu Huang, Yufei Wang 0006, Hanspeter Pfister, Bihan Wen
CVPR7
2023 Raw Image Reconstruction with Learned Compact Metadata
abstract
While raw images exhibit advantages over sRGB images (e.g., linearity and fine-grained quantization level), they are not widely used by common users due to the large storage requirements. Very recent works propose to compress raw images by designing the sampling masks in the raw image pixel space, leading to suboptimal image representations and redundant metadata. In this paper, we propose a novel framework to learn a compact representation in the latent space serving as the metadata in an end-to-end manner. Furthermore, we propose a novel sRGB-guided context model with the improved entropy estimation strategies, which leads to better reconstruction quality, smaller size of metadata, and faster speed. We illustrate how the proposed raw image compression scheme can adaptively allocate more bits to image regions that are important from a global perspective. The experimental results show that the proposed method can achieve superior raw image reconstruction results using a smaller size of the metadata on both uncompressed sRGB images and JPEG images. The code will be released at https://github.com/wyf0912/R2LCM
Yufei Wang 0006, Yi Yu 0011, Wenhan Yang, Lanqing Guo, Lap-Pui Chau, Alex Chichung Kot, Bihan Wen
CVPR7
2023 Towards Adversarially Robust Continual Learning
abstract
Recent studies show that models trained by continual learning can achieve the comparable performances as the standard supervised learning and the learning flexibility of continual learning models enables their wide applications in the real world. Deep learning models, however, are shown to be vulnerable to adversarial attacks. Though there are many studies on the model robustness in the context of standard supervised learning, protecting continual learning from adversarial attacks has not yet been investigated. To fill in this research gap, we are the first to study adversarial robustness in continual learning and propose a novel method called Task-Aware Boundary Augmentation (TABA) to boost the robustness of continual learning models. With extensive experiments on CIFAR-10 and CIFAR-100, we show the efficacy of adversarial training and TABA in defending adversarial attacks.
Chen Chen 0043, Lingjuan Lyu, Jun Zhao 0007, Bihan Wen
ICASSP5
2023 Benchmarking White Blood Cell Classification under Domain Shift
abstract
Recognizing the types of white blood cells (WBCs) in microscopic images of human blood smears is a fundamental task in the fields of pathology and hematology. Although previous studies have made significant contributions to the development of methods and datasets, few papers have investigated benchmarks or baselines that others can easily refer to. For instance, we observed notable variations in the reported accuracies of the same Convolutional Neural Network (CNN) model across different studies, yet no public implementation exists to reproduce these results. In this paper, we establish a benchmark for WBC recognition. Our results indicate that CNN-based models achieve high accuracy when trained and tested under similar imaging conditions. However, their performance drops significantly when tested under different conditions. Moreover, the ResNet classifier, which has been widely employed in previous work, exhibits an unreasonably poor generalization ability under domain shifts due to batch normalization. We investigate this issue and suggest some alternative normalization techniques that can mitigate it. We make fully-reproducible code publicly available1.
Satoshi Tsutsui, Bihan Wen
ICASSP3
2023 Unsupervised Deep Digital Staining for Microscopic Cell Images via Knowledge Distillation
abstract
Staining is critical to cell imaging and medical diagnosis, which is expensive, time-consuming, labor-intensive, and causes irreversible changes to cell tissues. Recent advances in deep learning enabled digital staining via supervised model training. However, it is difficult to obtain large-scale stained/unstained cell image pairs in practice, which need to be perfectly aligned with the supervision. In this work, we propose a novel unsupervised deep learning framework for the digital staining of cell images using knowledge distillation and generative adversarial networks (GANs). A teacher model is first trained mainly for the colorization of bright-field images. After that, a student GAN for staining is obtained by knowledge distillation with hybrid non-reference losses. We show that the proposed unsupervised deep staining method can generate stained images with more accurate positions and shapes of the cell targets. Compared with other unsupervised deep generative models for staining, our method achieves much more promising results both qualitatively and quantitatively.
Ziwang Xu, Lanqing Guo, Alex Chichung Kot, Bihan Wen
ICASSP5
2023 Hyperspectral Image Denoising Via Nonlocal Rank Residual Modeling
abstract
Nonlocal low-rank (LR) tensor modeling has shown great potential in hyperspectral image (HSI) denoising, which first uses the nonlocal self-similarity (NSS) prior to search for many similar full-band patches to form three-dimensional nonlocal full-band groups (tensors), and then usually enforces an LR penalty on each nonlocal full-band group. However, in most existing methods, the LR tensor is only approximated directly from the degraded nonlocal full-band tensor, which is subject to certain issues (e.g., in heavy noise environments) in obtaining a suboptimal tensor approximation, and thus leading to unsatisfactory denoising results. In this paper, we propose a novel nonlocal rank residual (NRR) approach for highly effective HSI denoising, which progressively approximates the underlying L-R tensor via minimizing the rank residual. Towards this end, we first obtain a good estimate of the original nonlocal full-band group by using the NSS prior, and then the rank residual between the de-graded nonlocal full-band group with the corresponding estimated nonlocal full-band group is minimized to achieve a more accurate LR tensor. Moreover, the global spectral LR prior is employed to reduce the spectral redundancy of HSI in the proposed denoising framework. Finally, we develop a simple yet effective alternating minimization algorithm to jointly refine global spectral information and nonlocal full-band groups. Experimental results clearly show that the proposed NRR algorithm outperforms many state-of-the-art HSI denoising methods. The source code of the proposed NRR algorithm for HSI denoising is available at: https://github.com/zhazhiyuan/NRR_HSI_Denoising_Demo.git.
Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu
ICASSP2
2023 Frequency Guidance Matters in Few-Shot Learning
abstract
Few-shot classification aims to learn a discriminative feature representation to recognize unseen classes with few labeled support samples. While most few-shot learning methods focus on exploiting the spatial information of image samples, frequency representation has also been proven essential in classification tasks. In this paper, we investigate the effect of different frequency components on the few-shot learning tasks. To enhance the performance and generalizability of few-shot methods, we propose a novel Frequency-Guided Few-shot Learning framework (dubbed FGFL), which leverages the task-specific frequency components to adaptively mask the corresponding image information, with a novel multi-level metric learning strategy including a triplet loss among original, masked and unmasked image as well as a contrastive loss between masked and original support and query sets to exploit more discriminative information. Extensive experiments on four benchmarks under several few-shot scenarios, i.e., standard, cross-dataset, cross-domain, and coarse-to-fine annotated classification, are conducted. Both qualitative and quantitative results show that our proposed FGFL scheme can attend to the class-discriminative frequency components, thus integrating those information towards more effective and generalizable few-shot learning.
Hao Cheng 0016, Siyuan Yang 0001, Joey Tianyi Zhou, Lanqing Guo, Bihan Wen
ICCV5
2023 Boundary-Aware Divide and Conquer: A Diffusion-based Solution for Unsupervised Shadow Removal
abstract
Recent deep learning methods have achieved superior results in shadow removal. However, most of these supervised methods rely on training over a huge amount of shadow and shadow-free image pairs, which require laborious annotations and may end up with poor model generalization. Shadows, in fact, only form partial degradation in images, while their non-shadow regions provide rich structural information potentially for unsupervised learning. In this paper, we propose a novel diffusion-based solution for unsupervised shadow removal, which separately modeling the shadow, non-shadow, and their boundary regions. We employ a pretrained unconditional diffusion model fused with non-corrupted information to generate the natural shadow-free image. While the diffusion model can restore the clear structure in the boundary region by utilizing its adjacent non-corrupted contextual information, it fails to address the inner shadow area due to the isolation of the non-corrupted contexts. Thus we further propose a Shadow-Invariant Intrinsic Decomposition module to exploit the underlying reflectance in the shadow region to maintain structural consistency during the diffusive sampling. Extensive experiments on the publicly available shadow removal datasets show that the proposed method achieves a significant improvement compared to existing unsupervised methods, and even is comparable with some existing supervised methods.
Lanqing Guo, Chong Wang 0011, Wenhan Yang, Yufei Wang 0006, Bihan Wen
ICCV5
2023 ExposureDiffusion: Learning to Expose for Low-light Image Enhancement
abstract
Previous raw image-based low-light image enhancement methods predominantly relied on feed-forward neural networks to learn deterministic mappings from low-light to normally-exposed images. However, they failed to capture critical distribution information, leading to visually undesirable results. This work addresses the issue by seamlessly integrating a diffusion model with a physics-based exposure model. Different from a vanilla diffusion model that has to perform Gaussian denoising, with the injected physics-based exposure model, our restoration process can directly start from a noisy image instead of pure noise. As such, our method obtains significantly improved performance and reduced inference time compared with vanilla diffusion models. To make full use of the advantages of different intermediate steps, we further propose an adaptive residual layer that effectively screens out the side-effect in the iterative refinement when the intermediate results have been already well-exposed. The proposed framework can work with both real-paired datasets, SOTA noise models, and different backbone networks. We evaluate the proposed method on various public benchmarks, achieving promising results with consistent improvements using different exposure models and backbones. Besides, the proposed method achieves better generalization capacity for unseen amplifying ratios and better performance than a larger feedforward neural model when few parameters are adopted. The code is released at https://github.com/wyf0912/ExposureDiffusion.
Yufei Wang 0006, Yi Yu 0011, Wenhan Yang, Lanqing Guo, Lap-Pui Chau, Alex Chichung Kot, Bihan Wen
ICCV7
2023 Enhancing Low-Light Images Using Infrared Encoded Images
abstract
Low-light image enhancement task is essential yet challenging as it is ill-posed intrinsically. Previous arts mainly focus on the low-light images captured in the visible spectrum using pixel-wise loss, which limits the capacity of recovering the brightness, contrast, and texture details due to the small number of income photons. In this work, we propose a novel approach to increase the visibility of images captured under low-light environments by removing the in-camera infrared (IR) cut-off filter, which allows for the capture of more photons and results in improved signal-to-noise ratio due to the inclusion of information from the IR spectrum. To verify the proposed strategy, we collect a paired dataset of low-light images captured without the IR cut-off filter, with corresponding long-exposure reference images with an external filter. The experimental results on the proposed dataset demonstrate the effectiveness of the proposed method, showing better performance quantitatively and qualitatively. The dataset and code are publicly available at https://wyf0912.github.io/ELIEI/
Shulin Tian, Yufei Wang 0006, Renjie Wan, Wenhan Yang, Alex Chichung Kot, Bihan Wen
ICIP6
2023 Removing Image Artifacts From Scratched Lens Protectors
abstract
A protector is placed in front of the camera lens for mobile devices to avoid damage, while the protector itself can be easily scratched accidentally, especially for plastic ones. The artifacts appear in a wide variety of patterns, making it difficult to see through them clearly. Removing image artifacts from the scratched lens protector is inherently challenging due to the occasional flare artifacts and the co-occurring interference within mixed artifacts. Though different methods have been proposed for some specific distortions, they seldom consider such inherent challenges. In our work, we consider the inherent challenges in a unified framework with two cooperative modules, which facilitate the performance boost of each other. We also collect a new dataset from the real world to facilitate training and evaluation purposes. The experimental results demonstrate that our method outperforms the baselines qualitatively and quantitatively. The code and datasets will be released at https://github.com/wyf0912/flare-removal
Yufei Wang 0006, Renjie Wan, Wenhan Yang, Bihan Wen, Lap-Pui Chau, Alex Chichung Kot
ISCAS4
2023 WBCAtt: A White Blood Cell Dataset Annotated with Detailed Morphological Attributes
abstract
The examination of blood samples at a microscopic level plays a fundamental role in clinical diagnostics. For instance, an in-depth study of White Blood Cells (WBCs), a crucial component of our blood, is essential for diagnosing blood-related diseases such as leukemia and anemia. While multiple datasets containing WBC images have been proposed, they mostly focus on cell categorization, often lacking the necessary morphological details to explain such categorizations, despite the importance of explainable artificial intelligence (XAI) in medical domains. This paper seeks to address this limitation by introducing comprehensive annotations for WBC images. Through collaboration with pathologists, a thorough literature review, and manual inspection of microscopic images, we have identified 11 morphological attributes associated with the cell and its components (nucleus, cytoplasm, and granules). We then annotated ten thousand WBC images with these attributes, resulting in 113k labels (11 attributes x 10.3k images). Annotating at this level of detail and scale is unprecedented, offering unique value to AI in pathology. Moreover, we conduct experiments to predict these attributes from cell images, and also demonstrate specific applications that can benefit from our detailed annotations. Overall, our dataset paves the way for interpreting WBC recognition models, further advancing XAI in the fields of pathology and hematology.
Satoshi Tsutsui, Winnie Pang, Bihan Wen
NeurIPS3
2023 Provenance of Training without Training Data: Towards Privacy-Preserving DNN Model Ownership Verification
abstract
In the era of deep learning, it is critical to protect the intellectual property of high-performance deep neural network (DNN) models. Existing proposals, however, are subject to adversarial ownership forgery (e.g., methods based on watermarks or fingerprints) or require full access to the original training dataset for ownership verification (e.g., methods requiring the replay of the learning process). In this paper, we propose a novel Provenance of Training (PoT) scheme, the first empirical study towards verifying DNN model ownership without accessing any original dataset while being robust against existing attacks. At its core, PoT relies on a coherent model chain built from the intermediate checkpoints saved during model training to serve as the ownership certificate. Through an in-depth analysis of model training, we propose six key properties that a legitimate model chain shall naturally hold. In contrast, it is difficult for the adversary to forge a model chain that satisfies these properties simultaneously without performing actual training. We systematically analyze PoT’s robustness against various possible attacks, including the adaptive attacks that are designed given the full knowledge of PoT’s design, and further perform extensive empirical experiments to demonstrate our security analysis.
Zhuotao Liu, Bihan Wen, Ke Xu 0002, Weiqiang Wang 0002, Wenbiao Zhao, Qi Li 0002
WWW4
2023 Reconciliation of statistical and spatial sparsity for robust visual classification
Hao Cheng 0016, Kim-Hui Yap, Bihan Wen
Neurocomputing3
2023 Efficient and Effective Nonconvex Low-Rank Subspace Clustering via SVT-Free Operators
abstract
With the growing interest in convex and nonconvex low-rank matrix learning problems, the widely used singular value thresholding (SVT) operators associated with rank relaxation functions often face higher computational complexity, particularly for large-scale data matrices. To improve the efficacy of low-rank subspace clustering and overcome the issue of high computational complexity, this work proposes an efficient and effective method that avoids the need for singular value decomposition (SVD) computations in the iteration scheme. This can be achieved through the use of a computationally efficient and compact formulation, as well as automatic removal of the optimal mean, which reduces time consumption and enhances evaluation performance. A unified clustering framework based on Schatten-$p$norm regularized by$\ell _{2,q}$-norm can be formulated using this processing way, where inner element suppression can be achieved by choosing appropriate$p$,$q \in (0,1)$. Additionally, calculating the optimal mean enhances the robustness of the proposed method in the presence of outliers. Unlike the general iteration scheme of the alternating direction method of multiplier (ADMM) algorithms that introduce auxiliary splitting variables, the proposed alternating re-weighted least square (ARwLS) algorithm uses matrix inverse and multiplication computations to obtain analytic solutions, resulting in faster processing speeds for each sub-problem. To further investigate, we provide the computational complexity of each iteration and the theoretical analysis of the convergence property, where the derived solution is a stationary point. Experimental results on synthetic data and several benchmark datasets demonstrate the promising efficiency and efficacy of the proposed clustering method compared to classical and competing algorithms.
Hengmin Zhang, Shuyi Li 0003, Jing Qiu 0002, Yang Tang 0001, Jie Wen 0001, Zhiyuan Zha, Bihan Wen
IEEE Trans. Circuits Syst. Video Technol.7
2023 Nonlocal Structured Sparsity Regularization Modeling for Hyperspectral Image Denoising
abstract
The non-local-based model for hyperspectral image (HSI) denoising first uses non-local self-similarity (NSS) prior to group similar full-band patches into three-dimensional non-local full-band groups (tensors) using a block matching (BM) operation, and then a low-rank (LR) penalty is typically applied to each non-local full-band group to reduce noise. While non-local-based methods have shown promising performance in HSI denoising, most existing methods have only considered the LR property of the non-local full-band group while ignoring the strong correlation between sparse coefficients. Moreover, such methods often result in unsatisfactory visual artifacts due to the noise sensitivity of BM operations, while requiring expensive computations. To address these limitations, this paper proposes a novel non-local structured sparsity regularization (NLSSR) approach for HSI denoising. First, to mitigate the noise sensitivity of the BM operation, we propose a graph-based domain distance scheme to index similar full-band patches to form the non-local full-band group. Second, we design an adaptive unidirectional low-rank (LR) dictionary with low complexity that takes into account the differences in intrinsic structure correlation among different modes of the non-local full-band tensor. Third, we utilize a global spectral LR prior to reduce spectral redundancy. Fourth, we develop a generalized soft-thresholding (GST) algorithm based on the alternating minimization framework to solve the NLSSR-based HSI denoising problem. We perform extensive experiments on both simulated and real data to show that the proposed NLSSR algorithm outperforms many popular or state-of-the-art HSI denoising methods in both quantitative and visual evaluations.
Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiachao Zhang, Jiantao Zhou 0001, Yilong Lu, Ce Zhu
IEEE Trans. Geosci. Remote. Sens.2
2023 Prototype Correction via Contrastive Augmentation for Few-Shot Unconstrained Palmprint Recognition
abstract
Unconstrained Palmprint Recognition (UPR) shows engaging potential owing to its high hygiene and privacy. The unconstrained acquisition usually produces wide variations, against which deep methods resort to large samples that are unavailable in practice, however. We focus on Few-Shot UPR (FS-UPR), a more general problem, recognizing query samples given a few support samples per class. Because scarce samples insufficiently represent potential variations, the augmentation methods train independent hallucinators on large samples to generate more ones. Whereas, the hallucinators trained independently of Few-Shot Learning (FSL) are blind of generating promising samples to boost the downstream FSL. Moreover, training hallucinators requires large samples per class, unavailable from unconstrained palmprint databases. We aim to address FS-UPR via contrastive augmentation merely on the support samples. Observing the variations to betransferableacross samples, we exploit low-rank representation to disentangle support samples intoprinciplesandvariationsin embedding space and augment features by variation transfer. To this end, we devise anend-to-endDeep Low-Rank Representation Feature Augmentation Network (DLRR-FAN) to simultaneously learn the embedding space and augmentation features with guaranteedrealityanddiversity. Furthermore, a Contrastive Recognition Regularizer (CRR) is tailored to secure thediscriminabilityof augmentation features. During each training episode, the task motivates DLRR-FAN to augment such features that correct the biased prototypes towards upcoming query samples with variations unseen in the support samples, namelytask-drivenprototype correction. Extensive experiments on both the typical and extended FS-UPR tasks demonstrate the efficacy of DLRR-FAN versus the state-of-the-art methods.
Kunlei Jing, Xinman Zhang, Chen Zhang 0013, Wanyu Lin, Hebo Ma, Bihan Wen
IEEE Trans. Inf. Forensics Secur.7
2023 Learning to Solve Multiple-TSP With Time Window and Rejections via Deep Reinforcement Learning
abstract
We propose a manager-worker framework (the implementation of our model is publically available at:https://github.com/zcaicaros/manager-worker-mtsptwr) based on deep reinforcement learning to tackle a hard yet nontrivial variant of Travelling Salesman Problem (TSP), i.e., multiple-vehicle TSP with time window and rejections (mTSPTWR), where customers who cannot be served before the deadline are subject to rejections. Particularly, in the proposed framework, a manager agent learns to divide mTSPTWR into sub-routing tasks by assigning customers to each vehicle via a Graph Isomorphism Network (GIN) based policy network. A worker agent learns to solve sub-routing tasks by minimizing the cost in terms of both tour length and rejection rate for each vehicle, the maximum of which is then fed back to the manager agent to learn better assignments. Experimental results demonstrate that the proposed framework outperforms strong baselines in terms of higher solution quality and shorter computation time. More importantly, the trained agents also achieve competitive performance for solving unseen larger instances.
Rongkai Zhang 0001, Zhiguang Cao, Wen Song 0004, Puay Siew Tan, Jie Zhang 0002, Bihan Wen, Justin Dauwels
IEEE Trans. Intell. Transp. Syst.7
2023 Graph Neural Networks With Triple Attention for Few-Shot Learning
abstract
Recent advances in Graph Neural Networks (GNNs) have achieved superior results in many challenging tasks, such as few-shot learning. Despite its capacity to learn and generalize a model from only a few annotated samples, GNN is limited in scalability, as deep GNN models usually suffer from severe over-fitting and over-smoothing. In this work, we propose a novel GNN framework with atriple-attention mechanism,i.e.node self-attention, neighbor attention, and layer memory attention, to tackle these challenges. We provide both theoretical analysis and illustrations to explain why the proposed attentive modules can improve GNN scalability for few-shot learning tasks. Our experiments show that the proposed Attentive GNN model outperforms the state-of-the-art few-shot learning methods using both GNN and non-GNN approaches. The improvement is consistent over the mini-ImageNet, tiered-ImageNet, CUB-200-2011, and Flowers-102 benchmarks, using both ConvNet-4 and ResNet-12 backbones, and under both the inductive and transductive settings. Furthermore, we demonstrate the superiority of our method for few-shot fine-grained and semi-supervised classification tasks with extensive experiments. The code for this work is publicly available athttps://github.com/chenghao-ch94/AGNN.
Hao Cheng 0016, Joey Tianyi Zhou, Wee-Peng Tay, Bihan Wen
IEEE Trans. Multim.4
2023 Hyper RPCA: Joint Maximum Correntropy Criterion and Laplacian Scale Mixture Modeling on-the-Fly for Moving Object Detection
abstract
Moving object detection is critical for automated video analysis in many vision-related tasks, such as surveillance tracking, video compression coding, etc. Robust Principal Component Analysis (RPCA), as one of the most popular moving object modelling methods, aims to separate the temporally-varying (i.e., moving) foreground objects from the static background in video, assuming the background frames to be low-rank while the foreground to be spatially sparse. Classic RPCA imposes sparsity of the foreground component using$\ell _1$-norm, and minimizes the modeling error via$\ell _2$-norm. We show that such assumptions can be too restrictive in practice, which limits the effectiveness of the classic RPCA, especially when processing videos with dynamic background, camera jitter, camouflaged moving object, etc. In this paper, we propose a novel RPCA-based model, called Hyper RPCA, to detect moving objects on the fly. Different from classic RPCA, the proposed Hyper RPCA jointly applies the maximum correntropy criterion (MCC) for the modeling error, and Laplacian scale mixture (LSM) model for foreground objects. Extensive experiments have been conducted, and the results demonstrate that the proposed Hyper RPCA has competitive performance for foreground detection to the state-of-the-art algorithms on several well-known benchmark datasets.
Zerui Shao, Yi-Fei Pu, Jiliu Zhou, Bihan Wen, Yi Zhang 0018
IEEE Trans. Multim.4
2023 Purifying Low-Light Images via Near-Infrared Enlightened Image
abstract
Cameras usually produce low-quality images under low-light conditions. Though many methods have been proposed to enhance the visibility of low-light images, they are mainly designed for illumination correction and less capable of sup-pressing the artifacts. In this paper, we propose to enhance the visibility and suppress artifacts by purifying low-light images under the guidance of the NIR enlightened image captured by using the near-infrared light as compensation. Specifically, we introduce a disentanglement framework to disentangle the structure and color components from the NIR enlightened and RGB images, respectively. Correspondingly, we introduce a new dataset with the RGB and NIR enlightened images for training and evaluation purposes. The experimental results show that our proposed method achieves promising results.
Renjie Wan, Boxin Shi, Wenhan Yang, Bihan Wen, Ling-Yu Duan, Alex Chichung Kot
IEEE Trans. Multim.4
2023 DisP+V: A Unified Framework for Disentangling Prototype and Variation From Single Sample per Person
abstract
Single sample per person face recognition (SSPP FR) is one of the most challenging problems in FR due to the extreme lack of enrolment data. To date, the most popular SSPP FR methods are the generic learning methods, which recognize query face images based on the so-called prototype plus variation (i.e., P+V) model. However, the classic P+V model suffers from two major limitations: 1) it linearly combines the prototype and variation images in the observational pixel-spatial space and cannot generalize to multiple nonlinear variations, e.g., poses, which are common in face images and 2) it would be severely impaired once the enrolment face images are contaminated by nuisance variations. To address the two limitations, it is desirable to disentangle the prototype and variation in a latent feature space and to manipulate the images in a semantic manner. To this end, we propose a novel disentangled prototype plus variation model, dubbed DisP+V, which consists of an encoder-decoder generator and two discriminators. The generator and discriminators play two adversarial games such that the generator nonlinearly encodes the images into a latent semantic space, where the more discriminative prototype feature and the less discriminative variation feature are disentangled. Meanwhile, the prototype and variation features can guide the generator to generate an identity-preserved prototype and the corresponding variation, respectively. Experiments on various real-world face datasets demonstrate the superiority of our DisP+V model over the classic P+V model for SSPP FR. Furthermore, DisP+V demonstrates its unique characteristics in both prototype recovery and face editing/interpolation.
Binghui Wang, Mang Ye, Yiu-Ming Cheung, Yiran Chen 0001, Bihan Wen
IEEE Trans. Neural Networks Learn. Syst.6
2023 Low-Rankness Guided Group Sparse Representation for Image Restoration
abstract
As a spotlighted nonlocal image representation model, group sparse representation (GSR) has demonstrated a great potential in diverse image restoration tasks. Most of the existing GSR-based image restoration approaches exploit the nonlocal self-similarity (NSS) prior by clustering similar patches into groups and imposing sparsity to each group coefficient, which can effectively preserve image texture information. However, these methods have imposed only plain sparsity over each individual patch of the group, while neglecting other beneficial image properties, e.g., low-rankness (LR), leads to degraded image restoration results. In this article, we propose a novel low-rankness guided group sparse representation (LGSR) model for highly effective image restoration applications. The proposed LGSR jointly utilizes the sparsity and LR priors of each group of similar patches under a unified framework. The two priors serve as the complementary priors in LGSR for effectively preserving the texture and structure information of natural images. Moreover, we apply an alternating minimization algorithm with an adaptively adjusted parameter scheme to solve the proposed LGSR-based image restoration problem. Extensive experiments are conducted to demonstrate that the proposed LGSR achieves superior results compared with many popular or state-of-the-art algorithms in various image restoration tasks, including denoising, inpainting, and compressive sensing (CS).
Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu, Alex Chichung Kot
IEEE Trans. Neural Networks Learn. Syst.2
2022 PIP: Physical Interaction Prediction via Mental Simulation with Span Selection
Jiafei Duan, Samson Yu Bai Jian, Soujanya Poria, Bihan Wen, Cheston Tan
ECCV (35)4
2022 DVS-Voltmeter: Stochastic Process-Based Event Simulator for Dynamic Vision Sensors
Songnan Lin, Zhenhua Guo 0001, Bihan Wen
ECCV (7)4
2022 Parameter-Free Style Projection for Arbitrary Image Style Transfer
abstract
Arbitrary image style transfer is a challenging task which aims to stylize a content image conditioned on arbitrary style images. In this task the feature-level content-style transformation plays a vital role for proper fusion of features. Existing feature transformation algorithms often suffer from loss of content or style details, non-natural stroke patterns, and unstable training. To mitigate these issues, this paper proposes a new feature-level style transformation technique, named Style Projection, for parameter-free, fast, and effective content-style transformation. This paper further presents a real-time feed-forward model to leverage Style Projection for arbitrary image style transfer, which includes a regularization term for matching the semantics between input contents and stylized outputs. Extensive qualitative analysis, quantitative evaluation, and user study have demonstrated the effectiveness and efficiency of the proposed methods.
Siyu Huang, Haoyi Xiong, Tianyang Wang 0004, Bihan Wen, Qingzhong Wang, Jun Huan, Dejing Dou
ICASSP4
2022 Feature Augmentation Learning for Few-Shot Palmprint Image Recognition With Unconstrained Acquisition
abstract
Few-shot learning is challenging in unconstrained palmprint recognition, where the palmprint images are collected by unconstrained acquisitions, i.e., different imaging sensors, backgrounds, palm postures, and illumination conditions. Furthermore, due to the lack of unconstrained palmprint databases and sufficient intra-class samples, it is difficult to apply the classic few-shot techniques, such as pre-training, fine-tuning, and sample augmentation, to generalize the model. In this work, we propose a novel feature augmentation network (FAN) for few-shot unconstrained palmprint recognition. Without any external databases, FAN aims to simultaneously remove the image variations caused by the unconstrained acquisitions and augment their feature representation from only a few support samples. To this end, the proposed deep self-expression module first decouples the support images into their principle and variation features. Assuming that the variations are translational across palm-print samples, the variation-sharing module achieves feature augmentations by swapping and combining all pairs of principle and variation features. The augmented palmprint features generated by FAN enable more general representations of categorical prototypes for few-shot unconstrained palmprint recognition. Experimental results on the standard palmprint databases show that FAN can effectively represent the prototypes of palmprint images from only a few available samples, thus outperforming the state-of-the-art methods in unconstrained palmprint recognition.
Kunlei Jing, Xinman Zhang, Bihan Wen
ICASSP4
2022 Simultaneous Nonlocal Low-Rank And Deep Priors For Poisson Denoising
abstract
Poisson noise is a common electronic noise, which has widely occurred in various photo-limited imaging systems. However, due to signal-dependent and multiplicative characteristics for Poisson noise, Poisson denoising is still an open problem. In this paper, we propose a novel approach using simultaneous nonlocal low-rank and deep priors (SNLDP) for Poisson denoising. The proposed SNLD-P simultaneously employs nonlocal self-similarity and deep image priors under the hybrid plug and play framework, which comprises multiple pairs of complementary priors, namely, nonlocal and local, shallow and deep, and internal and external. To make the optimization tractable, an effective alternating direction method of multiplier (ADMM) algorithm under the alternative minimization framework is provided to solve the proposed SNLDP-based Poisson denoising problem. Experimental results demonstrate the superiority of the proposed SNLDP over many popular or state-of-the-art Poisson denoising algorithms in terms of quantitative and visual perception.
Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu
ICASSP2
2022 REPNP: Plug-and-Play with Deep Reinforcement Learning Prior for Robust Image Restoration
abstract
Image restoration schemes based on the pre-trained deep models have received great attention due to their unique flexibility for solving various inverse problems. In particular, the Plug-and-Play (PnP) framework is a popular and powerful tool that can integrate an off-the-shelf deep denoiser for different image restoration tasks with known observation models. However, obtaining the observation model that exactly matches the actual one can be challenging in practice. Thus, the PnP schemes with conventional deep denoisers may fail to generate satisfying results in some real-world image restoration tasks. We argue that the robustness of the PnP framework is largely limited by using the off-the-shelf deep denoisers that are trained by deterministic optimization. To this end, we propose a novel deep reinforcement learning (DRL) based PnP framework, dubbed RePNP, by leveraging a light-weight DRL-based denoiser for robust image restoration tasks. Experimental results demonstrate that the proposed RePNP is robust to the observation model used in the PnP scheme deviating from the actual one. Thus, RePNP can generate more reliable restoration results for image deblurring and super resolution tasks. Compared with several state-of-the-art deep image restoration baselines, RePNP achieves better results subjective to model deviation with fewer model parameters.
Chong Wang 0011, Rongkai Zhang 0001, Saiprasad Ravishankar, Bihan Wen
ICIP4
2022 Nonconvex Structural Sparsity Residual Constraint for Image Restoration
abstract
This article proposes a novel nonconvex structural sparsity residual constraint (NSSRC) model for image restoration, which integrates structural sparse representation (SSR) with nonconvex sparsity residual constraint (NC-SRC). Although SSR itself is powerful for image restoration by combining the local sparsity and nonlocal self-similarity in natural images, in this work, we explicitly incorporate the novel NC-SRC prior into SSR. Our proposed approach provides more effective sparse modeling for natural images by applying a more flexible sparse representation scheme, leading to high-quality restored images. Moreover, an alternating minimizing framework is developed to solve the proposed NSSRC-based image restoration problems. Extensive experimental results on image denoising and image deblocking validate that the proposed NSSRC achieves better results than many popular or state-of-the-art methods over several publicly available datasets.
Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiachao Zhang, Ce Zhu
IEEE Trans. Cybern.3
2022 Make Web3.0 Connected
abstract
${\mathsf Web3.0}$, often cited to drastically shape our lives, is ubiquitous. However, few literatures have discussed the crucial differentiators that separate${\mathsf Web3.0}$from the era we are currently living in. Via a thorough analysis of the recent blockchain infrastructure evolution, we capture a key invariant featuring the evolution, based on which we provide the first academic definition for${\mathsf Web3.0}$. Our definition is not the only way of understanding${\mathsf Web3.0}$, yet, it captures the fundamental and defining trait of${\mathsf Web3.0}$, and meanwhile it is has two desirable properties. Under this definition, we articulate three key categories of infrastructural enablers for${\mathsf Web3.0}$: individual smart-contract capable blockchains, federated or centralized platforms capable of publishing verifiable states, and an interoperability platform to hyperconnect those state publishers to provide a unified and connected computing platform for${\mathsf Web3.0}$applications. While innovations in all categories are necessary to fully enable${\mathsf Web3.0}$, in this article, we present a design for the third enabler, i.e., the first interoperability platform, namely${\mathsf HyperService}$, that advances the state-of-the-art by simultaneously deliversinteroperabilityandprogrammabilityacrossheterogeneousblockchains and state publishers.${\mathsf HyperService}$is powered by two innovative designs:${\mathsf (i)}$a developer-facing programming framework that allows developers to build cross-chain applications in a unified programming model; and${\mathsf (ii)}$a secure blockchain-facing cryptography protocol that provably realizes those applications on blockchains. We implement a prototype of${\mathsf HyperService}$in approximately 62,000 lines of code to demonstrate its practicality, usability and scalability.
Zhuotao Liu, Yangxi Xiang, Peng Gao 0008, Haoyu Wang 0001, Xusheng Xiao, Bihan Wen, Qi Li 0002, Yih-Chun Hu
IEEE Trans. Dependable Secur. Comput.7
2022 A Unified Framework for Bidirectional Prototype Learning From Contaminated Faces Across Heterogeneous Domains
abstract
Existing heterogeneous face synthesis (HFS) methods focus on performing accurate image-to-image translation across domains, while they cannot effectively remove the nuisance facial variations such as poses, expressions or occlusions. To address such challenges, this paper studies a new practical heterogeneous prototype learning (HPL) problem. To be specific, given a face image contaminated by facial variations from a source domain, HPL aims to reconstruct the variation-free prototype in a specified target domain. To tackle HPL, we propose a unified and end-to-end framework named bidirectional heterogeneous prototype learning (BHPL). As a bidirectional learning framework, BHPL is able to simultaneously reconstruct the heterogeneous prototypes acrosssource-to-targetas well astarget-to-sourcedomains. Furthermore, BHPL is capable of learning the identity prototype features for the contaminated face images from both source and target domains in order to perform robust heterogeneous face recognition. BHPL consists of an encoder-decoder structural generator and two dual-task discriminators, which play an adversarial game such that the generator learns the identity prototype feature and generates the cross-domain identity-preserved prototype for each input face image from both domains, and the discriminators accurately predict face identity and distinguish real versus fake prototypes. Empirically studies on multiple heterogeneous face datasets containing facial variations demonstrate the effectiveness of BHPL.
Binghui Wang, Siyu Huang, Yiu-Ming Cheung, Bihan Wen
IEEE Trans. Inf. Forensics Secur.5
2022 Exploiting Non-Local Priors via Self-Convolution for Highly-Efficient Image Restoration
abstract
Constructing effective priors is critical to solving ill-posed inverse problems in image processing and computational imaging. Recent works focused on exploiting non-local similarity by grouping similar patches for image modeling, and demonstrated state-of-the-art results in many image restoration applications. However, compared to classic methods based on filtering or sparsity, non-local algorithms are more time-consuming, mainly due to the highly inefficient block matching step, i.e., distance between every pair of overlapping patches needs to be computed. In this work, we propose a novel Self-Convolution operator to exploit image non-local properties in a unified framework. We prove that the proposed Self-Convolution based formulation can generalize the commonly-used non-local modeling methods, as well as produce results equivalent to standard methods, but with much cheaper computation. Furthermore, by applying Self-Convolution, we propose an effective multi-modality image restoration scheme, which is much more efficient than conventional block matching for non-local modeling. Experimental results demonstrate that (1) Self-Convolution with fast Fourier transform implementation can significantly speed up most of the popular non-local image restoration algorithms, with two-fold to nine-fold faster block matching, and (2) the proposed online multi-modality image restoration scheme achieves superior denoising results than competing methods in both efficiency and effectiveness on RGB-NIR images. The code for this work is publicly available at https://github.com/GuoLanqing/Self-Convolution.
Lanqing Guo, Zhiyuan Zha, Saiprasad Ravishankar, Bihan Wen
IEEE Trans. Image Process.4
2022 FONT-SIR: Fourth-Order Nonlocal Tensor Decomposition Model for Spectral CT Image Reconstruction
abstract
Spectral computed tomography (CT) reconstructs images from different spectral data through photon counting detectors (PCDs). However, due to the limited number of photons and the counting rate in the corresponding spectral segment, the reconstructed spectral images are usually affected by severe noise. In this paper, we propose a fourth-order nonlocal tensor decomposition model for spectral CT image reconstruction (FONT-SIR). To maintain the original spatial relationships among similar patches and improve the imaging quality, similar patches without vectorization are grouped in both spectral and spatial domains simultaneously to form the fourth-order processing tensor unit. The similarity of different patches is measured with the cosine similarity of latent features extracted using principal component analysis (PCA). By imposing the constraints of the weighted nuclear and total variation (TV) norms, each fourth-order tensor unit is decomposed into a low-rank component and a sparse component, which can efficiently remove noise and artifacts while preserving the structural details. Moreover, the alternating direction method of multipliers (ADMM) is employed to solve the decomposition model. Extensive experimental results on both simulated and real data sets demonstrate that the proposed FONT-SIR achieves superior qualitative and quantitative performance compared with several state-of-the-art methods.
Xiang Chen 0015, Wenjun Xia, Yan Liu 0052, Hu Chen 0002, Jiliu Zhou, Zhiyuan Zha, Bihan Wen, Yi Zhang 0018
IEEE Trans. Medical Imaging7
2022 A Hybrid Structural Sparsification Error Model for Image Restoration
abstract
Recent works on structural sparse representation (SSR), which exploit image nonlocal self-similarity (NSS) prior by grouping similar patches for processing, have demonstrated promising performance in various image restoration applications. However, conventional SSR-based image restoration methods directly fit the dictionaries or transforms to the internal (corrupted) image data. The trained internal models inevitably suffer from overfitting to data corruption, thus generating the degraded restoration results. In this article, we propose a novel hybrid structural sparsification error (HSSE) model for image restoration, which jointly exploits image NSS prior using both the internal and external image data that provide complementary information. Furthermore, we propose a general image restoration scheme based on the HSSE model, and an alternating minimization algorithm for a range of image restoration applications, including image inpainting, image compressive sensing and image deblocking. Extensive experiments are conducted to demonstrate that the proposed HSSE-based scheme outperforms many popular or state-of-the-art image restoration methods in terms of both objective metrics and visual perception.
Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu, Alex Chichung Kot
IEEE Trans. Neural Networks Learn. Syst.2
2021 Self-Convolution: A Highly-Efficient Operator for Non-Local Image Restoration
abstract
Constructing effective image priors is critical to solving ill-posed inverse problems, such as image restoration. Recent works proposed to exploit image non-local similarity for inverse problems by grouping similar patches, and demonstrated state-of-the-art results in many applications. However, comparing to classic local methods based on filtering or sparsity, most of the non-local algorithms are time-consuming, mainly due to the highly inefficient and redundant block matching step, where the distance between each pair of overlapping patches needs to be computed. In this work, we propose a novel Self-Convolution operator to exploit image non-local similarity in a self-supervised way. The proposed Self-Convolution can generalize the commonly-used block matching step, and produce the equivalent results with much cheaper computation. Based on Self-Convolution, we propose an effective multi-modality image restoration scheme, which is much more efficient than conventional block matching for non-local modeling. Experimental results also demonstrate that Self-Convolution can significantly speed up most of the popular non-local image restoration algorithms, with two-fold to nine-fold faster block matching. The codes will be released on GitHub.
Lanqing Guo, Zhiyuan Zha, Saiprasad Ravishankar, Bihan Wen
ICASSP4
2021 Learning Sparsifying Transforms for Image Reconstruction in Electrical Impedance Tomography
abstract
Electrical Impedance Tomography (EIT) is a fast and non-invasive imaging technology that reconstructs the internal electrical properties of a subject. However, its functionality is limited by low spatial resolution arising from an ill-posed and ill-conditioned inverse problem. Several sparsity-promoting regularization methods have been applied to improve the quality of EIT image reconstruction, including various ℓ0and ℓ1-based analytical models (TV, TwIST, etc.), and a patch-based sparse representation via a learned dictionary (using the K-SVD algorithm), dubbed CS-EIT. To further exploit the potential of compressed sensing in Electrical Impedance Tomography, this paper incorporates the recent novel method of transform learning for EIT image reconstruction. We propose a blind compressed sensing algorithm, dubbed TL-EIT, which simultaneously optimizes the sparsifying transform and updates the reconstructed image. We demonstrate using both synthetic and in vivo data that the proposed TL-EIT is more effective than other sparsity-based algorithms for reconstructing high-quality EIT images. In addition, TL-EIT also accelerates the reconstruction process in comparison to other learning-based algorithms like CS-EIT.
Kaiyi Yang, Narong Borijindargoon, Boon Poh Ng, Saiprasad Ravishankar, Bihan Wen
ICASSP5
2021 R3L: Connecting Deep Reinforcement Learning To Recurrent Neural Networks For Image Denoising Via Residual Recovery
abstract
State-of-the-art image denoisers exploit various types of deep neural networks via deterministic training. Alternatively, very recent works utilize deep reinforcement learning for restoring images with diverse or unknown corruptions. Though deep reinforcement learning can generate effective policy networks for operator selection or architecture search in image restoration, how it is connected to the classic deterministic training in solving inverse problems remains unclear. In this work, we propose a novel image denoising scheme via Residual Recovery using Reinforcement Learning, dubbed R3L. We show that R3L is equivalent to a deep recurrent neural network that is trained using a stochastic reward, in contrast to many popular denoisers using supervised learning with deterministic losses. To benchmark the effectiveness of reinforcement learning in R3L, we train a recurrent neural network with the same architecture for residual recovery using the deterministic loss, thus to analyze how the two different training strategies affect the denoising performance. With such a unified benchmarking system, we demonstrate that the proposed R3L has better generalizability and robustness in image denoising when the estimated noise level varies, comparing to its counterparts using deterministic training, as well as various state-of the-art image denoising algorithms.
Rongkai Zhang 0001, Zhiyuan Zha, Justin Dauwels, Bihan Wen
ICIP5
2021 Multi-Scale Feature Guided Low-Light Image Enhancement
abstract
Low-light image enhancement aims at enlarging the intensity of image pixels to better match human perception and to improve the performance of subsequent vision tasks. While it is relatively easy to enlighten a globally low-light image, the lighting condition of realistic scenes is usually non-uniform and complex, e.g., some images may contain both bright and extremely dark regions, with or without rich features and information. Existing methods often generate abnormal light-enhancement results with over-exposure artifacts without proper guidance. To tackle this challenge, we propose a multi-scale feature guided attention mechanism in the deep generator, which can effectively perform a spatially-varying light enhancement. The attention map is fused by both the gray map and extracted feature map of the input image, to focus more on those dark and informative regions. Our baseline is an unsupervised generative adversarial network, which can be trained without any low/normal light image pair. Experimental results demonstrate the superiority in visual quality and performance of subsequent object detection over state-of-the-art alternatives.
Lanqing Guo, Renjie Wan, Guan-Ming Su, Alex Chichung Kot, Bihan Wen
ICIP5
2021 Joint Anomaly Detection and Inpainting for Microscopy Images Via Deep Self-Supervised Learning
abstract
While microscopy enables material scientists to view and analyze microstructures, the imaging results often include defects and anomalies with varied shapes and locations. The presence of such anomalies significantly degrades the quality of microscopy images and the subsequent analytical tasks. Comparing to classic feature-based methods, recent advancements in deep learning provide a more efficient, accurate, and scalable approach to detect and remove anomalies in microscopy images. However, most of the deep inpainting and anomaly detection schemes require a certain level of supervision, i.e., either annotation of the anomalies, or a corpus of purely normal data, which are limited in practice for supervision-starving microscopy applications. In this work, we propose a self-supervised deep learning scheme for joint anomaly detection and inpainting of microscopy images. The proposed anomaly detection model can be trained over a mixture of normal and abnormal microscopy images without any labeling. Instead of a two-stage scheme, our multi-task model can simultaneously detect abnormal regions and remove the defects via jointly training. To benchmark such microscopy application under the real-world setup, we propose a novel dataset of real microscopic images of integrated circuits, dubbed MIIC. The proposed dataset contains tens of thousands of normal microscopic images, while we labeled hundreds of them containing various imaging and manufacturing anomalies and defects for testing. Experiments show that the proposed model outperforms various popular or state-of-the-art competing methods for both microscopy image anomaly detection and inpainting.
Deruo Cheng, Xulei Yang, Tong Lin 0001, Yiqiong Shi, Kaiyi Yang, Bah-Hwee Gwee, Bihan Wen
ICIP8
2021 Labmat: Learned Feature-Domain Block Matching For Image Restoration
abstract
Grouping of similar patches, called block matching, has been widely used in image restoration applications. Popular block matching algorithms exploit image non-local similarities in spatial or a fixed transform domain, e.g., wavelets and DCT. However, applying these methods on corrupted patches usually leads to degraded matching accuracy, thus limiting the image restoration performance. In this work, we develop a novel methodology for performing block matching in a supervised way by learning multi-layer sparsifying transforms. The proposed learned transform-domain block matching method for image restoration, dubbed LABMAT, is shown to have better accuracy in terms of clustering similar blocks in the presence of noise, and it also achieves an improved denoising performance when it is incorporated into popular non-local denoising schemes.
Shijun Liang 0001, Berk Iskender, Bihan Wen, Saiprasad Ravishankar
ICIP3
2021 Systematic Analysis of Circular Artifacts for Stylegan
abstract
Recent research works have pointed out that the synthesized images by StyleGAN contain prominent circular artifacts which severely degrade the quality of generated images. In this work, we provide a systematic investigation on how those circular artifacts are formed by studying the functionalities of different modules that are used in the Style-GAN architecture. We present both analysis of the StyleGAN mechanism and extensive experiments to verify our claims. The key modules of StyleGAN that promote such undesired artifacts are highlighted based on the analysis. Besides, we propose a simple yet effective solution to remove the prominent circular artifacts for StyleGAN, by applying a simple but efficient pixel-instance normalization layer. The improved StyleGAN model trained via our proposed approach successfully prevents the appearance of circular artifacts in the generated images.
Way Tan, Bihan Wen, Cen Chen 0001, Zeng Zeng, Xulei Yang
ICIP2
2021 Low-Rank Regularized Joint Sparsity for Image Denoising
abstract
Nonlocal sparse representation models such as group sparse representation (GSR), low-rankness and joint sparsity (JS) have shown great potentials in image denoising studies, by effectively exploiting image nonlocal self-similarity (NSS) property. Popular dictionary-based JS algorithms apply convex JS penalties in their objective functions, which avoid NP-hard sparse coding step, but lead to only approximately sparse representation. Such approximated JS models fail to impose low-rankness of the underlying image data, resulting in degraded quality in image restoration. To simultaneously exploit the low-rank and JS priors, we propose a novel low-rank regularized joint sparsity model, dubbed LRJS, to enhance the dependency (i. e., low-rankness) of similar patches, thus better suppress independent noise. Moreover, to make the optimization tractable and robust, an alternating minimization algorithm with an adaptive parameter adjustment strategy is developed to solve the proposed LRJS-based image denoising problem. Experimental results demonstrate that the proposed LRJS outperforms many popular or state-of-the-art denoising algorithms in terms of both objective and visual perception met-
Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu
ICIP2
2021 Disentangling Prototype and Variation for Single Sample Face Recognition
abstract
Single sample per person face recognition (SSPP FR) is one of the most challenging problems in FR due to the extreme lack of enrolment data. State-of-the-art SSPP FR methods are based on the prototype plus variation (i.e., P+V) model. However, the classic P+V model has two major limitations: 1) It is a linear model and cannot generalize many non-linear variations; 2) It can be severely impaired once the enrolment face images are contaminated with variations. To this end, we propose a novel disentangled prototype plus variation model, dubbed DisP+V, to tackle such limitations. DisP+V consists of an encoder-decoder structural generator and two discriminators. The generator and discriminators play two adversarial games such that the generator nonlinearly encodes the images into a latent semantic space, where the more discriminative prototype feature and the less discriminative variation feature are disentangled. Meanwhile, the prototype and variation features in the latent space can guide the generator to generate an identity-preserved prototype and the corresponding variation, respectively. Experiments on various real-world face datasets demonstrate the superiority of our DisP+V model over the classic P+V model for SSPP FR. Furthermore, DisP+V demonstrates its unique characteristics in the challenging prototype recovery task.
Binghui Wang, Mang Ye, Yiran Chen 0001, Bihan Wen
ICME5
2021 Recent Advances in Adversarial Training for Adversarial Robustness
abstract
Adversarial training is one of the most effective approaches for deep learning models to defend against adversarial examples. Unlike other defense strategies, adversarial training aims to enhance the robustness of models intrinsically. During the past few years, adversarial training has been studied and discussed from various aspects, which deserves a comprehensive review. For the first time in this survey, we systematically review the recent progress on adversarial training for adversarial robustness with a novel taxonomy. Then we discuss the generalization problems in adversarial training from three perspectives and highlight the challenges which are not fully tackled. Finally, we present potential future directions.
Jinqi Luo, Jun Zhao 0007, Bihan Wen, Qian Wang 0002
IJCAI4
2021 ReLLIE: Deep Reinforcement Learning for Customized Low-Light Image Enhancement
abstract
Low-light image enhancement (LLIE) is a pervasive yet challenging problem, since: 1) low-light measurements may vary due to different imaging conditions in practice; 2) images can be enlightened subjectively according to diverse preference by each individual. To tackle these two challenges, this paper presents a novel deep reinforcement learning based method, dubbed ReLLIE, for customized low-light enhancement. ReLLIE models LLIE as a markov decision process, i.e., estimating the pixel-wise image-specific curves sequentially and recurrently. Given the reward computed from a set of carefully crafted non-reference loss functions, a lightweight network is proposed to estimate the curves for enlightening of a low-light image input. As ReLLIE learns a policy instead of one-one image translation, it can handle various low-light measurements and provide customized enhanced outputs by flexibly applying the policy different times. Furthermore, ReLLIE can enhance real-world images with hybrid corruptions, i.e., noise, by using a plug-and-play denoiser easily. Extensive experiments on various benchmarks demonstrate the advantages of ReLLIE, comparing to the state-of-the-art methods. (Code is available: https://github.com/GuoLanqing/ReLLIE.)
Rongkai Zhang 0001, Lanqing Guo, Siyu Huang, Bihan Wen
ACM Multimedia4
2021 VD-GAN: A Unified Framework for Joint Prototype and Representation Learning From Contaminated Single Sample per Person
abstract
Single sample per person (SSPP) face recognition with a contaminated biometric enrolment database (SSPP-ce FR) is an emerging practical FR problem, where the SSPP in the enrolment database is no longer standard but contaminated by nuisance facial variations such as expression, lighting, pose, and disguise. In this case, the conventional SSPP FR methods, including the patch-based and generic learning methods, will suffer from serious performance degradation. Few recent methods were proposed to tackle SSPP-ce FR by either performing prototype learning on the contaminated enrolment database or learning discriminative representations that are robust against variation. Despite that, most of these approaches can only handle a specified single variation, e.g., pose, but cannot be extended to multiple variations. To address these two limitations, we propose a novel Variation Disentangling Generative Adversarial Network (VDGAN) to jointly perform prototype learning and representation learning in a unified framework. The proposed VD-GAN consists of an encoder-decoder structural generator and a multi-task discriminator to handle universal variations including single, multiple, and even mixed variations in practice. The generator and discriminator play an adversarial game such that the generator learns a discriminative identity representation and generates an identity-preserved prototype for each face image, while the discriminator aims to predict face identity label, distinguish real vs. fake prototype, and disentangle target variations from the learned representations. Qualitative and quantitative evaluations on various real-world face datasets containing single/multiple and mixed variations demonstrate the effectiveness of VD-GAN.
Binghui Wang, Yiu-Ming Cheung, Yiran Chen 0001, Bihan Wen
IEEE Trans. Inf. Forensics Secur.5
2021 Image Restoration via Reconciliation of Group Sparsity and Low-Rank Models
abstract
Image nonlocal self-similarity (NSS) property has been widely exploited via various sparsity models such as joint sparsity (JS) and group sparse coding (GSC). However, the existing NSS-based sparsity models are either too restrictive, e.g., JS enforces the sparse codes to share the same support, or too general, e.g., GSC imposes only plain sparsity on the group coefficients, which limit their effectiveness for modeling real images. In this paper, we propose a novel NSS-based sparsity model, namely, low-rank regularized group sparse coding (LR-GSC), to bridge the gap between the popular GSC and JS. The proposed LR-GSC model simultaneously exploits the sparsity and low-rankness of the dictionary-domain coefficients for each group of similar patches. An alternating minimization with an adaptive adjusted parameter strategy is developed to solve the proposed optimization problem for different image restoration tasks, including image denoising, image deblocking, image inpainting, and image compressive sensing. Extensive experimental results demonstrate that the proposed LR-GSC algorithm outperforms many popular or state-of-the-art methods in terms of objective and perceptual metrics.
Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu
IEEE Trans. Image Process.2
2021 Triply Complementary Priors for Image Restoration
abstract
Recent works that utilized deep models have achieved superior results in various image restoration (IR) applications. Such approach is typically supervised, which requires a corpus of training images with distributions similar to the images to be recovered. On the other hand, the shallow methods, which are usually unsupervised remain promising performance in many inverse problems, e.g., image deblurring and image compressive sensing (CS), as they can effectively leverage nonlocal self-similarity priors of natural images. However, most of such methods are patch-based leading to the restored images with various artifacts due to naive patch aggregation in addition to the slow speed. Using either approach alone usually limits performance and generalizability in IR tasks. In this paper, we propose a joint low-rank and deep (LRD) image model, which contains a pair of triply complementary priors, namely, internal and external, shallow and deep, and non-local and local priors. We then propose a novel hybrid plug-and-play (H-PnP) framework based on the LRD model for IR. Following this, a simple yet effective algorithm is developed to solve the proposed H-PnP based IR problems. Extensive experimental results on several representative IR tasks, including image deblurring, image CS and image deblocking, demonstrate that the proposed H-PnP algorithm achieves favorable performance compared to many popular or state-of-the-art IR methods in terms of both objective and visual perception.
Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Joey Tianyi Zhou, Jiantao Zhou 0001, Ce Zhu
IEEE Trans. Image Process.2
2020 A Hybrid Structural Sparse Error Model for Image Deblocking
abstract
Inspired by the image nonlocal self-similarity (NSS) prior, structural sparse representation (SSR) models exploit each group as the basic unit for sparse representation, which have achieved promising results in various image restoration applications. However, conventional SSR models only exploited the group within the input degraded (internal) image for image restoration, which can be limited by over-fitting to data corruption. In this paper, we propose a novel hybrid structural sparse error (HSSE) model for image deblocking. The proposed HSSE model exploits image NSS prior over both the internal image and external image corpus, which can be complementary in both feature space and image plane. Moreover, we develop an alternating minimization with an adaptive parameter setting strategy to solve the proposed HSSE model. Experimental results demonstrate that the proposed HSSE-based image deblocking algorithm outperforms many state-of-the-art image deblocking methods in terms of objective and visual perception.
Zhiyuan Zha, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu, Bihan Wen
ICASSP5
2020 Joint Statistical and Spatial Sparse Representation for Robust Image and Image-Set Classification
abstract
Recent image classification schemes, by learning deep features from large-scale dataset, have achieved the significantly better results comparing to classic feature-based approaches. However, there are still challenges in practice, such as classifying noisy image-set queries and training over limited-scale dataset. Instead of applying generic deep features, the model-based approaches can be more effective for robust image and image-set classification tasks, as we need various image priors to exploit the inter- and intra-set data variations while prevent over-fitting. In this work, we propose a novel joint statistical and spatial sparse representation, dubbed J3S, to model the image or image-set data, by exploiting both their local patch structures and global Gaussian distribution into Riemannian manifold. To the best of our knowledge, no work to date utilized both global statistics and local patch structures jointly via sparse representation. We propose to solve a co-regularized sparse coding problem based on the J3S model, by coupling the local and global representations using joint sparsity. The learned J3S models are used for robust image and image-set classification. Experiments show that the proposed J3S-based image classification scheme outperforms the popular or state-of-the-art competing methods.
Hao Cheng 0016, Bihan Wen
ICIP2
2020 Opencc - an open Benchmark data set for Corpus Callosum Segmentation and Evaluation
abstract
Neuroimaging studies have revealed that the structural changes of the corpus callosum (CC) are evident in a variety of neurological diseases, such as epilepsy and autism. Segmentation of the CC from magnetic resonance images (MRI) of the brain is a crucial step in the diagnosis of various brain disorders. However, the lack of open benchmark CC datasets has hindered development of CC segmentation techniques. In this work, we present an open benchmark dataset - OpenCC - for CC segmentation and evaluation. The dataset was built through alternative application of automatic segmentation and manual refinement. The automatic segmentation is based on recent advances in deep learning - fully convolutional networks, specifically U-Net, while the manual refinement is done by domain radiologists. The resulting dataset consists of 4643 mid-sagittal (or near mid-sagittal) slices and their corresponding CC masks. Furthermore, we provided some baseline segmentation results on the OpenCC dataset by using two latest deep learning segmentation approaches. The OpenCC dataset can be used for comparison and evaluation of newly developed CC segmentation algorithms. We endeavor that, through the publishing of the OpenCC dataset and baseline segmentation results, we could promote further development of CC segmentation techniques.
Xulei Yang, Gabriel Tjio, Cen Chen 0001, Li Wang 0057, Bihan Wen, Yi Su 0001
ICIP6
2020 The Power Of Triply Complementary Priors For Image Compressive Sensing
abstract
Recent works that utilized deep models have achieved superior results in various image restoration applications. Such approach is typically supervised which requires a corpus of training images with distribution similar to the images to be recovered. On the other hand, the shallow methods which are usually unsupervised remain promising performance in many inverse problems, e.g., image compressive sensing (CS), as they can effectively leverage non-local self-similarity priors of natural images. However, most of such methods are patch-based leading to the restored images with various ringing artifacts due to naive patch aggregation. Using either approach alone usually limits performance and generalizability in image restoration tasks. In this paper, we propose a joint low-rank and deep (LRD) image model, which contains a pair of triply complementary priors, namely external and internal, deep and shallow, and local and nonlocal priors. We then propose a novel hybrid plug-and-play (H-PnP) framework based on the LRD model for image CS. To make the optimization tractable, a simple yet effective algorithm is proposed to solve the proposed H-PnP based image CS problem. Extensive experimental results demonstrate that the proposed H-PnP algorithm significantly outperforms the state-of-the-art techniques for image CS recovery such as SCSNet and WNNM.
Zhiyuan Zha, Xin Yuan 0002, Joey Tianyi Zhou, Jiantao Zhou 0001, Bihan Wen, Ce Zhu
ICIP5
2020 Reconciliation Of Group Sparsity And Low-Rank Models For Image Restoration
abstract
Image nonlocal self-similarity (NSS) property has been widely exploited via various sparsity models such as joint sparsity (JS) and group sparse coding (GSC). However, the existing NSS-based sparsity models are either too restrictive, i.e., JS enforces the sparse codes to share the same support, or too general, i.e., GSC imposes only plain sparsity on the group coefficients, which limit their effectiveness for modeling real images. In this paper, we propose a novel NSS-based sparsity model, namely low-rank regularized group sparse coding (LR-GSC), to bridge the gap between the popular GSC and JS. The proposed LR-GSC model simultaneously exploits the sparsity and low-rankness of the dictionary-domain coefficients for each group of similar patches. To make the proposed scheme tractable and robust, an alternating minimization with an adaptive adjusted parameter strategy is developed to solve the proposed optimization problem. Experimental results on both image deblocking and denoising demonstrate that the proposed LR-GSC image restoration algorithms outperform many popular or state-of-the-art methods, in terms of both the objective and perceptual quality.
Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu
ICME2
2020 Removing Backdoor-Based Watermarks in Neural Networks with Limited Data
abstract
Deep neural networks have been widely applied and achieved great success in various fields. As training deep models usually consumes massive data and computational resources, trading the trained deep models is highly-demanded and lucrative nowadays. Unfortunately, the naive trading schemes typically involves potential risks related to copyright and trustworthiness issues, e.g., a sold model can be illegally resold to others without further authorization to reap huge profits. To tackle this problem, various watermarking techniques are proposed to protect the model intellectual property, amongst which the backdoor-based watermarking is the most commonly-used one. However, the robustness of these watermarking approaches is not well evaluated under realistic settings, such as limited in-distribution data availability and agnostic of watermarking patterns. In this paper, we benchmark the robustness of watermarking, and propose a novel backdoor-based watermark removal framework using limited data, dubbed WILD. The proposed WILD removes the watermarks of deep models with only a small portion of training data, and the output model can perform the same as models trained from scratch without watermarks injected. In particular, a novel data augmentation method is utilized to mimic the behavior of watermark triggers. Combining with the distribution alignment between the normal and perturbed (e.g., occluded) data in the feature space, our approach generalizes well on all typical types of trigger contents. The experimental results demonstrate that our approach can effectively remove the watermarks without compromising the deep model performance for the original task with the limited access to training data.
Xuankai Liu, Fengting Li, Bihan Wen, Qi Li 0002
ICPR3
2020 Generating Person Images with Appearance-aware Pose Stylizer
abstract
Generation of high-quality person images is challenging, due to the sophisticated entanglements among image factors, e.g., appearance, pose, foreground, background, local details, global structures, etc. In this paper, we present a novel end-to-end framework to generate realistic person images based on given person poses and appearances. The core of our framework is a novel generator called Appearance-aware Pose Stylizer (APS) which generates human images by coupling the target pose with the conditioned person appearance progressively. The framework is highly flexible and controllable by effectively decoupling various complex person image factors in the encoding phase, followed by re-coupling them in the decoding phase. In addition, we present a new normalization method named adaptive patch normalization, which enables region-specific normalization and shows a good performance when adopted in person image generation model. Experiments on two benchmark datasets show that our method is capable of generating visually appealing and realistic-looking results using arbitrary image and pose inputs.
Siyu Huang, Haoyi Xiong, Zhi-Qi Cheng, Qingzhong Wang, Xingran Zhou, Bihan Wen, Jun Huan, Dejing Dou
IJCAI6
2020 Connecting Image Denoising and High-Level Vision Tasks via Deep Learning
abstract
Image denoising and high-level vision tasks are usually handled independently in the conventional practice of computer vision, and their connection is fragile. In this paper, we cope with the two jointly and explore the mutual influence between them with the focus on two questions, namely (1) how image denoising can help improving high-level vision tasks, and (2) how the semantic information from high-level vision tasks can be used to guide image denoising. First for image denoising we propose a convolutional neural network in which convolutions are conducted in various spatial resolutions via downsampling and upsampling operations in order to fuse and exploit contextual information on different scales. Second we propose a deep neural network solution that cascades two modules for image denoising and various high-level tasks, respectively, and use the joint loss for updating only the denoising network via backpropagation. We experimentally show that on one hand, the proposed denoiser has the generality to overcome the performance degradation of different high-level vision tasks. On the other hand, with the guidance of high-level vision information, the denoising network produces more visually appealing results. Extensive experiments demonstrate the benefit of exploiting image semantics simultaneously for image denoising and highlevel vision tasks via deep learning. The code is available online: https://github.com/Ding-Liu/DeepDenoising.
Ding Liu 0001, Bihan Wen, Jianbo Jiao, Xianming Liu 0005, Zhangyang Wang, Thomas S. Huang
IEEE Trans. Image Process.2
2020 Image Recovery via Transform Learning and Low-Rank Modeling: The Power of Complementary Regularizers
abstract
Recent works on adaptive sparse and on low-rank signal modeling have demonstrated their usefulness in various image/video processing applications. Patch-based methods exploit local patch sparsity, whereas other works apply low-rankness of grouped patches to exploit image non-local structures. However, using either approach alone usually limits performance in image reconstruction or recovery applications. In this work, we propose a simultaneous sparsity and low-rank model, dubbed STROLLR, to better represent natural images. In order to fully utilize both the local and non-local image properties, we develop an image restoration framework using a transform learning scheme with joint low-rank regularization. The approach owes some of its computational efficiency and good performance to the use of transform learning for adaptive sparse representation rather than the popular synthesis dictionary learning algorithms, which involve approximation of NP-hard sparse coding and expensive learning steps. We demonstrate the proposed framework in various applications to image denoising, inpainting, and compressed sensing based magnetic resonance imaging. Results show promising performance compared to state-of-the-art competing methods.
Bihan Wen, Yanjun Li 0001, Yoram Bresler
IEEE Trans. Image Process.1
2020 Group Sparsity Residual Constraint With Non-Local Priors for Image Restoration
abstract
Group sparse representation (GSR) has made great strides in image restoration producing superior performance, realized through employing a powerful mechanism to integrate the local sparsity and nonlocal self-similarity of images. However, due to some form of degradation (e.g., noise, down-sampling or pixels missing), traditional GSR models may fail to faithfully estimate sparsity of each group in an image, thus resulting in a distorted reconstruction of the original image. This motivates us to design a simple yet effective model that aims to address the above mentioned problem. Specifically, we propose group sparsity residual constraint with nonlocal priors (GSRC-NLP) for image restoration. Through introducing the group sparsity residual constraint, the problem of image restoration is further defined and simplified through attempts at reducing the group sparsity residual. Towards this end, we first obtain a good estimation of the group sparse coefficient of each original image group by exploiting the image nonlocal self-similarity (NSS) prior along with self-supervised learning scheme, and then the group sparse coefficient of the corresponding degraded image group is enforced to approximate the estimation. To make the proposed scheme tractable and robust, two algorithms, i.e., iterative shrinkage/thresholding (IST) and alternating direction method of multipliers (ADMM), are employed to solve the proposed optimization problems for different image restoration tasks. Experimental results on image denoising, image inpainting and image compressive sensing (CS) recovery, demonstrate that the proposed GSRC-NLP based image restoration algorithm is comparable to state-of-the-art denoising methods and outperforms several state-of-the-art image inpainting and image CS recovery methods in terms of both objective and perceptual quality metrics.
Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiantao Zhou 0001, Ce Zhu
IEEE Trans. Image Process.3
2020 From Rank Estimation to Rank Approximation: Rank Residual Constraint for Image Restoration
abstract
In this paper, we propose a novel approach for the rank minimization problem, termed rank residual constraint (RRC). Different from existing low-rank based approaches, such as the well-known nuclear norm minimization (NNM) and the weighted nuclear norm minimization (WNNM), which estimate the underlying low-rank matrix directly from the corrupted observation, we progressively approximate (approach) the underlying low-rank matrix via minimizing the rank residual. Through integrating the image nonlocal self-similarity (NSS) prior with the proposed RRC model, we apply it to image restoration tasks, including image denoising and image compression artifacts reduction. Toward this end, we first obtain a good reference of the original image groups by using the image NSS prior, and then the rank residual of the image groups between this reference and the degraded image is minimized to achieve a better estimate to the desired image. In this manner, both the reference and the estimated image in each iteration are improved gradually and jointly. Based on the group-based sparse representation model, we further provide a theoretical analysis on the feasibility of the proposed RRC model. Experimental results demonstrate that the proposed RRC model outperforms many state-of-the-art schemes in both the objective and perceptual qualities.
Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiantao Zhou 0001, Jiachao Zhang, Ce Zhu
IEEE Trans. Image Process.3
2020 A Benchmark for Sparse Coding: When Group Sparsity Meets Rank Minimization
abstract
Sparse coding has achieved a great success in various image processing tasks. However, a benchmark to measure the sparsity of image patch/group is missing since sparse coding is essentially an NP-hard problem. This work attempts to fill the gap from the perspective of rank minimization. We firstly design an adaptive dictionary to bridge the gap between group-based sparse coding (GSC) and rank minimization. Then, we show that under the designed dictionary, GSC and the rank minimization problems are equivalent, and therefore the sparse coefficients of each patch group can be measured by estimating the singular values of each patch group. We thus earn a benchmark to measure the sparsity of each patch group because the singular values of the original image patch groups can be easily computed by the singular value decomposition (SVD). This benchmark can be used to evaluate performance of any kind of norm minimization methods in sparse coding through analyzing their corresponding rank minimization counterparts. Towards this end, we exploit four well-known rank minimization methods to study the sparsity of each patch group and the weighted Schatten p-norm minimization (WSNM) is found to be the closest one to the real singular values of each patch group. Inspired by the aforementioned equivalence regime of rank minimization and GSC, WSNM can be translated into a non-convex weighted ℓp-norm minimization problem in GSC. By using the earned benchmark in sparse coding, the weighted ℓp-norm minimization is expected to obtain better performance than the three other norm minimization methods, i.e., ℓ1-norm, ℓp-norm and weighted ℓ1-norm. To verify the feasibility of the proposed benchmark, we compare the weighted ℓp-norm minimization against the three aforementioned norm minimization methods in sparse coding. Experimental results on image restoration applications, namely image inpainting and image compressive sensing recovery, demonstrate that the proposed scheme is feasible and outperforms many state-of-the-art methods.
Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiantao Zhou 0001, Jiachao Zhang, Ce Zhu
IEEE Trans. Image Process.3
2020 Image Restoration Using Joint Patch-Group-Based Sparse Representation
abstract
Sparse representation has achieved great success in various image processing and computer vision tasks. For image processing, typical patch-based sparse representation (PSR) models usually tend to generate undesirable visual artifacts, while group-based sparse representation (GSR) models lean to produce over-smooth effects. In this paper, we propose a new sparse representation model, termed joint patch-group based sparse representation (JPG-SR). Compared with existing sparse representation models, the proposed JPG-SR provides an effective mechanism to integrate the local sparsity and nonlocal self-similarity of images. We then apply the proposed JPG-SR to image restoration tasks, including image inpainting and image deblocking. An iterative algorithm based on the alternating direction method of multipliers (ADMM) framework is developed to solve the proposed JPG-SR based image restoration problems. Experimental results demonstrate that the proposed JPG-SR is effective and outperforms many state-of-the-art methods in both objective and perceptual quality.
Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiachao Zhang, Jiantao Zhou 0001, Ce Zhu
IEEE Trans. Image Process.3
2020 Image Restoration via Simultaneous Nonlocal Self-Similarity Priors
abstract
Through exploiting the image nonlocal self-similarity (NSS) prior by clustering similar patches to construct patch groups, recent studies have revealed that structural sparse representation (SSR) models can achieve promising performance in various image restoration tasks. However, most existing SSR methods only exploit the NSS prior from the input degraded (internal) image, and few methods utilize the NSS prior from external clean image corpus; how to jointly exploit the NSS priors of internal image and external clean image corpus is still an open problem. In this paper, we propose a novel approach for image restoration by simultaneously considering internal and external nonlocal self-similarity (SNSS) priors that offer mutually complementary information. Specifically, we first group nonlocal similar patches from images of a training corpus. Then a group-based Gaussian mixture model (GMM) learning algorithm is applied to learn an external NSS prior. We exploit the SSR model by integrating the NSS priors of both internal and external image data. An alternating minimization with an adaptive parameter adjusting strategy is developed to solve the proposed SNSS-based image restoration problems, which makes the entire algorithm more stable and practical. Experimental results on three image restoration applications, namely image denoising, deblocking and deblurring, demonstrate that the proposed SNSS produces superior results compared to many popular or state-of-the-art methods in both objective and perceptual quality measurements.
Zhiyuan Zha, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu, Bihan Wen
IEEE Trans. Image Process.5
2019 HyperService: Interoperability and Programmability Across Heterogeneous Blockchains
abstract
Blockchain interoperability, which allows state transitions across different blockchain networks, is critical functionality to facilitate major blockchain adoption. Existing interoperability protocols mostly focus on atomic token exchanges between blockchains. However, as blockchains have been upgraded from passive distributed ledgers into programmable state machines (thanks to smart contracts), the scope of blockchain interoperability goes beyond just token exchanges. In this paper, we present HyperService, the first platform that delivers interoperability and programmability across heterogeneous blockchains. HyperService is powered by two innovative designs: (i) a developer-facing programming framework that allows developers to build cross-chain applications in a unified programming model; and (ii) a secure blockchain-facing cryptography protocol that provably realizes those applications on blockchains. We implement a prototype of HyperService in approximately 35,000 lines of code to demonstrate its practicality. Our experiments show that (i) HyperService imposes reasonable latency, in order of seconds, on the end-to-end execution of cross-chain applications; (ii) the HyperService platform is scalable to continuously incorporate new large-scale production blockchains.
Zhuotao Liu, Yangxi Xiang, Peng Gao 0008, Haoyu Wang 0001, Xusheng Xiao, Bihan Wen, Yih-Chun Hu
CCS7
2019 A Comparative Study for the Nuclear Norms Minimization Methods
abstract
The nuclear norm minimization (NNM) is commonly used to approximate the matrix rank by shrinking all singular values equally. However, the singular values have clear physical meanings in many practical problems, and NNM may not be able to faithfully approximate the matrix rank. To alleviate the above-mentioned limitation of NNM, recent studies have suggested that the weighted nuclear norm minimization (WNNM) can achieve a better rank estimation than NNM, which heuristically set the weight being inverse to the singular values. However, it still lacks a rigorous explanation why WNNM is more effective than NMM in various applications. In this paper, we analyze NNM and WNNM from the perspective of group sparse representation (GSR). Concretely, an adaptive dictionary learning method is devised to connect the rank minimization and GSR models. Based on the proposed dictionary, we prove that NNM and WNNM are equivalent to ℓ1-norm minimization and the weighted ℓ1-norm minimization in GSR, respectively. Inspired by enhancing sparsity of the weighted ℓ1-norm minimization in comparison with ℓ1-norm minimization in sparse representation, we thus explain that WNNM is more effective than NMM. By integrating the image nonlocal self-similarity (NSS) prior with the WNNM model, we then apply it to solve the image denoising problem. Experimental results demonstrate that WNNM is more effective than NNM and outperforms several state-of-the-art methods in both objective and perceptual quality.
Zhiyuan Zha, Bihan Wen, Jiachao Zhang, Jiantao Zhou 0001, Ce Zhu
ICIP2
2019 Simultaneous Nonlocal Self-Similarity Prior for Image Denoising
abstract
Nonlocal image representation has achieved great success in various image processing tasks such as image denoising, image deblurring and image deblocking. Particularly, by exploiting the image nonlo-cal self-similarity (NSS) prior, many nonlocal similar patches can be searched across the whole image for a given patch, which has significantly boosted the performance of image restoration. To the best of our knowledge, most existing methods only consider the NSS prior of the input degraded image, while few methods exploit the NSS prior from external clean image corpus. However, how to utilize the NSS priors of input degraded image and external clean image corpus simultaneously is still an open problem. In this paper, we propose a novel approach for image denoising, which exploits simultaneous nonlocal self-similarity (SNSS) by integrating the NSS priors of both the input degraded image and external clean image corpus. Firstly, we search and group nonlocal similar patches from a clean image corpus, and a group-based Gaussian Mixture Model (GMM) learning algorithm is developed to learn an external NSS prior. Then, an optimal group is selected from the best suitable Gaussian component for a group of the noisy image. By integrating the group of the noisy image and the corresponding group of the Gaussian component with a low-rank constraint, an iterative algorithm is developed to solve the proposed SNSS model. Experimental results demonstrate that the proposed SNSS-based denoising method produces superior results compared with many state-of-the-art denoising methods in both objective and perceptual quality.
Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiachao Zhang, Jiantao Zhou 0001, Ce Zhu
ICIP3
2019 VIDOSAT: High-Dimensional Sparsifying Transform Learning for Online Video Denoising
abstract
Techniques exploiting the sparsity of images in a transform domain are effective for various applications in image and video processing. In particular, transform learning methods involve cheap computations and have been demonstrated to perform well in applications, such as image denoising and medical image reconstruction. Recently, we proposed methods for online learning of sparsifying transforms from streaming signals, which enjoy good convergence guarantees and involve lower computational costs than online synthesis dictionary learning. In this paper, we apply online transform learning to video denoising. We present a novel framework for online video denoising based on high-dimensional sparsifying transform learning for spatio-temporal patches. The patches are constructed either from corresponding 2D patches in successive frames or using an online block matching technique. The proposed online video denoising requires little memory and offers efficient processing. Numerical experiments evaluate the performance of the proposed video denoising algorithms on multiple video data sets. The proposed methods outperform several related and recent techniques, including denoising with 3D DCT, prior schemes based on dictionary learning, non-local means, background separation, and deep learning, as well as the popular VBM3D and VBM4D.
Bihan Wen, Saiprasad Ravishankar, Yoram Bresler
IEEE Trans. Image Process.1
2018 Joint Patch-Group Based Sparse Representation for Image Inpainting
abstract
Sparse representation has achieved great successes in various machine learning and image processing tasks. For image processing, typical patch-based sparse representation (PSR) models usually tend to generate undesirable visual artifacts, while group-based sparse representation (GSR) models produce over-smooth phenomena. In this paper, we propose a new sparse representation model, termed joint patch-group based sparse representation (JPG-SR). Compared with existing sparse representation models, the proposed JPG-SR provides a powerful mechanism to integrate the local sparsity and nonlocal self-similarity of images. We then apply the proposed JPG-SR model to a low-level vision problem, namely, image inpainting. To make the proposed scheme tractable and robust, an iterative algorithm based on the alternating direction method of multipliers (ADMM) framework is developed to solve the proposed JPG-SR model. Experimental results demonstrate that the proposed model is efficient and outperforms several state-of-the-art methods in both objective and perceptual quality.
Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiantao Zhou 0001, Ce Zhu
ACML3
2018 Deepcasd: An End-to-End Approach for Multi-Spectral Image Super-Resolution
abstract
Multi-spectral (MS) image super-resolution aims to reconstruct super-resolved multi-channel images from their low-resolution images by regularizing the image to be reconstructed. Recently data-driven regularization techniques based on sparse modeling and deep learning have achieved substantial improvements in single image reconstruction problems. Inspired by these data-driven methods, we develop a novel coupled analysis and synthesis dictionary (CASD) model for MS image super-resolution, by exploiting a regularizer that operates within, as well as across, multiple spectral channels using convolutional dictionaries. To learn the CASD model parameters, we propose a deep dictionary learning framework, named DeepCASD, by unfolding and training an end-to-end CASD based reconstruction network over an image data set. Experimental results show that the DeepCASD framework exhibits improved performance on multi-spectral image super-resolution compared to state-of-the-art learning based super-resolution algorithms.
Bihan Wen, Ulugbek Kamilov, Dehong Liu, Hassan Mansour, Petros Boufounos
ICASSP1
2018 Transim: Transfer Image Local Statistics Across EOTFS for HDR Image Applications
abstract
Despite the popularity of high dynamic range (HDR) technology in recent years, various algorithms for image and video applications are still designed and optimized for traditional standard dynamic range (SDR) data. Directly applying SDR-optimized algorithms to HDR images and video will result in significant artifacts or coding deficiency. In this work, we present a novel preprocessing method, dubbed TransIm, which transfers local statistics for the images from the desired domain (e.g. SDR) to the current domain (e.g., HDR), while maintaining its current visual presence. It is achieved by controlling the less perceivable “noise” that is orthogonal to the sparsifiable image content, using a unitary sparsifying transform. Numerical results show that the proposed TransIm can effectively transfer local patch variance from Gamma domain to Perceptual Quantizer (PQ) domain for HDR videos. We also demonstrate that the TransIm outputs are more robust to distortions and artifacts in seam carving applications.
Bihan Wen, Guan-Ming Su
ICME1
2018 When Image Denoising Meets High-Level Vision Tasks: A Deep Learning Approach
abstract
Conventionally, image denoising and high-level vision tasks are handled separately in computer vision. In this paper, we cope with the two jointly and explore the mutual influence between them. First we propose a convolutional neural network for image denoising which achieves the state-of-the-art performance. Second we propose a deep neural network solution that cascades two modules for image denoising and various high-level tasks, respectively, and use the joint loss for updating only the denoising network via back-propagation. We demonstrate that on one hand, the proposed denoiser has the generality to overcome the performance degradation of different high-level vision tasks. On the other hand, with the guidance of high-level vision information, the denoising network can generate more visually appealing results. To the best of our knowledge, this is the first work investigating the benefit of exploiting image semantics simultaneously for image denoising and high-level vision tasks via deep learning.
Ding Liu 0001, Bihan Wen, Xianming Liu 0005, Zhangyang Wang, Thomas S. Huang
IJCAI2
2018 Non-Local Recurrent Network for Image Restoration
abstract
Many classic methods have shown non-local self-similarity in natural images to be an effective prior for image restoration. However, it remains unclear and challenging to make use of this intrinsic property via deep networks. In this paper, we propose a non-local recurrent network (NLRN) as the first attempt to incorporate non-local operations into a recurrent neural network (RNN) for image restoration. The main contributions of this work are: (1) Unlike existing methods that measure self-similarity in an isolated manner, the proposed non-local module can be flexibly integrated into existing deep networks for end-to-end training to capture deep feature correlation between each location and its neighborhood. (2) We fully employ the RNN structure for its parameter efficiency and allow deep feature correlation to be propagated along adjacent recurrent states. This new design boosts robustness against inaccurate correlation estimation due to severely degraded images. (3) We show that it is essential to maintain a confined neighborhood for computing deep feature correlation given degraded images. This is in contrast to existing practice that deploys the whole image. Extensive experiments on both image denoising and super-resolution tasks are conducted. Thanks to the recurrent non-local operations and correlation propagation, the proposed NLRN achieves superior results to state-of-the-art methods with many fewer parameters.
Ding Liu 0001, Bihan Wen, Yuchen Fan 0001, Chen Change Loy, Thomas S. Huang
NeurIPS2
2017 When sparsity meets low-rankness: Transform learning with non-local low-rank constraint for image restoration
abstract
Recent works on adaptive sparse signal modeling have demonstrated their usefulness in various image/video processing applications. As the popular synthesis dictionary learning methods involve NP-hard sparse coding and expensive learning steps, transform learning has recently received more interest for its cheap computation. However, exploiting local patch sparsity alone usually limits performance in various image processing tasks. In this work, we propose a joint adaptive patch sparse and group low-rank model, dubbed STROLLR, to better represent natural images. We develop an image restoration framework based on the proposed model, which involves a simple and efficient alternating algorithm. We demonstrate applications, including image denoising and inpainting. Results show promising performance even when compared to state-of-the-art methods.
Bihan Wen, Yanjun Li 0001, Yoram Bresler
ICASSP1
2017 Joint Adaptive Sparsity and Low-Rankness on the Fly: An Online Tensor Reconstruction Scheme for Video Denoising
abstract
Recent works on adaptive sparse and low-rank signal modeling have demonstrated their usefulness, especially in image/video processing applications. While a patch-based sparse model imposes local structure, low-rankness of the grouped patches exploits non-local correlation. Applying either approach alone usually limits performance in various low-level vision tasks. In this work, we propose a novel video denoising method, based on an online tensor reconstruction scheme with a joint adaptive sparse and low-rank model, dubbed SALT. An efficient and unsupervised online unitary sparsifying transform learning method is introduced to impose adaptive sparsity on the fly. We develop an efficient 3D spatio-temporal data reconstruction framework based on the proposed online learning method, which exhibits low latency and can potentially handle streaming videos. To the best of our knowledge, this is the first work that combines adaptive sparsity and low-rankness for video denoising, and the first work of solving the proposed problem in an online fashion. We demonstrate video denoising results over commonly used videos from public datasets. Numerical experiments show that the proposed video denoising method outperforms competing methods.
Bihan Wen, Yanjun Li 0001, Luke Pfister, Yoram Bresler
ICCV1
2016 Learning flipping and rotation invariant sparsifying transforms
abstract
Adaptive sparse representation has been heavily exploited in signal processing and computer vision. Recently, sparsifying transform learning received interest for its cheap computation and optimal updates in the alternating algorithms. In this work, we develop a methodology for learning a Flipping and Rotation Invariant Sparsifying Transform, dubbed FRIST, to better represent natural images that contain textures with various geometrical directions. The proposed alternating learning algorithm involves efficient optimal updates. We demonstrate empirical convergence behavior of the proposed learning algorithm. Preliminary experiments show the usefulness of FRIST for image sparse representation, segmentation, robust inpainting, and MRI reconstruction with promising performances.
Bihan Wen, Saiprasad Ravishankar, Yoram Bresler
ICIP1
2016 COVERAGE - A novel database for copy-move forgery detection
abstract
We present COVERAGE - a novel database containing copy-move forged images and their originals with similar but genuine objects. COVERAGE is designed to highlight and address tamper detection ambiguity of popular methods, caused by self-similarity within natural images. In COVERAGE, forged-original pairs are annotated with (i) the duplicated and forged region masks, and (ii) the tampering factor/similarity metric. For benchmarking, forgery quality is evaluated using (i) computer vision-based methods, and (ii) human detection performance. We also propose a novel sparsity-based metric for efficiently estimating forgery quality. Experimental results show that (a) popular forgery detection methods perform poorly over COVERAGE, and (b) the proposed sparsity based metric best correlates with human detection performance. We release the COVERAGE database to the research community.
Bihan Wen, Subramanian Ramanathan, Tian-Tsong Ng, Xuanjing Shen, Stefan Winkler 0001
ICIP1
2016 Robust Single Image Super-Resolution via Deep Networks With Sparse Prior
abstract
Single image super-resolution (SR) is an ill-posed problem, which tries to recover a high-resolution image from its low-resolution observation. To regularize the solution of the problem, previous methods have focused on designing good priors for natural images, such as sparse representation, or directly learning the priors from a large data set with models, such as deep neural networks. In this paper, we argue that domain expertise from the conventional sparse coding model can be combined with the key ingredients of deep learning to achieve further improved results. We demonstrate that a sparse coding model particularly designed for SR can be incarnated as a neural network with the merit of end-to-end optimization over training data. The network has a cascaded structure, which boosts the SR performance for both fixed and incremental scaling factors. The proposed training and testing schemes can be extended for robust handling of images with additional degradation, such as noise and blurring. A subjective assessment is conducted and analyzed in order to thoroughly evaluate various SR techniques. Our proposed model is tested on a wide range of images, and it significantly outperforms the existing state-of-the-art methods for various scaling factors both quantitatively and perceptually.
Ding Liu 0001, Bihan Wen, Jianchao Yang, Wei Han 0002, Thomas S. Huang
IEEE Trans. Image Process.3
2015 Video denoising by online 3D sparsifying transform learning
abstract
Exploiting the sparsity of signals in an adaptive dictionary or transform domain benefits various applications in image/video processing. As opposed to synthesis dictionary learning, transform learning allows for cheap computations, and has been demonstrated to perform well in applications such as image denoising. Very recently, we proposed methods for online sparsifying transform learning, which are particularly useful for processing large-scale or streaming data. Online transform learning has good convergence guarantees and enjoys a much lower computational cost than online synthesis dictionary learning. In this work, we present a video denoising framework based on online 3D spatio-temporal sparsifying transform learning. The proposed scheme has low computational and memory costs, and can potentially handle streaming video. Our numerical experiments show promising performance for the proposed video denoising method compared to popular prior or state-of-the-art methods.
Bihan Wen, Saiprasad Ravishankar, Yoram Bresler
ICIP1
2015 Structured Overcomplete Sparsifying Transform Learning with Convergence Guarantees and Applications
Bihan Wen, Saiprasad Ravishankar, Yoram Bresler
Int. J. Comput. Vis.1
2014 Learning overcomplete sparsifying transforms with block cosparsity
abstract
The sparsity of images in a transform domain or dictionary has been widely exploited in image processing. Compared to the synthesis dictionary model, sparse coding in the (single) transform model is cheap. However, natural images typically contain diverse textures that cannot be sparsified well by a single transform. Hence, we propose a union of sparsifying transforms model, which is equivalent to an overcomplete transform model with block cosparsity (OC-TOBOS). Our alternating algorithm for transform learning involves simple closed-form updates. When applied to images, our algorithm learns a collection of well-conditioned transforms, and a good clustering of the patches or textures. Our learnt transforms provide better image representations than learned square transforms. We also show the promising denoising performance and speedups provided by the proposed method compared to synthesis dictionary-based denoising.
Bihan Wen, Saiprasad Ravishankar, Yoram Bresler
ICIP1