Shuhang Gu

dblp:126/1028 · DBLP profile ↗
← Back
73ranked-venue papers
11as first author
44since 2021 · last 2026
0000-0002-1394-6642ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 11 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 51 · 8 first-author · 28 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 DASR+: Training Domain Distance Aware Network for Unsupervised Image Super-Resolution
Xiaorui Zhao, Yunxuan Wei, Xin Deng 0002, Yawei Li 0001, Radu Timofte, Hengjie Song, Shuhang Gu
Int. J. Comput. Vis.7
2026 vProChain: Efficient Provenance Verification in Industrial Internet of Things (IIoT)
abstract
The Industrial Internet of Things (IIoT) has been widely deployed to enable real-time monitoring and automation. Within IIoT-driven production, supply chain management plays a critical role, necessitating verifiable provenance to ensure the authenticity and traceability of goods across multi-stakeholder networks. While blockchain provides a tamper-proof foundation, traditional storage structures suffer from unsecured data integrity, poor query efficiency, and scalability over provenance data. To address these challenges, we propose vProChain, an efficient provenance verification system to support verifiable and parallel queries over graph-structured provenance data. First, we design an Adaptive DAG Verkle Tree (ADVT) that deterministically maps supply chain dependencies into a graph-native authenticated data structure, enabling constant-size proofs and low-overhead verification. Second, we introduce the Merkle Inverted Patricia Trie (MIPT) to facilitate fast, verifiable multi-dimensional Boolean queries. Third, we develop a parallel provenance query algorithm that accelerates multi-hop path retrieval via consistent hashing and weighted bipartite matching. Finally, formal security analysis and extensive empirical evaluations demonstrate that vProChain can provide provable cryptographic guarantees for the soundness of provenance proofs and the completeness of query retrievals, while achieving high query efficiency in a large-scale IIoT environment.
Jiamin Deng, Zhe Peng, Chuan Zhang 0003, Shuhang Gu, Xin Xie 0001, Bin Xiao 0001
IEEE Internet Things J.4
2026 Noise-robust multivariate time series anomaly detection via importance awareness
Wuping Ke, Guowu Yang, Fumin Feng, Shuhang Gu, Desheng Zheng
Knowl. Based Syst.6
2026 DDNA: Dual-diffusion-enhanced network alignment
Shaobo Ren, Zhen Liu 0006, Haiyang Ren, Tongle Duan, Shuhang Gu
Knowl. Based Syst.5
2026 D2S-RSG-SSD: Dual Double-Sampling With Random Sub-Samples Generation for Self-Supervised Real Image Denoising
abstract
Recent advances in self-supervised image denoising have highlighted the potential of Blind-Spot Networks (BSNs). However, existing methods suffer from three major limitations: (1) Their effectiveness in real-world scenarios is limited by strong assumptions, such as noise independence, which rarely hold in practice. (2) While sampling-based strategies can partially improve performance, BSNs inherently suffer from information loss caused by centroid masking, and removing the blind spot leads to noise overfitting, both of which hinder denoising performance. (3) Sampling-based methods often introduce checkerboard artifacts, yet existing studies typically overlook the fundamental differences between these artifacts and real noise. To address these issues, we propose a novel self-supervised denoising framework, Dual Double-Sampling with Random Sub-samples Generation (D2S-RSG-SSD). To address Limitation 1, we introduce a sampling-based framework that breaks noise dependence by combining Random Sub-samples Generation (RSG) with a cross-paired loss $\mathcal {L}_{RSG}$LRSG. RSG generates diverse sub-samples with inherent variance, referred to as sampling differences, which serve as natural perturbations to augment training data and disrupt spatial noise correlations. The proposed loss function ensures full utilization of these sub-samples while stabilizing optimization. To address Limitation 2, we propose a Dual Double-Sampling (D2S) strategy with fixed sampling patterns and a dual-branch architecture. This design reduces reliance on pixel-level information and leverages complementary features to mitigate both noise overfitting and information loss. A key advantage is its compatibility with various advanced denoising networks, lifting the constraint of using BSNs in self-supervised settings. Additionally, we introduce a fixed sub-image sampling strategy to prevent pattern collapse during inference and ensure stability. To address Limitation 3, we explicitly differentiate checkerboard artifacts from real noise and develop a dedicated artifact remover to correct pixel discontinuities caused by sampling-based operations. This design preserves fine image details while reducing over-smoothing. Experiments on benchmark real-noise datasets and self-captured noisy images demonstrate the robustness and generalizability of our framework, achieving better performance over existing methods.
Xiao Liu 0022, Xiuya Shi, Yizhong Pan, Shuhang Gu, Wei Liu 0044, Chao Ren 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 ATD: Improved Transformer With Adaptive Token Dictionary for Image Restoration
abstract
Recently, Transformers have gained significant popularity in image restoration tasks such as image super-resolution and denoising, owing to their superior performance. However, balancing performance and computational burden remains a long-standing problem for transformer-based architectures. Due to the quadratic complexity of self-attention, existing methods often restrict attention to local windows, resulting in limited receptive field and suboptimal performance. To address this issue, we propose Adaptive Token Dictionary (ATD), a novel transformer-based architecture for image restoration that enables global dependency modeling with linear complexity relative to image size. The ATD model incorporates a learnable token dictionary, which summarizes external image priors (i.e., typical image structures) during the training process. To utilize this information, we introduce a token dictionary cross-attention (TDCA) mechanism that enhances the input features via interaction with the learned dictionary. Furthermore, we exploit the category information embedded in the TDCA attention maps to group input features into multiple categories, each representing a cluster of similar features across the image and serving as an attention group. We also integrate the learned category information into the feed-forward network to further improve feature fusion. ATD and its lightweight version ATD-light, achieve state-of-the-art performance on multiple image super-resolution benchmarks. Moreover, we develop ATD-U, a multi-scale variant of ATD, to address other image restoration tasks, including image denoising and JPEG compression artifacts removal. Extensive experiments demonstrate the superiority of out proposed models, both quantitatively and qualitatively.
Leheng Zhang, Yawei Li 0001, Xiaorui Zhao, Shuhang Gu
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Progressive Focused Transformer for Single Image Super-Resolution
abstract
Transformer-based methods have achieved remarkable results in image super-resolution tasks because they can capture non-local dependencies in low-quality input images. However, this feature-intensive modeling approach is computationally expensive because it calculates the similarities between numerous features that are irrelevant to the query features when obtaining attention weights. These unnecessary similarity calculations not only degrade the reconstruction performance but also introduce significant computational overhead. How to accurately identify the features that are important to the current query features and avoid similarity calculations between irrelevant features remains an urgent problem. To address this issue, we propose a novel and effective Progressive Focused Transformer (PFT) that links all isolated attention maps in the network through Progressive Focused Attention (PFA) to focus attention on the most important tokens. PFA not only enables the network to capture more critical similar features, but also significantly reduces the computational cost of the overall network by filtering out irrelevant features before calculating similarities. Extensive experiments demonstrate the effectiveness of the proposed method, achieving state-of-the-art performance on various single image super-resolution benchmarks.
Leheng Zhang, Shuhang Gu
CVPR4
2025 Learned Image Compression with Dictionary-based Entropy Model
abstract
Learned image compression methods have attracted great research interest and exhibited superior rate-distortion performance to the best classical image compression standards of the present. The entropy model plays a key role in learned image compression, which estimates the probability distribution of the latent representation for further entropy coding. Most existing methods employed hyper-prior and auto-regressive architectures to form their entropy models. However, they only aimed to explore the internal dependencies of latent representation while neglecting the importance of extracting prior from training data. In this work, we propose a novel entropy model named Dictionary-Based Cross Attention Entropy model, which introduces a learnable dictionary to summarize the typical structures occurring in the training dataset to enhance the entropy model. Extensive experimental results have demonstrated that the proposed model strikes a better balance between performance and latency, achieving state-of-the-art results on various benchmark datasets.
Jingbo Lu, Leheng Zhang, Mu Li 0005, Wen Li 0001, Shuhang Gu
CVPR6
2025 GeoDepth: From Point-to-Depth to Plane-to-Depth Modeling for Self-Supervised Monocular Depth Estimation
abstract
Self-supervised monocular depth estimation has long been treated as a point-wise prediction problem, where the depth of each pixel is usually estimated independently. However, artifacts are often observed in the estimated depth map, e.g., depth values for points located in the same region may jump dramatically. To address this issue, we propose a novel self-supervised monocular depth estimation framework called GeoDepth, where we explore the intrinsic geometric representation in 3D scenes for producing accurate and continuous depth maps. In particular, we model the complex 3D scene as a collection of planes with varying sizes, where each plane is characterized by a unique set of parameters, namely planar normal (indicating plane orientation) and planar offset (defining the perpendicular distance from the camera center to the plane). Under this modeling, points in the same plane are enforced to share a unique representation and their depth variations related only to pixel coordinates, thus this geometric relationship can be exploited to regularize the depth variations of these points. To this end, we design a structured plane generation module that introduces spatio-temporal geometric cues and the plane uniqueness principle to recover the correct scene plane representation. In addition, we develop a depth discontinuity module to identify depth discontinuity regions and subsequently optimize them. Our experiments on the KITTI and NYUv2 datasets demonstrate that GeoDepth achieves state-of-the-art performance, with additional tests on Make3D and ScanNet validating its generalization capabilities.
Haifeng Wu, Shuhang Gu, Lixin Duan, Wen Li 0001
CVPR2
2025 Robust Message Embedding via Attention Flow-Based Steganography
abstract
Image steganography can hide information in a host image and obtain a stego image that is perceptually indistinguishable from the original one. This technique has tremendous potential in scenarios like copyright protection and information retrospection. Some previous studies have proposed to enhance the robustness of the methods against image disturbances to increase their applicability. However, they generally cannot achieve a satisfying balance between the steganography quality and robustness. Instead of image-in-image steganography, we focus on the issue of message-in-image embedding that is robust to various real- world image distortions. This task aims to embed information into a natural image and the decoding result is required to be completely accurate, which increases the difficulty of data concealing and revealing. Inspired by the recent developments in transformer-based vision models, we discover that the tokenized representation of image is naturally suitable for steganography task. In this paper, we propose a novel message embedding framework, called Robust Message Steganography (RMSteg), which is competent to hide message via QR Code in a host image based on an normalizing flow-based model. The stego image derived by our method has imperceptible changes and the encoded message can be accurately restored even if the image is printed out and photographed. To our best knowledge, this is the first work that integrates the advantages of transformer models into normalizing flow. The code is available at https://github.com/huayuan4396/RMSteg.
Huayuan Ye, Shenzhuo Zhang, Shiqi Jiang 0001, Jing Liao 0001, Shuhang Gu, Dejun Zheng, Changbo Wang, Chenhui Li 0001
CVPR5
2025 Uncertainty-guided Perturbation for Image Super-Resolution Diffusion Model
abstract
Diffusion-based image super-resolution methods have demonstrated significant advantages over GAN-based approaches, particularly in terms of perceptual quality. Building upon a lengthy Markov chain, diffusion-based methods possess remarkable modeling capacity, enabling them to achieve outstanding performance in real-world scenarios. Unlike previous methods that focus on modifying the noise schedule or sampling process to enhance performance, our approach emphasizes the improved utilization of LR information. We find that different regions of the LR image can be viewed as corresponding to different timesteps in a diffusion process, where flat areas are closer to the target HR distribution but edge and texture regions are farther away. In these flat areas, applying a slight noise is more advantageous for the reconstruction. We associate this characteristic with uncertainty and propose to apply uncertainty estimate to guide region-specific noise level control, a technique we refer to as Uncertainty- guided Noise Weighting. Pixels with lower uncertainty (i.e., flat regions) receive reduced noise to preserve more LR information, therefore improving performance. Furthermore, we modify the network architecture of previous methods to develop our Uncertainty-guided Perturbation Super-Resolution (UPSR) model. Extensive experimental results demonstrate that, despite reduced model size and training overhead, the proposed UPSR method outperforms current state-of-the-art methods across various datasets, both quantitatively and qualitatively.
Leheng Zhang, Weiyi You, Kexuan Shi, Shuhang Gu
CVPR4
2025 Learning Pixel-Adaptive Multi-Layer Perceptrons for Real-Time Image Enhancement
abstract
Deep learning-based bilateral grid processing has emerged as a promising solution for image enhancement, inherently encoding spatial and intensity information while enabling efficient full-resolution processing through slicing operations. However, existing approaches are limited to linear affine transformations, hindering their ability to model complex color relationships. Meanwhile, while multi-layer perceptrons (MLPs) excel at non-linear mappings, traditional MLP-based methods employ globally shared parameters, which is hard to deal with localized variations. To overcome these dual challenges, we propose a Bilateral Grid-based Pixel-Adaptive Multi-layer Perceptron (BPAM) framework. Our approach synergizes the spatial modeling of bilateral grids with the non-linear capabilities of MLPs. Specifically, we generate bilateral grids containing MLP parameters, where each pixel dynamically retrieves its unique transformation parameters and obtain a distinct MLP for color mapping based on spatial coordinates and intensity values. In addition, we propose a novel grid decomposition strategy that categorizes MLP parameters into distinct types stored in separate subgrids. Multi-channel guidance maps are used to extract category-specific parameters from corresponding subgrids, ensuring effective utilization of color information during slicing while guiding precise parameter generation. Extensive experiments on public datasets demonstrate that our method outperforms state-of-the-art methods in performance while maintaining real-time processing capabilities.
Junyu Lou, Xiaorui Zhao, Kexuan Shi, Shuhang Gu
ICCV4
2025 Consistency Trajectory Matching for One-Step Generative Super-Resolution
Weiyi You, Leheng Zhang, Kexuan Shi, Shuhang Gu
ICCV6
2025 PerLDiff: Controllable Street View Synthesis Using Perspective-Layout Diffusion Model
Hualian Sheng, Sijia Cai, Bing Deng, Qiao Liang 0002, Wen Li 0001, Jieping Ye, Shuhang Gu
ICCV9
2025 Inductive Gradient Adjustment for Spectral Bias in Implicit Neural Representations
abstract
Implicit Neural Representations (INRs), as a versatile representation paradigm, have achieved success in various computer vision tasks. Due to the spectral bias of the vanilla multi-layer perceptrons (MLPs), existing methods focus on designing MLPs with sophisticated architectures or repurposing training techniques for highly accurate INRs. In this paper, we delve into the linear dynamics model of MLPs and theoretically identify the empirical Neural Tangent Kernel (eNTK) matrix as a reliable link between spectral bias and training dynamics. Based on this insight, we propose a practical **I**nductive **G**radient **A**djustment (**IGA**) method, which could purposefully improve the spectral bias via inductive generalization of eNTK-based gradient transformation matrix. Theoretical and empirical analyses validate impacts of IGA on spectral bias. Further, we evaluate our method on different INRs tasks with various INR architectures and compare to existing training techniques. The superior and consistent improvements clearly validate the advantage of our IGA. Armed with our gradient adjustment method, better INRs with more enhanced texture details and sharpened edges can be learned from data by tailored impacts on spectral bias. The codes are available at: [https://github.com/LabShuHangGU/IGA-INR](https://github.com/LabShuHangGU/IGA-INR).
Kexuan Shi, Hai Chen, Leheng Zhang, Shuhang Gu
ICML4
2025 Efficient Image Enhancement With a Diffusion-Based Frequency Prior
abstract
Due to the lack of appropriate priors, generating the content of dark regions remains a challenge in low-light image enhancement tasks. Currently, diffusion models employ robust image generation capabilities for enhancing low-light images. However, diffusion models require multiple iterations at the image feature level to generate details and content, which limits the speed. Moreover, the diffusion-based methods tend to generate unexpected artifacts in the degraded regions. To address these issues, we propose a Frequency Priors-guided Image Enhancement (FPIE) network, including a frequency prior generation network and an image restoration network. FPIE significantly accelerates inference by learning abstract prior with frequency domain constraints. Concretely, to learn compacted priors at the frequency domain, we introduce a joint training approach for the prior generation and restoration models to constrain the distribution of priors. Furthermore, to better utilize frequency-domain features for enhancing the network’s generation capabilities, a wavelet-based transformer block is introduced to produce intricate details and avoid the artifacts of the output. Extensive experimental results on the commonly used benchmarks demonstrate that our approach achieves state-of-the-art performances and well generalization to real-world images.
Qingsen Yan, Tao Hu 0013, Peng Wu 0015, Duwei Dai, Shuhang Gu, Wei Dong 0010, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Improved Implicit Neural Representation with Fourier Reparameterized Training
abstract
Implicit Neural Representation (INR) as a mighty representation paradigm has achieved success in various computer vision tasks recently. Due to the low-frequency bias issue of vanilla multi-layer perceptron (MLP), existing methods have investigated advanced techniques, such as positional encoding and periodic activation function, to improve the accuracy of INR. In this paper, we connect the network training bias with the reparameterization technique and theoretically prove that weight reparameterization could provide us a chance to alleviate the spectral bias of MLP. Based on our theoretical analysis, we propose a Fourier reparameterization method which learns coefficient matrix of fixed Fourier bases to compose the weights of MLP. We evaluate the proposed Fourier reparameterization method on different INR tasks with various MLP architectures, including vanilla MLP, MLP with positional encoding and MLP with advanced activation function, etc. The superiority approximation results on different MLP architectures clearly validate the advantage of our proposed method. Armed with our Fourier reparameterization method, better INR with more textures and less artifacts can be learned from the training data. The codes are available at https://github.com/LabShuHangGU/FR-INR.
Kexuan Shi, Shuhang Gu
CVPR3
2024 Transcending the Limit of Local Window: Advanced Super-Resolution Transformer with Adaptive Token Dictionary
abstract
Single Image Super-Resolution is a classic computer vision problem that involves estimating high-resolution (HR) images from low-resolution (LR) ones. Although deep neural networks (DNNs), especially Transformers for super-resolution, have seen significant advancements in recent years, challenges still remain, particularly in limited receptive field caused by window-based self-attention. To address these issues, we introduce a group of auxiliary Adaptive Token Dictionary to SR Transformer and establish an ATD-SR method. The introduced token dictionary could learn prior information from training data and adapt the learned prior to specific testing image through an adaptive refinement step. The refinement strategy could not only provide global information to all input tokens but also group image tokens into categories. Based on category partitions, we further propose a category-based self-attention mechanism designed to leverage distant but similar tokens for enhancing input features. The experimental results show that our method achieves the best performance on various single image super-resolution benchmarks.
Leheng Zhang, Yawei Li 0001, Xiaorui Zhao, Shuhang Gu
CVPR5
2024 Video Super-Resolution Transformer with Masked Inter&Intra-Frame Attention
abstract
Recently, Vision Transformer has achieved great success in recovering missing details in low-resolution sequences, i.e., the video super-resolution (VSR) task. Despite its su-periority in VSR accuracy, the heavy computational bur-den as well as the large memory footprint hinder the de-ployment of Transformer-based VSR models on constrained devices. In this paper, we address the above issue by proposing a novel feature-level masked processing frame-work: VSR with Masked Intra and inter-frame Attention (MIA-VSR). The core of MIA-VSR is leveraging feature-level temporal continuity between adjacent frames to re-duce redundant computations and make more rational use of previously enhanced SR features. Concretely, we propose an intra-frame and inter-frame attention block which takes the respective roles of past features and input features into consideration and only exploits previously enhanced fea-tures to provide supplementary information. In addition, an adaptive block-wise mask prediction module is developed to skip unimportant computations according to feature sim-ilarity between adjacent frames. We conduct detailed ab-lation studies to validate our contributions and compare the proposed method with recent state-of-the-art VSR approaches. The experimental results demonstrate that MIA-VSR improves the memory and computation efficiency over state-of-the-art methods, without trading off PSNR accuracy. The code is available at https://github.com/LabShuHangGU/MIA-VSR.
Leheng Zhang, Xiaorui Zhao, Keze Wang, Leida Li, Shuhang Gu
CVPR6
2024 Domain Prompt Learning Framework for Real Image Dehazing
abstract
Supervised dehazing models trained on synthetic datasets exhibit severe performance degradation in real-world scenarios due to the domain gap. Unsupervised methods are proposed to process real hazy images, while they suffer from the complex training procedure. In this paper, we present a universal domain prompt learning framework (DPLF) for boosting the performance of supervised dehazing models in real scenarios by introducing prompt learning and well-designed domain adapters. We train the learnable text prompts by CLIP feature alignment, which can discriminate between real hazy and clean images, and use these prompts as unsupervised text constraints. Notably, the distribution gap between synthetic and real haze can be regarded as the difference of the haze-relevant style domain. Motivated by this, we design the style domain prompt adapter to align features from synthetic and real haze domains. Extensive experiments on real-world datasets demonstrate significant performance improvement of the baseline dehazing models with our DPLF.
Kaihao Lin, Guoqing Wang 0001, Yuhui Wu 0001, Shuhang Gu, Xing Xu 0001, Yang Yang 0002
ICME4
2024 Stereo Risk: A Continuous Modeling Approach to Stereo Matching
abstract
We introduce Stereo Risk, a new deep-learning approach to solve the classical stereo-matching problem in computer vision. As it is well-known that stereo matching boils down to a per-pixel disparity estimation problem, the popular state-of-the-art stereo-matching approaches widely rely on regressing the scene disparity values, yet via discretization of scene disparity values. Such discretization often fails to capture the nuanced, continuous nature of scene depth. Stereo Risk departs from the conventional discretization approach by formulating the scene disparity as an optimal solution to a continuous risk minimization problem, hence the name "stereo risk". We demonstrate that $L^1$ minimization of the proposed continuous risk function enhances stereo-matching performance for deep networks, particularly for disparities with multi-modal probability distributions. Furthermore, to enable the end-to-end network training of the non-differentiable $L^1$ risk optimization, we exploited the implicit function theorem, ensuring a fully differentiable network. A comprehensive analysis demonstrates our method's theoretical soundness and superior performance over the state-of-the-art methods across various benchmark datasets, including KITTI 2012, KITTI 2015, ETH3D, SceneFlow, and Middlebury 2014.
Ce Liu 0004, Suryansh Kumar 0001, Shuhang Gu, Radu Timofte, Yao Yao 0008, Luc Van Gool
ICML3
2024 Causal Context Adjustment Loss for Learned Image Compression
abstract
In recent years, learned image compression (LIC) technologies have surpassed conventional methods notably in terms of rate-distortion (RD) performance. Most present learned techniques are VAE-based with an autoregressive entropy model, which obviously promotes the RD performance by utilizing the decoded causal context. However, extant methods are highly dependent on the fixed hand-crafted causal context. The question of how to guide the auto-encoder to generate a more effective causal context benefit for the autoregressive entropy models is worth exploring. In this paper, we make the first attempt in investigating the way to explicitly adjust the causal context with our proposed Causal Context Adjustment loss (CCA-loss). By imposing the CCA-loss, we enable the neural network to spontaneously adjust important information into the early stage of the autoregressive entropy model. Furthermore, as transformer technology develops remarkably, variants of which have been adopted by many state-of-the-art (SOTA) LIC techniques. The existing computing devices have not adapted the calculation of the attention mechanism well, which leads to a burden on computation quantity and inference latency. To overcome it, we establish a convolutional neural network (CNN) image compression model and adopt the unevenly channel-wise grouped strategy for high efficiency. Ultimately, the proposed CNN-based LIC network trained with our Causal Context Adjustment loss attains a great trade-off between inference latency and rate-distortion performance.
Minghao Han, Shiyin Jiang, Shengxi Li, Xin Deng 0002, Mai Xu, Ce Zhu, Shuhang Gu
NeurIPS7
2024 CrossHomo: Cross-Modality and Cross-Resolution Homography Estimation
abstract
Multi-modal homography estimation aims to spatially align the images from different modalities, which is quite challenging since both the image content and resolution are variant across modalities. In this paper, we introduce a novel framework namely CrossHomo to tackle this challenging problem. Our framework is motivated by two interesting findings which demonstrate the mutual benefits between image super-resolution and homography estimation. Based on these findings, we design a flexible multi-level homography estimation network to align the multi-modal images in a coarse-to-fine manner. Each level is composed of a multi-modal image super-resolution (MISR) module to shrink the resolution gap between different modalities, followed by a multi-modal homography estimation (MHE) module to predict the homography matrix. To the best of our knowledge, CrossHomo is the first attempt to address the homography estimation problem with both modality and resolution discrepancy. Extensive experimental results show that our CrossHomo can achieve high registration accuracy on various multi-modal datasets with different resolution gaps. In addition, the network has high efficiency in terms of both model complexity and running speed.
Xin Deng 0002, Enpeng Liu, Shengxi Li, Shuhang Gu, Mai Xu
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Towards a Flexible Semantic Guided Model for Single Image Enhancement and Restoration
abstract
Low-light image enhancement (LLIE) investigates how to improve the brightness of an image captured in illumination-insufficient environments. The majority of existing methods enhance low-light images in a global and uniform manner, without taking into account the semantic information of different regions. Consequently, a network may easily deviate from the original color of local regions. To address this issue, we propose a semantic-aware knowledge-guided framework (SKF) that can assist a low-light enhancement model in learning rich and diverse priors encapsulated in a semantic segmentation model. We concentrate on incorporating semantic knowledge from three key aspects: a semantic-aware embedding module that adaptively integrates semantic priors in feature representation space, a semantic-guided color histogram loss that preserves color consistency of various instances, and a semantic-guided adversarial loss that produces more natural textures by semantic priors. Our SKF is appealing in acting as a general framework in the LLIE task. We further present a refined framework SKF++ with two new techniques: (a) Extra convolutional branch for intra-class illumination and color recovery through extracting local information and (b) Equalization-based histogram transformation for contrast enhancement and high dynamic range adjustment. Extensive experiments on various benchmarks of LLIE task and other image processing tasks show that models equipped with the SKF/SKF++ significantly outperform the baselines and our SKF/SKF++ generalizes to different models and scenes well. Besides, the potential benefits of our method in face detection and semantic segmentation in low-light conditions are discussed.
Yuhui Wu 0001, Guoqing Wang 0001, Shaochong Liu, Yang Yang 0002, Wei Liu 0005, Xiongxin Tang, Shuhang Gu, Chongyi Li, Heng Tao Shen
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 Single Image Depth Prediction Made Better: A Multivariate Gaussian Take
abstract
Neural-network-based single image depth prediction (SIDP) is a challenging task where the goal is to predict the scene's per-pixel depth at test time. Since the problem, by definition, is illposed, the fundamental goal is to come up with an approach that can reliably model the scene depth from a set of training examples. In the pursuit of perfect depth estimation, most existing state-of-the-art learning techniques predict a single scalar depth value per-pixel. Yet, it is well-known that the trained model has accuracy limits and can predict imprecise depth. Therefore, an SIDP approach must be mindful of the expected depth variations in the model's prediction at test time. Accordingly, we introduce an approach that performs continuous modeling of per-pixel depth, where we can predict and reason about the per-pixel depth and its distribution. To this end, we model per-pixel scene depth using a multivariate Gaussian distribution. Moreover, contrary to the existing uncertainty modeling methods—in the same spirit, where per-pixel depth is assumed to be independent, we introduce per-pixel covariance modeling that encodes its depth dependency w.r.t. all the scene points. Unfortunately, per-pixel depth covariance modeling leads to a computationally expensive continuous loss function, which we solve efficiently using the learned low-rank approximation of the overall covariance matrix. Notably, when tested on benchmark datasets such as KITTI, NYU, and SUN-RGB-D, the SIDP model obtained by optimizing our loss function shows state-of-the-art results. Our method's accuracy (named MG) is among the top on the KITTI depth-prediction benchmark leaderboard11http://www.cvlibs.net/datasets/kitti/eval–depth.php?benchmark=depth–prediction.
Ce Liu 0004, Suryansh Kumar 0001, Shuhang Gu, Radu Timofte, Luc Van Gool
CVPR3
2023 VA-DepthNet: A Variational Approach to Single Image Depth Prediction
Ce Liu 0004, Suryansh Kumar 0001, Shuhang Gu, Radu Timofte, Luc Van Gool
ICLR3
2023 Learning Motion Refinement for Unsupervised Face Animation
abstract
Unsupervised face animation aims to generate a human face video based on the appearance of a source image, mimicking the motion from a driving video. Existing methods typically adopted a prior-based motion model (e.g., the local affine motion model or the local thin-plate-spline motion model). While it is able to capture the coarse facial motion, artifacts can often be observed around the tiny motion in local areas (e.g., lips and eyes), due to the limited ability of these methods to model the finer facial motions. In this work, we design a new unsupervised face animation approach to learn simultaneously the coarse and finer motions. In particular, while exploiting the local affine motion model to learn the global coarse facial motion, we design a novel motion refinement module to compensate for the local affine motion model for modeling finer face motions in local areas. The motion refinement is learned from the dense correlation between the source and driving images. Specifically, we first construct a structure correlation volume based on the keypoint features of the source and driving images. Then, we train a model to generate the tiny facial motions iteratively from low to high resolution. The learned motion refinements are combined with the coarse motion to generate the new image. Extensive experiments on widely used benchmarks demonstrate that our method achieves the best results among state-of-the-art baselines.
Jiale Tao, Shuhang Gu, Wen Li 0001, Lixin Duan
NeurIPS2
2023 Measuring Perceptual Color Differences of Smartphone Photographs
abstract
Measuring perceptual color differences (CDs) is of great importance in modern smartphone photography. Despite the long history, most CD measures have been constrained by psychophysical data of homogeneous color patches or a limited number of simplistic natural photographic images. It is thus questionable whether existing CD measures generalize in the age of smartphone photography characterized by greater content complexities and learning-based image signal processors. In this article, we put together so far the largest image dataset for perceptual CD assessment, in which the photographic images are 1) captured by six flagship smartphones, 2) altered by Photoshop, 3) post-processed by built-in filters of the smartphones, and 4) reproduced with incorrect color profiles. We then conduct a large-scale psychophysical experiment to gather perceptual CDs of 30,000 image pairs in a carefully controlled laboratory environment. Based on the newly established dataset, we make one of the first attempts to construct an end-to-end learnable CD formula based on a lightweight neural network, as a generalization of several previous metrics. Extensive experiments demonstrate that the optimized formula outperforms 33 existing CD measures by a large margin, offers reasonable local CD maps without the use of dense supervision, generalizes well to homogeneous color patch data, and empirically behaves as a proper metric in the mathematical sense. Our dataset and code are publicly available at https://github.com/hellooks/CDNet.
Zhihua Wang 0002, Keshuo Xu, Yang Yang 0201, Jianlei Dong, Shuhang Gu, Lihao Xu, Yuming Fang 0001, Kede Ma
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Revisiting Random Channel Pruning for Neural Network Compression
abstract
Channel (or 3D filter) pruning serves as an effective way to accelerate the inference of neural networks. There has been a flurry of algorithms that try to solve this practical problem, each being claimed effective in some ways. Yet, a benchmark to compare those algorithms directly is lacking, mainly due to the complexity of the algorithms and some custom settings such as the particular network configuration or training procedure. A fair benchmark is important for the further development of channel pruning. Meanwhile, recent investigations reveal that the channel configurations discovered by pruning algorithms are at least as important as the pre-trained weights. This gives channel pruning a new role, namely searching the optimal channel configuration. In this paper, we try to determine the channel configuration of the pruned models by random search. The proposed approach provides a new way to compare different methods, namely how well they behave compared with random pruning. We show that this simple strategy works quite well compared with other channel pruning methods. We also show that under this setting, there are surprisingly no clear winners among different channel importance evaluation methods, which then may tilt the research efforts into advanced channel configuration searching methods. Code will be released at https://github.com/ofsoundof/random_channel_pruning.
Yawei Li 0001, Kamil Adamczewski, Wen Li 0001, Shuhang Gu, Radu Timofte, Luc Van Gool
CVPR4
2022 Restore Globally, Refine Locally: A Mask-Guided Scheme to Accelerate Super-Resolution Networks
Xiaotao Hu, Jun Xu 0019, Shuhang Gu, Ming-Ming Cheng, Li Liu 0004
ECCV (19)3
2022 Exploiting Intra-Slice and Inter-Slice Redundancy for Learning-Based Lossless Volumetric Image Compression
abstract
3D volumetric image processing has attracted increasing attention in the last decades, in which one major research area is to develop efficient lossless volumetric image compression techniques to better store and transmit such images with massive amount of information. In this work, we propose the first end-to-end optimized learning framework for losslessly compressing 3D volumetric data. Our approach builds upon a hierarchical compression scheme by additionally introducing the intra-slice auxiliary features and estimating the entropy model based on both intra-slice and inter-slice latent priors. Specifically, we first extract the hierarchical intra-slice auxiliary features through multi-scale feature extraction modules. Then, an Intra-slice and Inter-slice Conditional Entropy Coding module is proposed to fuse the intra-slice and inter-slice information from different scales as the context information. Based on such context information, we can predict the distributions for both intra-slice auxiliary features and the slice images. To further improve the lossless compression performance, we also introduce two new gating mechanisms called Intra-Gate and Inter-Gate to generate the optimal feature representations for better information fusion. Eventually, we can produce the bitstream for losslessly compressing volumetric images based on the estimated entropy model. Different from the existing lossless volumetric image codecs, our end-to-end optimized framework jointly learns both intra-slice auxiliary features at different scales for each slice and inter-slice latent features from previously encoded slices for better entropy estimation. The extensive experimental results indicate that our framework outperforms the state-of-the-art hand-crafted lossless volumetric image codecs (e.g., JP3D) and the learning-based lossless image compression method on four volumetric image benchmarks for losslessly compressing both 3D Medical Images and Hyper-Spectral Images.
Shuhang Gu, Guo Lu, Dong Xu 0001
IEEE Trans. Image Process.2
2022 End-to-End Optimized 360° Image Compression
abstract
The 360° image that offers a 360-degree scenario of the world is widely used in virtual reality and has drawn increasing attention. In 360° image compression, the spherical image is first transformed into a planar image with a projection such as equirectangular projection (ERP) and then saved with the existing codecs. The ERP images that represent different circles of latitude with the same number of pixels suffer from the unbalance sampling problem, resulting in inefficiency using planar compression methods, especially for the deep neural network (DNN) based codecs. To tackle this problem, we introduce a latitude adaptive coding scheme for DNNs by allocating variant numbers of codes for different regions according to the latitude on the sphere. Specifically, taking both the number of allocated codes for each region and their entropy into consideration, we introduce a flexible regional adaptive rate loss for region-wise rate controlling. Latitude adaptive constraints are then introduced to prevent spending too many codes on the over-sampling regions. Furthermore, we introduce viewport-based distortion loss by calculating the average distortion on a set of viewports. We optimize and test our model on a large 360° dataset containing 19,790 images collected from the Internet. The experiment results demonstrate the superiority of the proposed latitude adaptive coding scheme. On the whole, our model outperforms the existing image compression standards, including JPEG, JPEG2000, HEVC Intra Coding, and VVC Intra Coding, and helps to save around 15% bits compared to the baseline learned image compression model for planar images.
Mu Li 0005, Jinxing Li 0003, Shuhang Gu, Feng Wu 0001, David Zhang 0001
IEEE Trans. Image Process.3
2022 VPU: A Video-Based Point Cloud Upsampling Framework
abstract
In this work, we propose a new patch-based framework called VPU for the video-based point cloud upsampling task by effectively exploiting temporal dependency among multiple consecutive point cloud frames, in which each frame consists of a set of unordered, sparse and irregular 3D points. Rather than adopting the sophisticated motion estimation strategy in video analysis, we propose a new spatio-temporal aggregation (STA) module to effectively extract, align and aggregate rich local geometric clues from consecutive frames at the feature level. By more reliably summarizing spatio-temporally consistent and complementary knowledge from multiple frames in the resultant local structural features, our method better infers the local geometry distributions at the current frame. In addition, our STA module can be readily incorporated with various existing single frame-based point upsampling methods (e.g., PU-Net, MPU, PU-GAN and PU-GCN). Comprehensive experiments on multiple point cloud sequence datasets demonstrate our video-based point cloud upsampling framework achieves substantial performance improvement over its single frame-based counterparts.
Kaisiyuan Wang, Lu Sheng, Shuhang Gu, Dong Xu 0001
IEEE Trans. Image Process.3
2021 Deep Line Encoding for Monocular 3D Object Detection and Depth Prediction
Ce Liu 0004, Shuhang Gu, Luc Van Gool, Radu Timofte
BMVC2
2021 The Heterogeneity Hypothesis: Finding Layer-Wise Differentiated Network Architectures
abstract
In this paper, we tackle the problem of convolutional neural network design. Instead of focusing on the design of the overall architecture, we investigate a design space that is usually overlooked, i.e. adjusting the channel configurations of predefined networks. We find that this adjustment can be achieved by shrinking widened baseline networks and leads to superior performance. Based on that, we articulate the "heterogeneity hypothesis": with the same training protocol, there exists a layer-wise differentiated net-work architecture (LW-DNA) that can outperform the original network with regular channel configurations but with a lower level of model complexity.The LW-DNA models are identified without extra computational cost or training time compared with the original network. This constraint leads to controlled experiments which direct the focus to the importance of layer-wise specific channel configurations. LW-DNA models come with advantages related to overfitting, i.e. the relative relationship between model complexity and dataset size. Experiments are conducted on various networks and datasets for image classification, visual tracking and image restoration. The resultant LW-DNA models consistently outperform the baseline models. Code is available at https://github.com/ofsoundof/Heterogeneity_Hypothesis.git.
Yawei Li 0001, Wen Li 0001, Martin Danelljan, Kai Zhang 0008, Shuhang Gu, Luc Van Gool, Radu Timofte
CVPR5
2021 Flow-Based Kernel Prior With Application to Blind Super-Resolution
abstract
Kernel estimation is generally one of the key problems for blind image super-resolution (SR). Recently, Double-DIP proposes to model the kernel via a network architecture prior, while KernelGAN employs the deep linear network and several regularization losses to constrain the kernel space. However, they fail to fully exploit the general SR kernel assumption that anisotropic Gaussian kernels are sufficient for image SR. To address this issue, this paper proposes a normalizing flow-based kernel prior (FKP) for kernel modeling. By learning an invertible mapping between the anisotropic Gaussian kernel distribution and a tractable latent distribution, FKP can be easily used to replace the kernel modeling modules of Double-DIP and KernelGAN. Specifically, FKP optimizes the kernel in the latent space rather than the network parameter space, which allows it to generate reasonable kernel initialization, traverse the learned kernel manifold and improve the optimization stability. Extensive experiments on synthetic and real-world images demonstrate that the proposed FKP can significantly improve the kernel estimation accuracy with less parameters, runtime and memory usage, leading to state-of-the-art blind SR results.
Jingyun Liang, Kai Zhang 0008, Shuhang Gu, Luc Van Gool, Radu Timofte
CVPR3
2021 Unsupervised Real-World Image Super Resolution via Domain-Distance Aware Training
abstract
These days, unsupervised super-resolution (SR) is soaring due to its practical and promising potential in real scenarios. The philosophy of off-the-shelf approaches lies in the augmentation of unpaired data, i.e. first generating synthetic low-resolution (LR) images ${\mathcal{Y}^g}$ corresponding to real-world high-resolution (HR) images ${\mathcal{X}^r}$ in the real-world LR domain ${\mathcal{Y}^r}$, and then utilizing the pseudo pairs $\left\{ {{\mathcal{Y}^g},{\mathcal{X}^r}} \right\}$ for training in a supervised manner. Unfortunately, since image translation itself is an extremely challenging task, the SR performance of these approaches is severely limited by the domain gap between generated synthetic LR images and real LR images. In this paper, we propose a novel domain-distance aware super-resolution (DASR) approach for unsupervised real-world image SR. The domain gap between training data (e.g. ${\mathcal{Y}^g}$) and testing data (e.g. ${\mathcal{Y}^r}$) is addressed with our domain-gap aware training and domain-distance weighted supervision strategies. Domain-gap aware training takes additional benefit from real data in the target domain while domain-distance weighted supervision brings forward the more rational use of labeled source domain data. The proposed method is validated on synthetic and real datasets and the experimental results show that DASR consistently outperforms state-of-the-art unsupervised SR approaches in generating SR outputs with more realistic and natural textures. Codes are available at https://github.com/ShuhangGu/DASR.
Yunxuan Wei, Shuhang Gu, Yawei Li 0001, Radu Timofte, Longcun Jin, Hengjie Song
CVPR2
2021 Improving Facial Attribute Recognition by Group and Graph Learning
abstract
Exploiting the relationships between attributes is a key challenge for improving multiple facial attribute recognition. In this work, we are concerned with two types of correlations that are spatial and non-spatial relationships. For the spatial correlation, we aggregate attributes with spatial similarity into a part-based group and then introduce a Group Attention Learning to generate the group attention and the part-based group feature. On the other hand, to discover the non-spatial relationship, we model a group-based Graph Correlation Learning to explore affinities of predefined part-based groups. We utilize such affinity information to control the communication between all groups and then refine the learned group features. Overall, we propose a unified network called Multi-scale Group and Graph Network. It incorporates these two newly proposed learning strategies and produces coarse-to-fine graph-based group features for improving facial attribute recognition. Comprehensive experiments demonstrate that our approach outperforms the state-of-the-art methods.
Shuhang Gu, Feng Zhu 0006, Rui Zhao 0001
ICME2
2021 Unsupervised Multimodal Video-to-Video Translation via Self-Supervised Learning
abstract
Existing unsupervised video-to-video translation methods fail to produce translated videos which are frame-wise realistic, semantic information preserving and video-level consistent. In this work, we propose UVIT, a novel unsupervised video-to-video translation model. Our model decomposes the style and the content, uses the specialized encoder-decoder structure and propagates the inter-frame information through bidirectional recurrent neural network (RNN) units. The style-content decomposition mechanism enables us to achieve style consistent video translation results as well as provides us with a good interface for modality flexible translation. In addition, by changing the input frames and style codes incorporated in our translation, we propose a video interpolation loss, which captures temporal information within the sequence to train our building blocks in a self-supervised manner. Our model can produce photo-realistic, spatio-temporal consistent translated videos in a multimodal way. Subjective and objective experimental results validate the superiority of our model over existing methods.
Kangning Liu, Shuhang Gu, Andrés Romero, Radu Timofte
WACV2
2021 You Only Look Yourself: Unsupervised and Untrained Single Image Dehazing Neural Network
Boyun Li, Yuanbiao Gou, Shuhang Gu, Zitao Liu 0001, Joey Tianyi Zhou, Xi Peng 0001
Int. J. Comput. Vis.3
2021 Learning Content-Weighted Deep Image Compression
abstract
Learning-based lossy image compression usually involves the joint optimization of rate-distortion performance, and requires to cope with the spatial variation of image content and contextual dependence among learned codes. Traditional entropy models can spatially adapt the local bit rate based on the image content, but usually are limited in exploiting context in code space. On the other hand, most deep context models are computationally very expensive and cannot efficiently perform decoding over the symbols in parallel. In this paper, we present a content-weighted encoder-decoder model, where the channel-wise multi-valued quantization is deployed for the discretization of the encoder features, and an importance map subnet is introduced to generate the importance masks for spatially varying code pruning. Consequently, the summation of importance masks can serve as an upper bound of the length of bitstream. Furthermore, the quantized representations of the learned code and importance map are still spatially dependent, which can be losslessly compressed using arithmetic coding. To compress the codes effectively and efficiently, we propose an upper-triangular masked convolutional network (triuMCN) for large context modeling. Experiments show that the proposed method can produce visually much better results, and performs favorably against deep and traditional lossy image compression approaches.
Mu Li 0005, Wangmeng Zuo, Shuhang Gu, Jane You, David Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Sequential Point Cloud Upsampling by Exploiting Multi-Scale Temporal Dependency
abstract
In this work, we propose a new sequential point cloud upsampling method called SPU, which aims to upsample sparse, non-uniform, and orderless point cloud sequences by effectively exploiting rich and complementary temporal dependency from multiple inputs. Specifically, these inputs include a set of multi-scale short-term features from the 3D points in three consecutive frames (i.e., the previous/current/subsequent frame) and a long-term latent representation accumulated throughout the point cloud sequence. Considering that these temporal clues are not well aligned in the coordinate space, we propose a new temporal alignment module (TAM) based on the cross-attention mechanism to transform each individual feature into the feature space of the current frame. We also propose a new gating mechanism to learn the optimal weights for these transformed features, based on which the transformed features can be effectively aggregated as the final fused feature. The fused feature can be readily fed into the existing single frame-based point cloud upsampling methods (e.g., PU-Net, MPU and PU-GAN) to generate the dense point cloud for the current frame. Comprehensive experiments on three benchmark datasets DYNA, COMA, and MSR Action3D demonstrate the effectiveness of our method for upsampling point cloud sequences.
Kaisiyuan Wang, Lu Sheng, Shuhang Gu, Dong Xu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 Deep Coupled Feedback Network for Joint Exposure Fusion and Image Super-Resolution
abstract
Nowadays, people are getting used to taking photos to record their daily life, however, the photos are actually not consistent with the real natural scenes. The two main differences are that the photos tend to have low dynamic range (LDR) and low resolution (LR), due to the inherent imaging limitations of cameras. The multi-exposure image fusion (MEF) and image super-resolution (SR) are two widely-used techniques to address these two issues. However, they are usually treated as independent researches. In this paper, we propose a deep Coupled Feedback Network (CF-Net) to achieve MEF and SR simultaneously. Given a pair of extremely over-exposed and under-exposed LDR images with low-resolution, our CF-Net is able to generate an image with both high dynamic range (HDR) and high-resolution. Specifically, the CF-Net is composed of two coupled recursive sub-networks, with LR over-exposed and under-exposed images as inputs, respectively. Each sub-network consists of one feature extraction block (FEB), one super-resolution block (SRB) and several coupled feedback blocks (CFB). The FEB and SRB are to extract high-level features from the input LDR image, which are required to be helpful for resolution enhancement. The CFB is arranged after SRB, and its role is to absorb the learned features from the SRBs of the two sub-networks, so that it can produce a high-resolution HDR image. We have a series of CFBs in order to progressively refine the fused high-resolution HDR image. Extensive experimental results show that our CF-Net drastically outperforms other state-of-the-art methods in terms of both SR accuracy and fusion performance. The software code is available here https://github.com/ytZhang99/CF-Net.
Xin Deng 0002, Mai Xu, Shuhang Gu, Yiping Duan
IEEE Trans. Image Process.4
2021 Layer-Output Guided Complementary Attention Learning for Image Defocus Blur Detection
abstract
Defocus blur detection (DBD), which has been widely applied to various fields, aims to detect the out-of-focus or in-focus pixels from a single image. Despite the fact that the deep learning based methods applied to DBD have outperformed the hand-crafted feature based methods, the performance cannot still meet our requirement. In this paper, a novel network is established for DBD. Unlike existing methods which only learn the projection from the in-focus part to the ground-truth, both in-focus and out-of-focus pixels, which are completely and symmetrically complementary, are taken into account. Specifically, two symmetric branches are designed to jointly estimate the probability of focus and defocus pixels, respectively. Due to their complementary constraint, each layer in a branch is affected by an attention obtained from another branch, effectively learning the detailed information which may be ignored in one branch. The feature maps from these two branches are then passed through a unique fusion block to simultaneously get the two-channel output measured by a complementary loss. Additionally, instead of estimating only one binary map from a specific layer, each layer is encouraged to estimate the ground truth to guide the binary map estimation in its linked shallower layer followed by a top-to-bottom combination strategy, gradually exploiting the global and local information. Experimental results on released datasets demonstrate that our proposed method remarkably outperforms state-of-the-art algorithms.
Jinxing Li 0003, Lingxiao Yang, Shuhang Gu, Guangming Lu 0002, Yong Xu 0001, David Zhang 0001
IEEE Trans. Image Process.4
2020 Group Sparsity: The Hinge Between Filter Pruning and Decomposition for Network Compression
abstract
In this paper, we analyze two popular network compression techniques, i.e. filter pruning and low-rank decomposition, in a unified sense. By simply changing the way the sparsity regularization is enforced, filter pruning and low-rank decomposition can be derived accordingly. This provides another flexible choice for network compression because the techniques complement each other. For example, in popular network architectures with shortcut connections (e.g. ResNet), filter pruning cannot deal with the last convolutional layer in a ResBlock while the low-rank decomposition methods can. In addition, we propose to compress the whole network jointly instead of in a layer-wise manner. Our approach proves its potential as it compares favorably to the state-of-the-art on several benchmarks. Code is available at https://github.com/ofsoundof/group_sparsity.
Yawei Li 0001, Shuhang Gu, Christoph Mayer 0007, Luc Van Gool, Radu Timofte
CVPR2
2020 Multi-Domain Learning for Accurate and Few-Shot Color Constancy
abstract
Color constancy is an important process in camera pipeline to remove the color bias of captured image caused by scene illumination. Recently, significant improvements in color constancy accuracy have been achieved by using deep neural networks (DNNs). However, existing DNNbased color constancy methods learn distinct mappings for different cameras, which require a costly data acquisition process for each camera device. In this paper, we start a pioneer work to introduce multi-domain learning to color constancy area. For different camera devices, we train a branch of networks which share the same feature extractor and illuminant estimator, and only employ a camera-specific channel re-weighting module to adapt to the camera-specific characteristics. Such a multi-domain learning strategy enables us to take benefit from crossdevice training data. The proposed multi-domain learning color constancy method achieved state-of-the-art performance on three commonly used benchmark datasets. Furthermore, we also validate the proposed method in a fewshot color constancy setting. Given a new unseen device with limited number of training samples, our method is capable of delivering accurate color constancy by merely learning the camera-specific parameters from the few-shot dataset. Our project page is publicly available at https://github.com/msxiaojin/MDLCC.
Shuhang Gu, Lei Zhang 0006
CVPR2
2020 Improving Deep Video Compression by Resolution-Adaptive Flow Coding
Dong Xu 0001, Guo Lu, Wanli Ouyang, Shuhang Gu
ECCV (2)6
2020 Video Super-Resolution with Recurrent Structure-Detail Network
Takashi Isobe, Xu Jia 0012, Shuhang Gu, Songjiang Li, Shengjin Wang, Qi Tian 0001
ECCV (12)3
2020 DHP: Differentiable Meta Pruning via HyperNetworks
Yawei Li 0001, Shuhang Gu, Kai Zhang 0008, Luc Van Gool, Radu Timofte
ECCV (8)2
2020 Learned Dynamic Guidance for Depth Image Reconstruction
abstract
The depth images acquired by consumer depth sensors (e.g., Kinect and ToF) usually are of low resolution and insufficient quality. One natural solution is to incorporate a high resolution RGB camera and exploit the statistical correlation of its data and depth. In recent years, both optimization-based and learning-based approaches have been proposed to deal with the guided depth reconstruction problems. In this paper, we introduce a weighted analysis sparse representation (WASR) model for guided depth image enhancement, which can be considered a generalized formulation of a wide range of previous optimization-based models. We unfold the optimization by the WASR model and conduct guided depth reconstruction with dynamically changed stage-wise operations. Such a guidance strategy enables us to dynamically adjust the stage-wise operations that update the depth image, thus improving the reconstruction quality and speed. To learn the stage-wise operations in a task-driven manner, we propose two parameterizations and their corresponding methods: dynamic guidance with Gaussian RBF nonlinearity parameterization (DG-RBF) and dynamic guidance with CNN nonlinearity parameterization (DG-CNN). The network structures of the proposed DG-RBF and DG-CNN methods are designed with the the objective function of our WASR model in mind and the optimal network parameters are learned from paired training data. Such optimization-inspired network architectures enable our models to leverage the previous expertise as well as take benefit from training data. The effectiveness is validated for guided depth image super-resolution and for realistic depth image reconstruction tasks using standard benchmarks. Our DG-RBF and DG-CNN methods achieve the best quantitative results (RMSE) and better visual quality than the state-of-the-art approaches at the time of writing. The code is available at https://github.com/ShuhangGu/GuidedDepthSR.
Shuhang Gu, Shi Guo, Wangmeng Zuo, Yunjin Chen, Radu Timofte, Luc Van Gool, Lei Zhang 0006
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Self-Guided Network for Fast Image Denoising
abstract
During the past years, tremendous advances in image restoration tasks have been achieved using highly complex neural networks. Despite their good restoration performance, the heavy computational burden hinders the deployment of these networks on constrained devices, \eg smart phones and consumer electronic products. To tackle this problem, we propose a self-guided network (SGN), which adopts a top-down self-guidance architecture to better exploit image multi-scale information. SGN directly generates multi-resolution inputs with the shuffling operation. Large-scale contextual information extracted at low resolution is gradually propagated into the higher resolution sub-networks to guide the feature extraction processes at these scales. Such a self-guidance strategy enables SGN to efficiently incorporate multi-scale information and extract good local features to recover noisy images. We validate the effectiveness of SGN through extensive experiments. The experimental results demonstrate that SGN greatly improves the memory and runtime efficiency over state-of-the-art efficient methods, without trading off PSNR accuracy.
Shuhang Gu, Yawei Li 0001, Luc Van Gool, Radu Timofte
ICCV1
2019 Fast Image Restoration With Multi-Bin Trainable Linear Units
abstract
Tremendous advances in image restoration tasks such as denoising and super-resolution have been achieved using neural networks. Such approaches generally employ very deep architectures, large number of parameters, large receptive fields and high nonlinear modeling capacity. In order to obtain efficient and fast image restoration networks one should improve upon the above mentioned requirements. In this paper we propose a novel activation function, the multi-bin trainable linear unit (MTLU), for increasing the nonlinear modeling capacity together with lighter and shallower networks. We validate the proposed fast image restoration networks for image denoising (FDnet) and super-resolution (FSRnet) on standard benchmarks. We achieve large improvements in both memory and runtime over current state-of-the-art for comparable or better PSNR accuracies.
Shuhang Gu, Wen Li 0001, Luc Van Gool, Radu Timofte
ICCV1
2019 Learning Filter Basis for Convolutional Neural Network Compression
abstract
Convolutional neural networks (CNNs) based solutions have achieved state-of-the-art performances for many computer vision tasks, including classification and super-resolution of images. Usually the success of these methods comes with a cost of millions of parameters due to stacking deep convolutional layers. Moreover, quite a large number of filters are also used for a single convolutional layer, which exaggerates the parameter burden of current methods. Thus, in this paper, we try to reduce the number of parameters of CNNs by learning a basis of the filters in convolutional layers. For the forward pass, the learned basis is used to approximate the original filters and then used as parameters for the convolutional layers. We validate our proposed solution for multiple CNN architectures on image classification and image super-resolution benchmarks and compare favorably to the existing state-of-the-art in terms of reduction of parameters and preservation of accuracy. Code is available at https://github.com/ofsoundof/learning_filter_basis.
Yawei Li 0001, Shuhang Gu, Luc Van Gool, Radu Timofte
ICCV2
2018 Learning Latent Patterns in Molecular Data for Explainable Drug Side Effects Prediction
Pengwei Hu 0001, Zhu-Hong You, Tiantian He 0001, Shaochun Li, Shuhang Gu, Keith C. C. Chan
BIBM5
2018 Video Rain Streak Removal by Multiscale Convolutional Sparse Coding
abstract
Videos captured by outdoor surveillance equipments sometimes contain unexpected rain streaks, which brings difficulty in subsequent video processing tasks. Rain streak removal from a video is thus an important topic in recent computer vision research. In this paper, we raise two intrinsic characteristics specifically possessed by rain streaks. Firstly, the rain streaks in a video contain repetitive local patterns sparsely scattered over different positions of the video. Secondly, the rain streaks are with multiscale configurations due to their occurrence on positions with different distances to the cameras. Based on such understanding, we specifically formulate both characteristics into a multiscale convolutional sparse coding (MS-CSC) model for the video rain streak removal task. Specifically, we use multiple convolutional filters convolved on the sparse feature maps to deliver the former characteristic, and further use multiscale filters to represent different scales of rain streaks. Such a new encoding manner makes the proposed method capable of properly extracting rain streaks from videos, thus getting fine video deraining effects. Experiments implemented on synthetic and real videos verify the superiority of the proposed method, as compared with the state-of-the-art ones along this research line, both visually and quantitatively.
Minghan Li 0001, Qi Xie 0002, Qian Zhao 0002, Wei Wei 0006, Shuhang Gu, Deyu Meng
CVPR5
2018 Learning Convolutional Networks for Content-Weighted Image Compression
abstract
Lossy image compression is generally formulated as a joint rate-distortion optimization problem to learn encoder, quantizer, and decoder. Due to the non-differentiable quantizer and discrete entropy estimation, it is very challenging to develop a convolutional network (CNN)-based image compression system. In this paper, motivated by that the local information content is spatially variant in an image, we suggest that: (i) the bit rate of the different parts of the image is adapted to local content, and (ii) the content-aware bit rate is allocated under the guidance of a content-weighted importance map. The sum of the importance map can thus serve as a continuous alternative of discrete entropy estimation to control compression rate. The binarizer is adopted to quantize the output of encoder and a proxy function is introduced for approximating binary operation in backward propagation to make it differentiable. The encoder, decoder, binarizer and importance map can be jointly optimized in an end-to-end manner. And a convolutional entropy encoder is further presented for lossless compression of importance map and binary codes. In low bit rate image compression, experiments show that our system significantly outperforms JPEG and JPEG 2000 by structural similarity (SSIM) index, and can produce the much better visual result with sharp edges, rich textures, and fewer artifacts.
Mu Li 0005, Wangmeng Zuo, Shuhang Gu, Debin Zhao, David Zhang 0001
CVPR3
2018 Integrating Local and Non-local Denoiser Priors for Image Restoration
abstract
Image local structural prior and non-local self-similarity (NSS) prior are two categories of priors which have been commonly used for solving the ill-posed image restoration problem. As they exploit different properties of the natural images, it is interesting to investigate whether the two categories of priors can be combined to achieve better restoration performance. Inspired by recently proposed Regularization by denoising [1] idea, we propose LNIR which incorporates a Local CNN denoiser prior and a NSS-based denoiser prior implicitly for Image Restoration. Our experimental results on the image deblurring and super-resolution tasks demonstrate the effectiveness of the proposed method. The proposed LNIR algorithm can not only flexibly adapt to different restoration tasks, but also delivers state-of-the-art restoration results.
Shuhang Gu, Radu Timofte, Luc Van Gool
ICPR1
2018 Online multi-layer dictionary pair learning for visual classification
Shuhang Gu, Jianrui Cai
Expert Syst. Appl.3
2018 Learning a Deep Single Image Contrast Enhancer from Multi-Exposure Images
abstract
Due to the poor lighting condition and limited dynamic range of digital imaging devices, the recorded images are often under-/over-exposed and with low contrast. Most of previous single image contrast enhancement (SICE) methods adjust the tone curve to correct the contrast of an input image. Those methods, however, often fail in revealing image details because of the limited information in a single image. On the other hand, the SICE task can be better accomplished if we can learn extra information from appropriately collected training data. In this work, we propose to use the convolutional neural network (CNN) to train a SICE enhancer. One key issue is how to construct a training dataset of low-contrast and high-contrast image pairs for end-to-end CNN learning. To this end, we build a large-scale multi-exposure image dataset, which contains 589 elaborately selected high-resolution multi-exposure sequences with 4,413 images. Thirteen representative multi-exposure image fusion and stack-based high dynamic range imaging algorithms are employed to generate the contrast enhanced images for each sequence, and subjective experiments are conducted to screen the best quality one as the reference image of each scene. With the constructed dataset, a CNN can be easily trained as the SICE enhancer to improve the contrast of an under-/over-exposure image. Experimental results demonstrate the advantages of our method over existing SICE methods with a significant margin.
Jianrui Cai, Shuhang Gu, Lei Zhang 0006
IEEE Trans. Image Process.2
2017 Learning Dynamic Guidance for Depth Image Enhancement
abstract
The depth images acquired by consumer depth sensors (e.g., Kinect and ToF) usually are of low resolution and insufficient quality. One natural solution is to incorporate with high resolution RGB camera for exploiting their statistical correlation. However, most existing methods are intuitive and limited in characterizing the complex and dynamic dependency between intensity and depth images. To address these limitations, we propose a weighted analysis representation model for guided depth image enhancement, which advances the conventional methods in two aspects: (i) task driven learning and (ii) dynamic guidance. First, we generalize the analysis representation model by including a guided weight function for dependency modeling. And the task-driven learning formulation is introduced to obtain the optimized guidance tailored to specific enhancement task. Second, the depth image is gradually enhanced along with the iterations, and thus the guidance should also be dynamically adjusted to account for the updating of depth image. To this end, stage-wise parameters are learned for dynamic guidance. Experiments on guided depth image upsampling and noisy depth image restoration validate the effectiveness of our method.
Shuhang Gu, Wangmeng Zuo, Shi Guo, Yunjin Chen, Chongyu Chen, Lei Zhang 0006
CVPR1
2017 Learning Deep CNN Denoiser Prior for Image Restoration
abstract
Model-based optimization methods and discriminative learning methods have been the two dominant strategies for solving various inverse problems in low-level vision. Typically, those two kinds of methods have their respective merits and drawbacks, e.g., model-based optimization methods are flexible for handling different inverse problems but are usually time-consuming with sophisticated priors for the purpose of good performance, in the meanwhile, discriminative learning methods have fast testing speed but their application range is greatly restricted by the specialized task. Recent works have revealed that, with the aid of variable splitting techniques, denoiser prior can be plugged in as a modular part of model-based optimization methods to solve other inverse problems (e.g., deblurring). Such an integration induces considerable advantage when the denoiser is obtained via discriminative learning. However, the study of integration with fast discriminative denoiser prior is still lacking. To this end, this paper aims to train a set of fast and effective CNN (convolutional neural network) denoisers and integrate them into model-based optimization method to solve other inverse problems. Experimental results demonstrate that the learned set of denoisers can not only achieve promising Gaussian denoising results but also can be used as prior to deliver good performance for various low-level vision applications.
Kai Zhang 0008, Wangmeng Zuo, Shuhang Gu, Lei Zhang 0006
CVPR3
2017 Joint Convolutional Analysis and Synthesis Sparse Representation for Single Image Layer Separation
abstract
Analysis sparse representation (ASR) and synthesis sparse representation (SSR) are two representative approaches for sparsity-based image modeling. An image is described mainly by the non-zero coefficients in SSR, while is mainly characterized by the indices of zeros in ASR. To exploit the complementary representation mechanisms of ASR and SSR, we integrate the two models and propose a joint convolutional analysis and synthesis (JCAS) sparse representation model. The convolutional implementation is adopted to more effectively exploit the image global information. In JCAS, a single image is decomposed into two layers, one is approximated by ASR to represent image large-scale structures, and the other by SSR to represent image fine-scale textures. The synthesis dictionary is adaptively learned in JCAS to describe the texture patterns for different single image layer separation tasks. We evaluate the proposed JCAS model on a variety of applications, including rain streak removal, high dynamic range image tone mapping, etc. The results show that our JCAS method outperforms state-of-the-arts in these applications in terms of both quantitative measure and visual perception quality.
Shuhang Gu, Deyu Meng, Wangmeng Zuo, Lei Zhang 0006
ICCV1
2017 Weighted Nuclear Norm Minimization and Its Applications to Low Level Vision
Shuhang Gu, Qi Xie 0002, Deyu Meng, Wangmeng Zuo, Xiangchu Feng, Lei Zhang 0006
Int. J. Comput. Vis.1
2017 Joint Image Denoising and Disparity Estimation via Stereo Structure PCA and Noise-Tolerant Cost
Jianbo Jiao, Qingxiong Yang, Shengfeng He, Shuhang Gu, Lei Zhang 0006, Rynson W. H. Lau
Int. J. Comput. Vis.4
2016 Dictionary Pair Classifier Driven Convolutional Neural Networks for Object Detection
abstract
Feature representation and object category classification are two key components of most object detection methods. While significant improvements have been achieved for deep feature representation learning, traditional SVM/softmax classifiers remain the dominant methods for the final object category classification. However, SVM/softmax classifiers lack the capacity of explicitly exploiting the complex structure of deep features, as they are purely discriminative methods. The recently proposed discriminative dictionary pair learning (DPL) model involves a fidelity term to minimize the reconstruction loss and a discrimination term to enhance the discriminative capability of the learned dictionary pair, and thus is appropriate for balancing the representation and discrimination to boost object detection performance. In this paper, we propose a novel object detection system by unifying DPL with the convolutional feature learning. Specifically, we incorporate DPL as a Dictionary Pair Classifier Layer (DPCL) into the deep architecture, and develop an end-to-end learning algorithm for optimizing the dictionary pairs and the neural networks simultaneously. Moreover, we design a multi-task loss for guiding our model to accomplish the three correlated tasks: objectness estimation, categoryness computation, and bounding box regression. From the extensive experiments on PASCAL VOC 2007/2012 benchmarks, our approach demonstrates the effectiveness to substantially improve the performances over the popular existing object detection frameworks (e.g., R-CNN [13] and FRCN [12]), and achieves new state-of-the-arts.
Keze Wang, Liang Lin 0004, Wangmeng Zuo, Shuhang Gu, Lei Zhang 0006
CVPR4
2016 Multispectral Images Denoising by Intrinsic Tensor Sparsity Regularization
abstract
Multispectral images (MSI) can help deliver more faithful representation for real scenes than the traditional image system, and enhance the performance of many computer vision tasks. In real cases, however, an MSI is always corrupted by various noises. In this paper, we propose a new tensor-based denoising approach by fully considering two intrinsic characteristics underlying an MSI, i.e., the global correlation along spectrum (GCS) and nonlocal self-similarity across space (NSS). In specific, we construct a new tensor sparsity measure, called intrinsic tensor sparsity (ITS) measure, which encodes both sparsity insights delivered by the most typical Tucker and CANDECOMP/ PARAFAC (CP) low-rank decomposition for a general tensor. Then we build a new MSI denoising model by applying the proposed ITS measure on tensors formed by non-local similar patches within the MSI. The intrinsic GCS and NSS knowledge can then be efficiently explored under the regularization of this tensor sparsity measure to finely rectify the recovery of a MSI from its corruption. A series of experiments on simulated and real MSI denoising problems show that our method outperforms all state-of-the-arts under comprehensive quantitative performance measures.
Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Zongben Xu, Shuhang Gu, Wangmeng Zuo, Lei Zhang 0006
CVPR5
2016 Weighted Schatten p-Norm Minimization for Image Denoising and Background Subtraction
abstract
Low rank matrix approximation (LRMA), which aims to recover the underlying low rank matrix from its degraded observation, has a wide range of applications in computer vision. The latest LRMA methods resort to using the nuclear norm minimization (NNM) as a convex relaxation of the nonconvex rank minimization. However, NNM tends to over-shrink the rank components and treats the different rank components equally, limiting its flexibility in practical applications. We propose a more flexible model, namely, the weighted Schatten p-norm minimization (WSNM), to generalize the NNM to the Schatten p-norm minimization with weights assigned to different singular values. The proposed WSNM not only gives better approximation to the original low-rank assumption, but also considers the importance of different rank components. We analyze the solution of WSNM and prove that, under certain weights permutation, WSNM can be equivalently transformed into independent non-convex lp-norm subproblems, whose global optimum can be efficiently solved by generalized iterated shrinkage algorithm. We apply WSNM to typical low-level vision problems, e.g., image denoising and background subtraction. Extensive experimental results show, both qualitatively and quantitatively, that the proposed WSNM can more effectively remove noise, and model the complex and dynamic scenes compared with state-of-the-art methods.
Yuan Xie 0006, Shuhang Gu, Yan Liu 0004, Wangmeng Zuo, Wensheng Zhang 0002, Lei Zhang 0006
IEEE Trans. Image Process.2
2016 Learning Iteration-wise Generalized Shrinkage-Thresholding Operators for Blind Deconvolution
abstract
Salient edge selection and time-varying regularization are two crucial techniques to guarantee the success of maximum a posteriori (MAP)-based blind deconvolution. However, the existing approaches usually rely on carefully designed regularizers and handcrafted parameter tuning to obtain satisfactory estimation of the blur kernel. Many regularizers exhibit the structure-preserving smoothing capability, but fail to enhance salient edges. In this paper, under the MAP framework, we propose the iteration-wise ℓp-norm regularizers together with data-driven strategy to address these issues. First, we extend the generalized shrinkage-thresholding (GST) operator for ℓp-norm minimization with negative p value, which can sharpen salient edges while suppressing trivial details. Then, the iteration-wise GST parameters are specified to allow dynamical salient edge selection and time-varying regularization. Finally, instead of handcrafted tuning, a principled discriminative learning approach is proposed to learn the iterationwise GST operators from the training dataset. Furthermore, the multi-scale scheme is developed to improve the efficiency of the algorithm. Experimental results show that, negative p value is more effective in estimating the coarse shape of blur kernel at the early stage, and the learned GST operators can be well generalized to other dataset and real world blurry images. Compared with the state-of-the-art methods, our method achieves better deblurring results in terms of both quantitative metrics and visual quality, and it is much faster than the state-of-the-art patch-based blind deconvolution method.
Wangmeng Zuo, Dongwei Ren, David Zhang 0001, Shuhang Gu, Lei Zhang 0006
IEEE Trans. Image Process.4
2015 Discriminative learning of iteration-wise priors for blind deconvolution
abstract
The maximum a posterior (MAP)-based blind deconvolution framework generally involves two stages: blur kernel estimation and non-blind restoration. For blur kernel estimation, sharp edge prediction and carefully designed image priors are vital to the success of MAP. In this paper, we propose a blind deconvolution framework together with iteration specific priors for better blur kernel estimation. The family of hyper-Laplacian (Pr(d) ∝ e-∥d∥pp/λ) is adopted for modeling iteration-wise priors of image gra- dients, where each iteration has its own model parameters {λ(t), p(t)}. To avoid heavy parameter tuning, all iteration-wise model parameters can be learned using our principled discriminative learning model from a training set, and can be directly applied to other dataset and real blurry images. Interestingly, with the generalized shrinkage / thresholding operator, negative p value (p <;0) is allowable and we find that it contributes more in estimating the coarse shape of blur kernel. Experimental results on synthetic and real world images demonstrate that our method achieves better deblurring results than the existing gradient prior-based methods. Compared with the state-of-the-art patch prior-based method, our method is competitive in restoration results but is much more efficient.
Wangmeng Zuo, Dongwei Ren, Shuhang Gu, Liang Lin 0004, Lei Zhang 0006
CVPR3
2015 Convolutional Sparse Coding for Image Super-Resolution
abstract
Most of the previous sparse coding (SC) based super resolution (SR) methods partition the image into overlapped patches, and process each patch separately. These methods, however, ignore the consistency of pixels in overlapped patches, which is a strong constraint for image reconstruction. In this paper, we propose a convolutional sparse coding (CSC) based SR (CSC-SR) method to address the consistency issue. Our CSC-SR involves three groups of parameters to be learned: (i) a set of filters to decompose the low resolution (LR) image into LR sparse feature maps, (ii) a mapping function to predict the high resolution (HR) feature maps from the LR ones, and (iii) a set of filters to reconstruct the HR images from the predicted HR feature maps via simple convolution operations. By working directly on the whole image, the proposed CSC-SR algorithm does not need to divide the image into overlapped patches, and can exploit the image global correlation to produce more robust reconstruction of image local structures. Experimental results clearly validate the advantages of CSC over patch based SC in SR application. Compared with state-of-the-art SR methods, the proposed CSC-SR method achieves highly competitive PSNR results, while demonstrating better edge and texture preservation performance.
Shuhang Gu, Wangmeng Zuo, Qi Xie 0002, Deyu Meng, Xiangchu Feng, Lei Zhang 0006
ICCV1
2014 Weighted Nuclear Norm Minimization with Application to Image Denoising
abstract
As a convex relaxation of the low rank matrix factorization problem, the nuclear norm minimization has been attracting significant research interest in recent years. The standard nuclear norm minimization regularizes each singular value equally to pursue the convexity of the objective function. However, this greatly restricts its capability and flexibility in dealing with many practical problems (e.g., denoising), where the singular values have clear physical meanings and should be treated differently. In this paper we study the weighted nuclear norm minimization (WNNM) problem, where the singular values are assigned different weights. The solutions of the WNNM problem are analyzed under different weighting conditions. We then apply the proposed WNNM algorithm to image denoising by exploiting the image nonlocal self-similarity. Experimental results clearly show that the proposed WNNM algorithm outperforms many state-of-the-art denoising algorithms such as BM3D in terms of both quantitative measure and visual perception quality.
Shuhang Gu, Lei Zhang 0006, Wangmeng Zuo, Xiangchu Feng
CVPR1
2014 Projective dictionary pair learning for pattern classification
Shuhang Gu, Lei Zhang 0006, Wangmeng Zuo, Xiangchu Feng
NIPS1
2012 Fast image super resolution via local regression
Shuhang Gu, Nong Sang, Fan Ma
ICPR1