Yuanchao Bai

dblp:137/5961 · DBLP profile ↗
← Back
28ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0003-3449-6537ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 8 first-author · 15 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Practical Lossless Volumetric Medical Image Compression via Tri-Plane Context Tree Learning
abstract
Lossless compression of volumetric medical images is of paramount importance for clinical and research applications where data fidelity is essential. Traditional compression methods are often limited in efficiency due to rigid, handcrafted models. Conversely, deep neural network (DNN)-based compression methods, while effective, demand substantial computational resources, hindering deployment in resource-constrained settings. To address these challenges, we propose a novel tri-plane context tree (TCT)-based method for lossless volumetric medical image compression that delivers high performance without relying on DNNs or external training data. To exploit intra-slice and inter-slice redundancies, we introduce a compact tri-plane context representation that decomposes complex 3D context modeling into efficient 2D modeling on three orthogonal planes. By integrating this representation with a context tree framework, we develop an input-specific TCT model employing an adaptive binary tree structure. At each tree node, the model dynamically selects from a suite of tri-plane based predictors and contextual feature extractors, enabling data-adaptive context modeling tailored to local structural characteristics. Instead of offline training, we sample a subset of the input volume to learn the TCT model by optimizing the minimum description length (MDL) through iterative construction and pruning. With the learned TCT model, each pixel retrieves its corresponding context, computes the prediction residual using the predictor dictated by the context, and performs entropy encoding based on the associated histograms. Experimental results demonstrate that the proposed method achieves compression performance on par with recent DNN-based methods on multiple datasets, while maintaining low computational cost and fast coding speeds, making it highly applicable in practice.
Yuanchao Bai, Kai Wang 0070, Yuanbo Du, Jie Chen 0001, Teng Fang, Xianming Liu 0005, Wen Gao 0001
IEEE Trans. Image Process.1
2026 3D-SLARM: Practical Lossless Volumetric Image Compression via a 3D-Scanning Lightweight Autoregressive Model
abstract
Volumetric images often encapsulate critical information, making it essential to employ lossless compression to preserve data integrity. Although various learned methods have demonstrated effective lossless compression for volumetric images, balancing high compression ratios with rapid coding speeds and lightweight architectures remains challenging. In this paper, we propose a 3D-scanning lightweight autoregressive model (3D-SLARM) for practical lossless volumetric image compression. 3D-SLARM integrates a novel 3D plane scanning module, a lightweight feature extraction (FE) module, and a lightweight distribution parameter and adaptive range predictor (DPARP) module. Initially, 3D-SLARM leverages a 3D plane scanning module to determine the scanning order of each voxel, allowing parallel coding of voxels within the same plane. Next, the lightweight FE module captures both intra-slice and inter-slice dependencies in the receptive field defined by the 3D plane scanning module. By incorporating our proposed serial re-parameterization (SerRep) technology alongside non-centric masked convolution (NCMC), the FE module attains a lightweight design while effectively capturing complex dependencies. Finally, 3D-SLARM employs a lightweight DPARP module to compute distribution parameters for both 8-bit and high bit-depth volumetric images. For high bit-depth images, the module further generates an adaptive probability range for each voxel, resulting in compact, voxel-specific PMF tables that facilitate efficient compression. Extensive experiments demonstrate that our 3D-SLARM achieves state-of-the-art lossless compression performance on majority volumetric image datasets and maintains fast coding speed with a lightweight design, underscoring its practical applicability.
Kai Wang 0070, Yuanchao Bai, Daxin Li, Deming Zhai, Junjun Jiang, Xianming Liu 0005
IEEE Trans. Image Process.2
2025 CALLIC: Content Adaptive Learning for Lossless Image Compression
abstract
Learned lossless image compression has achieved significant advancements in recent years. However, existing methods often rely on training amortized generative models on massive datasets, resulting in sub-optimal probability distribution estimation for specific testing images during encoding process. To address this challenge, we explore the connection between the Minimum Description Length (MDL) principle and Parameter-Efficient Transfer Learning (PETL), leading to the development of a novel content-adaptive approach for learned lossless image compression, dubbed CALLIC. Specifically, we first propose a content-aware autoregressive self-attention mechanism by leveraging convolutional gating operations, termed Masked Gated ConvFormer (MGCF), and pretrain MGCF on training dataset. Cache then Crop Inference (CCI) is proposed to accelerate the coding process. During encoding, we decompose pretrained layers, including depth-wise convolutions, using low-rank matrices and then adapt the incremental weights on testing image by Rate-guided Progressive Fine-Tuning (RPFT). RPFT fine-tunes with gradually increasing patches that are sorted in descending order by estimated entropy, optimizing learning process and reducing adaptation time. Extensive experiments across diverse datasets demonstrate that CALLIC sets a new state-of-the-art (SOTA) for learned lossless image compression.
Daxin Li, Yuanchao Bai, Kai Wang 0070, Junjun Jiang, Xianming Liu 0005, Wen Gao 0001
AAAI2
2025 A Wavelet-based Image Coding Framework for Data Storage on DNA
abstract
In the face of the exponential growth of digital data, DNA is expected to become a new storage medium. Image data makes up a large proportion of digital data. However, existing DNA data storage models are mainly designed for general files. To address this issue, we propose a novel image encoding method for DNA data storage. We employ discrete wavelet transform to decompose the image and utilize an improved exponent-mantissa representation for numerical data. Subsequently, we achieve enhanced compression performance through context-adaptive arithmetic coding. Additionally, we construct a dictionary between ternary sequences and oligonucleotides to generate nucleotide sequences that meet the specified constraints. Experiments show that our method outperforms JPEG-DNA and BioCoder in compression performance and generates higher-quality nucleotide sequences.
Chen Qin, Yuanchao Bai, Wenbo Zhao 0004, Xianming Liu 0005
VCIP2
2025 Learning Lossless Compression for High Bit-Depth Volumetric Medical Image
abstract
Recent advances in learning-based methods have markedly enhanced the capabilities of image compression. However, these methods struggle with high bit-depth volumetric medical images, facing issues such as degraded performance, increased memory demand, and reduced processing speed. To address these challenges, this paper presents the Bit-Division based Lossless Volumetric Image Compression (BD-LVIC) framework, which is tailored for high bit-depth medical volume compression. The BD-LVIC framework skillfully divides the high bit-depth volume into two lower bit-depth segments: the Most Significant Bit-Volume (MSBV) and the Least Significant Bit-Volume (LSBV). The MSBV concentrates on the most significant bits of the volumetric medical image, capturing vital structural details in a compact manner. This reduction in complexity greatly improves compression efficiency using traditional codecs. Conversely, the LSBV deals with the least significant bits, which encapsulate intricate texture details. To compress this detailed information effectively, we introduce an effective learning-based compression model equipped with a Transformer-Based Feature Alignment Module, which exploits both intra-slice and inter-slice redundancies to accurately align features. Subsequently, a Parallel Autoregressive Coding Module merges these features to precisely estimate the probability distribution of the least significant bit-planes. Our extensive testing demonstrates that the BD-LVIC framework not only sets new performance benchmarks across various datasets but also maintains a competitive coding speed, highlighting its significant potential and practical utility in the realm of volumetric medical image compression.
Kai Wang 0070, Yuanchao Bai, Daxin Li, Deming Zhai, Junjun Jiang, Xianming Liu 0005
IEEE Trans. Image Process.2
2025 Self-Supervised Multi-Camera Collaborative Depth Prediction With Latent Diffusion Models
abstract
Depth map estimation from images is a crucial task in self-driving applications. Existing methods can be categorized into two groups: multi-view stereo and monocular depth estimation. The former requires cameras to have large overlapping areas and a sufficient baseline between them, while the latter that processes each image independently can hardly guarantee the structure consistency between cameras. In this paper, we propose a novel self-supervised multi-camera collaborative depth prediction method with latent diffusion models, which does not require large overlapping areas while maintaining structure consistency between cameras. Specifically, we introduce MCDP, a new generative foundation model for estimating depth attributes for multi-cameras. We formulate the depth estimation as a weighted combination of depth bases, in which the weights are updated iteratively by the recurrent refinement strategy. During the iterative update, the results of depth estimation are compared across cameras, and the information of overlapping areas is propagated to the whole depth maps with the help of basis formulation in diffusion process. We integrate the GRU-based Weight Net into the diffusion process, allowing the refined hidden state to serve as a conditional input to accurately control the next iterative denoising step. Furthermore, by incorporating the proposed depth consistency loss, we ensure structural consistency across cameras, even in regions with minimal overlap. Experimental results on DDAD, NuScenes, Cityscapes, and Waymo Open Datasets demonstrate the superior performance of our method, and show great help for the downstream task.
Jialei Xu, Xianming Liu 0005, Yuanchao Bai, Junjun Jiang, Xiangyang Ji
IEEE Trans. Intell. Transp. Syst.3
2024 Parameter Efficient Adaptation for Image Restoration with Heterogeneous Mixture-of-Experts
abstract
Designing single-task image restoration models for specific degradation has seen great success in recent years. To achieve generalized image restoration, all-in-one methods have recently been proposed and shown potential for multiple restoration tasks using one single model. Despite the promising results, the existing all-in-one paradigm still suffers from high computational costs as well as limited generalization on unseen degradations. In this work, we introduce an alternative solution to improve the generalization of image restoration models. Drawing inspiration from recent advancements in Parameter Efficient Transfer Learning (PETL), we aim to tune only a small number of parameters to adapt pre-trained restoration models to various tasks. However, current PETL methods fail to generalize across varied restoration tasks due to their homogeneous representation nature. To this end, we propose AdaptIR, a Mixture-of-Experts (MoE) with orthogonal multi-branch design to capture local spatial, global spatial, and channel representation bases, followed by adaptive base combination to obtain heterogeneous representation for different degradations. Extensive experiments demonstrate that our AdaptIR achieves stable performance on single-degradation tasks, and excels in hybrid-degradation tasks, with training only 0.6% parameters for 8 hours.
Hang Guo 0002, Tao Dai 0001, Yuanchao Bai, Bin Chen 0011, Xudong Ren, Zexuan Zhu 0001, Shutao Xia
NeurIPS3
2024 Semantic Ensemble Loss and Latent Refinement for High-Fidelity Neural Image Compression
abstract
Recent advancements in neural compression have surpassed traditional codecs in PSNR and MS-SSIM measurements. However, at low bit-rates, these methods can introduce visually displeasing artifacts, such as blurring, color shifting, and texture loss, thereby compromising perceptual quality of images. To address these issues, this study presents an enhanced neural compression method designed for optimal visual fidelity. We have trained our model with a sophisticated semantic ensemble loss, integrating Charbonnier loss, perceptual loss, style loss, and a non-binary adversarial loss, to enhance the perceptual quality of image reconstructions. Additionally, we have implemented a latent refinement process to generate content-aware latent codes. These codes adhere to bit-rate constraints, and prioritize bit allocation to regions of greater importance. Our empirical findings demonstrate that this approach significantly improves the statistical fidelity of neural image compression.
Daxin Li, Yuanchao Bai, Kai Wang 0070, Junjun Jiang, Xianming Liu 0005
VCIP2
2024 Enhancing Privacy-Utility Tradeoff with Few-Round Strategy in Heterogeneous Federated Learning
abstract
Federated learning inherently provides a certain level of privacy protection, which however is often inadequate in many real-world scenarios. Existing privacy-preserving methods frequently incur unbearable time overheads or result in non-negligible deterioration to model performance, thus suffering from the tradeoff between performance and privacy. In this work, we propose a novel Federated Privacy-Preserving Knowledge Transfer framework, namely FedPPKT, which employs data-free knowledge distillation in a meta-learning manner to rapidly generates pseudo data and performs privacy-preserving knowledge transfer. FedPPKT establishes a protective barrier between the original private data and the federated model, thereby ensuring user privacy. Furthermore, leveraging the few-round strategy of FedPPKT, it has the capability to reduce the number of communication rounds, further mitigating the risk of privacy exposure for user data. With the help of the meta generator, the problem of uneven local label distribution on clients is alleviated, mitigating data heterogeneity and improving model performance. Experiments show that FedPPKT outperforms the state-of-the-art privacy-preserving federated learning methods. Our code is publicly available at https://github.com/HIT-weiqb/FedPPKT.
Qingbin Wei, Feilong Zhang 0002, Yuanchao Bai, Deming Zhai, Junjun Jiang, Xianming Liu 0005
VCIP3
2024 Deep Lossy Plus Residual Coding for Lossless and Near-Lossless Image Compression
abstract
Lossless and near-lossless image compression is of paramount importance to professional users in many technical fields, such as medicine, remote sensing, precision engineering and scientific research. But despite rapidly growing research interests in learning-based image compression, no published method offers both lossless and near-lossless modes. In this paper, we propose a unified and powerful deep lossy plus residual (DLPR) coding framework for both lossless and near-lossless image compression. In the lossless mode, the DLPR coding system first performs lossy compression and then lossless coding of residuals. We solve the joint lossy and residual compression problem in the approach of VAEs, and add autoregressive context modeling of the residuals to enhance lossless compression performance. In the near-lossless mode, we quantize the original residuals to satisfy a given ℓ∞error bound, and propose a scalable near-lossless compression scheme that works for variable ℓ∞bounds instead of training multiple networks. To expedite the DLPR coding, we increase the degree of algorithm parallelization by a novel design of coding context, and accelerate the entropy coding with adaptive residual interval. Experimental results demonstrate that the DLPR coding system achieves both the state-of-the-art lossless and near-lossless image compression performance with competitive coding speed.
Yuanchao Bai, Xianming Liu 0005, Kai Wang 0070, Xiangyang Ji, Xiaolin Wu 0001, Wen Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 GroupedMixer: An Entropy Model With Group-Wise Token-Mixers for Learned Image Compression
abstract
Transformer-based entropy models have gained prominence in recent years due to their superior ability to capture long-range dependencies in probability distribution estimation compared to convolution-based methods. However, previous transformer-based entropy models suffer from a sluggish coding process due to pixel-wise autoregression or duplicated computation during inference. In this paper, we propose a novel transformer-based entropy model called GroupedMixer, which enjoys both faster coding speed and better compression performance than previous transformer-based methods. Specifically, our approach builds upon group-wise autoregression by first partitioning the latent variables into groups along spatial-channel dimensions, and then entropy coding the groups with the proposed transformer-based entropy model. The global causal self-attention is decomposed into more efficient group-wise interactions, implemented using inner-group and cross-group token-mixers. The inner-group token-mixer incorporates contextual elements within a group while the cross-group token-mixer interacts with previously decoded groups. Alternate arrangement of two token-mixers enables global contextual reference. To further expedite the network inference, we introduce context cache optimization to GroupedMixer, which caches attention activation values in cross-group token-mixers and avoids complex and duplicated computation. Experimental results demonstrate that the proposed GroupedMixer yields the state-of-the-art rate-distortion performance with fast compression speed.
Daxin Li, Yuanchao Bai, Kai Wang 0070, Junjun Jiang, Xianming Liu 0005, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Illumination-Aware Low-Light Image Enhancement with Transformer and Auto-Knee Curve
abstract
Images captured under low-light conditions suffer from several combined degradation factors, including low brightness, low contrast, noise, and color bias. Many learning-based techniques attempt to learn the low-to-clear mapping between low-light and normal-light images. However, they often fall short when applied to low-light images taken in wide-contrast scenes because uneven illumination brings illumination-varying noise and the enhanced images are easily over-saturated in highlight areas. In this article, we present a novel two-stage method to tackle the problem of uneven illumination distribution in low-light images. Under the assumption that noise varies with illumination, we design an illumination-aware transformer network for the first stage of image restoration. In this stage, we introduce the Illumination-aware Attention Block featured with Illumination-aware Multi-head Self-attention, which incorporates different scales of illumination features to guide the attention module, thereby enhancing the denoising and reconstruction capabilities of the restoration network. In the second stage, we innovatively introduce a cubic auto-knee curve transfer with a global parameter predictor to alleviate the over-exposure caused by uneven illumination. We also adopt a white balance correction module to address color bias issues at this stage. Extensive experiments on various benchmarks demonstrate the advantages of our method over state-of-the-art methods qualitatively and quantitatively.
Jinwang Pan, Xianming Liu 0005, Yuanchao Bai, Deming Zhai, Junjun Jiang, Debin Zhao
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Learning Lossless Compression for High Bit-Depth Medical Imaging
abstract
We propose a learned lossless image compression method for high bit-depth medical imaging (up to 16 bit-depths). Instead of compressing a high bit-depth medical image as a whole, we split it into two low bit-depth subimages, i.e., the most significant bytes (MSB) subimage and the least significant bytes (LSB) subimage, respectively. The MSB subimage depicts piece-wise smooth structure information that is relatively easy to compress. We thus use traditional lossless codecs for low complexity. The LSB subimage depicts the complementary texture information that is more challenging to compress. We design an autoregressive entropy model conditioned on the MSB subimage that models the probability distribution of the LSB subimage and effectively reduces the redundancy between the MSB and LSB subimages. We then encode the LSB subimage to bitstreams based on the learned entropy model. The compressed high bit-depth medical image is finally stored including the bitstreams of the MSB and LSB subimages. Experimental results demonstrate the state-of-the-art compression performance of the proposed method on high bit-depth medical images, compared with both existing traditional and learned lossless image codecs.
Kai Wang 0070, Yuanchao Bai, Deming Zhai, Daxin Li, Junjun Jiang, Xianming Liu 0005
ICME2
2023 Learning Spatial-Frequency Transformer for Visual Object Tracking
abstract
Recently, some researchers have begun to adopt the Transformer to combine or replace the widely used ResNet as their new backbone network. As the Transformer captures the long-range relations between pixels well using the self-attention scheme, which complements the issues caused by the limited receptive field of CNN. Although their trackers work well in regular scenarios, they simply flatten the 2D features into a sequence to better match the Transformer. We believe these operations ignore the spatial prior of the target object, which may lead to sub-optimal results only. In addition, many works demonstrate that self-attention is actually a low-pass filter, which is independent of input features or keys/queries. That is to say, it may suppress the high-frequency component of the input features and preserve or even amplify the low-frequency information. To handle these issues, in this paper, we propose a unified Spatial-Frequency Transformer that models the Gaussian spatial Prior and High-frequency emphasis Attention (GPHA) simultaneously. To be specific, Gaussian spatial prior is generated using dual Multi-Layer Perceptrons (MLPs) and injected into the similarity matrix produced by multiplying Query and Key features in self-attention. The output will be fed into a softmax layer and then decomposed into two components, i.e., the direct and high-frequency signal. The low- and high-pass branches are rescaled and combined to achieve all-pass, therefore, the high-frequency features will be protected well in stacked self-attention layers. We further integrate the Spatial-Frequency Transformer into the Siamese tracking framework and propose a novel tracking algorithm termed SFTransT. The cross-scale fusion based SwinTransformer is adopted as the backbone, and also a multi-head cross-attention module is used to boost the interaction between search and template features. The output will be fed into the tracking head for target localization. Extensive experiments on short-term and long-term tracking benchmarks all demonstrate the effectiveness of our proposed framework. Source code will be released athttps://github.com/Tchuanm/SFTransT.git.
Chuanming Tang, Xiao Wang 0014, Yuanchao Bai, Zhe Wu 0006, Jianlin Zhang 0001, Yongmei Huang
IEEE Trans. Circuits Syst. Video Technol.3
2022 Towards End-to-End Image Compression and Analysis with Transformers
abstract
We propose an end-to-end image compression and analysis model with Transformers, targeting to the cloud-based image classification application. Instead of placing an existing Transformer-based image classification model directly after an image codec, we aim to redesign the Vision Transformer (ViT) model to perform image classification from the compressed features and facilitate image compression with the long-term information from the Transformer. Specifically, we first replace the patchify stem (i.e., image splitting and embedding) of the ViT model with a lightweight image encoder modelled by a convolutional neural network. The compressed features generated by the image encoder are injected convolutional inductive bias and are fed to the Transformer for image classification bypassing image reconstruction. Meanwhile, we propose a feature aggregation module to fuse the compressed features with the selected intermediate features of the Transformer, and feed the aggregated features to a deconvolutional neural network for image reconstruction. The aggregated features can obtain the long-term information from the self-attention mechanism of the Transformer and improve the compression performance. The rate-distortion-accuracy optimization problem is finally solved by a two-step training strategy. Experimental results demonstrate the effectiveness of the proposed model in both the image compression and the classification tasks.
Yuanchao Bai, Xianming Liu 0005, Junjun Jiang, Yaowei Wang 0001, Xiangyang Ji, Wen Gao 0001
AAAI1
2022 ChebyLighter: Optimal Curve Estimation for Low-light Image Enhancement
abstract
Low-light enhancement aims to recover a high contrast normal light image from a low-light image with bad exposure and low contrast. Inspired by curve adjustment in photo editing software and Chebyshev approximation, this paper presents a novel model for brightening low-light images. The proposed model, ChebyLighter, learns to estimate pixel-wise adjustment curves for a low-light image recurrently to reconstruct an enhanced output. In ChebyLighter, Chebyshev image series are first generated. Then pixel-wise coefficient matrices are estimated with Triple Coefficient Estimation (TCE) modules and the final enhanced image is recurrently reconstructed by Chebyshev Attention Weighted Summation (CAWS). The TCE module is specifically designed based on dual attention mechanism with three necessary inputs. Our method can achieve ideal performance because adjustment curves can be obtained with numerical approximation by our model. With extensive quantitative and qualitative experiments on diverse test images, we demonstrate that the proposed method performs favorably against state-of-the-art low-light image enhancement algorithms.
Jinwang Pan, Deming Zhai, Yuanchao Bai, Junjun Jiang, Debin Zhao, Xianming Liu 0005
ACM Multimedia3
2022 Multi-Camera Collaborative Depth Prediction via Consistent Structure Estimation
abstract
Depth map estimation from images is an important task in robotic systems. Existing methods can be categorized into two groups including multi-view stereo and monocular depth estimation. The former requires cameras to have large overlapping areas and sufficient baseline between cameras, while the latter that processes each image independently can hardly guarantee the structure consistency between cameras. In this paper, we propose a novel multi-camera collaborative depth prediction method that does not require large overlapping areas while maintaining structure consistency between cameras. Specifically, we formulate the depth estimation as a weighted combination of depth basis, in which the weights are updated iteratively by a refinement network driven by the proposed consistency loss. During the iterative update, the results of depth estimation are compared across cameras and the information of overlapping areas is propagated to the whole depth maps with the help of basis formulation. Experimental results on DDAD and NuScenes datasets demonstrate the superior performance of our method.
Jialei Xu, Xianming Liu 0005, Yuanchao Bai, Junjun Jiang, Xiaozhi Chen, Xiangyang Ji
ACM Multimedia3
2021 Learning Scalable lY=-Constrained Near-Lossless Image Compression via Joint Lossy Image and Residual Compression
abstract
We propose a novel joint lossy image and residual compression framework for learning ℓ∞-constrained near-lossless image compression. Specifically, we obtain a lossy reconstruction of the raw image through lossy image compression and uniformly quantize the corresponding residual to satisfy a given tight ℓ∞error bound. Suppose that the error bound is zero, i.e., lossless image compression, we formulate the joint optimization problem of compressing both the lossy image and the original residual in terms of variational auto-encoders and solve it with end-to-end training. To achieve scalable compression with the error bound larger than zero, we derive the probability model of the quantized residual by quantizing the learned probability model of the original residual, instead of training multiple networks. We further correct the bias of the derived probability model caused by the context mismatch between training and inference. Finally, the quantized residual is encoded according to the bias-corrected probability model and is concatenated with the bitstream of the compressed lossy image. Experimental results demonstrate that our near-lossless codec achieves the state-of-the-art performance for lossless and near-lossless image compression, and achieves competitive PSNR while much smaller ℓ∞error compared with lossy image codecs at high bit rates.
Yuanchao Bai, Xianming Liu 0005, Wangmeng Zuo, Yaowei Wang 0001, Xiangyang Ji
CVPR1
2020 FFA-Net: Feature Fusion Attention Network for Single Image Dehazing
abstract
In this paper, we propose an end-to-end feature fusion at-tention network (FFA-Net) to directly restore the haze-free image. The FFA-Net architecture consists of three key components:1) A novel Feature Attention (FA) module combines Channel Attention with Pixel Attention mechanism, considering that different channel-wise features contain totally different weighted information and haze distribution is uneven on the different image pixels. FA treats different features and pixels unequally, which provides additional flexibility in dealing with different types of information, expanding the representational ability of CNNs. 2) A basic block structure consists of Local Residual Learning and Feature Attention, Local Residual Learning allowing the less important information such as thin haze region or low-frequency to be bypassed through multiple local residual connections, let main network architecture focus on more effective information. 3) An Attention-based different levels Feature Fusion (FFA) structure, the feature weights are adaptively learned from the Feature Attention (FA) module, giving more weight to important features. This structure can also retain the information of shallow layers and pass it into deep layers.The experimental results demonstrate that our proposed FFA-Net surpasses previous state-of-the-art single image dehazing methods by a very large margin both quantitatively and qualitatively, boosting the best published PSNR metric from 30.23 dB to 36.39 dB on the SOTS indoor test dataset. Code has been made available at GitHub.
Xu Qin, Zhilin Wang, Yuanchao Bai, Huizhu Jia
AAAI3
2020 Single-Image Blind Deblurring Using Multi-Scale Latent Structure Prior
abstract
Blind image deblurring is a challenging problem in computer vision, which aims to restore both the blur kernel and the latent sharp image from only a blurry observation. Inspired by the prevalent self-example prior in image super-resolution, in this paper, we observe that a coarse enough image down-sampled from a blurry observation is approximately a low-resolution version of the latent sharp image. We prove this phenomenon theoretically and define the coarse enough image as a latent structure prior of the unknown sharp image. Starting from this prior, we propose to restore sharp images from the coarsest scale to the finest scale on a blurry image pyramid and progressively update the prior image using the newly restored sharp image. These coarse-to-fine priors are referred to as multi-scale latent structures (MSLSs). Leveraging the MSLS prior, our algorithm comprises two phases: 1) we first preliminarily restore sharp images in the coarse scales and 2) we then apply a refinement process in the finest scale to obtain the final deblurred image. In each scale, to achieve lower computational complexity, we alternately perform a sharp image reconstruction with fast local self-example matching, an accelerated kernel estimation with error compensation, and a fast non-blind image deblurring, instead of computing any computationally expensive non-convex priors. We further extend the proposed algorithm to solve more challenging non-uniform blind image deblurring problem. The extensive experiments demonstrate that our algorithm achieves the competitive results against the state-of-the-art methods with much faster running speed.
Yuanchao Bai, Huizhu Jia, Ming Jiang 0001, Xianming Liu 0005, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 Contrast Enhancement via Dual Graph Total Variation-Based Image Decomposition
abstract
Images captured in low lighting environment suffer from both low luminance contrast and noise corruption. However, most existing contrast enhancement algorithms only consider contrast boosting, which tends to reveal or amplify noise that is originally not visible in the dark areas. In this paper, we propose a joint contrast enhancement and denoising algorithm, which is based on structure/texture layer decomposition via minimization of dual forms of graph total variation (GTV). Specifically, the structure layer is expected to be generally smoothing but with sharp edges at the foreground background boundaries, for which we propose a quadratic form of GTV (QGTV) as the prior that promotes signal smoothness along graph structure. For the texture layer, a re-weighted GTV (RGTV) is tailored to noise removal while preserving true image details. We provide theoretical analysis about the filtering behavior of these two priors. Furthermore, a boost factor is derived per patch via optimal contrast-tone mapping to improve the overall brightness level of the patch. Finally, an optimization objective function is formulated, which casts image decomposition, brightness boosting, and noise reduction into a unified optimization framework. We further propose a fast approach to efficiently solve the optimization and provide analysis about the convergency. The experimental results show that the proposed method outperforms the state-of-the-art works in subjective, objective, and statistical quality evaluation.
Xianming Liu 0005, Deming Zhai, Yuanchao Bai, Xiangyang Ji, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2019 Reconstruction-cognizant Graph Sampling Using Gershgorin Disc Alignment
abstract
Graph sampling with noise is a fundamental problem in graph signal processing (GSP). Previous works assume an unbiased least square (LS) signal reconstruction scheme and select samples greedily via expensive extreme eigenvector computation. A popular biased scheme using graph Laplacian regularization (GLR) solves a system of linear equations for its reconstruction. Assuming this GLR-based scheme, we propose a reconstruction-cognizant sampling strategy to maximize the numerical stability of the linear system-i.e., minimize the condition number of the coefficient matrix. Specifically, we maximize the eigenvalue lower bounds of the matrix, represented by left-ends of Gershgorin discs of the coefficient matrix. To accomplish this efficiently, we propose an iterative algorithm to traverse the graph nodes via Breadth First Search (BFS) and align the left-ends of all corresponding Gershgorin discs at lower-bound threshold T using two basic operations: disc shifting and scaling. We then perform binary search to maximize T given a sample budget K. Experiments on real graph data show that the proposed algorithm can effectively promote large eigenvalue lower bounds, and the reconstruction MSE is the same or smaller than existing sampling methods for different budget K at much lower complexity.
Yuanchao Bai, Gene Cheung, Xianming Liu 0005, Wen Gao 0001
ICASSP1
2019 Graph-Based Blind Image Deblurring From a Single Photograph
abstract
Blind image deblurring, i.e., deblurring without knowledge of the blur kernel, is a highly ill-posed problem. The problem can be solved in two parts: i) estimate a blur kernel from the blurry image, and ii) given an estimated blur kernel, de-convolve the blurry input to restore the target image. In this paper, we propose a graph-based blind image deblurring algorithm by interpreting an image patch as a signal on a weighted graph. Specifically, we first argue that a skeleton image-a proxy that retains the strong gradients of the target but smooths out the details-can be used to accurately estimate the blur kernel and has a unique bi-modal edge weight distribution. Then, we design a reweighted graph total variation (RGTV) prior that can efficiently promote a bi-modal edge weight distribution given a blurry patch. Further, to analyze RGTV in the graph frequency domain, we introduce a new weight function to represent RGTV as a graph l1-Laplacian regularizer. This leads to a graph spectral filtering interpretation of the prior with desirable properties, including robustness to noise and blur, strong piecewise smooth (PWS) filtering and sharpness promotion. Minimizing a blind image deblurring objective with RGTV results in a non-convex non-differentiable optimization problem. Leveraging the new graph spectral interpretation for RGTV, we design an efficient algorithm that solves for the skeleton image and the blur kernel alternately. Specifically for Gaussian blur, we propose a further speedup strategy for blind Gaussian deblurring using accelerated graph spectral filtering. Finally, with the computed blur kernel, recent non-blind image deblurring algorithms can be applied to restore the target image. Experimental results demonstrate that our algorithm successfully restores latent sharp images and outperforms state-of-the-art methods quantitatively and qualitatively.
Yuanchao Bai, Gene Cheung, Xianming Liu 0005, Wen Gao 0001
IEEE Trans. Image Process.1
2018 FPGA-Based Real-Time Super-Resolution System for Ultra High Definition Videos
abstract
The market benefits from a barrage of Ultra High Definition (Ultra-HD) displays, yet most extant cameras are barely equipped with Full-HD video capturing. In order to upgrade existing videos without extra storage costs, we propose an FPGA-based super-resolution system that enables real-time Ultra-HD upscaling in high quality. Our super-resolution system crops each frame into blocks, measures their total variation values, and dispatches them accordingly to a neural network or an interpolation module for upscaling. This approach balances the FPGA resource utilization, the attainable frame rate, and the image quality. Evaluations demonstrate that the proposed system achieves superior performance in both throughput and reconstruction quality, comparing to current approaches.
Zhuolun He, Hanxian Huang, Ming Jiang 0001, Yuanchao Bai, Guojie Luo
FCCM4
2018 Blind Image Deblurring Via Reweighted Graph Total Variation
abstract
Blind image deblurring, i.e., deblurring without knowledge of the blur kernel, is a highly ill-posed problem. The problem can be solved in two parts: i) estimate a blur kernel from the blurry image, and ii) given estimated blur kernel, de-convolve blurry input to restore the target image. In this paper, by interpreting an image patch as a signal on a weighted graph, we first argue that a skeleton image-a proxy that retains the strong gradients of the target but smooths out the details-can be used to accurately estimate the blur kernel and has a unique bi-modal edge weight distribution. We then design a reweighted graph total variation (RGTV) prior that can efficiently promote bi-modal edge weight distribution given a blurry patch. However, minimizing a blind image deblurring objective with RGTV results in a non-convex non-differentiable optimization problem. We propose a fast algorithm that solves for the skeleton image and the blur kernel alternately. Finally with the computed blur kernel, recent non-blind image deblurring algorithms can be applied to restore the target image. Experimental results show that our algorithm can robustly estimate the blur kernel with large kernel size, and the reconstructed sharp image is competitive against the state-of-the-art methods.
Yuanchao Bai, Gene Cheung, Xianming Liu 0005, Wen Gao 0001
ICASSP1
2018 Robust Contrast Enhancement via Graph-Based Cartoon-Texture Decomposition
abstract
In this paper, we propose a robust contrast enhancement algorithm based on cartoon and texture layer decomposition. Specifically, the cartoon layer is expected to be generally smoothing but with sharp edges at the foreground and background boundaries, for which we propose a quadratic form of graph total variation (GTV) as the prior to promote signal smoothness along graph structure. For the texture layer, a re-weighted GTV is tailored to remove noises while preserving true image details. Finally, an optimization objective function is formulated, which casts image decomposition, contrast enhancement and noise reduction into a unified framework. We propose an efficient algorithm to solve it. Experimental results show that our generated images outperform state-of-the-art schemes noticeably in subjective quality evaluation.
Deming Zhai, Xianming Lu, Xiangyang Ji, Yuanchao Bai, Debin Zhao, Wen Gao 0001
ICME4
2015 A fast super-resolution method based on sparsity properties
abstract
Super-resolution enhancement is a kind of promising approach to enhance the spatial resolution of images. To super-resolve a satisfying result, regularization term design and blur kernel estimation are two important aspects which need to be carefully considered. In this paper, we propose a robust regularized super-resolution reconstruction approach based on two sparsity properties to deal with these two aspects. Firstly, we design a sparse reweighted TV L1 prior to restrict the first derivative of the upsampled image. Then, noticing that only deblurring sparse high gradient areas can sharpen the super-resolution result, we design an over-deblurring control method to decrease the artifacts caused by inaccurate blur kernel estimation. We also design a fast optimization algorithm to solve our model. The experimental results show that the proposed approach achieves a remarkable performance both in visual quality and run time.
Yuanchao Bai, Huizhu Jia, Rui Chen 0006, Ming Jiang 0001, Wen Gao 0001
VCIP1
2014 Layer-based image completion by poisson surface reconstruction
abstract
Image completion has been widely used to repair damaged regions of a given digital image in a visually plausible way. However, it is difficult to infer appropriate information, meanwhile keep globally coherent just from the origin image when its critical parts are missing. To address this problem, we propose a novel layer-divided image completion scheme, which contains two major steps. First, we extract foregrounds of both target image and source image, and then we apply a guided Poisson surface reconstruction technique to complete the target foreground according to parameters obtained from optimal-matching calculation. Second, to fill the remaining damaged part, a related exemplar-based image completion algorithm is further devised. Several experiments and comparisons show the effectiveness and robustness of our proposed algorithm.
Hengjin Liu, Huizhu Jia, Yuanchao Bai, Wen Gao 0001
VCIP5