Zhi Jin 0002

dblp:22/3510-2 · DBLP profile ↗
← Back
52ranked-venue papers
9as first author
40since 2021 · last 2026
0000-0001-9670-7366ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 41 · 6 first-author · 30 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SparseGS-W: Sparse-View Gaussian Splatting for Unconstrained Image Collections With Diffusion Priors
abstract
Synthesizing novel views from unconstrained image collections is an important but challenging task in computer vision. Existing methods, which optimize per-image appearance and transient occlusion through implicit neural networks from dense training views (approximately 1000 images), struggle to perform effectively with sparse inputs, resulting in noticeable artifacts. In this work, we introduce SparseGS-W, a novel framework designed to boost the reconstruction of unconstrained scenes and novel view synthesis using as few as five training images. Motivated by the observation that diffusion prior constrained by limited sparse inputs can remove artifacts through fast and efficient fine-tuning, we propose a plug-and-play Constrained Novel-View Enhancement module to iteratively enhance the quality of rendered novel views. We further present an Occlusion Handling scheme, which flexibly removes occlusions utilizing the inherent inpainting capability of constrained diffusion priors. Both components are capable of extracting appearance features from any user-provided reference image, enabling flexible modeling of illumination-consistent scenes. Extensive experiments demonstrate that SparseGS-W achieves superior performance not only in full-reference metrics, but also in commonly used non-reference metrics such as FID, ClipIQA and MUSIQ.
Jiawei Wu 0001, Yikun Ma, Zhi Jin 0002
IEEE Trans. Circuits Syst. Video Technol.5
2026 Control-Lit: Illumination Controllable Backlit Image Enhancement
abstract
Backlit image enhancement (BIE) aims to address image degradation caused by challenging lighting conditions. By enhancing the illumination of underexposed areas and restoring image details while avoiding overexposure, it achieves an overall harmonious luminance. In contrast to traditional BIE methods that apply global enhancements with limited effectiveness, we propose a controllable Mamba-based enhancement method, termed Control-Lit. Our method not only delivers effective backlit image enhancement but also allows users to interactively adjust the illumination of specific regions. Control-Lit achieves superior global enhancement performance by employing the Dark Channel Prior (DCP)-based Finite Scalar Quantization (DFSQ) module that provides a high-quality image prior. Additionally, the Dark Channel Prior Enhancement (DCPE) module is designed to guide the network in pixel-wise illumination adjustment at the feature level, thereby achieving enhancement of backlit regions. Furthermore, to address the loss of content and details in backlit regions, the Global-Local Vision State Space (GLVSS) module is incorporated to extract both global and local features for BIE. To enable customizable, controllable enhancement, we introduce an illumination control vector. By adjusting the coefficient map elements that are multiplied by this vector, users can achieve precise regional illumination adjustments. To further validate the generalization of different light enhancement methods, we contribute a synthetic backlit image dataset using a relighting generative model. Along with current widely used datasets, the experimental results demonstrate that our method achieves state-of-the-art performance on all datasets while enabling high-quality and controllable image illumination adjustment. The code is available at https://github.com/wuhj43/Control-Lit.
Hongjun Wu 0003, Yi Tang 0008, Chongyi Li, Zhi Jin 0002
IEEE Trans. Circuits Syst. Video Technol.4
2026 FreeDehaze: Towards Training-Free Real-World Image Dehazing via Diffusion Degradation Prior
abstract
Restoring high-quality images from degraded hazy images is a challenging task, particularly in real-world scenarios. Recent investigations seek to address this limitation by exploring advanced methods for synthesizing haze and incorporating real-world hazy images. Due to the inherent diversity and complexity of real-world haze, these methods struggle to accurately model haze representations. Based on our observation that the hazy images generated by advanced text-to-image diffusion models exhibit a remarkable resemblance to real-world haze, it suggests that these diffusion models effectively internalize haze representations. Hence, we propose FreeDehaze, a novel training-free diffusion method for real-world image dehazing. FreeDehaze is a posterior-based framework capable of addressing non-linear dehazing challenges without relying on additional degradation estimation networks. It follows the human cognition for image restoration, beginning with perception and subsequently enhancing the image. The core method initially generates pseudo-clean images based on abstract textual descriptions. Subsequently, optimal transport aligns the denoising network output with the pseudo-clean image within a PCA-based haze subspace, facilitating high-fidelity dehazing. Extensive experiments demonstrate that FreeDehaze outperforms comparative methods in subjective metrics on challenging datasets (e.g., RTTS, URHI, and O-HAZE) and achieves competitive objective metrics, demonstrating strong generalization although without additional training.
Jiawei Wu 0001, Yikun Ma, Wenqi Ren, Zhi Jin 0002, Xiaochun Cao
IEEE Trans. Image Process.4
2026 Virtual Consistency Model for All-in-One Image Restoration
abstract
All-in-one Image Restoration (AIR) seeks to address diverse degradations using a unified model trained only once. Existing methods often rely on degradation-specific guidance, leading to conflicting gradients during training. In contrast, diffusion models offer a promising alternative by operating in a high-noise space where diverse degradations exhibit a homogeneous Gaussian distribution. This characteristic alleviates gradient conflicts associated with task-specific degradations. However, existing diffusion-based AIR methods often suffer from a lack of direct supervision in the image space, leading to error accumulation during the iterative denoising process and image fidelity compromisation. This highlights a fundamental dilemma for AIR: the optimal space for modeling degradations is inherently suboptimal for preserving image fidelity. To address this issue, we propose a Virtual Consistency Model for AIR (VCMAIR), which restores images in the high-noise space while employing a novel consistency function to enforce accurate supervision in the image space. Extensive experiments demonstrate that the proposed method outperforms existing state-of-the-art methods across a comprehensive benchmark of diverse degradation scenarios, including both standard AIR tasks and challenging real-world image restoration tasks.
Jiawei Wu 0001, Luwei Tu, Zhi Jin 0002, Kaihao Zhang, Wenqi Ren, Xiaochun Cao
IEEE Trans. Image Process.4
2025 MotionDiff: Training-Free Zero-Shot Interactive Motion Editing via Flow-Assisted Multi-View Diffusion
Yikun Ma, Jiawei Wu 0001, Zhi Jin 0002
ICCV5
2025 Accelerating Learned Video Compression via Low-Resolution Representation Learning
Zidian Qiu, Zongyao He, Zhi Jin 0002
ICIG (1)3
2025 FDG-Diff: Frequency-Domain-Guided Diffusion Framework for Compressed Hazy Image Restoration
abstract
In this study, we reveal that the interaction between haze degradation and JPEG compression introduces complex joint loss effects, which significantly complicate image restoration. Existing dehazing models often neglect compression effects, which limits their effectiveness in practical applications. To address these challenges, we introduce three key contributions. First, we design FDG-Diff, a novel frequency-domain-guided dehazing framework that improves JPEG image restoration by leveraging frequency-domain information. Second, we introduce the High-Frequency Compensation Module (HFCM), which enhances spatial-domain detail restoration by incorporating frequency-domain augmentation techniques into a diffusion-based restoration framework. Lastly, the introduction of the Degradation-Aware Denoising Timestep Predictor (DADTP) module further enhances restoration quality by enabling adaptive region-specific restoration, effectively addressing regional degradation inconsistencies in compressed hazy images. Experimental results across multiple compressed dehazing datasets demonstrate that our method consistently outperforms the latest state-of-the-art approaches.
Ruicheng Zhang, Kanghui Tian, Qixiang Liu, Zhi Jin 0002
ICME5
2025 MB-TaylorFormer V2: Improved Multi-Branch Linear Transformer Expanded by Taylor Formula for Image Restoration
abstract
Recently, Transformer networks have demonstrated outstanding performance in the field of image restoration due to the global receptive field and adaptability to input. However, the quadratic computational complexity of Softmax-attention poses a significant limitation on its extensive application in image restoration tasks, particularly for high-resolution images. To tackle this challenge, we propose a novel variant of the Transformer. This variant leverages the Taylor expansion to approximate the Softmax-attention and utilizes the concept of norm-preserving mapping to approximate the remainder of the first-order Taylor expansion, resulting in a linear computational complexity. Moreover, we introduce a multi-branch architecture featuring multi-scale patch embedding into the proposed Transformer, which has four distinct advantages: 1) various sizes of the receptive field; 2) multi-level semantic information; 3) flexible shapes of the receptive field; 4) accelerated training and inference speed. Hence, the proposed model, named the second version of Taylor formula expansion-based Transformer (for short MB-TaylorFormer V2) has the capability to concurrently process coarse-to-fine features, capture long-distance pixel interactions with limited computational cost, and improve the approximation of the Taylor expansion remainder. Experimental results across diverse image restoration benchmarks demonstrate that MB-TaylorFormer V2 achieves state-of-the-art performance in multiple image restoration tasks, such as image dehazing, deraining, desnowing, motion deblurring, and denoising, with very little computational overhead.
Zhi Jin 0002, Yuwei Qiu, Kaihao Zhang, Hongdong Li, Wenhan Luo
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Joint Resources Optimization for Soft Video Transmission Over IRS-Assisted SR Network
abstract
Intelligent reflective surface (IRS) assisted symbiotic radio (SR) network has been proposed as a promising solution for the sixth generation (6G) mobile wireless system, which achieves mutualistic spectrum sharing and highly reliable backscattering communication with extremely low energy cost. On the other hand, exponential growth in video traffic makes wireless video transmission more challenging in the 6G era. With the assistance of IRS based secondary link in SR network, an efficient soft video transmission scheme (IRSCast) is proposed to achieve linear quality transition under the drastically varying wireless channel. To minimize the transmission distortion of the video signal, a multivariable optimization problem is formulated to jointly optimize the wireless resources, including transmission power, active beamforming of the primary transmitter (PTx), and passive beamforming of the secondary transmitter (STx). Then, an alternating optimization method is utilized to decouple the multivariate optimization problem into multiple univariate sub-problems that are finally solved by semi-positive definite relaxation and Lagrange multiplier methods. The simulation results demonstrated that the proposed IRSCast method significantly improves the objective and subjective quality of the received video.
Lei Luo 0003, Zhi Jin 0002, Hongwei Guo 0001, Ce Zhu
IEEE Trans. Circuits Syst. Video Technol.3
2025 Fourier-Based Decoupling Network for Joint Low-Light Image Enhancement and Deblurring
abstract
Nighttime handheld photography is often simultaneously affected by low light and blur degradations due to object motion and camera shake. Previous methods typically design specific modules to restore the degradations in the spatial domain independently. However, the interdependence of low light and blur degradations in the spatial domain makes it difficult for these approaches to effectively decouple the degradations, limiting the performance of the designed modules. In this paper, we observe that in the Fourier domain, low light and blur degradations can be represented independently in the amplitude and phase of the image. Through an in-depth analysis of the underlying physical degradation process, we discover that low light degradation exhibits distinct characteristics across different frequency bands in amplitude, while blur degradation is characterized by phase correlation. Leveraging these insights, we mathematically derive a frequency attention mechanism and a filtering mechanism for learning decoupled representations of these degradations, proposing a Fourier-based Decoupling Network for joint low-light image enhancement and deblurring. Experimental results demonstrate that our method achieves the state-of-the-art performance on both synthetic and real-world datasets and exhibits significantly sharper edges. Code is available at https://github.com/Jabruson/FDN-TIP2025.
Luwei Tu, Jiawei Wu 0001, Deyu Meng, Zhi Jin 0002
IEEE Trans. Image Process.5
2025 Eliminating Moiré Patterns Across Diverse Image Resolutions via DMMNet
abstract
The occurrence of frequency aliasing between the camera and high-frequency scene elements causes moiré patterns in images, leading to color distortions and a loss of fine details, thereby reducing image quality. The intricate frequency characteristics and diverse appearances inherent in moiré patterns render their removal, commonly referred to as demoiréing, particularly challenging. Recent advancements in deep learning-based demoiréing methods have showcased notable efficacy. However, prevailing techniques often specialize in mitigating moiré patterns exclusively within either the frequency or spatial domains. Additionally, these methods generally perform well at specific image resolutions, but struggle to maintain effectiveness across different resolutions due to less generalized architectures. To address these issues, we propose a Dual-domain Multi-level Multi-scale Network DMMNet, working in both spatial and frequency domains sequentially. The Multi-scale Multi-level Demoire Stage (MMDS) in our framework focuses on moiré patterns removal in the spatial domain. To adeptly integrate features from various semantic levels, we introduce a pioneering plug-and-play Adjacent Cross Attention (ACA) module within the MMDS. Subsequently, the Frequency Separation and Reconstruction Stage (FSRS) restores high-frequency texture details, reconstructs color information, and eliminates residual moiré patterns in the wavelet frequency domain. Ultimately, the clean image is obtained by converting it back to the spatial domain. Extensive experimental assessments, spanning both quantitative metrics and qualitative visual evaluations, attest to the superior efficacy of DMMNet to State-Of-The-Art (SOTA) demoiréing methods, concurrently exhibiting enhanced generalization for demoiréing across diverse image resolutions. We posit that the proposed methodology presents a viable solution for broader applications in the realm of demoiréing. Code will be available onhttps://github.com/Mr-Ma-yikun/DMMNet.
Yikun Ma, Zhi Jin 0002
IEEE Trans. Multim.3
2025 Colorization-Inspired Customized Low-Light Image Enhancement by a Decoupled Network
abstract
Recently, numerous inspirational approaches have been proposed to enhance the visual quality of the images captured under poor lighting conditions. Simultaneously, in order to accommodate diverse user esthetics, researchers have explored customized operations within the enhancement process. However, most existing studies ignore the significance of the chrominance component, which often leads to unsatisfactory results in terms of color. To address this issue, we novelly decompose the low-light image enhancement (LLIE) task into the brightening and colorization subtasks and develop a decoupled network called CCNet for colorization-inspired customized enhancement. Specifically, the brightening subtask aims to restore images with normal contrast, less noise, and sharper details. While the colorization subtask utilizes the chrominance information from low-light images as color guidance to predict rich chrominance in enhanced images. Then, in the inference stage, users can adjust the color style or the saturation of color guidance to obtain customized results. Extensive experiments demonstrate that our proposed method achieves superior performance in both general and customized LLIE tasks-particularly in terms of improving chrominance components. Code is available at: https://github.com/FVL2020/CCNet.
Zhi Jin 0002
IEEE Trans. Neural Networks Learn. Syst.1
2024 Latent Modulated Function for Computational Optimal Continuous Image Representation
abstract
The recent work Local Implicit Image Function (LIIF) and subsequent Implicit Neural Representation (INR) based works have achieved remarkable success in Arbitrary-Scale Super-Resolution (ASSR) by using MLP to decode Low-Resolution (LR) features. However, these continuous image representations typically implement decoding in High-Resolution (HR) High-Dimensional (HD) space, leading to a quadratic increase in computational cost and seriously hindering the practical applications of ASSR. To tackle this problem, we propose a novel Latent Modulated Function (LMF), which decouples the HR-HD decoding process into shared latent decoding in LR-HD space and independent rendering in HR Low-Dimensional (LD) space, thereby realizing the first computational optimal paradigm of continuous image representation. Specifically, LMF utilizes an HD MLP in latent space to generate latent modulations of each LR feature vector. This enables a modulated LD MLP in render space to quickly adapt to any input feature vector and perform rendering at arbitrary resolution. Further-more, we leverage the positive correlation between modulation intensity and input image complexity to design a Controllable Multi-Scale Rendering (CMSR) algorithm, offering the flexibility to adjust the decoding efficiency based on the rendering precision. Extensive experiments demonstrate that converting existing INR-based ASSR methods to LMF can reduce the computational cost by up to 99.9%, accelerate inference by up to 57×, and save up to 76% of parameters, while maintaining competitive performance. The code is available at https://github.com/HeZongyao/LMF.
Zongyao He, Zhi Jin 0002
CVPR2
2024 Dynamic Implicit Image Function for Efficient Arbitrary-Scale Super-Resolution
abstract
Implicit Neural Representation (INR)-based methods have achieved remarkable success in Arbitrary-Scale Super-Resolution (ASSR). However, these continuous image representations, where a decoder infers pixel values across a continuous spatial domain, suffer from rapidly increasing computational cost as the scale factor increases, limiting the practical applications of ASSR. To address this problem, we propose a Dynamic Implicit Image Function (DIIF) for efficient ASSR. Instead of independently using each image coordinate and its nearby 2D features as decoder inputs, DIIF introduces a coordinate grouping and slicing strategy to decode pixel value slices from coordinate slices. To perform efficient arbitrary-scale decoding, we further introduce a dynamic coordinate slicing strategy empowered by our Coarse-to-Fine MLP (C2F-MLP), which allows adjusting the number of coordinates in each slice as the scale factor varies. Extensive experiments demonstrate that DIIF can seamlessly integrate with INR-based ASSR methods, significantly reducing computational cost and runtime, while maintaining State-Of-The-Art (SOTA) SR performance.
Zongyao He, Zhi Jin 0002
ICME2
2024 CAPformer: Compression-Aware Pre-trained Transformer for Low-Light Image Enhancement
abstract
Low-Light Image Enhancement (LLIE) has advanced with the surge in phone photography demand, yet many existing methods neglect compression, a crucial concern for resource-constrained phone photography. Most LLIE methods overlook this, hindering their effectiveness. In this study, we investigate the effects of JPEG compression on low-light images and reveal substantial information loss caused by JPEG due to widespread low pixel values in dark areas. Hence, we propose the Compression-Aware Pre-trained Transformer (CAPformer) network, employing a novel pre-training strategy to learn lossless information from uncompressed low-light images. Additionally, the proposed Brightness-Guided Self-Attention (BGSA) mechanism enhances rational information gathering. Experiments demonstrate the superiority of our approach in mitigating compression effects on LLIE, showcasing its potential for improving LLIE in resource-constrained scenarios.
Zhi Jin 0002
ICME2
2024 Understanding and improving zero-reference deep curve estimation for low-light image enhancement
Dandan Zhan, Zhi Jin 0002
Appl. Intell.3
2024 CSPN: A Category-Specific Processing Network for Low-Light Image Enhancement
abstract
Images captured in low-light conditions usually suffer from degradation problems. Recently, numerous deep learning-based methods are proposed for low-light image enhancement. They either focus on performance improvement with negligence of computational complicity, or are extremely computationally efficient networks with poor performance. In this work, we intend to figure out a solution, which strikes a balance between computational cost and performance. Moreover, we observe that different regions of an image contain different amounts of information, where the region with less information is easier to restore than that with more information. Hence, we propose to crop a low-light image into patches and classify these patches into “simple”, “medium” and “hard” categories based on their involved information. Then, we enhance different patch categories with different network complexities, therefore, a Category-specific Processing Network (CSPN) is proposed to achieve the computational complexity and performance balance. The patch classification is implemented by the proposed Grey-Level Co-occurrence Matrix (GLCM) entropy-based algorithm, which measures the content complexity of an image by analyzing the statistics of the difference between pixels. As the frequency domain contains exclusive feature information that is beneficial for improving image quality, the wavelet transform is introduced during the enhancement. Extensive experimental results demonstrate the superiority of our proposed CSPN over other state-of-the-art methods in various datasets with the least amount of computational cost.
Hongjun Wu 0003, Luwei Tu, Constantin Patsch, Zhi Jin 0002
IEEE Trans. Circuits Syst. Video Technol.5
2024 Learning From Text: A Multimodal Face Inpainting Network for Irregular Holes
abstract
Irregular hole face inpainting is a challenging task, since the appearance of faces varies greatly (e.g., different expressions and poses) and the human vision is more sensitive to subtle blemishes in the inpainted face images. Without external information, most existing methods struggle to generate new content containing semantic information for face components in the absence of sufficient contextual information. As it is known that text can be used to describe the content of an image in most cases, and is flexible and user-friendly. In this work, a concise and effective Multimodal Face Inpainting Network (MuFIN) is proposed, which simultaneously utilizes the information of the known regions and the descriptive text of the input image to address the problem of irregular hole face inpainting. To fully exploit the rest parts of the corrupted face images, a plug-and-play Multi-scale Multi-level Skip Fusion Module (MMSFM), which extracts multi-scale features and fuses shallow features into deep features at multiple levels, is illustrated. Moreover, to bridge the gap between textual and visual modalities and effectively fuse cross-modal features, a Multi-scale Text-Image Fusion Block (MTIFB), which incorporates text features into image features from both local and global scales, is developed. Extensive experiments conducted on two commonly used datasets CelebA and Multi-Modal-CelebA-HQ demonstrate that our method outperforms state-of-the-art methods both qualitatively and quantitatively, and can generate realistic and controllable results.
Dandan Zhan, Zhi Jin 0002
IEEE Trans. Circuits Syst. Video Technol.4
2024 Keypoints Filtrating Nonlinear Refinement in Spatial Target Pose Estimation with Deep Learning
abstract
Spatial target pose estimation with deep learning has garnered increasing attention in recent years. However, the existing methods in this field suffer from poor generalization. In this study, we propose a robust and reliable pose estimation method for spatial targets. The method aims to achieve keypoints filtrating. It involves a detection network tasked with identifying the target area, while the subsequent stage employs a classification network to regress keypoints from the detected target area. To improve the accuracy of pose estimation, we leverage spatial target geometric constraints to formulate 2-D–3-D keypoints equations for an initial pose. Then, we create a nonlinear optimization equation based on the confidence of 2-D keypoints and accomplish nonlinear refinement. We conduct extensive experiments on commonly used datasets and demonstrate the effectiveness of the proposed method. Furthermore, thanks to the effectiveness of keypoints filtrating and nonlinear refinement, the proposed method is robust with challenging scenarios and domain bias.
Lijun Zhong, Shengpeng Chen, Zhi Jin 0002, Pengyu Guo
IEEE Trans. Ind. Informatics3
2024 An Efficient Latent Style Guided Transformer-CNN Framework for Face Super-Resolution
abstract
In the Face Super-Resolution (FSR) task, it is important to precisely recover facial textures while maintaining facial contours for realistic high resolution faces. Although several CNN-based FSR methods have achieved great performance, they fail in restoring the facial contours due to the limitation of local convolutions. In contrast, Transformer-based methods which use self-attention as the basic component, are expert in modeling long-range dependencies between image patches. However, learning long-range dependencies often deteriorates facial textures due to the lack of locality. Therefore, a question is naturally raised:how to effectively combine the superiority of CNN and Transformer for better reconstructing faces?To address this issue, we propose an Efficient Latent Style guided Transformer-CNN framework for FSR calledELSFace, which can sufficiently integrate the advantages of CNN and Transformer. The framework consists of a Feature Preparation Stage and a Feature Carving Stage. Basic facial contours and textures are generated in the Feature Preparation Stage, and separately guided by latent styles, so that facial details are better represented in reconstruction. CNN and Transformer streams in the Feature Carving Stage are used to individually restore facial textures and facial contours, respectively in a parallel recursive way. Considering the negligence of high-frequency features when learning the long-range dependencies, we design the High-Frequency Enhancement Block (HFEB) in the Transformer stream. The Sharp Loss is also proposed for better perceptual quality in optimization. Extensive experimental results demonstrate that our ELSFace can achieve the best results among all metrics compared to the state-of-the-art CNN and Transformer-based methods on commonly used datasets and real-world tasks. Meanwhile, our ELSFace method has the least model parameters and running time. The codes are released athttps://github.com/FVL2020/ELSFace.
Yuwei Qiu, Zhi Jin 0002
IEEE Trans. Multim.4
2023 Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video Grounding
abstract
Spatio-Temporal Video Grounding (STVG) aims to localize the target object spatially and temporally according to the given language query. It is a challenging task in which the model should well understand dynamic visual cues (e.g., motions) and static visual cues (e.g., object appearances) in the language description, which requires effective joint modeling of spatiotemporal visuallinguistic dependencies. In this work, we propose a novel framework in which a static vision-language stream and a dynamic vision-language stream are developed to collaboratively reason the target tube. The static stream performs cross-modal understanding in a single frame and learns to attend to the target object spatially according to intraframe visual cues like object appearances. The dynamic stream models visual-linguistic dependencies across multiple consecutive frames to capture dynamic cues like motions. We further design a novel cross-stream collaborative block between the two streams, which enables the static and dynamic streams to transfer useful and complementary information from each other to achieve collaborative reasoning. Experimental results show the effectiveness of the collaboration of the two streams and our overall frame-work achieves new state-of-the-art performance on both HCSTVG and VidSTG datasets.
Zihang Lin, Chaolei Tan, Jianfang Hu, Zhi Jin 0002, Tiancai Ye, Wei-Shi Zheng 0001
CVPR4
2023 MB-TaylorFormer: Multi-branch Efficient Transformer Expanded by Taylor Formula for Image Dehazing
abstract
In recent years, Transformer networks are beginning to replace pure convolutional neural networks (CNNs) in the field of computer vision due to their global receptive field and adaptability to input. However, the quadratic computational complexity of softmax-attention limits the wide application in image dehazing task, especially for high-resolution images. To address this issue, we propose a new Transformer variant, which applies the Taylor expansion to approximate the softmax-attention and achieves linear computational complexity. A multi-scale attention refinement module is proposed as a complement to correct the error of the Taylor expansion. Furthermore, we introduce a multi-branch architecture with multi-scale patch embedding to the proposed Transformer, which embeds features by overlapping deformable convolution of different scales. The design of multi-scale patch embedding is based on three key ideas: 1) various sizes of the receptive field; 2) multi-level semantic information; 3) flexible shapes of the receptive field. Our model, named Multi-branch Transformer expanded by Taylor formula (MB-TaylorFormer), can em-bed coarse to fine features more flexibly at the patch embedding stage and capture long-distance pixel interactions with limited computational cost. Experimental results on several dehazing benchmarks show that MB-TaylorFormer achieves state-of-the-art (SOTA) performance with a light computational burden. The source code and pre-trained models are available at https://github.com/FVL2020/ICCV-2023-MB-TaylorFormer.
Yuwei Qiu, Kaihao Zhang, Wenhan Luo, Hongdong Li, Zhi Jin 0002
ICCV6
2023 Brighten-and-Colorize: A Decoupled Network for Customized Low-Light Image Enhancement
abstract
Low-Light Image Enhancement (LLIE) aims to improve the perceptual quality of an image captured in low-light conditions. Generally, a low-light image can be divided into lightness and chrominance components. Recent advances in this area mainly focus on the refinement of the lightness, while ignoring the role of chrominance. It easily leads to chromatic aberration and, to some extent, limits the diverse applications of chrominance in customized LLIE. In this work, a "brighten-and-colorize'' network (called BCNet), which introduces image colorization to LLIE, is proposed to address the above issues. BCNet can accomplish LLIE with accurate color and simultaneously enables customized enhancement with varying saturations and color styles based on user preferences. Specifically, BCNet regards LLIE as a multi-task learning problem: brightening and colorization. The brightening sub-task aligns with other conventional LLIE methods to get a well-lit lightness. The colorization sub-task is accomplished by regarding the chrominance of the low-light image as color guidance like the user-guide image colorization. Upon completion of model training, the color guidance (i.e., input low-light chrominance) can be simply manipulated by users to acquire customized results. This customized process is optional and, due to its decoupled nature, does not compromise the structural and detailed information of lightness. Extensive experiments on the commonly used LLIE datasets show that the proposed method achieves both State-Of-The-Art (SOTA) performance and user-friendly customization.
Zhi Jin 0002
ACM Multimedia2
2023 FourLLIE: Boosting Low-Light Image Enhancement by Fourier Frequency Information
abstract
Recently, Fourier frequency information has attracted much attention in Low-Light Image Enhancement (LLIE). Some researchers noticed that, in the Fourier space, the lightness degradation mainly exists in the amplitude component and the rest exists in the phase component. By incorporating both the Fourier frequency and the spatial information, these researchers proposed remarkable solutions for LLIE. In this work, we further explore the positive correlation between the magnitude of amplitude and the magnitude of lightness, which can be effectively leveraged to improve the lightness of low-light images in the Fourier space. Moreover, we find that the Fourier transform can extract the global information of the image, and does not introduce massive neural network parameters like Multi-Layer Perceptrons (MLPs) or Transformer. To this end, a two-stage Fourier-based LLIE network (FourLLIE) is proposed. In the first stage, we improve the lightness of low-light images by estimating the amplitude transform map in the Fourier space. In the second stage, we introduce the Signal-to-Noise-Ratio (SNR) map to provide the prior for integrating the global Fourier frequency and the local spatial information, which recovers image details in the spatial space. With this ingenious design, FourLLIE outperforms the existing state-of-the-art (SOTA) LLIE methods on four representative datasets while maintaining good model efficiency. Notably, compared with a recent Transformer-based SOTA method SNR-Aware, FourLLIE reaches superior performance with only 0.31% parameters. Code is available at https://github.com/wangchx67/FourLLIE
Hongjun Wu 0003, Zhi Jin 0002
ACM Multimedia3
2023 Quality of Task Perception based Performance Optimization of Time-delayed Teleoperation
abstract
This paper proposes a Quality-of-Task-Perception (QoTP) based performance optimization approach for bilateral haptic teleoperation. For time-delayed teleoperation, stabilizing control schemes are combined with communication and data reduction algorithms to ensure stability, transparency, and Quality of Experience (QoE). An adaptive control scheme switching strategy to improve the QoE of teleoperation considering network quality of service (QoS) and quality of control (QoC) is proposed in our previous work. In this paper, we introduce a novel concept named quality of task perception (QoTP) to optimize teleoperation from another dimension in addition to QoS and QoC. QoTP represents the pre-cognition of the task and the accuracy of the environment restoration. The proposed optimization approach is applied to a haptic teleoperation system with switchable control schemes (prediction-based or passivity-based). An environment restoration model is set on the leader side using the least squares method (LSM) to fit different environment models and provide force feedback without the influence of round-trip delay. We also evaluate the system performance with different delays, control schemes, and model complexities both objectively and subjectively. Our experiments validate the proposed approach and show that the QoE performance increases when selecting the more accurate environment restoration model in the QoTP dimension considering the system’s computing power.
Xiao Xu 0001, Zican Wang, Zhi Jin 0002, Eckehard G. Steinbach
RO-MAN5
2023 Reconstruction with robustness: A semantic prior guided face super-resolution framework for multiple degradations
Hongjun Wu 0003, Huanrong Zhang, Zhi Jin 0002, Driton Salihu, Jianfang Hu
Image Vis. Comput.4
2023 Estimating Human Weight From a Single Image
abstract
Body weight, as one of the biometric traits, has been studied in both the forensic and medical domains. However, estimating weight directly from 2-D images is particularly challenging since visual inspection is rather sensitive to the distance between the subject and camera, even for frontal view images. In this case, the widely used body mass index (BMI), which is associated with body height and weight, can be employed as a measure of weight to indicate health conditions. Previous works on the estimation of BMI have predominantly focused on using multiple 2-D images, 3-D images, or facial images; however, these cues are not always available. To address this issue, we explore the feasibility of obtaining BMI from a single 2-D body image with the dual-branch regression framework proposed in this work. More specifically, the framework comprises an anthropometric feature computation branch and a deep learning-based feature extraction branch. One aggregation layer maps all the features to an estimated BMI value. In addition, a new public 2-D image-to-BMI dataset, which contains 4189 images (1477 males and 2712 females) from approximately 3000 subjects with attributes including gender, age, height, and weight, was collected and released to facilitate the study. Extensive experiments confirm that the proposed framework combining anthropometric features and deep features outperforms the single-type feature approaches to BMI estimation in most cases.
Zhi Jin 0002, Junjia Huang, Wenjin Wang 0002, Aolin Xiong, Xiaojun Tan
IEEE Trans. Multim.1
2022 A Lightweight Image Entropy-Based Divide-and-Conquer Network for Low-Light Image Enhancement
abstract
Images captured in low-light conditions usually suffer from degradation problems. Based on the observation, we found that different image regions have different enhancement difficulties and can be processed by networks with different capacities. Hence, in this work, we propose a lightweight image entropy-based divide-and-conquer network called IEDCN for low-light image enhancement. Our network consists of Pre-processing, Enhancement, and Refinement three stages. In the Pre-processing Stage, we crop the low-light image into patches, and classify them into “simple”, “medium” and “hard” groups according to their image entropy. Then patches in each group are enhanced separately by corresponding branches with the divide-and-conquer strategy in the Enhancement Stage. Finally, the combined segments from the branches are refined by the last stage as the final output. Compared with other state-of-the-art methods, our IEDCN with only 0.73M parameters can effectively improve the quality of enhanced images, while saving up to 53% Flops on the LOL dataset.
Hongjun Wu 0003, Jingzhou Luo, Zhi Jin 0002
ICME5
2022 You Never Stop Dancing: Non-freezing Dance Generation via Bank-constrained Manifold Projection
abstract
One of the most overlooked challenges in dance generation is that the auto-regressive frameworks are prone to freezing motions due to noise accumulation. In this paper, we present two modules that can be plugged into the existing models to enable them to generate non-freezing and high fidelity dances. Since the high-dimensional motion data are easily swamped by noise, we propose to learn a low-dimensional manifold representation by an auto-encoder with a bank of latent codes, which can be used to reduce the noise in the predicted motions, thus preventing from freezing. We further extend the bank to provide explicit priors about the future motions to disambiguate motion prediction, which helps the predictors to generate motions with larger magnitude and higher fidelity than possible before. Extensive experiments on AIST++, a public large-scale 3D dance motion benchmark, demonstrate that our method notably outperforms the baselines in terms of quality, diversity and time length.
Jiangxin Sun, Huang Hu, Hanjiang Lai, Zhi Jin 0002, Jianfang Hu
NeurIPS5
2022 Self-distillation framework for indoor and outdoor monocular depth estimation
Meng Pan, Huanrong Zhang, Zhi Jin 0002
Multim. Tools Appl.4
2022 Depth-guided asymmetric CycleGAN for rain synthesis and image deraining
Yinhe Qi, Huanrong Zhang, Zhi Jin 0002, Wanquan Liu
Multim. Tools Appl.3
2022 Attention guided deep features for accurate body mass index estimation
Zhi Jin 0002, Junjia Huang, Aolin Xiong, Yuxian Pang, Wenjin Wang 0002, Beichen Ding
Pattern Recognit. Lett.1
2022 SRDRL: A Blind Super-Resolution Framework With Degradation Reconstruction Loss
abstract
Recent years have witnessed the remarkable success of deep learning-based single image super-resolution (SISR) methods. However, most of the existing SISR methods assume that low-resolution (LR) images are purely bicubic downsampled from high-resolution (HR) images. Once the actual degradation is not bicubic, their outstanding performance is hard to maintain. Since the real-world image degradation process can be modeled by a combination of downsampling, blurring, and noise, several SR methods have been proposed to super-resolve LR images with multiple blur kernels and noise levels. However, these SR methods require prior knowledge of the degradation process, which is difficult to obtain in practical applications. To address these issues, we propose a degradation reconstruction loss (DRL), which captures the degradation-wise differences between SR images and HR images via a degradation simulator. Empowered by the degradation simulator, the proposed loss, and an efficient SR network, a blind SR framework (SRDRL) without prior knowledge that can handle multiple degradations is formed. Extensive experimental results demonstrate that the proposed SRDRL outperforms the state-of-the-art blind SR methods and denosing+SR methods on multi-degraded datasets. The degradation reconstruction loss can be a plug-and-play loss for existing SR methods to handle multiple degradations. The source code can be found athttps://github.com/FVL2020/SRDRL.
Zongyao He, Zhi Jin 0002, Yao Zhao 0001
IEEE Trans. Multim.2
2021 Degradation Reconstruction Loss: A Perceptual-Oriented Super-Resolution Framework for Multi-downsampling Degradations
Zongyao He, Zhi Jin 0002, Xiao Xu 0001, Lei Luo 0003
ICIG (3)2
2021 Seeing Health with Eyes: Feature Combination for Image-Based Human BMI Estimation
abstract
Body Mass Index (BMI) is an important measurement of human obesity and health, which can provide useful information for plenty of practical purposes, such as monitoring, re-identification, and health care. Recently, some data-driven advances have been proposed to estimate BMI by 2D or 3D features from face images, frontal-body images and RGB-D images. However, due to the privacy issue or limitations of 3D cameras, the required data is hard to be obtained. More importantly, each of the previous works has only studied for a single type of features, hence it is worth investigating whether combinations of different features are more effective. To address this issue, we analyze the correlation of various features extracted from 2D body images with the estimated BMI, and then propose an accurate BMI estimation method with the optimal feature combination. Extensive experiments demonstrate that the proposed method outperforms these image-based BMI estimation methods which only utilize a single type feature in most cases. Code has been made available at : https://github.com/FVL2020/Features_for_BMI_estimation.
Junjia Huang, Chenming Shang, Aolin Xiong, Yuxian Pang, Zhi Jin 0002
ICME5
2021 SemFSR: An Unsupervised Face SR with Semantic Features for Multiple Degradations
abstract
Face Super-Resolution (FSR) field has witnessed significant progress with the development of deep learning, which is also widely applied in high-level vision tasks as the preprocessing step. Recently, FSR methods are exploring to utilize facial priors in the reconstruction of High-Resolution (HR) faces. However, facial priors directly extracted from Low-Resolution (LR) faces are less accurate or even unavailable. Meanwhile, the generalization ability of FSR methods can be further improved across different degradations, e.g., manual interpolated degradations, and real-world degradation with noise and blur. To tackle these problems, we propose a coarse-to-fine unsupervised FSR method based on semantic features called SemFSR under multiple degradations. The SemFSR contains the Degradation Stage and the Generation Stage. The Degradation Stage learns to generate degraded LR faces interfered with noise and blur. At the Generation Stage, we firstly reconstruct the "Coarse-SR" face from the degraded LR face for more accurate semantic features. Then we further propose the Channel Attention Block with Semantic features (CAB-S) and Semantic Loss to reconstruct the "Fine-SR" face. Both quantitative and qualitative experiments demonstrate the superiority of our SemFSR when encountering multiple degraded LR faces, compared with the state-of-the-art supervised and unsupervised FSR methods.
Huanrong Zhang, Zhi Jin 0002
ICTAI3
2021 When Face Completion Meets Irregular Holes: An Attributes Guided Deep Inpainting Network
abstract
Lots of convolutional neural network (CNN)-based methods have been proposed to implement face completion with regular holes. However, in practical applications, irregular holes are more common to see. Moreover, due to the distinct attributes and large variation of appearance for human faces, it is more challenging to fill irregular holes in face images while keeping content consistent with the rest region. Since facial attributes (e.g., gender, smiling, pointy nose, etc.) allow for a more understandable description of one face, they can provide some hints that benefit the face completion task. In this work, we propose a novel attributes-guided face completion network (AttrFaceNet), which comprises a facial attribute prediction subnet and a face completion subnet. The attribute prediction subnet predicts facial attributes from the rest parts of the corrupted images and guides the face completion subnet to fill the missing regions. The proposed AttrFaceNet is evaluated in an end-to-end way on commonly used datasets CelebA and Helen. Extensive experimental results show that our method outperforms state-of-the-art methods qualitatively and quantitatively especially in large mask size cases. Code is available at https://github.com/FVL2020/AttrFaceNet.
Dandan Zhan, Zhi Jin 0002
ACM Multimedia4
2021 A general model compression method for image restoration network
Zhi Jin 0002, Huanrong Zhang
Signal Process. Image Commun.2
2021 Dual-Stream Multi-Path Recursive Residual Network for JPEG Image Compression Artifacts Reduction
abstract
JPEG is the most widely used lossy image compression standard. When using JPEG with high compression ratios, visual artifacts cannot be avoided. These artifacts not only degrade the user experience but also negatively affect many low-level image processing tasks. Recently, convolutional neural network (CNN)-based compression artifact removal approaches have achieved significant success, however, at the cost of high computational complexity due to an enormous number of parameters. To address this issue, we propose a dual-stream recursive residual network (STRRN) which consists of structure and texture streams for separately reducing the specific artifacts related to high-frequency or low-frequency image components. The outputs of these streams are combined and fed into an aggregation network to further enhance the restored images. By using parameter sharing, the proposed network reduces the total number of training parameters significantly. Moreover, experiments conducted on five commonly used datasets confirm that the proposed STRRN can efficiently reduce the compression artifacts, while using up to 4.6 times less training parameters and 5 times less running time compared to the state-of-the-art approaches.
Zhi Jin 0002, Wenbin Zou, Xia Li 0006, Eckehard G. Steinbach
IEEE Trans. Circuits Syst. Video Technol.1
2021 A Visually Interpretable Deep Learning Framework for Histopathological Image-Based Skin Cancer Diagnosis
abstract
Owing to the high incidence rate and the severe impact of skin cancer, the precise diagnosis of malignant skin tumors is a significant goal, especially considering treatment is normally effective if the tumor is detected early. Limited published histopathological image sets and the lack of an intuitive correspondence between the features of lesion areas and a certain type of skin cancer pose a challenge to the establishment of high-quality and interpretable computer-aided diagnostic (CAD) systems. To solve this problem, a light-weight attention mechanism-based deep learning framework, namely, DRANet, is proposed to differentiate 11 types of skin diseases based on a real histopathological image set collected by us during the last 10 years. The CAD system can output not only the name of a certain disease but also a visualized diagnostic report showing possible areas related to the disease. The experimental results demonstrate that the DRANet obtains significantly better performance than baseline models (i.e., InceptionV3, ResNet50, VGG16, and VGG19) with comparable parameter size and competitive accuracy with fewer model parameters. Visualized results produced by the hidden layers of the DRANet actually highlight part of the class-specific regions of diagnostic points and are valuable for decision making in the diagnosis of skin diseases.
Shancheng Jiang, Huichuan Li, Zhi Jin 0002
IEEE J. Biomed. Health Informatics3
2020 Towards Lighter and Faster: Learning Wavelets Progressively for Image Super-Resolution
abstract
Due to the significant development of deep learning (DL) techniques, recent advances in the super-resolution (SR) field have achieved a great performance. While seeking for better performance, the later proposed networks prone to be deeper and heavier, which limits the applications of SR algorithms in the resource-constrain devices. Some advances rely on recurrent/recursive learning to reduce the number of network parameters, however, they ignore the caused long inference time, since the more recurrences/recursions are involved, the longer inference time the network needs. To address this trade-off issue between reconstruction performance, the number of network parameters, and inference time, we propose a lightweight and fast network (WSR) to learn wavelet coefficients of the target image progressively for single image super-resolution. More specifically, the network comprises two main branches. One is used for predicting the second level low-frequency wavelet coefficients, and the other one is designed in a recurrent way for predicting the rest wavelet coefficients at the first and second levels. Finally, an inverse wavelet transformation is adopted to reconstruct the SR images from these coefficients. In addition, we propose a deformable convolution kernel (side window) to construct the side-information multi-distillation block (S-IMDB), which is the basic unit of the recurrent blocks (RBs). We train the WSR with loss constraints at wavelet and spatial domains. Comprehensive experiments demonstrate that our WSR achieves a better trade-off than most of the state-of-the-art approaches. Code is available at https://github.com/FVL2020/WSR.
Huanrong Zhang, Zhi Jin 0002, Xiaojun Tan
ACM Multimedia2
2020 Video salient object detection via spatiotemporal attention neural networks
Yi Tang 0008, Wenbin Zou, Yang Hua 0001, Zhi Jin 0002, Xia Li 0006
Neurocomputing4
2020 A Flexible Deep CNN Framework for Image Restoration
abstract
Image restoration is a long-standing problem in image processing and low-level computer vision. Recently, discriminative convolutional neural network (CNN)-based approaches have attracted considerable attention due to their superior performance. However, most of these frameworks are designed for one specific image restoration task; hence, they seldom show high performance on other image restoration tasks. To address this issue, we propose a flexible deep CNN framework that exploits the frequency characteristics of different types of artifacts. Hence, the same approach can be employed for a variety of image restoration tasks by adjusting the architecture. For reducing the artifacts with similar frequency characteristics, a quality enhancement network that adopts residual and recursive learning is proposed. Residual learning is utilized to speed up the training process and boost the performance; recursive learning is adopted to significantly reduce the number of training parameters as well as boost the performance. Moreover, lateral connections transmit the extracted features between different frequency streams via multiple paths. One aggregation network combines the outputs of these streams to further enhance the restored images. We demonstrate the capabilities of the proposed framework with three representative applications: image compression artifacts reduction (CAR), image denoising, and single image super-resolution (SISR). Extensive experiments confirm that the proposed framework outperforms the state-of-the-art approaches on benchmark datasets for these applications.
Zhi Jin 0002, Dmytro Bobkov, Wenbin Zou, Xia Li 0006, Eckehard G. Steinbach
IEEE Trans. Multim.1
2019 An Efficient Quality Enhancement Solution for Stereo Images
Yingqing Peng, Zhi Jin 0002, Wenbin Zou, Yi Tang 0008, Xia Li 0006
ICIG (3)2
2019 Robust Plane Detection Using Depth Information From a Consumer Depth Camera
abstract
The emerging of depth-camera technology is paving the way for a variety of new applications and it is believed that plane detection is one of them. In fact, planes are common in man-made living structures, thus their accurate detection can benefit many visual-based applications. The use of depth information allows detecting planes characterized by complex pattern and texture, where the texture-based plane detection algorithms usually fail. In this paper, we propose a robust depth-driven plane detection (DPD) algorithm which consists of two parts: the growing-based plane detection and a two-stage refinement. The proposed approach starts from the seed patch with the highest planarity and uses the estimated equation of the growing plane and a dynamic threshold function to steer the growing process. Aided with this mechanism, each seed patch can grow to its maximum extent, and then the next seed patch starts to grow. This process is iteratively repeated so as to detect all the planes. Moreover, the refinement is proposed to tackle two common problems suffered by growing-based approaches, the over-growing problem, and the under-growing problem. Validated by extensive experiments, the proposed DPD algorithm is able to accurately detect planes and robust to various testing conditions. In terms of applications, it can be used as the pre-processing step for a variety of applications, such as, planar object recognition, super-resolution of the time-of-flight depth images with intrinsically low resolution.
Zhi Jin 0002, Tammam Tillo, Wenbin Zou, Yao Zhao 0001, Xia Li 0006
IEEE Trans. Circuits Syst. Video Technol.1
2019 Weakly Supervised Salient Object Detection With Spatiotemporal Cascade Neural Networks
abstract
Recently, deep learning techniques have substantially boosted the performance of salient object detection in still images. However, the salient object detection in videos by using traditional handcrafted features or deep learning features is not fully investigated, probably due to the lack of sufficient manually labeled video data for saliency modeling, especially for the data-driven deep learning. This paper proposes a novel weakly supervised approach to the salient object detection in a video, which can learn a robust saliency prediction model by using very limited manually labeled data and a large amount of weakly labeled data that could be easily generated in a supervised approach. Furthermore, we propose a spatiotemporal cascade neural network architecture for saliency modeling, in which two fully convolutional networks are cascaded to evaluate the visual saliency from both spatial and temporal cues to lead the optimal video saliency prediction. The proposed approach is extensively evaluated on the widely used challenging data sets, and the experiments demonstrate that our proposed approach substantially outperforms the state-of-the-art salient object detection models.
Yi Tang 0008, Wenbin Zou, Zhi Jin 0002, Yuhuan Chen, Yang Hua 0001, Xia Li 0006
IEEE Trans. Circuits Syst. Video Technol.3
2019 Joint Texture/Depth Power Allocation for 3-D Video SoftCast
abstract
Recently, a novel uncoded (pseudoanalog) scheme called SoftCast is proposed for wireless video transmission, which eliminates the cliff effect of the state-of-the-art source-channel coding based schemes and achieves linear quality transition within a wide range of channel signal-to-noise ratio. Therefore, SoftCast-like uncoded and hybrid transmission has become an attractive research issue for natural 2-D video. However, very few studies focus on the SoftCast-based wireless transmission of the 3-D video (3DV) currently. One critical issue of 3DV SoftCast is how to allocate the limited power budget of the transmitter to the texture videos and depth maps of the 3DV to achieve the optimal overall quality on the receiver side, including the transmission quality of the reference views and the synthesis quality of the virtual views. This paper attempts to solve the optimal joint power allocation problem in an efficient way. First, we formulate the target problem as a constrained power-distortion optimization (PDO) problem mathematically. Then, each part of the distortion is analyzed and formulated in a closed form. Finally, the PDO problem is mapped to an unconstrained convex optimization problem and solved by the Lagrangian multiplier method. Simulation results demonstrate that the performance of the proposed method is close to that of the full search method, which can provide the best performance theoretically. Nevertheless, the complexity of the proposed method is negligible compared with that of the full search method. In addition, as compared with the fixed ratio (e.g., 1:1) power allocation between texture and depth, the proposed method can achieve a PNSR gain up to 1.8 dB.
Lei Luo 0003, Taihai Yang, Ce Zhu, Zhi Jin 0002, Shu Tang
IEEE Trans. Multim.4
2018 Color Image Demosaicking Using a 3-Stage Convolutional Neural Network Structure
abstract
Color demosaicking (CDM) is a critical first step for the acquisition of high-quality RGB images with single chip cameras. Conventional CDM approaches are mostly based on interpolation schemes and hand-crafted image priors, which result in unpleasant visual artifacts in some cases. Motivated by the special characteristics of inter-channel correlations (higher correlations for R/G and G/B channels than that for R/B), in this paper, a 3-stage convolutional neural network (CNN) structure for CDM is proposed. In the first stage, the G channel is reconstructed independently. Then, by using the reconstructed G channel as guidance, the R and B channels are recovered in the second stage. Finally, high-quality RGB color images are reconstructed in the third stage. The objective and visual quality evaluation results show that the proposed structure achieves noticeable quality improvements in comparison to the state-of-the-art approaches.
Kai Cui 0003, Zhi Jin 0002, Eckehard G. Steinbach
ICIP2
2018 Multi-Scale Spatiotemporal Conv-LSTM Network for Video Saliency Detection
abstract
Recently, deep neural networks have been crucial techniques for image salient detection. However, two difficulties prevent the development of deep learning in video saliency detection. The first one is that the traditional static network cannot conduct a robust motion estimation in videos. The other is that the data-driven deep learning is in lack of sufficient manually annotated pixel-wise ground truths for video saliency network training. In this paper, we propose a multi-scale spatiotemporal convolutional LSTM network (MSST-ConvLSTM) to incorporate spatial and temporal cues for video salient objects detection. Furthermore, as manually pixel-wised labeling is very time-consuming, we sign lots of coarse labels, which are mixed with fine labels to train a robust saliency prediction model. Experiments on the widely used challenging benchmark datasets (e.g., FBMS and DAVIS) demonstrate that the proposed approach has competitive performance of video saliency detection compared with the state-of-the-art saliency models.
Yi Tang 0008, Wenbin Zou, Zhi Jin 0002, Xia Li 0006
ICMR3
2017 Multi-modal metric learning for vehicle re-identification in traffic surveillance environment
abstract
Vehicle re-identification (Re-Id) aims to retrieve the same vehicle captured by disjoint cameras at different time instants from different locations, and is a challenging task mainly due to the high similarity among the captured vehicle images in surveillance environment. With the rapid development of Convolutional Neural Network (CNN), learning-based deep features have been adopted to combine with hand-crafted features to re-identify vehicles in traffic surveillance environment. However, the two kinds of features are in different feature space, and if they are fused directly together, their complementary correlation is not able to be fully explored. To address such an issue, this paper proposes a multi-modal metric learning architecture to fuse deep features and hand-crafted ones in an end-to-end optimization network, which achieves a more robust and discriminative feature representation for vehicle re-identification. The extensive experiments on a large-scale traffic surveillance vehicle dataset demonstrate that our proposed approach substantially outperforms the state-of-the-art methods on vehicle Re-Id.
Yi Tang 0008, Di Wu 0009, Zhi Jin 0002, Wenbin Zou, Xia Li 0006
ICIP3
2017 A CNN cascade for quality enhancement of compressed depth images
abstract
Transmitting depth images along with the corresponding textures enables a wide range of receiver-side 3D applications. Since each pixel on the depth images represents a corresponding 3D scene geometric information, when compressed during transmission the compression artifacts will lead to severe geometry distortions and visual perceptual degradation. To solve this problem, in this paper we proposed a convolutional neural network (CNN) cascade for suppressing the compression artifacts on depth images. According to the feature of depth images, we furthermore, adopt a weighted loss function for network training which can adaptively improve the learning efficiency and accuracy. Meanwhile, in order to overcome the limited training data problem, we audaciously trained our network on textures first and then finetune on the target depth images. To our best knowledge, few works have applied CNN on depth images targeting for compression artifacts reduction (CAR). Through extensive experiments, our proposed solution achieves higher quality for both reconstructed depth images and synthesized virtual views than the state-of-the-art methods.
Zhi Jin 0002, Lei Luo 0003, Yi Tang 0008, Wenbin Zou, Xia Li 0006
VCIP1
2016 Virtual-View-Assisted Video Super-Resolution and Enhancement
abstract
A 3-D multiview video gives users an experience that is different from that provided by a traditional video; however, it puts a huge burden on limited bandwidth resources. Mixed-resolution video in a multiview system can alleviate this problem by using different video resolutions for different views. However, to reduce visual uncomfortableness and to make this video format more suitable for free-viewpoint television, the low-resolution (LR) views need to be super-resolved to the target full resolution. In this paper, we propose a virtual-view-assisted super-resolution algorithm, where the inter-view similarity is used to determine whether the missing pixels in the super-resolved frame need to be filled by virtual-view pixels or by spatial interpolated pixels. The decision mechanism is steered by the texture characteristics of the neighbors of each missing pixel. Furthermore, the inter-view similarity is used, on the one hand, to enhance the quality of the virtual-view-copied pixels by compensating the luminance difference between different views and, on the other hand, to enhance the original LR pixels in the super-resolved frame by reducing their compression distortion. Thus, the proposed method can recover the details in regions with edges while maintaining good quality at smooth areas by properly exploiting the high-quality virtual-view pixels and the directional correlation of pixels. The experimental results demonstrate the effectiveness of the proposed approach with a peak signal-to-noise ratio gain of up to 3.85 dB.
Zhi Jin 0002, Tammam Tillo, Jimin Xiao, Yao Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.1