Li Chen 0021

dblp:c/LiChen21 · DBLP profile ↗
← Back
44ranked-venue papers
0as first author
12since 2021 · last 2025
0000-0001-9899-2535ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 9 since 2021Artificial intelligence and machine learning · 7 · 3 since 2021Systems, architecture and hardware · 2
YearPublicationVenuePosition
2025 StreamLIC: A Lightweight Learned Image Compression Model for Stream Processing On FPGA
abstract
Learned image compression offers superior rate-distortion performance over conventional compression methods, but its high computational complexity remains a major barrier to hardware deployment. This paper tackles this challenge by introducing a lightweight learned image compression model, co-designed with a dedicated hardware architecture for efficient FPGA inference. By adopting a stream processing perspective, we propose several methods to mitigate inference bottlenecks while maintaining competitive performance: a block-sparse, row-wise autoregressive prediction model for low-latency entropy modeling, a hardware-friendly nonlinear module for efficient analysis and synthesis transforms, and multi-stage training to enhance model performance without inference overhead. Experimental results show our model achieves competitive compression performance with an exceptionally low operation count. Further-more, hardware simulations confirm that the proposed codec can achieve real-time processing on FPGA.
Hanlun Zhang, Li Chen 0021
VCIP2
2025 Exploring Bidirectional Bounds for Minimax-Training of Energy-Based Models
Cong Geng, Jia Wang 0004, Li Chen 0021, Jes Frellsen, Søren Hauberg
Int. J. Comput. Vis.3
2024 A Coding Framework and Benchmark Towards Low-Bitrate Video Understanding
abstract
Video compression is indispensable to most video analysis systems. Despite saving the transportation bandwidth, it also deteriorates downstream video understanding tasks, especially at low-bitrate settings. To systematically investigate this problem, we first thoroughly review the previous methods, revealing that three principles, i.e., task-decoupled, label-free, and data-emerged semantic prior, are critical to a machine-friendly coding framework but are not fully satisfied so far. In this paper, we propose a traditional-neural mixed coding framework that simultaneously fulfills all these principles, by taking advantage of both traditional codecs and neural networks (NNs). On one hand, the traditional codecs can efficiently encode the pixel signal of videos but may distort the semantic information. On the other hand, highly non-linear NNs are proficient in condensing video semantics into a compact representation. The framework is optimized by ensuring that a transportation-efficient semantic representation of the video is preserved w.r.t. the coding procedure, which is spontaneously learned from unlabeled data in a self-supervised manner. The videos collaboratively decoded from two streams (codec and NN) are of rich semantics, as well as visually photo-realistic, empirically boosting several mainstream downstream video analysis task performances without any post-adaptation procedure. Furthermore, by introducing the attention mechanism and adaptive modeling scheme, the video semantic modeling ability of our approach is further enhanced. Fianlly, we build a low-bitrate video understanding benchmark with three downstream tasks on eight datasets, demonstrating the notable superiority of our approach. All codes, data, and models will be open-sourced for facilitating future research.
Yuan Tian 0017, Guo Lu, Yichao Yan, Guangtao Zhai, Li Chen 0021
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Audio-Visual Saliency for Omnidirectional Videos
Xilei Zhu, Huiyu Duan, Kaiwei Zhang, Yucheng Zhu, Li Chen 0021, Xiongkuo Min, Guangtao Zhai
ICIG (5)7
2023 CLSA: A Contrastive Learning Framework With Selective Aggregation for Video Rescaling
abstract
Video rescaling has recently drawn extensive attention for its practical applications such as video compression. Compared to video super-resolution, which focuses on upscaling bicubic-downscaled videos, video rescaling methods jointly optimize a downscaler and a upscaler. However, the inevitable loss of information during downscaling makes the upscaling procedure still ill-posed. Furthermore, the network architecture of previous methods mostly relies on convolution to aggregate information within local regions, which cannot effectively capture the relationship between distant locations. To address the above two issues, we propose a unified video rescaling framework by introducing the following designs. First, we propose to regularize the information of the downscaled videos via a contrastive learning framework, where, particularly, hard negative samples for learning are synthesized online. With this auxiliary contrastive learning objective, the downscaler tends to retain more information that benefits the upscaler. Second, we present a selective global aggregation module (SGAM) to efficiently capture long-range redundancy in high-resolution videos, where only a few representative locations are adaptively selected to participate in the computationally-heavy self-attention (SA) operations. SGAM enjoys the efficiency of the sparse modeling scheme while preserving the global modeling capability of SA. We refer to the proposed framework as Contrastive Learning framework with Selective Aggregation (CLSA) for video rescaling. Comprehensive experimental results show that CLSA outperforms video rescaling and rescaling-based video compression methods on five datasets, achieving state-of-the-art performance.
Yuan Tian 0017, Yichao Yan, Guangtao Zhai, Li Chen 0021
IEEE Trans. Image Process.4
2022 Enhanced Deep Animation Video Interpolation
abstract
Existing learning-based frame interpolation algorithms extract consecutive frames from high-speed natural videos to train the model. Compared to natural videos, cartoon videos are usually in a low frame rate. Besides, the motion between consecutive cartoon frames is typically nonlinear, which breaks the linear motion assumption of interpolation algorithms. Thus, it is unsuitable for generating a training set directly from cartoon videos. For better adapting frame interpolation algorithms from nature video to animation video, we present AutoFI, a simple and effective method to automatically render training data for deep animation video interpolation. AutoFI takes a layered architecture to render synthetic data, which ensures the assumption of linear motion. Experimental results show that AutoFI performs favorably in training both DAIN and ANIN. However, most frame interpolation algorithms will still fail in error-prone areas, such as fast motion or large occlusion. Besides AutoFI, we also propose a plug-and-play sketch-based post-processing module, named SktFI, to refine the final results using user-provided sketches manually. With AutoFI and SktFI, the interpolated animation frames show high perceptual quality.
Wang Shen, Wenbo Bao, Guangtao Zhai, Li Chen 0021
ICIP5
2022 An Efficient Content-aware Downsampling-based Video Compression Framework
abstract
Recently, deep learning-based video compression algorithms have achieved competitive performance in Bjøntegaard delta (BD) rate, especially those adopting super-resolution networks as post-processing modules in downsampling-based video compression (DBC) frameworks. However, limited by the non-differentiable characteristics of traditional codecs, DBC frameworks mainly focus on improving the performance of super-resolution modules while ignoring optimizing downscaling modules. It is crucial to improve video compression performance without introducing additional modifications to the decoder client in practical application scenarios. We propose a context-aware processing network (CPN) compatible with standard codecs with no computational burden introduced to the client, which preserves the critical information and essential structures during downscaling. The proposed CPN works as a precoder cascaded by standard codecs to improve the compression performance on the server before encoding and transmission. Besides, a surrogate codec is employed to simulate the degradation process of the standard codecs and backpropagate the gradient to optimize the CPN. Experimental results show that the proposed method outperforms latest pre-processing networks and achieves considerable performance compared with the latest DBC frameworks.
Li Chen 0021
VCIP2
2022 Spatial Temporal Video Enhancement Using Alternating Exposures
abstract
High-speed video acquisition under poor illumination conditions is a challenging task. Imaging using long exposure can ensure brightness and suppress noise. However, the captured images may be blurry due to fast object movements or camera shakes. Imaging with short exposure can record sharp textures, but the high camera gain may cause noticeable noise. To alleviate this dilemma, we design a camera system using alternating exposures, where frames expose cyclically in a short-long way. The system consists of restoration and interpolation modules to reconstruct sharp, noise-reduced, high-frame-rate frames from low-frame-rate alternate-exposed input images. We design an optical-flow-based alternate-complementary alignment architecture for spatial enhancement, which effectively aligns the short-exposed and long-exposed images in a two-stage progressive way. Moreover, it explores complementary information from short-exposed and long-exposed inputs to ensure consistency between outputs. We propose a flow-enhanced frame interpolation module for temporal enhancement, which refines the intermediate flows and reconstructs the intermediate images based on the restored images of the alignment network and warped input neighboring frames. The whole network with two modules is end-to-end jointly learnable. We first evaluate the algorithm on simulation data. To demonstrate practicality, we then test it on real data by setting up a prototype camera. We propose an effective spatial degradation regularization strategy to reduce the domain gap between simulation and real data. Besides, we extend our method by integrating multi-frame exposure fusion technology to reduce overexposure areas in real scenarios. Experimental results show that our method performs favorably against state-of-the-art methods on both synthetic data and real-world data.
Wang Shen, Guo Lu, Guangtao Zhai, Li Chen 0021, Muhammad Salman Asif
IEEE Trans. Circuits Syst. Video Technol.5
2021 Video Compression based on Jointly Learned Down-Sampling and Super-Resolution Networks
abstract
With the blooming of deep learning technology in computer vision, the integration of deep learning and the traditional video coding has made significant improvements, especially applying the super-resolution neural network as the post-processing module in the down-sampling-based video compression framework. However, the pre-processing module lacks back-propagated gradients for jointly considering down-sampling and up-sampling due to the non-differentiability of the traditional video codec. In this paper, we propose an end- to-end down-sampling-based video compression framework applying convolutional neural networks both as down-sampling and upsampling. We use a virtual codec neural network to approximate the actual video codec so that the gradient can be effectively back-propagated for joint training. Experimental results show the superiority of our proposed framework compared with the predefined down-sampling-based video compression and various methods of joint training.
Yuzhuo Wei, Li Chen 0021, Li Song 0001
VCIP2
2021 Fast and Context-Aware Framework for Space-Time Video Super-Resolution
abstract
Increasing the spatial resolution and frame rate of a video simultaneously has attracted attention in recent years. The current one-stage space-time video super-resolution (STVSR) methods are difficult to deal with large motion and complex scenes, and are time-consuming and memory intensive. We propose an efficient STVSR framework, which can correctly handle complicated scenes such as occlusion and large motion and generate results with clearer texture. In REDS dataset, our method outperforms all existing one-stage methods. Our method is lightweight and can generate 720p frames at 16fps on a NVIDIA GTX 1080 Ti GPU.
Xueheng Zhang, Li Chen 0021, Li Song 0001
VCIP2
2021 An End-to-End Learning Framework for Video Compression
abstract
Traditional video compression approaches build upon the hybrid coding framework with motion-compensated prediction and residual transform coding. In this paper, we propose the first end-to-end deep video compression framework to take advantage of both the classical compression architecture and the powerful non-linear representation ability of neural networks. Our framework employs pixel-wise motion information, which is learned from an optical flow network and further compressed by an auto-encoder network to save bits. The other compression components are also implemented by the well-designed networks for high efficiency. All the modules are jointly optimized by using the rate-distortion trade-off and can collaborate with each other. More importantly, the proposed deep video compression framework is very flexible and can be easily extended by using lightweight or advanced networks for higher speed or better efficiency. We also propose to introduce the adaptive quantization layer to reduce the number of parameters for variable bitrate coding. Comprehensive experimental results demonstrate the effectiveness of the proposed framework on the benchmark datasets.
Guo Lu, Xiaoyun Zhang 0001, Wanli Ouyang, Li Chen 0021, Dong Xu 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Video Frame Interpolation and Enhancement via Pyramid Recurrent Framework
abstract
Video frame interpolation aims to improve users' watching experiences by generating high-frame-rate videos from low-frame-rate ones. Existing approaches typically focus on synthesizing intermediate frames using high-quality reference images. However, the captured reference frames may suffer from inevitable spatial degradations such as motion blur, sensor noise, etc. Few studies have approached the joint video enhancement problem, namely synthesizing high-frame-rate and high-quality results from low-frame-rate degraded inputs. In this paper, we propose a unified optimization framework for video frame interpolation with spatial degradations. Specifically, we develop a frame interpolation module with a pyramid structure to cyclically synthesize high-quality intermediate frames. The pyramid module features adjustable spatial receptive field and temporal scope, thus contributing to controllable computational complexity and restoration ability. Besides, we propose an inter-pyramid recurrent module to connect sequential models to exploit the temporal relationship. The pyramid module integrates the recurrent module, thus can iteratively synthesize temporally smooth results. And the pyramid modules share weights across iterations, thus it does not expand the model's parameter size. Our model can be generalized to several applications such as up-converting the frame rate of videos with motion blur, reducing compression artifacts, and jointly super-resolving low-resolution videos. Extensive experimental results demonstrate that our method performs favorably against state-of-the-art methods on various video frame interpolation and enhancement tasks.
Wang Shen, Wenbo Bao, Guangtao Zhai, Li Chen 0021, Xiongkuo Min
IEEE Trans. Image Process.4
2020 Blurry Video Frame Interpolation
abstract
Existing works reduce motion blur and up-convert frame rate through two separate ways, including frame deblurring and frame interpolation. However, few studies have approached the joint video enhancement problem, namely synthesizing high-frame-rate clear results from low-frame-rate blurry inputs. In this paper, we propose a blurry video frame interpolation method to reduce motion blur and up-convert frame rate simultaneously. Specifically, we develop a pyramid module to cyclically synthesize clear intermediate frames. The pyramid module features adjustable spatial receptive field and temporal scope, thus contributing to controllable computational complexity and restoration ability. Besides, we propose an inter-pyramid recurrent module to connect sequential models to exploit the temporal relationship. The pyramid module integrates a recurrent module, thus can iteratively synthesize temporally smooth results without significantly increasing the model size. Extensive experimental results demonstrate that our method performs favorably against state-of-the-art methods. The source code and pre-trained model are available at https://github.com/laomao0/BIN.
Wang Shen, Wenbo Bao, Guangtao Zhai, Li Chen 0021, Xiongkuo Min
CVPR4
2020 Content Adaptive and Error Propagation Aware Deep Video Compression
Guo Lu, Chunlei Cai, Xiaoyun Zhang 0001, Li Chen 0021, Wanli Ouyang, Dong Xu 0001
ECCV (2)4
2020 Adversarial Text Image Super-Resolution using Sinkhorn Distance
abstract
Convolutional neural network-based methods have demonstrated promising results for single image super-resolution. However, existing methods usually approach the problem on natural scenes rather than texts, whereas the latter can provide more informative messages to viewers. In this paper, instead of using the Lp-norm as the supervision metric, we propose a novel one for better preserving semantic information in text images. Our new metric combines optimal transport in a primal form with Sinkhorn distance defined in an adversarially learned feature space. Since the Sinkhorn distance measures the similarity between two features in terms of both feature components and spatial locations, our metric can maintain the spatial structure of texts during network optimization. Experimental results on text datasets show that our method performs favorably against state-of-the-art approaches in both quantitative and qualitative evaluations. We will publish the code, datasets, and models upon acceptance.
Cong Geng, Li Chen 0021, Xiaoyun Zhang 0001
ICASSP2
2020 End-to-End Optimized ROI Image Compression
abstract
Compressing an image with more bits automatically allocated to the region of interest (ROI) than to the background can both protect key information and reduce substantial redundancy. This paper models ROI image compression as an optimization problem of minimizing a weighted sum of the rate of the image and distortion of the ROI. The traditional framework solves this problem by cascading ROI prediction and ROI coding, through which achieving the optimized solution is impossible. To improve coding performance, we propose a novel deep-learning-based unified framework that can achieve rate distortion optimization for ROI compression. Specifically, the proposed framework includes a pair of ROI encoder and decoder convolutional neural networks and a learned entropy codec. The encoder network simultaneously generates multiscale representations that support efficient rate allocation and an implicit ROI mask that guides rate allocation. The proposed framework can automatically complete ROI image compression, and it can be optimized from data in an end-to-end manner. To effectively train the framework by back propagation, we develop a soft-to-hard ROI prediction scheme to make the entire framework differential. To improve visual quality, we propose a hierarchical distortion loss function to protect both pixel-level fidelity for ROI and structural similarity for the entire image. The proposed framework is implemented in two scenarios: salient-target and face-target ROI compression. Comparative experiments demonstrate the advantages of the proposed framework over the traditional framework, including considerably better subjective visual quality, significantly higher objective ROI compression performance and execution efficiency.
Chunlei Cai, Li Chen 0021, Xiaoyun Zhang 0001
IEEE Trans. Image Process.2
2020 Deep Non-Local Kalman Network for Video Compression Artifact Reduction
abstract
Video compression algorithms are widely used to reduce the huge size of video data, but they also introduce unpleasant visual artifacts due to the lossy compression. In order to improve the quality of the compressed videos, we proposed a deep non-local Kalman network for compression artifact reduction. Specifically, the video restoration is modeled as a Kalman filtering procedure and the decoded frames can be restored from the proposed deep Kalman model. Instead of using the noisy previous decoded frames as temporal information, the less noisy previous restored frame is employed in a recursive way, which provides the potential to generate high quality restored frames. In the proposed framework, several deep neural networks are utilized to estimate the corresponding states in the Kalman filter and integrated together in the deep Kalman filtering network. More importantly, we also exploit the non-local prior information by incorporating the spatial and temporal non-local networks for better restoration. Our approach takes the advantages of both the model-based methods and learning-based methods, by combining the recursive nature of the Kalman model and powerful representation ability of neural networks. Extensive experimental results on the Vimeo-90k and HEVC benchmark datasets demonstrate the effectiveness of our proposed method.
Guo Lu, Xiaoyun Zhang 0001, Wanli Ouyang, Dong Xu 0001, Li Chen 0021
IEEE Trans. Image Process.5
2019 A Novel Deep Progressive Image Compression Framework
abstract
In Internet applications, compressing the image without perceptually distinguishable distortions and loading the images without notable delays in the client end can significantly improve the user experience. Compressing the image at high bit rates can maintain the high quality of the decoded image but in cost of long transmitting and decoding time, resulting in bad user experience. The progressive coding scheme can resolve the conflict between the high quality requirement and the large loading delay. This paper proposes a novel efficient progressive image coding framework based on deep convolutional neural networks. The proposed framework is composed of a uniform encoder network and two progressive decoder networks. The encoder network decomposes the input image into two scales of representations, that can be transmitted and reconstructed progressively into a basic quality preview image and a high-quality image by two individual decoder networks respectively. All the networks are jointly learned when achieving the rate distortion optimization of both scales. Experiments results show that the proposed method has much better coding performance than the commercial codecs WebP and JPEG, which are commonly used in Internet applications. Meanwhile, the proposed codec consumes much less time to load the image compared with WebP.
Chunlei Cai, Li Chen 0021, Xiaoyun Zhang 0001, Guo Lu
PCS2
2019 Identifying and Pruning Redundant Structures for Deep Neural Networks
abstract
Deep convolutional neural networks have achieved considerable success in the field of computer vision. However, it is difficult to deploy state-of-the-art models on resource-constrained platforms due to their high storage, memory bandwidth, and computational costs. In this paper, we propose a structured pruning method which employs a three-step process to reduce the resource consumption of neural networks. First, we train an initial network on the training set and evaluate it on the validation set. Next, we introduce an iterative pruning and fine-tuning algorithm to identify and prune redundant structures, which results in a pruned network with a compact architecture. Finally, we train the pruned network from scratch on both the training set and validation set to obtain the final accuracy on the test set. In the experiments, our pruning method significantly reduces the model size (by 87.2% on CIFAR-10), saves inference time (53.3% on CIFAR-10), and achieves better performance as compared to recent state-of-the-art methods.
Wenyao Gan, Li Song 0001, Li Chen 0021, Rong Xie 0004, Xiao Gu 0001
VCIP3
2019 FPGA Based Video Transcoding System with 2K-4K Super-Resolution Conversion
abstract
We present a FPGA-based system supporting video stream transcoding with 2k full high-definition (FHD) video to 4k ultra high-definition (UHD) video super- resolution(SR) conversion. Our system focuses on building a functional pipeline with convolutional neural network (CNN) accelerator and real-time video codec unit for converting H.264 video stream to H.265/HEVC video stream. The overall video processing system can be used as an important plug-in module in the video streaming network to improve the video stream service quality.
Yuzhuo Wei, Li Chen 0021, Rong Xie 0004, Li Song 0001, Xiaoyun Zhang 0001
VCIP2
2019 Efficient Variable Rate Image Compression With Multi-Scale Decomposition Network
abstract
While deep learning image compression methods have shown an impressive coding performance, most of them output a single-optimized-compression rate using a trained-specific network. However, in practice, it is essential to support the variable rate compression or meet a target rate with a high-coding performance. This paper proposes a novel image compression method, making it possible for a single convolutional neural network (CNN) model to generate the variable rate efficiently with an optimized rate-distortion (RD) performance. The method consists of CNN-based multi-scale decomposition transform and content adaptive rate allocation. Specifically, the transform network is learned to decompose the input image into several scales of representations while optimizing the RD performance for all scales. Rate allocation algorithms for two typical scenarios are provided to determine the optimal scale of each image block for a given target rate or quality factor. For a target rate, the allocation is adaptive based on content complexity. In addition, for a target quality factor which indicates a tradeoff between the rate and the quality, the optimal scale is determined by minimizing the RD cost. The experimental results have shown that our method has outperformed the JPEG2000 and BPG standards with high efficiency and the state-of-the-art RD performance as measured by the multi-scale structural similarity index metric. Moreover, our method can strictly control the rate to generate the target compression result.
Chunlei Cai, Li Chen 0021, Xiaoyun Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2019 KalmanFlow 2.0: Efficient Video Optical Flow Estimation via Context-Aware Kalman Filtering
abstract
Recent studies on optical flow typically focus on the estimation of the single flow field in between a pair of images but pay little attention to the multiple consecutive flow fields in a longer video sequence. In this paper, we propose an efficient video optical flow estimation method by exploiting the temporal coherence and context dynamics under a Kalman filtering system. In this system, pixel's motion flow is first formulated as a second-order time-variant state vector and then optimally estimated according to the measurement and system noise levels within the system by maximum a posteriori criteria. Specifically, we evaluate the measurement noise according to the flow's temporal derivative, spatial gradient, and warping error. We determine the system noise based on the similarity of contextual information, which is represented by the compact features learned by pre-trained convolutional neural networks. The context-aware Kalman filtering helps improve the robustness of our method against abrupt change of light and occlusion/dis-occlusion in complicated scenes. The experimental results and analyses on the MPI Sintel, Monkaa, and Driving video datasets demonstrate that the proposed method performs favorably against the state-of-the-art approaches.
Wenbo Bao, Xiaoyun Zhang 0001, Li Chen 0021
IEEE Trans. Image Process.3
2018 Rcdfnn: Robust Change Detection Based on Convolutional Fusion Neural Network
abstract
Video change detection, which plays an important role in computer vision, is far from being well resolved due to the complexity of diverse scenes in real world. Most of the current methods are designed based on hand-crafted features and perform well in some certain scenes but may fail on others. This paper puts up forward a deep learning based method to automatically fuse multiple basic detections into an optimal one. Specifically, a convolutional fusion neural network is designed to obtain an adaptive fusion strategy based on features extracted from video content. Limited by the amount of available labeled dataset for change detection, this paper leverages an extractor that well trained on external dataset to improve generalization. Experiments show that the proposed method generates state-of-the-art result compared with nine recent outstanding algorithms and it performs well for diverse scenarios such as dynamic background, camera jitter and night videos.
Chunlei Cai, Li Chen 0021, Xiaoyun Zhang 0001
ICASSP2
2018 KalmanFlow: Efficient Kalman Filtering for Video Optical Flow
abstract
This paper proposes an efficient optical flow filtering method for video sequences. Motivated by the observation that motions in videos have strong temporal coherence, we use Kalman filtering to exploit this characteristic for more accurate flow fields. In the proposed system, pixel's motion flow is formulated as a time-variant state vector and optimally estimated by Kalman filter according to the noise level, which is evaluated using flow's temporal derivative, spatial gradient and matching error. Experiments on MPI Sintel video dataset demonstrate that the temporal coherence employed during Kalman filtering has the advantage of more consistent results, and can contribute to the state-of-the-art methods.
Wenbo Bao, Xiaoyun Zhang 0001, Li Chen 0021
ICIP3
2018 Frame Interpolation via Refined Deep Voxel Flow
abstract
Traditional frame interpolation methods first estimate motion between two consecutive frames and then synthesize intermediate frames. This problem is challenging because of complex motion and video scenes. In this paper, we present an end-to-end deep network for frame interpolation problem. Based on a video synthesis method deep voxel flow (DVF), refinement modules are designed to increase the accuracy of voxel flow, which we call Refined DVF (RDVF). A deeper architecture with more convolution and deconvolution layers is also utilized to help extract motion. Our results greatly improve the performance of original DVF and compare favorably to state-of-the-art methods both quantitatively and qualitatively.
Zhifeng Zhang 0003, Li Chen 0021, Rong Xie 0004, Li Song 0001
ICIP2
2018 A Wavelet-based Learning for Face Hallucination with Loop Architecture
abstract
Face hallucination is a specific super-resolution problem which aims to generate high-resolution(HR) faces from low-resolution(LR) input. Recently, deep learning methods have been widely applied in single-image super resolution. Considering face images have great similarities in both pixel value and global structure, we propose a wavelet-based deep learning method with loop architecture for face hallucination. In contrast to existing wavelet-based methods that generate wavelet coefficients independently without considering relationships between them, we propose a three-stage method with loop architecture. This alternately updated loop structure explores the statistical relationships among wavelet coefficients and has a maximum use of information flow with a small number of parameters. Because of multi-resolution property of wavelet transform, we adopt a mixed input strategy to train images with different sizes to realize multi-scale face hallucination without retraining and adding extra sub-networks. Experiments demonstrate that our method can get a robust performance with multi-scale face hallucination.
Cong Geng, Li Chen 0021, Xiaoyun Zhang 0001
VCIP2
2018 High-Order Model and Dynamic Filtering for Frame Rate Up-Conversion
abstract
This paper proposes a novel frame rate up-conversion method through high-order model and dynamic filtering (HOMDF) for video pixels. Unlike the constant brightness and linear motion assumptions in traditional methods, the intensity and position of the video pixels are both modeled with high-order polynomials in terms of time. Then, the key problem of our method is to estimate the polynomial coefficients that represent the pixel's intensity variation, velocity, and acceleration. We propose to solve it with two energy objectives: one minimizes the auto-regressive prediction error of intensity variation by its past samples, and the other minimizes video frame's reconstruction error along the motion trajectory. To efficiently address the optimization problem for these coefficients, we propose the dynamic filtering solution inspired by video's temporal coherence. The optimal estimation of these coefficients is reformulated into a dynamic fusion of the prior estimate from pixel's temporal predecessor and the maximum likelihood estimate from current new observation. Finally, frame rate up-conversion is implemented using motion-compensated interpolation by pixel-wise intensity variation and motion trajectory. Benefited from the advanced model and dynamic filtering, the interpolated frame has much better visual quality. Extensive experiments on the natural and synthesized videos demonstrate the superiority of HOMDF over the state-of-the-art methods in both subjective and objective comparisons.
Wenbo Bao, Xiaoyun Zhang 0001, Li Chen 0021, Lianghui Ding
IEEE Trans. Image Process.3
2018 Novel Integration of Frame Rate Up Conversion and HEVC Coding Based on Rate-Distortion Optimization
abstract
Frame rate up conversion (FRUC) can improve the visual quality by interpolating new intermediate frames. However, high frame rate videos by FRUC are confronted with more bitrate consumption or annoying artifacts of interpolated frames. In this paper, a novel integration framework of FRUC and high efficiency video coding (HEVC) is proposed based on rate-distortion optimization, and the interpolated frames can be reconstructed at encoder side with low bitrate cost and high visual quality. First, joint motion estimation (JME) algorithm is proposed to obtain robust motion vectors, which are shared between FRUC and video coding. What's more, JME is embedded into the coding loop and employs the original motion search strategy in HEVC coding. Then, the frame interpolation is formulated as a rate-distortion optimization problem, where both the coding bitrate consumption and visual quality are taken into account. Due to the absence of original frames, the distortion model for interpolated frames is established according to the motion vector reliability and coding quantization error. Experimental results demonstrate that the proposed framework can achieve 21% ~ 42% reduction in BDBR, when compared with the traditional methods of FRUC cascaded with coding.
Guo Lu, Xiaoyun Zhang 0001, Li Chen 0021
IEEE Trans. Image Process.3
2017 Spatiotemporal salient object detection based on distance transform and energy optimization
Bing Yang 0003, Xiaoyun Zhang 0001, Li Chen 0021
Neurocomputing3
2017 Edge guided salient object detection
Bing Yang 0003, Xiaoyun Zhang 0001, Li Chen 0021, Hua Yang 0001
Neurocomputing3
2016 Principal components analysis-based visual saliency detection
abstract
In this paper, a novel patch-wise saliency detection algorithm is proposed based on Principal Component Analysis (PCA). As a powerful statistical procedure in data analysis, PCA are fully exploited to convert color space and produce compact patch representation. Specifically, images are first converted to linearly uncorrelated channels and divided into non-overlapped patches. Then the patches are represented by the coefficients of principal components using PCA analysis. Based on the compact representation of patches, two types of distinctiveness are introduced: center-surround contrast and global rarity. Experimental results demonstrate that the PCA-based color space conversion and patch representation can improve the accuracy of human fixations prediction, and the proposed algorithm outperforms the mainstream algorithms on predicting human fixations.
Bing Yang 0003, Xiaoyun Zhang 0001, Jing Liu 0002, Li Chen 0021
ICASSP4
2016 Frame rate up-conversion based on motion-region segmentation
abstract
The key problem of frame rate up-conversion (FRUC) is to obtain true motion vectors (MV), especially for the motion boundaries. In this paper, we propose a novel FRUC algorithm based on motion-region segmentation. According to region's temporal consistency, motion-regions are determined by a categorization of detected feature points' true MVs. Then, constrained by MV's spatial smoothness within a region, true motions are propagated to the entire frame. This motion-region segmentation based method achieves truthful motion vector field and preferable interpolated frames. Experiments show that comparing to the state-of-art methods, the proposed algorithm produces videos with better quality in terms of objective and subjective evaluation.
Wenbo Bao, Xiaoyun Zhang 0001, Li Chen 0021
VCIP4
2015 Spatial Error Concealment With an Adaptive Linear Predictor
abstract
In this paper, a novel spatial error concealment (EC) algorithm is proposed. Under the sequential recovery framework, pixels in missing blocks are successively reconstructed based on adaptive linear predictor. The predictor automatically tunes its order and support shape according to local contexts. The predictor order and support shape are determined using Bayesian information criterion, which is able to strike a balance between the bias and variance of the prediction errors. The flexibility of the order-adaptive predictor is able to recover more important features or structures. A novel scan order based on the uncertainty of each pixel is also proposed to alleviate error propagation problem. Compared with the state-of-the-art EC algorithms, experimental results show that the proposed method gives better reconstruction performance in terms of objective and subjective evaluations.
Jing Liu 0002, Guangtao Zhai, Xiaokang Yang 0001, Bing Yang 0003, Li Chen 0021
IEEE Trans. Circuits Syst. Video Technol.5
2014 Lossless Predictive Coding for Images With Bayesian Treatment
abstract
Adaptive predictor has long been used for lossless predictive coding of images. Most of existing lossless predictive coding techniques mainly focus on suitability of prediction model for training set with the underlying assumption of local consistency, which may not hold well on object boundaries and cause large predictive error. In this paper, we propose a novel approach based on the assumption that local consistency and patch redundancy exist simultaneously in natural images. We derive a family of linear models and design a new algorithm to automatically select one suitable model for prediction. From the Bayesian perspective, the model with maximum posterior probability is considered as the best. Two types of model evidence are included in our algorithm. One is traditional training evidence, which represents the models’ suitability for current pixel under the assumption of local consistency. The other is target evidence, which is proposed to express the preference for different models from the perspective of patch redundancy. It is shown that the fusion of training evidence and target evidence jointly exploits the benefits of local consistency and patch redundancy. As a result, our proposed predictor is more suitable for natural images with textures and object boundaries. Comprehensive experiments demonstrate that the proposed predictor achieves higher efficiency compared with the state-of-the-art lossless predictors.
Jing Liu 0002, Guangtao Zhai, Xiaokang Yang 0001, Li Chen 0021
IEEE Trans. Image Process.4
2013 Occlusion handling frame rate up-conversion
abstract
Motion-compensated frame interpolation (MCFI) is a technique used extensively to enhance the temporal resolution of video sequences. In order to obtain a high quality interpolation, the motion vector field (MVF) between frames must be well-estimated. However, many current techniques for determining the MVF are prone to errors in occlusion regions. In this work, we propose an improved algorithm for improving the quality of MCFI by restoring the unreliable MVF and pixels in occlusion regions. We first utilize a dual motion estimation (DME) scheme which performs better in occlusion regions. Occlusion regions are determined by the ratio of two directional matching errors. Then, MVs in occlusion regions are refined using an orientation-based refinement (OBR) method, which promotes occluded MVs with its orthogonal neighboring MVF. Finally, regional blending (RB) is proposed to restore the unreliable pixels in occlusion regions for further error concealment. Experimental results demonstrate that the proposed algorithm provides a better quality than previous benchmark frame rate up-conversion (FRUC) methods both objectively and subjectively.
Li Chen 0021, Xiaoyun Zhang 0001
ICASSP3
2013 Sample-based image completion using structure synthesis
abstract
Image completion technique is widely used in image processing applications such as textural recovery, object removal, image edit, etc. When filling in the missing areas of an image, it is often a challenge to keep local consistency of image structures while avoiding ambiguity and visual artifacts. To tackle with this problem, we propose a robust sample-based image completion scheme which is a cascade of two major procedures. First, we extract structural information from both source and sample images and then perform template matching under constraints of boundary band map and contour consistency to reconstruct the damaged structures. Second, a weighted exemplar-based image synthesis algorithm is further devised taking the previous structural information and matching results into account. Extensive experiments and comparative study show the reliability and superiority of our image completion algorithm.
Chongwu Tang, Li Chen 0021, Guangtao Zhai, Xiaokang Yang 0001
ICIP3
2013 Effective early termination using adaptive search order for frame rate up-conversion
abstract
Motion Estimation (ME) palys a crucial part in frame rate up-conversion (FRUC) and video compression. One popular chip estimator, 3-D Recursive Search (3DRS), has achieved success in nowadays high definition televisions(HDTV). However, for future ultra high definition television (UHDTV) applications, the amount of computation increases exponentially. Therefore, there is an urgent need for further computational reduction. In this paper, we propose an effective early termination (ET) technique to reduce the calculation of absolute difference (AD), which is the most computation-intensive process in ME. First, motion vector (MV) candidates are examined in an adaptive search order (ASO) of their reliability to be true estimator. Then, the ET method dynamically changes the threshold for different motions to terminate accurately among reliable candidates. From the simulation, we can achieve an average computation reduction by 42% (up to 65%) with a slight PSNR degradation of 0.16dB on average.
Li Chen 0021, Xiaoyun Zhang 0001
ISCAS3
2013 Hybrid image interpolation with soft-decision kernel regression
abstract
Parametric linear autoregressive (AR) model has been widely used in image processing but is known to induce unstable results. The recently emerged nonparametric kernel regression is an effective structural method for forestalling outliers but often brings over-smoothed output. This paper introduces a hybrid algorithm for image interpolation through combining the strength of parametric and nonparametric modeling techniques. More specifically, it is a soft-decision kernel regression (SKR) method in which the soft-decision AR model is embedded into the adaptive kernel regression framework. Compared with the state-of-the-art interpolation methods, simulation results show that the proposed SKR algorithm achieves comparative or better results in terms of objective and subjective quality.
Jing Liu 0002, Xiaokang Yang 0001, Guangtao Zhai, Li Chen 0021
ISCAS4
2013 Lossless predictive coding with Bayesian treatment
abstract
Natural image statistics have been widely exploited for lossless predictive coding and other applications. However, traditional adaptive techniques always focus on the local consistency of training set regardless of what the predicted target looks like. We investigate the problem of introducing the model evidence of predicted target since self-similarity inherent in natural images gives some kind of prior information for the distribution of predicted result. The proposed Bayesian model integrated with both training evidence and target evidence takes full advantages of local structure as well as self-similarity. Experimental results demonstrate that the proposed context model achieves best results compared with the state-of-the-art lossless predictors.
Jing Liu 0002, Xiaokang Yang 0001, Guangtao Zhai, Li Chen 0021, Xianghui Sun, Wanhong Chen, Ying Zuo
VCIP4
2013 Sample-based image completion using structure synthesis
Chongwu Tang, Li Chen 0021, Guangtao Zhai, Xiaokang Yang 0001
J. Vis. Commun. Image Represent.3
2012 Nonlinear additive model based saliency map weighting strategy for image quality assessment
abstract
Most state-of-the-art image quality metrics are based on the two-step approach: local distortion/fidelity measurement and pooling. During the pooling stage, many weighting strategies have been proposed incorporating properties of the distortion itself, various masking effects and visual attention. Recently, researchers have devoted great enthusiasm and effort to the improvement of image quality assessment using visual saliency models. In this research, it is noticed that visual saliency features of both the original image and the distorted one have impacts on the process of image quality assessment. To reduce the overlapping effects, a nonlinear additive model is proposed to integrate saliency features from the original and distorted images towards improved error weighting results. Our extensive experimental studies on four publicly available image databases (LIVE, TID2008, CSIQ and A57) indicate that the proposed improved nonlinear additive model based saliency map weighting strategy constantly leads to higher prediction accuracy for image quality assessment than traditional methods.
Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Li Chen 0021, Wenjun Zhang 0001
MMSP4
2011 A fast video stabilization algorithm based on block matching and edge completion
abstract
Purpose of video stabilization is to register the frames of a video sequence with relative motions between each other to yield a stable video of higher perceptual quality. In this paper we focus on the problem of fast and robust video stabilization for the same scene based on temporal block matching. We use VoD principle to find the local motion vectors between adjacent frames, and then use statistical analysis to generate the global vibrant motion vector. After motion compensation, we further design an edge completion algorithm incorporating mosaicking and inpainting of neighbour frames, so as to reduce the impact of error propagation. Experimental results and comparative studies will be provided to justify the effectiveness of the proposed algorithm.
Chongwu Tang, Xiaokang Yang 0001, Li Chen 0021, Guangtao Zhai
MMSP3
2011 Example-based image contrast enhancement
abstract
In this paper, a novel example-based contrast enhancement algorithm is proposed. The proposed approach enhances the contrast by learning some important informative priors from the histogram of the example image. The experimental results indicate that the proposed Example-based Dist-Stretched (ExDS) contrast enhancement algorithm can boost the image contrast effectively. And thanks to the example-based learning process, the output images from the ExDS algorithm have more natural looking than those of traditional histogram equalization based methods. The proposed ExDS algorithm can also be extended to the applications of contrast correction for old film restoration as well as tone mapping for image and video post-productions.
Xiaokang Yang 0001, Li Chen 0021, Guangtao Zhai, Wenjun Zhang 0001
MMSP3
2010 Color transfer via local binary patterns mapping
abstract
Color transfer is a process of carrying over image colors from one image to another. Since images have diverse texture, color, content and other features, key challenge for color transfer is to find a correct mapping between image and target image. In this paper, a new color transfer method based on feature points extraction and local binary patterns(LBP) mapping is proposed. We construct a framework for feature points extraction and mapping. In the framework, we use SIFT descriptor for feature points selection and we design a LBP mapping method for feature points matching between source and target image. Final feature points in source image are obtained using this mapping. Greyscale source image and final feature points are used as inputs of a scribble-based colorization method to get the final color transfer result. Experimental results show the benefits of our method.
Xiaokang Yang 0001, Li Chen 0021
ICIP3