VLDB 2026 Research / reviewers in the wild / expert
Jie Liang 0001
dblp:51/239-1
· DBLP profile ↗
144ranked-venue papers
9as first author
23since 2021 · last 2026
0000-0003-3003-4343ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 107 · 9 first-author · 18 since 2021Computer networks · 17 · 1 since 2021Artificial intelligence and machine learning · 8 · 5 since 2021Databases, data management, data science and information retrieval · 7 · 2 since 2021Systems, architecture and hardware · 4Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spike-hammer: An efficient spike-driven hybrid architecture for multi-modal emotion recognition with physiological signals
Haisheng Fu, Yuchen Zou, Guohe Zhang, Jie Liang 0001 |
Neural Networks | 5 |
| 2025 | SELIC: Semantic-Enhanced Learned Image Compression via High-Level Textual GuidanceabstractLearned image compression (LIC) techniques have achieved remarkable progress; however, effectively integrating high-level semantic information remains challenging. In this work, we present a Semantic-Enhanced Learned Image Compression framework, termed SELIC, which leverages high-level textual guidance to improve rate-distortion performance. Specifically, SELIC employs a text encoder to extract rich semantic descriptions from the input image. These textual features are transformed into fixed-dimension tensors and seamlessly fused with the image-derived latent representation. By embedding the SELIC tensor directly into the compression pipeline, our approach enriches the bitstream without requiring additional inputs at the decoder, thereby maintaining fast and efficient decoding. Extensive experiments on benchmark datasets (e.g., Kodak) demonstrate that integrating semantic information substantially enhances compression quality. Our SELIC-guided method outperforms a baseline LIC model without semantic integration by approximately 0.1-0.15 dB across a wide range of bit rates in PSNR and achieves a 4.9% BD-rate improvement over VVC. Moreover, this improvement comes with minimal computational overhead, making the proposed SELIC framework a practical solution for advanced image compression applications. Haisheng Fu, Jie Liang 0001, Zhenman Fang, Jingning Han |
ICME | 2 |
| 2024 | Learned Image Compression with Dual-Branch Encoder and Conditional Information CodingabstractRecent advancements in deep learning-based image compression are notable. However, prevalent schemes that employ a serial context-adaptive entropy model to enhance rate-distortion (R-D) performance are markedly slow. Furthermore, the complexities of the encoding and decoding networks are substantially high, rendering them unsuitable for some practical applications. In this paper, we propose two techniques to balance the trade-off between complexity and performance. First, we introduce two branching coding networks to independently learn a low-resolution latent representation and a high-resolution latent representation of the input image, discriminatively representing the global and local information therein. Second, we utilize the high-resolution latent representation as conditional information for the low-resolution latent representation, furnishing it with global information, thus aiding in the reduction of redundancy between low-resolution information. We do not utilize any serial entropy models. Instead, we employ a parallel channel-wise auto-regressive entropy model for encoding and decoding low-resolution and high-resolution latent representations. Experiments demonstrate that our method is approximately twice as fast in both encoding and decoding compared to the parallelizable checkerboard context model, and it also achieves a 1.2% improvement in R-D performance compared to state-of-the-art learned image compression schemes. Our method also outperforms classical image codecs including H.266/VVC-intra (4:4:4) and some recent learned methods in rate-distortion performance, as validated by both PSNR and MS-SSIM metrics on the Kodak dataset. Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Zhenman Fang, Guohe Zhang, Jingning Han |
DCC | 3 |
| 2024 | WeConvene: Learned Image Compression with Wavelet-Domain Convolution and Entropy Model
Haisheng Fu, Jie Liang 0001, Zhenman Fang, Jingning Han, Feng Liang 0001, Guohe Zhang |
ECCV (50) | 2 |
| 2024 | Efficient Learned Image Compression with Selective Kernel Residual Module and Channel-Wise Causal Context ModelabstractRecently, learning-based image compression approaches have achieved superior performance over classical image compression methods. However, their complexities remain quite high. In this paper, we propose two efficient modules to reduce the complexity. First, we introduce a selective kernel residual module into the core network, which effectively expands the receptive field and captures global information. Second, we present an improved channel-wise causal context model, designed to not only reduce encoding and decoding time but also ensure rate-distortion performance. Experimental results demonstrate that our proposed method achieves better tradeoff than recent leading learned image compression methods, and also outperforms the latest H.266/VVC (4:4:4) in terms of PSNR and MS-SSIM metrics. Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Zhenman Fang, Guohe Zhang, Jingning Han |
ICASSP | 3 |
| 2024 | Fast and High-Performance Learned Image Compression With Improved Checkerboard Context Model, Deformable Residual Module, and Knowledge DistillationabstractDeep learning-based image compression has made great progresses recently. However, some leading schemes use serial context-adaptive entropy model to improve the rate-distortion (R-D) performance, which is very slow. In addition, the complexities of the encoding and decoding networks are quite high and not suitable for many practical applications. In this paper, we propose four techniques to balance the trade-off between the complexity and performance. We first introduce the deformable residual module to remove more redundancies in the input image, thereby enhancing compression performance. Second, we design an improved checkerboard context model with two separate distribution parameter estimation networks and different probability models, which enables parallel decoding without sacrificing the performance compared to the sequential context-adaptive model. Third, we develop a three-pass knowledge distillation scheme to retrain the decoder and entropy coding, and reduce the complexity of the core decoder network, which transfers both the final and intermediate results of the teacher network to the student network to improve its performance. Fourth, we introduce$L_{1}$regularization to make the numerical values of the latent representation more sparse, and we only encode non-zero channels in the encoding and decoding process to reduce the bit rate. This also reduces the encoding and decoding time. Experiments show that compared to the state-of-the-art learned image coding scheme, our method can be about 20 times faster in encoding and 70-90 times faster in decoding, and our R-D performance is also 2.3% higher. Our method achieves better rate-distortion performance than classical image codecs including H.266/VVC-intra (4:4:4) and some recent learned methods, as measured by both PSNR and MS-SSIM metrics on the Kodak and Tecnick-40 datasets. Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Zhenman Fang, Guohe Zhang, Jingning Han |
IEEE Trans. Image Process. | 3 |
| 2023 | ROI-Based Deep Image Compression with Swin TransformersabstractEncoding the Region Of Interest (ROI) with better quality than the background has many applications including video conferencing systems, video surveillance and object-oriented vision tasks. In this paper, we propose a ROI-based image compression framework with Swin transformers as main building blocks for the autoencoder network. The binary ROI mask is integrated into different layers of the network to provide spatial information guidance. Based on the ROI mask, we can control the relative importance of the ROI and non-ROI by modifying the corresponding Lagrange multiplier λ for different regions. Experimental results show our model achieves higher ROI PSNR than other methods and modest average PSNR for human evaluation. When tested on models pre-trained with original images, it has superior object detection and instance segmentation performance on the COCO validation dataset. Jie Liang 0001, Haisheng Fu, Jingning Han |
ICASSP | 2 |
| 2023 | Asymmetric Learned Image Compression With Multi-Scale Residual Block, Importance Scaling, and Post-Quantization FilteringabstractRecently, deep learning-based image compression has made significant progresses, and has achieved better rate-distortion (R-D) performance than the latest traditional method, H.266/VVC, in both MS-SSIM metric and the more challenging PSNR metric. However, a major problem is that the complexities of many leading learned schemes are too high. In this paper, we propose an efficient and effective image coding framework, which achieves similar R-D performance with lower complexity than the state of the art. First, we develop an improved multi-scale residual block (MSRB) that can expand the receptive field and capture global information more efficiently, which further reduces the spatial correlation of the latent representations. Second, an importance scaling network is introduced to directly scale the latents to achieve content-adaptive bit allocation without sending side information, which is more flexible than previous importance map methods. Third, we apply a post-quantization filter (PQF) to reduce the quantization error, motivated by the Sample Adaptive Offset (SAO) filter in video coding. Moreover, our experiments show that the performance of the system is less sensitive to the complexity of the decoder. Therefore, we design an asymmetric paradigm, in which the encoder employs three stages of MSRBs to improve the learning capacity, whereas the decoder only uses one stage of MSRB, which reduces the decoder complexity and still yields satisfactory performance. Experimental results show that compared to the state-of-the-art method, the encoding and decoding time of the proposed method are about 17 times faster, and the R-D performance is only reduced by about 1% on both Kodak and Tecnick-40 datasets, which is still better than H.266/VVC(4:4:4) and other leading learning-based methods. Our source code is publicly available athttps://github.com/fengyurenpingsheng. Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Guohe Zhang, Jingning Han |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Learned Image Compression With Gaussian-Laplacian-Logistic Mixture Model and Concatenated Residual ModulesabstractRecently deep learning-based image compression methods have achieved significant achievements and gradually outperformed traditional approaches including the latest standard Versatile Video Coding (VVC) in both PSNR and MS-SSIM metrics. Two key components of learned image compression are the entropy model of the latent representations and the encoding/decoding network architectures. Various models have been proposed, such as autoregressive, softmax, logistic mixture, Gaussian mixture, and Laplacian. Existing schemes only use one of these models. However, due to the vast diversity of images, it is not optimal to use one model for all images, even different regions within one image. In this paper, we propose a more flexible discretized Gaussian-Laplacian-Logistic mixture model (GLLMM) for the latent representations, which can adapt to different contents in different images and different regions of one image more accurately and efficiently, given the same complexity. Besides, in the encoding/decoding network design part, we propose a concatenated residual blocks (CRB), where multiple residual blocks are serially connected with additional shortcut connections. The CRB can improve the learning ability of the network, which can further improve the compression performance. Experimental results using the Kodak, Tecnick-100 and Tecnick-40 datasets show that the proposed scheme outperforms all the leading learning-based methods and existing compression standards including VVC intra coding (4:4:4 and 4:2:0) in terms of the PSNR and MS-SSIM. The source code is available at https://github.com/fengyurenpingsheng. Haisheng Fu, Feng Liang 0001, Bing Li 0022, Jie Liang 0001, Guohe Zhang, Dong Liu 0002, Chengjie Tu, Jingning Han |
IEEE Trans. Image Process. | 6 |
| 2023 | Auto-Weighted Layer Representation Based View Synthesis Distortion Estimation for 3-D Video CodingabstractRecently, various view synthesis distortion estimation models have been studied to better serve 3-D video coding. However, they can hardly model the relationship quantitatively among different levels of depth changes, texture degeneration, and view synthesis distortion (VSD), which is crucial for rate-distortion optimization and rate allocation. In this paper, an auto-weighted layer representation based view synthesis distortion estimation model is developed. Firstly, sub-VSD (S-VSD) is defined according to the level of depth changes and their associated texture degeneration. After that, a set of theoretical derivations demonstrate that the VSD can be approximately decomposed into the S-VSDs multiplied by their associated weights. To obtain the S-VSDs efficiently, a layer-based representation method is developed, where all the pixels with the same level of depth changes are represented with a layer. It enables the S-VSD calculation at the layer level. Meanwhile, a nonlinear mapping function is learnt to accurately represent the relationship between the VSD and S-VSDs, automatically providing weights for the S-VSDs during VSD estimation. To learn such a function, a dataset of the VSD and its associated S-VSDs are built, termed as VSDSet. Experimental results show that the VSD can be accurately estimated with the weights learnt by the nonlinear mapping function once its associated S-VSDs are available. The proposed method outperforms the relevant state-of-the-art methods in both accuracy and efficiency. The VSDSet and source code of the proposed method will be available athttps://github.com/jianjin008/. Xingxing Zhang 0001, Lili Meng, Weisi Lin, Jie Liang 0001, Huaxiang Zhang 0001, Yao Zhao 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | Learned Image Compression with Inception Residual Blocks and Multi-Scale Attention Module
Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Guohe Zhang, Jiangning Han |
PCS | 3 |
| 2022 | Region-of-interest and channel attention-based joint optimization of image compression and computer vision
Linwei Ye, Jie Liang 0001, Yang Wang 0003, Jingning Han |
Neurocomputing | 3 |
| 2022 | Outage Analysis of NOMA-Enabled Backscatter Communications With Intelligent Reflecting SurfacesabstractIntelligent reflecting surface (IRS) has emerged as a potential technology to achieve smart wireless communications and high energy efficiency. On the other hand, nonorthogonal multiple access (NOMA)-enabled backscatter communications have shown a great potential in large-scale Internet of Things (IoT) networks. In this article, we consider a downlink IRS-assisted backscatter communication with NOMA. We further consider a two-user scenario with channel disparity from the base station. We first derive the probability density function of the sum of the modulus of reflected channels, where each channel follows the Rayleigh distribution with dissimilar variances. The respective and generalized closed-form outage probability expressions are derived for the considered scenario. Simulation results validate the accuracy of the analytical outage probability expressions. We demonstrate that the far user can achieve a superior performance with the increase of reflecting elements or the reflection coefficients. Suyue Li, Lina Bariah, Sami Muhaidat, Anhong Wang, Jie Liang 0001 |
IEEE Internet Things J. | 5 |
| 2022 | AS-Net: An attention-aware downsampling network for point clouds oriented to classification tasks
Yakun Yang, Anhong Wang, Donghan Bu, Zewen Feng, Jie Liang 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2022 | Color-Sensitivity-Based Rate-Distortion Optimization for H.265/HEVCabstractRate-Distortion Optimization (RDO) is an important step in video coding to achieve the best quality under a certain compression ratio constraint. The traditional RDO assigns equal importance to different color components. However, Human Visual System (HVS) has different sensitivities to different components. In this paper, the color-sensitivity-based combined PSNR (CSPSNR) is utilized as the distortion measurement in the process of RDO, where the characteristics of the color sensitivities of HVS are taken into account. Firstly, the distortion weights of luma and chroma components are derived from the criterion of maximizing CSPSNR. Then Lagrange multiplier and quantization parameter (QP) are adjusted according to the variation of distortion weights among different components. Finally, the CSPSNR-based RDO (CSRDO) adaptively calculates the RD costs of luma and chroma components under different sampling rates to improve the coding efficiency of the whole sequence. Experimental results in H.265/HEVC demonstrate that the proposed method can achieve 3.11% and 3.58% BD-RATE gain for AI and RA configurations in terms of CSPSNR on average. Xiwu Shang, Jie Liang 0001, Xiaoli Zhao 0003, Hua Han 0002, Yifan Zuo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Laplacian Pyramid Dense Network for Hyperspectral PansharpeningabstractHyperspectral (HS) pansharpening aims to create a pansharpened image that integrates the spatial details of the panchromatic (PAN) image and the spectral content of the HS image. In this article, we present a deep convolutional network within the mature Gaussian–Laplacian pyramid for pansharpening (LPPNet). The overall structure of LPPNet is a cascade of the Laplacian pyramid dense network with a similar structure at each pyramid level. Following the general idea of multiresolution analysis (MRA), the subband residuals of the desired HS images are extracted from the PAN image and injected into the upsampled HS image to reconstruct the high-resolution HS images level by level. Applying the mature Laplace pyramid decomposition technique to the convolution neural network (CNN) can simplify the pansharpening problem into several pyramid-level learning problems so that the pansharpening problem can be solved with a shallow CNN with fewer parameters. Specifically, the Laplacian pyramid technology is used to decompose the image into different levels that can differentiate large- and small-scale details, and each level is handled by a spatial subnetwork in a divide-and-conquer way to make the network more efficient. Experimental results show that the proposed LPPNet method performs favorably against some state-of-the-art pansharpening methods in terms of objective indexes and subjective visual appearance. Wenqian Dong, Tongzhen Zhang, Jiahui Qu, Song Xiao 0001, Jie Liang 0001, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Learned Bi-Resolution Image Coding using Generalized Octave ConvolutionsabstractLearned image compression has recently shown the potential to outperform the standard codecs. State-of-the-art rate-distortion (R-D) performance has been achieved by context-adaptive entropy coding approaches in which hyperprior and autoregressive models are jointly utilized to effectively capture the spatial dependencies in the latent representations. However, the latents are feature maps of the same spatial resolution in previous works, which contain some redundancies that affect the R-D performance. In this paper, we propose a learned bi-resolution image coding approach that is based on the recently developed octave convolutions to factorize the latents into high and low resolution components. Therefore, the spatial redundancy is reduced, which improves the R-D performance. Novel generalized octave convolution and octave transposed-convolution architectures with internal activation layers are also proposed to preserve more spatial structure of the information. Experimental results show that the proposed scheme outperforms all existing learned methods as well as standard codecs such as the next-generation video coding standard VVC (4:2:0) in both PSNR and MS-SSIM. We also show that the proposed generalized octave convolution can improve the performance of other auto-encoder-based schemes such as semantic segmentation and image denoising. Jie Liang 0001, Jingning Han, Chengjie Tu |
AAAI | 2 |
| 2021 | Modulated Variable-Rate Deep Video CompressionabstractRate adaption is one of the decisive factors for the applications of video compression. However, previous deep video compression methods are usually optimized for a single fixed rate-distortion (R-D) tradeoff. While they can achieve multiple bitrates by training multiple independent models, the realized bitrates are limited to several discrete points on the R-D curve and the storage cost increases proportionally to the number of models. In this paper, we propose a variable-rate scheme for deep video compression, which can achieve continuously variable rate by a single model, i.e., it can reach any point on the R-D curve. In our scheme, two deep auto-encoders are used to compress the residual and the motion vector field respectively, which directly generate the final bitstream. The basic rate adaptation can be achieved by using the R-D tradeoff parameter to deeply modulate all the internal feature maps of the auto-encoders. However, other modules in our scheme, notably motion estimation and motion compensation, also affect the final bitrate indirectly. We further use the R-D tradeoff parameter to modulate them via a conditional map, which effectively improves the compression efficiency. We use a multi-rate-distortion loss function together with a step-by-step training strategy to optimize the entire scheme. Our experiments show that the proposed scheme achieves continuously variable rate by a single model with almost the same compression efficiency as multiple fixed-rate models. The additional parameters and computation of our model are negligible when compared with a single fixed-rate model. Dong Liu 0002, Jie Liang 0001, Houqiang Li, Feng Wu 0001 |
DCC | 3 |
| 2021 | A Deeply Modulated Scheme for Variable-Rate Video CompressionabstractRate adaption is one of the decisive factors for the applications of video compression. Previous deep video compression methods are usually optimized for a single fixed rate-distortion (R-D) tradeoff. While they can achieve multiple bitrates by training multiple independent models, the achievable bitrates are limited to several discrete points on the R-D curve and the storage cost increases proportionally to the number of models. We propose a variable-rate scheme for deep video compression, which can achieve continuously variable rate by a single model, i.e., reaching any point on the R-D curve. In our scheme, two deep auto-encoders are used to compress the residual and the motion vector field respectively, which directly generate the final bitstream. The basic rate adaptation can be achieved by using the R-D tradeoff parameter to deeply modulate all the internal feature maps of the auto-encoders. In addition, other modules in our scheme, notably motion estimation and motion compensation, also affect the final bitrate indirectly. We further use the R-D tradeoff parameter to modulate them via a conditional map, thereby effectively improving the compression efficiency. We use a multi-rate-distortion loss function together with a step-by-step training strategy to optimize the entire scheme. The experimental results show the proposed scheme achieves continuously variable rate by a single model with almost the same compression efficiency as multiple fixed-rate models. The additional parameters and computation of our model are negligible when compared with a single fixed-rate model. Dong Liu 0002, Jie Liang 0001, Houqiang Li, Feng Wu 0001 |
ICIP | 3 |
| 2021 | Unsupervised stereoscopic image retargeting via view synthesis and stereo cycle consistency losses
Xiaoting Fan, Jianjun Lei 0001, Jie Liang 0001, Yuming Fang 0001, Xiaochun Cao, Nam Ling |
Neurocomputing | 3 |
| 2021 | An extended context-based entropy hybrid modeling for image compression
Haisheng Fu, Feng Liang 0001, Qian Zhang 0082, Jie Liang 0001, Chengjie Tu, Guohe Zhang |
Signal Process. Image Commun. | 5 |
| 2021 | Stereoscopic Image Retargeting Based on Deep Convolutional Neural NetworkabstractStereoscopic image retargeting aims at converting stereoscopic images to the target resolution adaptively. Different from 2D image retargeting, stereoscopic image retargeting needs to preserve both the shape structure of salient objects and depth consistency of 3D scenes. In this paper, we present a stereoscopic image retargeting method based on deep convolutional neural network to obtain high-quality retargeted images with both object shape preservation and scene depth preservation. First, a cross-attention extraction mechanism is constructed to generate attention map, which contains the valuable attention features of the left and right images and the common attention features between them. Second, since the disparity map can provide accurate depth information of objects in 3D scenes, a disparity-assisted 3D significance map generation module is utilized to further preserve the valuable depth information of stereoscopic images. Finally, in order to predict the retargeted stereoscopic images accurately, an image consistency loss is developed to preserve the geometric structure of salient objects, and a disparity consistency loss is introduced to eliminate depth distortions. Experimental results demonstrate that the proposed deep convolutional neural network can provide favorable stereoscopic image retargeting results. Xiaoting Fan, Jianjun Lei 0001, Jie Liang 0001, Yuming Fang 0001, Nam Ling, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Learned Multi-Resolution Variable-Rate Image Compression With Octave-Based Residual BlocksabstractRecently deep learning-based image compression has shown the potential to outperform traditional codecs. However, most existing methods train multiple networks for multiple bit rates, which increase the implementation complexity. In this paper, we propose a new variable-rate image compression framework, which employs generalized octave convolutions (GoConv) and generalized octave transposed-convolutions (GoTConv) with built-in generalized divisive normalization (GDN) and inverse GDN (IGDN) layers. Novel GoConv- and GoTConv-based residual blocks are also developed in the encoder and decoder networks. Our scheme also uses a stochastic rounding-based scalar quantization. To further improve the performance, we encode the residual between the input and the reconstructed image from the decoder network as an enhancement layer. To enable a single model to operate with different bit rates and to learn multi-rate image features, a new objective function is introduced. Experimental results show that the proposed framework trained with variable-rate objective function outperforms the standard codecs such as H.265/HEVC-based BPG and state-of-the-art learning-based variable-rate methods. Jie Liang 0001, Jingning Han, Chengjie Tu |
IEEE Trans. Multim. | 2 |
| 2020 | Deep Learning-Based Image Compression with Trellis Coded QuantizationabstractRecently many works attempt to develop image compression models based on deep learning architectures, where the uniform scalar quantizer (SQ) is commonly applied to the feature maps between the encoder and decoder. In this paper, we propose to incorporate trellis coded quantizer (TCQ) into a deep learning based image compression framework. A soft-to-hard strategy is applied to allow for back propagation during training. We develop a simple image compression model that consists of three subnetworks (encoder, decoder and entropy estimation), and optimize all of the components in an end-to-end manner. We experiment on two high resolution image datasets and both show that our model can achieve superior performance at low bit rates. We also show the comparisons between TCQ and SQ based on our proposed baseline model and demonstrate the advantage of TCQ. Jie Liang 0001, Yang Wang 0003 |
DCC | 3 |
| 2020 | Learned Variable-Rate Image Compression With Residual Divisive NormalizationabstractRecently deep learning-based image compression has shown the potential to outperform traditional codecs. However, most existing methods train multiple networks for multiple bit rates, which increases the implementation complexity. In this paper, we propose a variable-rate image compression framework, which employs more Generalized Divisive Normalization (GDN) layers than previous GDN-based methods. Novel GDN-based residual sub-networks are also developed in the encoder and decoder networks. Our scheme also uses a stochastic rounding-based scalar quantization. To further improve the performance, we encode the residual between the input and the reconstructed image from the decoder network as an enhancement layer. To enable a single model to operate with different bit rates and to learn multi-rate image features, a new objective function is introduced. Experimental results show that the proposed framework trained with variable-rate objective function outperforms all standard codecs such as H.265/HEVC-based BPG and state-of-the-art learning-based variable-rate methods. Jie Liang 0001, Jingning Han, Chengjie Tu |
ICME | 2 |
| 2020 | Variable-Rate Multi-Frequency Image Compression using Modulated Generalized Octave ConvolutionabstractIn this proposal, we design a learned multi-frequency image compression approach that uses generalized octave convolutions to factorize the latent representations into high-frequency (HF) and low-frequency (LF) components, and the LF components have lower resolution than HF components, which can improve the rate-distortion performance, similar to wavelet transform. Moreover, compared to the original octave convolution, the proposed generalized octave convolution (GoConv) and octave transposed-convolution (GoTConv) with internal activation layers preserve more spatial structure of the information, and enable more effective filtering between the HF and LF components, which further improve the performance. In addition, we develop a variable-rate scheme using the Lagrangian parameter to modulate all the internal feature maps in the autoencoder, which allows the scheme to achieve the large bitrate range of the JPEG AI with only three models. Experiments show that the proposed scheme achieves much better Y MS-SSIM than VVC. In terms of YUV PSNR, our scheme is very similar to HEVC. Haisheng Fu, Qian Zhang 0082, Shang Wang 0006, Jie Liang 0001, Dong Liu 0002, Feng Liang 0001, Guohe Zhang, Chengjie Tu |
MMSP | 6 |
| 2020 | Fusion of hyperspectral and panchromatic images using structure tensor and matting model
Wenqian Dong, Song Xiao 0001, Jie Liang 0001, Jiahui Qu |
Neurocomputing | 3 |
| 2020 | Light field all-in-focus image fusion based on spatially-guided angular information
Yingchun Wu, Yumei Wang, Jie Liang 0001, Ivan V. Bajic, Anhong Wang |
J. Vis. Commun. Image Represent. | 3 |
| 2020 | Super resolution of single depth image based on multi-dictionary learning with edge feature regularization
Anhong Wang, Hong Shangguan, Yingchun Wu, Donghong Li, Youcheng Wu, Jie Liang 0001 |
Multim. Tools Appl. | 7 |
| 2020 | Improved hybrid layered image compression using deep learning and traditional codecs
Haisheng Fu, Feng Liang 0001, Nai Bian, Qian Zhang 0082, Jie Liang 0001, Chengjie Tu |
Signal Process. Image Commun. | 7 |
| 2020 | EAAT: Environment-Aware Adaptive Transmission for Split-Screen Video StreamingabstractWith the tremendous growth of video contents and mobility demands, there is a need to develop more personalized video services. Split-screen services, such as picture in picture become more and more popular. Furthermore, the user's viewing environment affects the user's quality of experience (QoE). Therefore, video transmission of split-screen services face several major challenges, such as to quantify the impact of environmental factors on user's QoE; how to assess the user's QoE of the split-screen services; how to choose the bit-rate of each video stream to maximize user's QoE of the split-screen services. To address these challenges, in the paper, an environment-aware adaptive transmission (EAAT) scheme for split-screen video streaming is first presented. Then, we introduce a mathematical model for characterizing user's QoE to be affected by environmental factors in the proposed EAAT. In the model, the QoE of user's relationship with the viewing environment is proposed. Based on the model, a problem of maximizing user's QoE is formulated, and we develop a heuristic algorithm to solve the optimization problem. In addition, we conduct various trace-bandwidth experiments to rigorously evaluate the proposed EAAT scheme in different network environments, and show that EAAT can enrich the video quality while saving network resources. Xiangyang Gong, Jie Liang 0001, Wendong Wang 0003, Xirong Que |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Pixel-Level View Synthesis Distortion Estimation for 3D Video CodingabstractRecently, region-based 3D video coding has been proposed. However, existing view synthesis distortion estimation (VSDE) methods are performed at the frame level. To guide the rate-distortion optimization process of region-based 3D video coding schemes, this paper proposes the first pixel-level VSDE (PL-VSDE) method. We first give the definition of the pixel-level view synthesis distortion. To estimate it, a backward prediction method is then developed, which starts from the pixels of interest (POIs) in the virtual view and finds their corresponding pixels in the reference view via a coarse-to-fine approach, denoted as coarse-to-fine backward prediction (CFBP) method. Additionally, the CFBP fully considers the details of 3D warping, the rounding operation and the warping competition in view synthesis, leading to improve accuracy of the prediction. Besides, a table-lookup method and a warping property are introduced to speed up the CFBP. After integrating the CFBP into the PL-VSDE, we can estimate the view synthesis distortion at the pixel level. Our method is carried out pixel-by-pixel independently, which is friendly for parallel processing. The experimental results demonstrate that our proposed method has significant advantages in both accuracy and efficiency compared with the state-of-the-art frame-level VSDE methods. Jie Liang 0001, Yao Zhao 0001, Chunyu Lin, Lili Meng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Saliency Analysis and Gaussian Mixture Model-Based Detail Extraction Algorithm for Hyperspectral PansharpeningabstractThe purpose of hyperspectral (HS) pansharpening is to improve the spatial resolution of HS images using panchromatic (PAN) images so as to obtain the pansharpened images with both high spectral diversity and high spatial resolution. The classical component substitution (CS)-based pansharpening approaches can be decomposed into two sequential phases: detail extraction and detail injection. In general, the detail extraction is performed by computing the difference between the PAN image and a weighted average of the HS bands, whereas the detail injection depends on the injection gain which is defined locally or globally. In this article, we introduce a novel pansharpening algorithm in which the extracted details and the injection gain are estimated over salient and nonsalient regions achieved via saliency analysis and Gaussian mixture model. The proposed method is applied to four credible CS-based pansharpening methods and also compared to other state-of-the-art methods. Experimental results show that the modified CS methods have better performance than the original methods. In addition, our method also achieves comparable or better performance than the other state-of-the-art methods. Wenqian Dong, Jie Liang 0001, Song Xiao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Censor-Based Cooperative Multi-Antenna Spectrum Sensing with Imperfect Reporting ChannelsabstractThe present contribution proposes a spectrally efficient censor-based cooperative spectrum sensing (C-CSS) approach in a sustainable cognitive radio network that consists of multiple antenna nodes and experiences imperfect sensing and reporting channels. In this context, exact analytic expressions are first derived for the corresponding probability of detection, probability of false alarm, and secondary throughput, assuming that each secondary user (SU) sends its detection outcome to a fusion center only when it has detected a primary signal. Capitalizing on the findings of the analysis, the effects of critical measures, such as the detection threshold, the number of SUs, and the number of employed antennas, on the overall system performance are also quantified. In addition, the optimal detection threshold for each antenna based on the Neyman-Pearson criterion is derived and useful insights are developed on how to maximize the system throughput with a reduced number of SUs. It is shown that the C-CSS approach provides two distinct benefits compared with the conventional sensing approach, i.e., without censoring: i) the sensing tail problem, which exists in imperfect sensing environments, can be mitigated; and ii) less SUs are ultimately required to obtain higher secondary throughput, rendering the system more sustainable. Omar Alhussein, Paschalis C. Sofotasios, Sami Muhaidat, Paul D. Yoo, Jie Liang 0001, Anhong Wang |
IEEE Trans. Sustain. Comput. | 6 |
| 2019 | Error Analysis of NOMA-Based User Cooperation with SWIPTabstractThe present contribution analyzes the performance of non-orthogonal multiple access (NOMA)-based user cooperation with simultaneous wireless information and power transfer (SWIPT). In particular, we consider a two-user NOMA-based cooperative SWIPT scenario, in which the near user acts as a SWIPT-enabled relay that assists the farthest user. In this context, we derive analytic expressions for the pairwise error probability (PEP) of both users assuming the both amplify-and-forward (AF) and decode-and-forward (DF) relay protocols. The derived expressions are expressed in closed-form and have a tractable algebraic representation which renders them convenient to handle both analytically and numerically. In addition to this, we derive a simple asymptotic closed-form expression for the PEP in the high signal-to-noise ratio (SNR) regime which provide useful insights on the impact of the involved parameters on the overall system performance. Capitalizing on this, we subsequently quantify the maximum achievable diversity order of both users. It is shown that numerical and simulation results corroborate the derived analytic expressions. Furthermore, the offered results provide interesting insights into the error rate performance of each user, which are expected to be useful in future designs and deployments of NOMA based SWIPT systems. Suyue Li, Lina Bariah, Sami Muhaidat, Paschalis C. Sofotasios, Jie Liang 0001, Anhong Wang |
DCOSS | 5 |
| 2019 | DSSLIC: Deep Semantic Segmentation-based Layered Image CompressionabstractDeep learning has revolutionized many computer vision fields in the last few years, including learning-based image compression. In this paper, we propose a deep semantic segmentation-based layered image compression (DSSLIC) framework in which the segmentation map of the input image is obtained and encoded as the base layer of the bit-stream. A compact representation of the input image is also generated and encoded as the first enhancement layer. The segmentation map and the compact version of the image are then employed to obtain a coarse reconstruction of the image. The residual between the input and the coarse reconstruction is additionally encoded as another enhancement layer. Experimental results show that the proposed framework outperforms the H.265/HEVC-based BPG and other codecs in both PSNR and MS-SSIM metrics in RGB domain. Besides, since semantic map is included in the bit-stream, the proposed scheme can facilitate many other tasks such as image search and object-based adaptive image compression1. Jie Liang 0001, Jingning Han |
ICASSP | 2 |
| 2019 | Compression Artifact Removal with Stacked Multi-Context Channel-Wise Attention NetworkabstractImage compression plays an important role in saving disk storage and transmission bandwidth. Among traditional compression standards, JPEG is one of the commonly used standards in lossy image compression. However, the decompressed JPEG images usually have inevitable artifacts due to the quantization step, especially at low bitrate. Many recent works leverage deep learning networks to remove the JPEG artifacts and have achieved notable progress. In this paper, we propose a stacked multi-context channel-wise attention model. The channel-wise attention adaptively integrates features along the channel dimension given a set of feature maps. We apply multiple context-based channel attentions to enable the network to capture features from different resolutions. The entire architecture is trained progressively from the image space of low quality factor to that of high quality factor. Experiments show that we can achieve the state-of-the-art performance with lower complexity. Jie Liang 0001, Yang Wang 0003 |
ICIP | 2 |
| 2019 | Censor-Based Multi-Antenna Cooperative Spectrum Sensing over Erroneous Feedback ChannelsabstractWe propose a spectrally efficient censor-based cooperative spectrum sensing (C-CSS) approach for a sustainable cognitive radio network that consists of multiple antenna nodes and experiences imperfect sensing and reporting channels. First, analytic expressions are derived for the corresponding probabilities of detection and false alarm, assuming that each secondary user sends its detection outcome to a fusion center only when it believes to have detected a primary user's signal. Second, we derive lower bounds for the probability of false alarm, where we show that a sensing tail problem, which exist in the conventional (non-censor-based) scheme, can be effectively mitigated with the aid of the proposed C-CSS scheme. Simulation results are presented to corroborate the derived analytic results, and to provide theoretical and technical insights that are useful for the design of cognitive radio networks. Omar Alhussein, Paschalis C. Sofotasios, Sami Muhaidat, Paul D. Yoo, Jie Liang 0001, Anhong Wang |
WCNC | 6 |
| 2019 | Simultaneous color-depth super-resolution with conditional generative adversarial networks
Lijun Zhao 0002, Huihui Bai 0001, Jie Liang 0001, Bing Zeng 0001, Anhong Wang, Yao Zhao 0001 |
Pattern Recognit. | 3 |
| 2019 | Local activity-driven structural-preserving filtering for noise removal and image smoothing
Lijun Zhao 0002, Huihui Bai 0001, Jie Liang 0001, Anhong Wang, Bing Zeng 0001, Yao Zhao 0001 |
Signal Process. | 3 |
| 2019 | DCSN-Cast: Deep compressed sensing network for wireless video multicast
Hehe Wu, Anhong Wang, Jie Liang 0001, Suyue Li |
Signal Process. Image Commun. | 3 |
| 2019 | A Depth-Bin-Based Graphical Model for Fast View Synthesis Distortion EstimationabstractDuring 3-D video communication, transmission errors, such as packet loss, could happen to the texture and depth sequences. View synthesis distortion will be generated when these sequences are used to synthesize virtual views according to the depth-image-based rendering method. A depth-value-based graphical model (DVGM) has been employed to achieve the accurate packet-loss-caused view synthesis distortion estimation (VSDE). However, the DVGM models the complicated view synthesis processes at depth-value level, which costs too much computation and is difficult to be applied in practice. In this paper, a depth-bin-based graphical model (DBGM) is developed, in which the complicated view synthesis processes are modeled at depth-bin level so that it can be used for the fast VSDE with 1-D parallel camera configuration. To this end, several depth values are fused into one depth bin, and a depth-bin-oriented rule is developed to handle the warping competition process. Then, the properties of the depth bin are analyzed and utilized to form the DBGM. Finally, a conversion algorithm is developed to convert the per-pixel input depth value probability distribution into the depth-bin format. Experimental results verify that our proposed method is 8-$32\times $ faster and requires 17%-60% less memory than the DVGM, with exactly the same accuracy. Jie Liang 0001, Yao Zhao 0001, Chunyu Lin, Anhong Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Color-Sensitivity-Based Combined PSNR for Objective Video Quality AssessmentabstractThe peak signal-to-noise ratio (PSNR) has been widely employed as an objective video quality assessment (VQA) metric. Usually, videos are represented in the YCbCr color space, which results in three PSNR values for each video frame. Several VQA metrics have been proposed to measure the video quality with a single combined PSNR. However, these metrics are derived heuristically without theoretical justification. In this paper, based on our extensive subjective tests on the sensitivity of the human visual system to different color components, we derive the optimal weighting coefficients of a color-sensitivity-based combined PSNR (CSPSNR). Moreover, to verify the performance of the combined PSNR, test sequences with different levels of combined PSNRs are used to evaluate the quality of the videos. However, no such database is currently available for measuring the effectiveness of different methods regarding combined PSNRs. In this paper, we design a novel coding scheme to produce sequences whose PSNRs are the combinations of different levels of PSNRs of YCbCr, with which the correlation between the subjective score and the combined PSNR is analyzed. Experiment results and statistical analysis demonstrate that the proposed CSPSNR correlates better with the mean opinion score than the existing methods. Xiwu Shang, Jie Liang 0001, Haiwu Zhao, Chengjia Wu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Distortion Estimation-Based Adaptive Power Allocation for Hybrid Digital-Analog Video TransmissionabstractHybrid digital–analog (HDA) video transmission schemes have shown advantages in avoiding thecliff effect. However, most current HDA schemes assume perfect transmission of the digital signal, which is hardly the case in practice. In this paper, we propose an adaptive recursive distortion estimate for the HDA system (ARDE-HDA), which recursively estimates the decoder-side distortion from the encoder and adaptively allocates the transmission power between the digital and analog signals in HDA. First, we derive the closed-form expression of a recursive distortion estimation (RDE) method, which does not require the digital part of the HDA output to be decoded perfectly, and both transmission error and superposition process of digital and analog parts are taken into consideration. Then, based on the deduced RDE model, an adaptive power allocation is proposed for the digital and analog parts to minimize the decoder-side distortion. Finally, simulation results are presented, which show the accuracy of the proposed RDE model and the ARDE-HDA method. Our method can achieve an average of 4.10 dB gain over existing HDA methods and 15.14-dB gain over the Softcast method in terms of the peak signal-to-noise ratio. Anhong Wang, Jie Liang 0001, Suyue Li, Xiong Zhang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Semi-Recurrent Cnn-Based Vae-Gan for Sequential Data GenerationabstractA semi-recurrent hybrid VAE-GAN model for generating sequential data is introduced. In order to consider the spatial correlation of the data in each frame of the generated sequence, CNNs are utilized in the encoder, generator, and discriminator. The subsequent frames are sampled from the latent distributions obtained by encoding the previous frames. As a result, the dependencies between the frames are maintained. Two testing frameworks for synthesizing a sequence with any number of frames are also proposed. The promising experimental results on piano music generation indicates the potential of the proposed framework in modelling other sequential data such as video. Jie Liang 0001 |
ICASSP | 2 |
| 2018 | Recursive Distortion Estimation for Hybrid Digital-Analog Video TransmissionabstractRecently, hybrid digital-analog (HDA) video transmission scheme has shown advantages in avoiding the cliff effect. However, most HDA schemes assume perfect transmission of digital signals, which is hardly the case in practice. This paper proposes a scheme named RDE-HDA that recursively estimates the decoder-side distortion from the encoder-side for HDA system, and both transmission error and superposition process of digital and analog parts are taken into consideration. Therefore, our method does not require the digital part of the HDA output to be transmitted losslessly, making the HDA more practical. We derive the closed-form expression of the recursive distortion estimation for HDA system. The accuracy of our method is verified by simulation results. Anhong Wang, Jie Liang 0001 |
ICASSP | 3 |
| 2018 | Privacy-Preserving Age Estimation for Content RatingabstractContent rating (aka. maturity rating) rates the suitability of kinds of media (e.g., movies and video games) to its audience. It is essential to prevent a specific age group of people such as children from inappropriate information. However, in practice the administration of content rating system is usually suggestion-based declaration by media sources or key-based password which can easily fail if someone ignores the suggestions or somehow knows the keys. In this paper, we propose to estimate user's age in a privacy-preserving manner for automatic content rating. Several privacy-preserving approaches on facial images with different degree of privacy are proposed and evaluated on a deep neural network architecture for age estimation accuracy. We also introduce an attention mechanism which can adaptively learn discriminative features from the processed facial images. Experiments show that the proposed attention-based model performs better than the baseline model and achieves a reasonable performance to that with raw images in testing. Linwei Ye, Noman Mohammed, Yang Wang 0003, Jie Liang 0001 |
MMSP | 5 |
| 2018 | A real-time system for online learning-based visual transcription of piano music
Jie Liang 0001, Howard Cheng |
Multim. Tools Appl. | 2 |
| 2017 | Convolutional neural network-based depth image artifact removalabstractIn 3D video coding and depth-based image rendering, the distortion of the compressed depth image often leads to wrong 3D warpping. In this paper, by generalizing the recent work of convolutional neural network (CNN)-based depth image up-sampling, we propose a CNN-based depth image artifact removal scheme, where both the compressed depth and color images are used to enhance the depth accuracy. The proposed CNN has two sub-networks: joint depth-color sub-network and joint depth sub-network. During the depth and color feature extraction, the gradient of the depth image is used as the input to color image, while the gradient of color image is used as the input of depth feature extraction. Such an exchange of gradient information improves the learned features. Experimental results in terms of both objective and subjective quality of the depth and color images verify the efficiency of the proposed method. Lijun Zhao 0002, Jie Liang 0001, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001 |
ICIP | 2 |
| 2017 | A new combined PSNR for objective video quality assessmentabstractIn video coding, quality evaluation is important for improving the coding efficiency. Usually Peak Signal-to-Noise Ratio (PSNR) is utilized to measure the performance of different coding techniques. During the video coding process in YCbCr color space, there are three PSNRs, one for each color component. Sometimes they may contradict to each other, which poses a problem for evaluating the coding performance. Several video quality assessment (VQA) metrics have been proposed to measure the video quality with a combined PSNR. However, these combined PSNRs are obtained heuristically without theoretical justification. In this paper, we propose a color-sensitivity-based combined PSNR (CSP-SNR) based on extensive subjective tests on the sensitivity of human visual system (HVS) to different color components. Subjective experiment results demonstrate that the proposed combined PSNR correlates well with the mean opinion score (MOS) than existing methods. Xiwu Shang, Haiwu Zhao, Jie Liang 0001, Chengjia Wu |
ICME | 4 |
| 2017 | Single depth image super-resolution with multiple residual dictionary learning and refinementabstractLearning-based image super-resolution methods often use large datasets to learn texture features. When these methods are applied to depth images, emphasis should be given on learning the geometrical structures at object boundaries, since depth images do not have much texture information. In this paper, we develop a scheme to learn multiple residual dictionaries from only one external image. After depth image super-resolution, some artifacts may appear. An adaptive depth map refinement method is then proposed to remove these artifacts along the depth edges, based on the shape-adaptive weighted median filtering method. Experimental results demonstrate the advantage of the proposed method over many other methods. Lijun Zhao 0002, Huihui Bai 0001, Jie Liang 0001, Anhong Wang, Yao Zhao 0001 |
ICME | 3 |
| 2017 | Depth map up-sampling with fractal dimension and texture-depth boundary consistencies
Meiqin Liu 0002, Yao Zhao 0001, Jie Liang 0001, Chunyu Lin, Huihui Bai 0001 |
Neurocomputing | 3 |
| 2017 | Adaptive residual-based distributed compressed sensing for soft video multicasting over wireless networks
Anhong Wang, Suyue Li, Jie Liang 0001 |
Multim. Tools Appl. | 6 |
| 2017 | Power and sub-channel optimization of JPEG 2000 image transmission over OFDM-based cognitive radio networks
Golara Javadi, Atousa Hajshirmohammadi, Jie Liang 0001 |
Signal Process. Image Commun. | 3 |
| 2017 | QoE-driven optimization for cloud-assisted DASH-based scalable interactive multiview video streaming over wireless network
Mincheng Zhao, Xiangyang Gong, Jie Liang 0001, Wendong Wang 0003, Xirong Que, Yihua Guo, Shiduan Cheng |
Signal Process. Image Commun. | 3 |
| 2017 | Joint Source-Channel Coding of JPEG 2000 Image Transmission Over Two-Way Multi-Relay NetworksabstractIn this paper, we develop a two-way multi-relay scheme for JPEG 2000 image transmission. We adopt a modified time-division broadcast cooperative protocol, and derive its power allocation and relay selection under a fairness constraint. The symbol error probability of the optimal system configuration is then derived. After that, a joint source-channel coding (JSCC) problem is formulated to find the optimal number of JPEG 2000 quality layers for the image and the number of channel coding packets for each JPEG 2000 codeblock that can minimize the reconstructed image distortion for the two users, subject to a rate constraint. Two fast algorithms based on dynamic programming and branch and bound are then developed. Simulation demonstrates that the proposed JSCC scheme achieves better performance and lower complexity than other similar transmission systems. Chongyuan Bi, Jie Liang 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Graph-Based Transform for 2D Piecewise Smooth Signals With Random Discontinuity LocationsabstractThe graph-based block transform recently emerged as an effective tool for compressing some special signals such as depth images in 3D videos. However, in existing methods, overheads are required to describe the graph of the block, from which the decoder has to calculate the transform via time-consuming eigendecomposition. To address these problems, in this paper, we aim to develop a single graph-based transform for a class of 2D piecewise smooth signals with similar discontinuity patterns. We first consider the deterministic case with a known discontinuity location in each row. We propose a 2D first-order autoregression (2D AR1) model and a 2D graph for this type of signals. We show that the closed-form expression of the inverse of a biased Laplacian matrix of the proposed 2D graph is exactly the covariance matrix of the proposed 2D AR1 model. Therefore, the optimal transform for the signal are the eigenvectors of the proposed graph Laplacian. Next, we show that similar results hold in the random case, where the locations of the discontinuities in different rows are randomly distributed within a confined region, and we derive the closed-form expression of the corresponding optimal 2D graph Laplacian. The theory developed in this paper can be used to design both pre-computed transforms and signal-dependent transforms with low complexities. Finally, depth image coding experiments demonstrate that our methods can achieve similar performance to the state-of-the-art method, but our complexity is much lower. Jie Liang 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | Graph-Based Transform for 2D Piecewise Smooth Signals with Random DiscontinuitiesabstractThe graph-based transform has recently emerged as an effective tool for compressing some special signals such as depth images in 3D videos. However, one limitation of this approach is that it needs to apply eigen-decomposition to each block. To reduce the complexity, in this paper, we develop a systematic approach to find a universal optimal graph-based transform for a class of 2D piecewise smooth signals. Each block in the class can include a discontinuity whose locations in different rows are randomly distributed within a confined region. We first define a special 2D graph model for this class of signals. Our derivation then reveals that the inverse of the covariance matrix of this class of signals is equal to its graph Laplacian with a bias value added to the first diagonal element. Furthermore, the edge values within the confined region have a closed-form expression. If the bias value is assumed negligible then the approximation of the optimal transform for the class of signals is given by the eigenvectors of the true graph Laplacian and can be pre-computed. Therefore online eigen-decomposition for the class of signals can be avoided, and the complexity of the encoder and decoder can thus be reduced. The feasibility of the proposed scheme is demonstrated via depth image coding examples. Jie Liang 0001 |
DCC | 2 |
| 2016 | Scalable Compression of Deep Neural NetworksabstractDeep neural networks generally involve some layers with millions of parameters, making them difficult to be deployed and updated on devices with limited resources such as mobile phones and other smart embedded systems. In this paper, we propose a scalable representation of the network parameters, so that different applications can select the most suitable bit rate of the network based on their own storage constraints. Moreover, when a device needs to upgrade to a high-rate network, the existing low-rate network can be reused, and only some incremental data are needed to be downloaded. We first hierarchically quantize the weights of a pre-trained deep neural network to enforce weight sharing. Next, we adaptively select the bits assigned to each layer given the total bit budget. After that, we retrain the network to fine-tune the quantized centroids. Experimental results show that our method can achieve scalable compression with graceful degradation in the performance. Xing Wang 0003, Jie Liang 0001 |
ACM Multimedia | 2 |
| 2016 | Encoder-Driven Inpainting Strategy in Multiview Video CompressionabstractIn free viewpoint video systems, a user has the freedom to select a virtual view from which an image of the 3D scene is rendered, and the scene is commonly represented by color and depth images of multiple nearby viewpoints. In such representation, there exists data redundancy across multiple dimensions: 1) a 3D voxel may be represented by pixels in multiple viewpoint images (inter-view redundancy); 2) a pixel patch may recur in a distant spatial region of the same image due to self-similarity (inter-patch redundancy); and 3) pixels in a local spatial region tend to be similar (inter-pixel redundancy). It is important to exploit these redundancies during inter-view prediction toward effective multiview video compression. In this paper, we propose an encoder-driven inpainting strategy for inter-view predictive coding, where explicit instructions are transmitted minimally, and the decoder is left to independently recover remaining missing data via inpainting, resulting in lower coding overhead. In particular, after pixels in a reference view are projected to a target view via depth-image-based rendering at the decoder, the remaining holes in the target view are filled via an inpainting process in a block-by-block manner. First, blocks are ordered in terms of difficulty-to-inpaint by the decoder. Then, explicit instructions are only sent for the reconstruction of the most difficult blocks. In particular, the missing pixels are explicitly coded via a graph Fourier transform or a sparsification procedure using discrete cosine transform, leading to low coding cost. For blocks that are easy to inpaint, the decoder independently completes missing pixels via template-based inpainting. We apply our proposed scheme to frames in a prediction structure defined by JCT-3V where inter-view prediction is dominant, and experimentally we show that our scheme achieves up to 3-dB gain in peak-signal-to-noise-ratio in reconstructed image quality over a comparable 3D-High Efficiency Video Coding implementation using fixed 16 $\times $ 16 block size. Yu Gao 0003, Gene Cheung, Thomas Maugey, Pascal Frossard, Jie Liang 0001 |
IEEE Trans. Image Process. | 5 |
| 2015 | A Hybrid Transmission Approach for DASH over MBMS in LTE NetworkabstractDynamic adaptive streaming over HTTP (DASH) has been a research hotspot, and the current studies focus on the DASH transmission optimization using unicast mode. However, DASH streaming also could be transmitted by multicast mode, which could effectively reduce transmission resource consumption especially when multiple DASH clients request the same video program in parallel. In this paper, we propose the Hybrid Transmission strategies for DASH (HTD) in LTE network, which is considered both unicast and multicast modes for DASH. The optimization problem is formulated as a Mixed Binary Integer Programming (MBIP) problem, and a two-level greedy algorithm is proposed, which could improve the quality of experience (QoE) of wireless DASH users, and save the wireless resources in LTE network. Simulation results demonstrate that our scheme achieves better performance than traditional single transmission mode in the literature. Xiangyang Gong, Jie Liang 0001, Shiju Zhang, Mincheng Zhao, Wendong Wang 0003 |
GLOBECOM | 3 |
| 2015 | Multi-resolution compressed sensing reconstruction via approximate message passingabstractWe consider multi-resolution (MR) compressed sensing reconstruction, where instead of always reconstructing the signal at the original high resolution (HR), we enable the reconstruction of a better-quality low-resolution (LR) signal when the sampling rate is too low. We propose an approximate message passing (AMP)-based solution (MR-AMP). Theoretical analyses show that in addition to reduced complexity, our method can produce a LR signal with bounded mean squared error (MSE) even when the MSE of the conventional HR reconstruction is unbounded. The performance of the proposed scheme is verified using both synthetic data and natural images. Xing Wang 0003, Jie Liang 0001 |
ICIP | 2 |
| 2015 | Quantized dictionary for sparse representationabstractDictionary learning for sparse representation has drawn considerable attention in recent years. In particular, the K-SVD algorithm is an efficient approach, and various modifications of the K-SVD have been developed for applications such as face recognition. However, the efficient storage of the dictionary has not been studied. Currently, the dictionary is simply normalized and saved as floating-point numbers, which could be quite large and lead to excessive cost and delay if the dictionary needs to be transmitted, e.g., to mobile users. In this paper, we develop a quantized K-SVD (Q-KSVD) to reduce the storage of the dictionary. We compress each basis image in the dictionary by the conventional image coding method. Moreover, we integrate the image compression step into various modified K-SVD optimization schemes, and develop an algorithm to find the optimal dictionary when there is a constraint on the total bits of the compressed dictionary. Our algorithm selects dictionary bases by ranking the contribution-rate slopes of all bases. This method also serves as an efficient approach to find the optimal number of bases of the dictionary at each rate constraint. Face recognition experiments using four K-SVD-based methods show that our method can achieve different tradeoffs between the dictionary storage space and the recognition accuracy. It can achieve comparable performance with as little as 3% of the original storage space. It can even yield higher accuracy than the uncompressed dictionary in some cases. Jie Liang 0001, Yao Zhao 0001, Chunyu Lin, Huihui Bai 0001 |
MMSP | 2 |
| 2015 | A cloud-assisted DASH-based Scalable Interactive Multiview Video Streaming frameworkabstractInteractive multiview video streaming (IMVS) allows viewers to periodically switch viewpoint. Its user experience can be further enhanced by creating virtual views from neighboring coded views using view synthesis techniques. Dynamic adaptive streaming over HTTP (DASH) is a new standard that can adjust the quality of video streaming according to the network condition. In this paper, we propose an improved DASH-based IMVS scheme over wireless networks. The main contributions are twofold. First, our scheme allows virtual views to be generated at either the cloud-based server or the client, and can adaptively select the optimal approach based on the network condition and the cost of the cloud. Second, scalable video coding is used in our system. Simulations with the NS3 tool demonstrate the advantage of our proposed scheme over the existing approach with client-based view synthesis and single-layer video coding. Mincheng Zhao, Xiangyang Gong, Jie Liang 0001, Wendong Wang 0003, Xirong Que, Shiduan Cheng |
PCS | 3 |
| 2015 | A Generalized Mixture of Gaussians for Fading ChannelsabstractThe analysis of composite fading channels, which are typically encountered in wireless channels due to multipath and shadowing is quite involved, as the underlying fading distributions do not lend themselves to analysis. An example of such channels are the Nakagami/Rayleigh-Lognormal fading channels. Several simplified expressions have been proposed in the literature. In this paper, a generalized fading model for composite and non-composite fading models, based on the so-called Mixture of Gaussians (MoG) distribution, is proposed. The well-known expectation-maximization algorithm is utilized to estimate the parameters of the MoG model. Furthermore, relying on the proposed MoG model, we derive closed form expressions for several performance metrics used in wireless communication systems, including the raw moments, the amount of fading, the outage probability, the average channel capacity, and the moment generating function. In addition, the symbol error rate of L-branch maximum ratio combining diversity receiver is studied for linear coherent signaling schemes. Monte Carlo simulations are presented to corroborate the analytical results and to assess the accuracy of the MoG model. Omar Alhussein, Bassant Selim, Tasneem Assaf, Sami Muhaidat, Jie Liang 0001, George K. Karagiannidis |
VTC Spring | 5 |
| 2015 | Performance analysis of energy detection over mixture gamma based fading channels with diversity receptionabstractThe present paper is devoted to the evaluation of energy detection based spectrum sensing over different multipath fading and shadowing conditions. This is realized by means of a unified and versatile approach that is based on the particularly flexible mixture gamma distribution. To this end, novel analytic expressions are firstly derived for the probability of detection over MG fading channels for the conventional single-channel communication scenario. These expressions are subsequently employed in deriving closed-form expressions for the case of square-law combining and square-law selection diversity methods. The validity of the offered expressions is verified through comparisons with results from respective computer simulations. Furthermore, they are employed in analyzing the performance of energy detection over multipath fading, shadowing and composite fading conditions, which provides useful insighs on the performance and design of future cognitive radio based communication systems. Omar Alhussein, Ahmed Y. Al Hammadi, Paschalis C. Sofotasios, Sami Muhaidat, Jie Liang 0001, Mahmoud Al-Qutayri, George K. Karagiannidis |
WiMob | 5 |
| 2015 | Approximate message passing-based compressed sensing reconstruction with generalized elastic net prior
Xing Wang 0003, Jie Liang 0001 |
Signal Process. Image Commun. | 2 |
| 2015 | View Synthesis Distortion Estimation With a Graphical Model and Recursive Calculation of Probability DistributionabstractDepth-image-based rendering (DIBR) is frequently used in multiview video applications such as free-viewpoint television. In this paper, we consider the two DIBR algorithms used in the Moving Picture Experts Group view synthesis reference software, and develop a scheme for the encoder to estimate the distortion of the synthesized virtual view at the decoder when the reference texture and depth sequences experience transmission errors such as packet loss. We first develop a graphical model to analyze how random errors in the reference depth image affect the synthesized virtual view. The warping competition rule adopted in the DIBR algorithms is explicitly represented by the graphical model. We then consider the case where packet loss occurs to both the encoded texture and depth images during transmission and develop a recursive optimal distribution estimation (RODE) method to calculate the per-pixel texture and depth probability distributions in each frame of the reference views. The RODE is then integrated with the graphical model method to estimate the distortion in the synthesized view caused by packet loss. Experimental results verify the accuracy of the graphical model method, the RODE, and the combined estimation scheme. Jie Liang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | QoE-Driven Cross-Layer Optimization for Wireless Dynamic Adaptive Streaming of Scalable Videos Over HTTPabstractRecently, Dynamic Adaptive Streaming over HTTP (DASH) has attracted significant attention. In this paper, we consider DASH-based transmission of scalable videos in wireless broadband access networks (e.g., long-term evolution and WiMAX), and propose three methods to enhance the quality of experience of wireless DASH users. First, we design an improved mapping scheme from scalable video coding layers to DASH layers that can provide the desired bitrates, enhance the video end-to-end throughput, and reduce the HTTP communication overhead. Second, we develop a DASH-friendly scheduling and resource allocation algorithm by integrating the DASH-based media delivery and the radio-level adaptation via a cross-layer approach. It utilizes the characteristics of video content and scalable video coding, and greatly reduces the possibility of video playback interruption by considering the client buffer status. The optimization problem is formulated as a mixed binary integer programming problem, and is solved by a subgradient method. Finally, a DASH proxy-based bitrate stabilization algorithm is proposed to improve the video playback smoothness that can achieve the desired tradeoff between playback quality and stability. Simulations with the Qualnet tool demonstrate that our schemes achieve better performances than other methods in the literature. Mincheng Zhao, Xiangyang Gong, Jie Liang 0001, Wendong Wang 0003, Xirong Que, Shiduan Cheng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | 3D geometry representation using multiview coding of image tilesabstractCompression of dynamic 3D geometry obtained from depth sensors is challenging, because noise and temporal inconsistency inherent in acquisition of depth data means there is no one-to-one correspondence between sets of 3D points in consecutive time instants. In this paper, instead of coding 3D points (or meshes) directly, we propose to represent an object's 3D geometry as a collection of tile images. Specifically, we first place a set of image tiles around an object. Then, we project the object's 3D geometry onto the tiles that are interpreted as 2D depth images, which we subsequently encode using a modified multiview image codec tuned for piecewise smooth signals. The crux of the tile image framework is the “optimal” placement of image tiles - one that yields the best tradeoff in rate and distortion. We show that if only planar and cylindrical tiles are considered, then the optimal placement problem for K tiles can be mapped to a tractable piece-wise linear approximation problem. We propose an efficient dynamic programming algorithm to find an optimal solution to the piecewise linear approximation problem. Experimental results show that optimal tiling outperforms naïve tiling by up to 35% in rate reduction, and graph transform can further exploit the smoothness of the tile images for coding gain. Yu Gao 0003, Gene Cheung, Thomas Maugey, Pascal Frossard, Jie Liang 0001 |
ICASSP | 5 |
| 2014 | Side information-aided compressed sensing reconstruction via approximate message passingabstractIn this paper, the side information (SI)-aided compressed sensing reconstruction is considered, where a sparse signal is observed via a noisy underdetermined linear system, and a SI is available during the reconstruction. We develop a SI-aided approximate message passing (SI-AMP) algorithm to solve the problem. Based on the corresponding state evolution formula, the asymptotic prediction performance and noise-sensitivity analysis of the scheme are derived. Simulation results are presented to verify the efficiency of the proposed method. Xing Wang 0003, Jie Liang 0001 |
ICASSP | 2 |
| 2014 | Scheduling and resource allocation for wireless dynamic adaptive streaming of scalable videos over HTTPabstractRecently dynamic adaptive streaming over HTTP (DASH) has gained significant attentions. In this paper, we study DASH-based transmission of scalable videos in wireless broadband access networks (e.g. LTE, WiMAX), and propose a DASH-friendly scheduling and resource allocation scheme (DFSRA) to enhance the Quality-of-Experience (QoE) of wireless DASH users. By integrating DASH-based media delivery and radio-level adaptation in a cross-layer manner, the scheme can determine the optimal wireless resource allocation. A gradient-based algorithm is proposed to solve the optimization problem, which has the following advantages over traditional algorithms: (1) The characteristics of video content and scalable video coding are utilized. (2) The client buffer status is considered to reduce video playback interruption and buffer overflow. Simulations with the Qualnet tool demonstrate that our scheme achieves better performance than existing methods in the literature. Mincheng Zhao, Xiangyang Gong, Jie Liang 0001, Wendong Wang 0003, Xirong Que, Shiduan Cheng |
ICC | 3 |
| 2014 | Multiple Description Coding With Randomly and Uniformly Offset QuantizersabstractIn this paper, two multiple description coding schemes are developed, based on prediction-induced randomly offset quantizers and unequal-deadzone-induced near-uniformly offset quantizers, respectively. In both schemes, each description encodes one source subset with a small quantization stepsize, and other subsets are predictively coded with a large quantization stepsize. In the first method, due to predictive coding, the quantization bins that a coefficient belongs to in different descriptions are randomly overlapped. The optimal reconstruction is obtained by finding the intersection of all received bins. In the second method, joint dequantization is also used, but near-uniform offsets are created among different low-rate quantizers by quantizing the predictions and by employing unequal deadzones. By generalizing the recently developed random quantization theory, the closed-form expression of the expected distortion is obtained for the first method, and a lower bound is obtained for the second method. The schemes are then applied to lapped transform-based multiple description image coding. The closed-form expressions enable the optimization of the lapped transform. An iterative algorithm is also developed to facilitate the optimization. Theoretical analyzes and image coding results show that both schemes achieve better performance than other methods in this category. Lili Meng, Jie Liang 0001, Upul Samarawickrama, Yao Zhao 0001, Huihui Bai 0001, André Kaup |
IEEE Trans. Image Process. | 2 |
| 2013 | M-channel multiple description coding based on uniformly offset quantizers with optimal deadzoneabstractThis paper proposes an improved source-splitting-based two-rate M-channel multiple description coding scheme, where the source is split into M subsets. In each description, one subset is coded at a high rate, and others are predictively coded at a low rate. Uniform offsets among low-rate quantizers of different descriptions are achieved by employing unequal deadzones and by quantizing the predictions. When several descriptions are received, the optimal reconstruction of each subset is achieved by finding the intersection of all received quantization bins. The closed-form expression of the expected distortion is obtained. The proposed scheme is applied to lapped transform-based multiple description image coding and achieves improved performance. The optimal deadzone selection and its impact are also given in this paper. Lili Meng, Jie Liang 0001, Yao Zhao 0001, Huihui Bai 0001, Chunyu Lin, André Kaup |
ICASSP | 2 |
| 2013 | View interpolation confidence-aided compressed sensing of multiview imagesabstractIn this paper, a hybrid multiview imaging system is considered, where traditional cameras and compressed sensing (CS) cameras are interleavingly placed. To improve the reconstruction quality of the CS cameras, the interpolated image from the two neighboring traditional cameras is used as side information. Different from existing CS-based multiview imaging systems, we incorporate in the CS reconstruction both the frame-level and pixel-level confidences of the interpolated view, based on the knowledge of occluded pixels and holes in it. Simulation results demonstrate the flexibility and superior performance of the proposed framework. Xing Wang 0003, Jie Liang 0001 |
ICASSP | 2 |
| 2013 | Rate-complexity tradeoff for client-side free viewpoint image renderingabstractFree viewpoint video enables a client to interactively choose a viewpoint from which to synthesize an image via depth-image-based rendering (DIBR). However, synthesizing a novel viewpoint image using texture and depth maps from two nearby views entails a sizable computation overhead. Further, to reduce transmission rate, recent proposals synthesize the second reference view itself using texture and depth maps of the first reference view via a complex inpainting algorithm to complete large disocclusion holes in the second reference image-a small amount of auxiliary information (AI) is transmitted by sender to aid the inpainting process-resulting in an even higher computation cost. In this paper, we study the optimal tradeoff between transmission rate and client-side complexity, so that in the event that a client device is computation-constrained, complexity of DIBR-based view synthesis can be scalably reduced at the expense of a controlled increase in transmission rate. Specifically, for standard view synthesis paradigm that requires texture and depth maps of two neighboring reference views, we design a dynamic programming algorithm to select the optimal subset of intermediate virtual views for rendering and encoding at server, so that a client performs only video decoding of these views, reducing overall view synthesis complexity. For new view synthesis paradigm that synthesizes the second reference view itself from the first, we optimize the transmission of AI used to assist inpainting of large disocclusion holes, so that some computation-expensive exemplar block search operations are avoided, reducing inpainting complexity. Experimental results show that the proposed schemes can scalably and gracefully reduce client-side complexity, and the proposed optimizations achieve better rate-complexity tradeoff than competing schemes. Yu Gao 0003, Gene Cheung, Jie Liang 0001 |
ICIP | 3 |
| 2013 | Multiple description coding with randomly offset quantizersabstractA multiple description coding scheme based on prediction-induced randomly offset quantizers is proposed, where each description encodes one source subset with a small quantization stepsize, and other subsets are predictively coded with a large quantization stepsize. Due to the prediction, the quantization bins that a coefficient belongs to in different descriptions are randomly overlapped with each others. The optimal reconstruction is obtained by finding the intersection of all received quantization bins. Using the recently developed random quantization theory, the closed-form expression of the expected distortion is obtained. The proposed scheme is then applied to lapped transform-based multiple-description image coding, and an iterative optimization scheme is developed to find the optimal lapped transform. Experimental results show that the proposed scheme achieves better performance than other methods in this category. Lili Meng, Jie Liang 0001, Upul Samarawickrama, Yao Zhao 0001, Huihui Bai 0001, André Kaup |
ISCAS | 2 |
| 2013 | A cooperation incentive scheme based on coalitional game theory for sparse and dense VANETs
Di Wu 0007, Yanrong Gao, Guozhen Tan, Limin Sun 0001, Jie Liang 0001, Jiangchuan Liu |
IWCMC | 5 |
| 2013 | Distortion estimation for two-step view synthesisabstractIn depth-image-based-rendering (DIBR), the quality of the synthesized virtual view depends on that of the depth maps in the reference views. In this paper, we develop a framework to estimate the distortion of the synthesized view when a simplified two-step warping algorithm is used and when there are random errors in the reference depth maps. A graph-based method is developed to obtain the depth and texture distributions in the synthesized view. Experimental results demonstrate the accuracy of the estimated distortion. Jie Liang 0001, Xiangyang Gong |
PCS | 2 |
| 2013 | Layered Source-Channel Coding over Two-Way Relay NetworksabstractIn this paper, we study the performances of various layered source-channel coding schemes over two-way relay networks, where the sources are coded into different layers and transmitted by progressive coding or superposition coding. The two-way relay network employs multiple-access-broadcast or time- division-broadcast protocol, and the relay uses decode-and-forward (DF) protocol. We first derive the optimal distortion exponents of different schemes at high signal-to-noise ratios (SNRs). Fast resource allocation algorithms are then developed to optimize the expected distortion of the source at finite SNRs. Simulation results show that layered coding achieves better performance than single rate coding, and the proposed fast resource allocation scheme is near-optimal. Chongyuan Bi, Jie Liang 0001 |
VTC Fall | 2 |
| 2013 | Relay selection in cognitive radio networks with interference constraintsabstractIn this study, the authors investigate the outage probability of underlay cognitive radio systems with relay selection. In particular, they consider a secondary multi‐relay network operating in the amplify‐and‐forward (AF) mode and only the ‘best’ relay is selected, which satisfies an index of merit. The proposed selection strategy takes into consideration the effect of primary user (PU) interference. That is, the authors assume that the secondary multi‐relay network is exposed to unwanted interference from a neighboring PU network. They derive a closed‐form outage probability expression and further present a thorough asymptotic diversity order analysis of the underlying scenario. A simulation study is presented to corroborate the analytical results and to have further insight into the performance of the proposed selection strategy. Mehdi Seyfi, Sami Muhaidat, Jie Liang 0001 |
IET Commun. | 3 |
| 2013 | Fast synthesized and predicted just noticeable distortion maps for perceptual multiview video coding
Yu Gao 0003, Xiaoyu Xiu, Jie Liang 0001, Weisi Lin |
J. Vis. Commun. Image Represent. | 3 |
| 2013 | Fast transmission distortion estimation and adaptive error protection for H.264/AVC-based embedded video conferencing systems
Jie Liang 0001, Inder Singh |
Signal Process. Image Commun. | 2 |
| 2012 | Noncausal directional intra prediction: Theoretical analysis and simulationabstractThis paper presents the theoretical analysis and simulation of noncausal directional intra prediction for image and video coding, where noncausal pixels, that is, pixels inside, below, or to the right of the target block, are used to predict the block, at the cost of extra bits to code those noncausal reference pixels. The proposed method generalizes the conventional causal intra prediction. The optimal number and locations of noncausal reference pixels are determined by minimizing the total differential entropies of the prediction residuals and the noncausal pixels. In order to obtain the differential entropy, a statistical image model is used to derive the autocorrelation of the residuals. In addition, the optimal sinusoidal approximations to transform the residuals are obtained by maximizing the coding gain. In the simulation, the optimal noncausal reference pixels and transforms of 4 × 4 and 8 × 8 blocks are identified for up to 33 directions. The results could provide insights for the design of more advanced directional intra prediction for the High Efficiency Video Coding (HEVC). Yu Gao 0003, Jie Liang 0001 |
ICIP | 2 |
| 2012 | Just noticeable distortion map prediction for perceptual multiview video codingabstractThe just noticeable distortion (JND) map is a useful tool for perceptual video coding. However, direct calculation of the JND map incurs high complexity, and the problem is aggravated in multiview video coding. In this paper, the motion and disparity vectors obtained during the video coding are employed to predict the JND maps in order to reduce the complexity. The error propagation of the prediction is studied and a JND block refreshing approach is proposed, when the prediction is not satisfactory, to alleviate the influence of the error propagation. The performance of the proposed JND prediction method is evaluated in a perceptual MVC framework, where the prediction residuals are tuned according to the JND thresholds to save the bits without affecting the perceptual quality. Experimental results show that the JND prediction method has better accuracy and lower complexity than an existing JND synthesis method. In addition, the proposed method leads to negligible degradation of the coding performance, compared to the direct JND method. Yu Gao 0003, Xiaoyu Xiu, Jie Liang 0001, Weisi Lin |
ICIP | 3 |
| 2012 | End-to-end distortion estimation for H.264 with unconstrained intra predictionabstractThis paper presents a macroblock-level algorithm to estimate the end-to-end distortion of H.264/AVC-based video transmission that allows unrestricted intra prediction. Some fast approximations are developed to reduce its complexity so that it can be used in real-time applications. Experimental results demonstrate the performance of the algorithm and the impact of the unconstrained intra prediction. Jie Liang 0001, Inder Singh |
ICIP | 2 |
| 2012 | Ztitch: A mobile phone application for immersive panorama creation, navigation, and social sharingabstractThis paper presents the design of Ztitch, a mobile phone application that can create immersive panoramic scenes in real time, where multiple photos are arranged in the 3D space according to their camera poses, maintaining a realistic perspective of the scene. By taking full advantage of the touchscreen, accelerometer, and gyroscope in the phone, Ztitch allows users to easily fine-tune each photo's position, navigate the scene, and edit the scenes created by others. When there is a undesired gap or overlap between the two ends of a 360° immersive panorama due to unknown camera focal length or accumulated errors, a fast and accurate algorithm is developed to jointly adjust the entire sequence to achieve the desired panorama. This is again achieved by leveraging the touchscreen of the phone. In addition, a fast color-balancing and exposure-compensation technique is developed to blend neighboring images. Andrew Au, Jie Liang 0001 |
MMSP | 2 |
| 2012 | Relay selection in underlay cognitive radio networksabstractIn this paper, we investigate the performance of relay selection in an underlay cognitive radio system in the presence of primary user (PU) interference. In particular, we consider a secondary multi-relay network operating in the amplify-and-forward (AF) mode and only the “best” relay which satisfies an index of merit is selected. The proposed selection strategy takes into consideration the effect of PU interference, i.e., we assume that a secondary relay network is exposed to unwanted interference from a neighboring PU network. We derive a closed-form outage probability expression for the secondary multi-relay network and further present a thorough asymptotical diversity order analysis. A simulation study is presented to corroborate the analytical results and to have further insight into the performance the proposed selection strategy. Mehdi Seyfi, Sami Muhaidat, Jie Liang 0001 |
WCNC | 3 |
| 2012 | Max-min relay selection in bidirectional cooperative networks with imperfect channel estimationabstractThe authors study the performance of wireless bidirectional relay-assisted networks in the presence of imperfect channel state information, where two end-source terminals S1 and S2 communicate with the assistance of M relay terminals Rj's. The max–min relay selection criterion is used to select the best relay that maximises the minimum signal-to-noise ratio of the links S1 → Rj → S2 and S2 → Rj → S1 over all relay terminals. The authors investigate the impact of imperfect channel estimation on the outage probability Pout of the system by means of the correlation coefficient pSi of the estimated channel gains and their actual values. Furthermore, the authors show that in a bidirectional relay-assisted network neither of the links S1 → Rj → S2 and S2 → Rj → S1 dominates the performance of the system. Instead, the performance is determined by the average performance of the two links, based on that the authors then discuss the power allocation in such networks. The authors demonstrate that in order to minimise Pout of the entire system, increasing the transmission power of the link with better estimation cannot compensate for the effect of the worse link and therefore the optimum power allocation with the least complexity is to transmit at each source terminals S1 and S2 with equal powers. Numerical results are also presented to corroborate the analytical expressions. M. Jafar Taghiyar, Sami Muhaidat, Jie Liang 0001 |
IET Commun. | 3 |
| 2012 | Delay-Cognizant Interactive Streaming of Multiview Video With Free Viewpoint SynthesisabstractIn interactive multiview video streaming (IMVS), a client receives and observes one of many available viewpoints of the same scene and periodically requests from the server view switches to neighboring views, as the video is played back in time uninterruptedly. One key technical challenge is to design a frame coding structure that facilitates periodic view switching and achieves an optimal tradeoff between storage cost and expected transmission rate. In this paper, we first propose three significant improvements over existing IMVS systems and then study the corresponding frame structure optimization. First, using depth-image-based rendering, the new IMVS system enables free viewpoint switching, i.e., by encoding and transmitting both texture and depth maps of captured views, a client can select and synthesize any virtual view from an almost continuum of viewpoints between the left-most and right-most captured views. Second, the IMVS system adopts a more realistic Markovian view-switching model with memory that more accurately captures user behaviors than previous memoryless models . A view-switching model is used in predicting client's future view-switching patterns. Third, assuming that the round-trip-time (RTT) delay during server-client communication is nonnegligible, during an IMVS session, the IMVS system additionally transmits redundant frames RTT into future playback, so that zero-delay view switching can be achieved. Given these improvements, we formalize a new joint optimization of the frame coding structure, transmission schedule, and quantization parameters of the texture and depth maps of multiple camera views. We propose an iterative algorithm to achieve fast and near-optimal solutions. The convergence of the algorithm is also demonstrated. Experimental results show that the proposed optimized rate-allocation method requires 38% lower transmission rate than the fixed rate-allocation scheme. In addition, with the same storage, the transmission rate of the optimized frame structure can be up to 55% lower than that of an I-frame-only structure and 27% lower than that of the structure without distributed source coding frames. Xiaoyu Xiu, Gene Cheung, Jie Liang 0001 |
IEEE Trans. Multim. | 3 |
| 2012 | Amplify-and-Forward Selection Cooperation over Rayleigh Fading Channels with Imperfect CSIabstractIn this paper, we investigate the performance of selection cooperation in the presence of imperfect channel estimation. In particular, we consider a cooperative scenario with multiple relays and amplify-and-forward protocol over frequency flat fading channels. In the selection scheme, only the "best" relay which maximizes the effective signal-to-noise ratio (SNR) at the receiver end is selected. We present lower and upper bounds on the effective SNR and further we provide closed-form expressions for the bounds on average symbol error rate (ASER), outage probability and average capacity per bandwidth of the received signal in the presence of channel estimation errors. A simulation study is presented to corroborate the analytical results and to demonstrate the performance of relay selection with imperfect channel estimation. Mehdi Seyfi, Sami Muhaidat, Jie Liang 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2011 | System Distortion Exponents of Two-Way Relay NetworksabstractDistortion exponent (DE) is the asymptotic exponential decay rate of the expected distortion between a source and its reconstruction in the high-SNR regime. This paper presents two improvements to the existing DE analyses of two-way relaying networks, where two users communicate in both directions with the help of a relay. First, we allow the users to have different bandwidth ratios, and derive the corresponding achievable DE regions of the system with different source coding, transmission, and relaying strategies. Secondly, we propose a new performance measure called the system distortion exponent (SDE), which characterizes the weighted average of the two distortions. The SDEs of the aforementioned scenarios are also derived. Yu Gao 0003, Jing Wang 0029, Jie Liang 0001 |
GLOBECOM | 3 |
| 2011 | JPEG XR optimization with graph-based soft decision quantizationabstractJPEG XR is the latest image compression standard. In this paper, two graph-based soft decision quantization (SDQ) methods are developed to optimize the rate-distortion performance of JPEG XR. The first approach uses a full graph, whose number of states is determined by the block size. The second method employs a fast and adaptive event-based graph, where the number of states depends on the number of nonzero indices in a normalized block, which is usually much less than the block size. Experimental results show that the fast method performs as good as the full graph method, and both methods can achieve up to 0.5 dB gain over JPEG XR. Yu Gao 0003, Duncan Chan, Jie Liang 0001 |
ICIP | 3 |
| 2011 | Frame structure optimization for interactive multiview video streaming with bounded network delayabstractInteractive multiview video streaming (IMVS) is an application that streams to a client one out of N available video views for observation, but client can periodically request switches to neighboring views as the video is played back uninterrupted in time. Previous IMVS works focused on the design of a frame structure at encoding time, trading off expected transmission rate with storage, without knowing the exact view trajectory a client may select at stream time. None of the existing IMVS schemes, however, explicitly addressed the network delay problem, and so a client will suffer a round trip time (RTT) delay for each requested view-switch. In this pa- per, we optimize frame structure for a bounded RTT, so that a client can switch to neighboring views as the video is played back without view-switching delay. The key idea is to send additional views likely to be requested by a client within one RTT beyond the current requested view. Each required set of contiguous views (corresponding to a given current requested single view) are pre-encoded using frames of previously transmitted set of views as predictors to lower transmission rate. Using I-, P- and distributed source coding (DSC) frames, we first formulate the structure design problem as a Lagrangian minimization for a desired bandwidth/storage tradeoff. We then develop a low-complexity greedy algorithm to automatically generate a good structure. Experimental results show that for the same storage cost, the transmission rate of the proposed structure can be 42% lower than that of I-frame-only structure, and 8% lower than that of the structure without DSC frames. Xiaoyu Xiu, Gene Cheung, Jie Liang 0001 |
ICIP | 3 |
| 2011 | Optimizing frame structure for interactive multiview video streaming with viewsynthesisabstractTraditional multiview video coding schemes compress all captured video frames exploiting all possible inter-view and temporal frame correlation for coding gain, creating complex inter-frame dependencies in the process. In contrast, interactive multiview video streaming (IMVS) demands data navigation flexibility in the frame structure design, so that server can send only a single periodically selected video view for decoding and display at client, saving transmission bandwidth. In this paper, we generalize previous IMVS frame structure optimization to allow a client to request an arbitrary virtual view; i.e., the server sends two adjacent coded views for the client to synthesize the desired virtual view. Since existing IMVS schemes transmit only one view at a time, they employ only cross-time pre diction; i.e., the frame of previous time instant from which the client switches is used as predictor for the requested view. In our new scenario, two coded views are transmitted, thus within-time prediction can also be used, where the coded frame of one transmitted view is used to predict the frame of the other view of same time instant. Using I-frames, P-frames and Merge (M-) frames as building blocks, we formulate a Lagrangian problem to find the optimal frame structure for a desired storage/streaming rate tradeoff, with the right mixture of cross-time / within-time prediction types. Experiments show that for the same storage cost, the expected streaming rate of the proposed structure can be 40% lower than that of the I-frame-only structure, and 9% lower than that of the structure using M-frames but with cross-time prediction only. Xiaoyu Xiu, Gene Cheung, Antonio Ortega, Jie Liang 0001 |
ICME | 4 |
| 2011 | Perceptual multiview video coding using synthesized Just Noticeable Distortion mapsabstractIn this paper, a perceptual multiview video coding scheme is proposed, based on the synthesized Just Noticeable Distortion (JND) maps. In JND-based perceptual video coding, the residues after intra or inter prediction are tuned according to the corresponding JND thresholds to save the bits without affecting the perceptual quality. In our scheme, to reduce the computational cost of generating the multiview JND maps, only the JND maps of some anchor views are calculated directly. These maps are then used to synthesize the JND maps of other views via the block-based Depth Image Based Rendering (DIBR) method, which can be more than 30 times faster than direct computation with reasonable error. Experimental results show that the proposed scheme can improve the perceptual performance of the JMVC by more than 1 dB in terms of the Peak Signal Perceptual Noise Ratio (PSPNR). Yu Gao 0003, Xiaoyu Xiu, Jie Liang 0001, Weisi Lin |
ISCAS | 3 |
| 2011 | Ztitch: a mobile phone application for 3D scene creation, navigation, and sharingabstractModern smartphones provide an excellent platform for creating 3D scenes from photos. While there already exists many mobile applications that can stitch a set of photos to create a single, panoramic landscape photo, this paper proposes the creation of panoramic scenes where multiple photos are projected in a 3D space using the pinhole camera model, so that a realistic perspective of the scene is maintained. Our application allows users to automatically create a panoramic 3D scene in real-time using the live images from the phone's camera, and to manually fine-tune each photo's position via the touchscreen if the default position is inaccurate. The application enables scenes to be easily shared to other users, and was developed using the Silverlight framework so that it can run across multiple platforms. Andrew Au, Jie Liang 0001 |
ACM Multimedia | 2 |
| 2011 | Distortion exponents of two-way relaying networks with Multiple-Access Broadcast protocolabstractIn this paper, we study the end-to-end distortions when transmitting two Gaussian signals in a three-node, half-duplex, and two-way relaying network, where two users communicate in both directions with the help of one relay, and can transmit at different rates. We focus on the two-phase Multiple-Access Broadcast (MABC) cooperation protocol with decode-and-forward (DF), amplify-and-forward (AF), or compress-and-forward (CF) strategy at the relay. In each case, we analyze the achievable distortion exponents of the two sources at high signal-to-noise ratio (SNR) regime. The results illustrate the effects of the bandwidth ratio and cooperation strategies on the optimized distortion exponents. We also derive the achievable diversity-multiplexing gain tradeoffs of these two-way relaying protocols. Jing Wang 0029, Jie Liang 0001 |
WCNC | 2 |
| 2011 | Average capacity performance of opportunistic relay selection with outdated CSIabstractThe authors investigate the effect of feedback delay on the average capacity of a decode-and-forward cooperative network with relay selection. In particular, a multi-relay cooperative scenario is considered, where the best relay is selected from a subset of relays that are able to decode the source information correctly. In this selection scenario, the authors assume that the destination terminal estimates the relay-to-destination (R→D) channel-state-information perfectly and sends the index of the best relay to the relay terminals via a delayed feedback link. The authors investigate the performance of the considered scenario in terms of average capacity. Simulation results are presented to corroborate the analytical results. Mehdi Seyfi, Sami Muhaidat, Jie Liang 0001 |
IET Commun. | 3 |
| 2011 | A three-layer scheme for M-channel multiple description image coding
Upul Samarawickrama, Jie Liang 0001, Chao Tian 0002 |
Signal Process. | 2 |
| 2011 | Performance Analysis of Relay Selection With Feedback Delay and Channel Estimation ErrorsabstractIn this letter, we investigate the effect of feedback delay and channel estimation errors in a decode-and-forward (DF) cooperative network with relay selection. In particular, we consider a multirelay cooperative scenario, where the best relay is selected from a subset of relays that are able to decode the source information correctly. In the selection scenario, the destination terminal estimates the relay-to-destination (R → D) channel state information (CSI) and sends the index of the best relay to the relay terminals via a delayed feedback link. We investigate the performance of the considered scenario in terms of average symbol error rate (ASER) and asymptotic diversity order. Simulation results are presented to corroborate the analytical results. Mehdi Seyfi, Sami Muhaidat, Jie Liang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2011 | Relay Selection in Dual-Hop Vehicular NetworksabstractIn this letter, we investigate cooperative diversity with relay selection over cascaded Rayleigh fading channels. In particular, we analyze the performance of a relay selection scheme for cooperative vehicular networks with the decode-and-forward (DF) protocol. Only the “best” relay, which satisfies an index of merit, is selected. We ignore the direct transmission between the source (S) and its destination (D), and assume that the destination has perfect knowledge of theS→RandR→Dchannel gains. We study the performance of the underlying scheme in terms of outage probability and investigate its achievable diversity order. Mehdi Seyfi, Sami Muhaidat, Jie Liang 0001, Murat Uysal |
IEEE Signal Process. Lett. | 3 |
| 2011 | Rectification-Based View Interpolation and Extrapolation for Multiview Video CodingabstractIn this paper, we first develop improved projective rectification-based view interpolation and extrapolation methods, and apply them to view synthesis prediction-based multiview video coding (MVC). A geometric model for these view synthesis methods is then developed. We also propose an improved model to study the rate-distortion (R-D) performances of various practical MVC schemes, including the current joint multiview video coding standard. Experimental results show that our schemes achieve superior view synthesis results, and can lead to better R-D performance in MVC. Simulation results with the theoretical models help explaining the experimental results. Xiaoyu Xiu, Derek Pang, Jie Liang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | Effect of Feedback Delay on the Performance of Cooperative Networks with Relay SelectionabstractIn this paper, we analyze the effect of feedback delay and channel estimation errors on the performance of a decode-and-forward (DF) cooperative transmission scenario with relay selection. In our relay selection scheme, only one relay with the best relay-to-destination (R → D) channel quality is selected among the set of relays that decode the source information correctly. Specifically, the destination terminal first estimates the channel state information (CSI) of all active R → D links and then sends the index of the best relay to the relay terminals via a delayed feedback link. Due to the time varying nature of the fading channels, selection is performed based on the old version of the channel estimate. Closed-form expressions for the outage probability, average capacity and average symbol error rate (ASER) are derived. Through asymptotic diversity order analysis, we show that the presence of feedback delay reduces the asymptotic diversity order to one, while the effect of channel estimation errors reduces it to zero. Finally, simulation results are presented to corroborate the analytical results. Mehdi Seyfi, Sami Muhaidat, Jie Liang 0001, Mehrdad Dianati |
IEEE Trans. Wirel. Commun. | 3 |
| 2010 | Amplify-and-Forward Selection Cooperation with Channel Estimation ErrorabstractIn this paper, we investigate the performance of selection cooperation in the presence of imperfect channel estimation. In particular, we consider a cooperative scenario with multiple relays and amplify-and-forward protocol over frequency flat fading channels. In the selection scheme, only the "best" relay which maximizes the effective signal-to-noise ratio (SNR) at the receiver end is selected. We present lower and upper bounds on the effective SNR and derive closed-form expressions for the average symbol error rate (ASER), in the presence of channel estimation errors. A simulation study is presented to corroborate the analytical results and to demonstrate the performance of relay selection with imperfect channel estimation. Mehdi Seyfi, Sami Muhaidat, Jie Liang 0001 |
GLOBECOM | 3 |
| 2010 | Distortion exponents of source transmission over two-way relaying cooperative networksabstractIn this paper, we consider the source transmission in a three-node, half-duplex, and two-way relaying network, where two users communicate with the help of one relay. The relay employs the decode-and-forward (DF) based relaying protocol. We study the distortion exponent that characterizes the high signal-to-noise ratio (SNR) behavior of the end-to-end distortion of the reconstructed signal at each user node. We first provide an upper bound on the achievable distortion exponents of two-way relaying communications, which is tight at large bandwidth ratio. We then investigate the performance of various coding and transmission schemes, including the conventional one-way relaying strategies and source-channel coding in two-way relaying with single-rate coding or limited channel state feedback. We derive the achievable distortion exponents of all these schemes and illustrate the effect of the bandwidth ratio, feedback resolution, and relaying strategies on the optimal distortion exponent. Jing Wang 0029, Jie Liang 0001 |
ICASSP | 2 |
| 2010 | A three-layer algorithm for M-channel multiple description image codingabstractIn this paper, a three-layer scheme is developed for M-channel multiple description image coding. In each description, a subset of the source samples is encoded in the first layer. In the second layer, the remaining subsets are encoded sequentially by predicting from the already encoded subsets. The third layer encoding is designed to refine the reconstruction when only one description is lost, which is the dominant loss scenario in practice. We first derive the closed-form expressions of the expected distortion of the system for 1-D sources when different numbers of descriptions are received. The scheme is then applied to lapped transform based image coding. Simulation results show that the method outperforms some competing schemes. Upul Samarawickrama, Jie Liang 0001, Chao Tian 0002 |
ICIP | 2 |
| 2010 | An improved rate-distortion model for multiview video codingabstractAn improved rate-distortion model is developed to compare the performances of different multiview video coding (MVC) schemes. We first formulate the coding efficiency of one frame with a given temporal and inter-view prediction configuration. The overall performance of a MVC scheme is then obtained by averaging the coding efficiencies of all frames in all views, with different prediction configurations. We then apply the model to evaluate the theoretical performances of various MVC schemes, including the motion and disparity estimation-based MVC, the rectified view interpolation-based MVC, and the rectified view extrapolation-based MVC. The findings of these theoretical studies agree well with the experimental results. Xiaoyu Xiu, Jie Liang 0001 |
ICIP | 2 |
| 2010 | Lloyd-Max quantization-based priority index assignment for the scalable extension of H.264/AVCabstractA fast priority index (PID) assignment algorithm is developed for the MGS (Medium Grain Scalability) packets in the scalable extension of the H.264/AVC. The contributions of the paper are threefold. First, we formulate the index assignment problem as the quantization of the rate-distortion (R-D) slopes of MGS packets, and use the Lloyd-Max algorithm to find the optimal solution. The slope quantization index of a packet is used as its PID. The complexity of our method is much lower than existing method. Secondly, the quantization-based PIDs facilitate the comparisons of packets from different video streams. Video multiplexing results show that the overall PSNR can be improved up to 1 dB. Finally, we propose some real-time and adaptive implementations of the proposed method, which have the same performance as the offline method. Xiaozheng Huang, Jie Liang 0001, Jiangchuan Liu |
ISCAS | 2 |
| 2010 | Distortion exponents of two-way relaying cooperative networksabstractIn this paper, we study the transmission of Gaussian signals in a three-node half-duplex bidirectional relaying network, where two users communicate in both directions with the help of one relay, and can transmit at different rates. The relay employs amplify-and-forward (AF) or decode-and-forward (DF) based cooperation protocols. We analyze the distortion exponent that characterizes the high signal-to-noise ratio (SNR) behavior of the end-to-end distortion of the reconstructed signal at each user node. The different rates of the two users necessitate the study of a new concept - the achievable distortion exponent region of the system. We first derive an outer bound on the distortion exponent region of two-way relaying communications, which is tight at large bandwidth ratio. We then obtain the optimal distortion exponent pairs of conventional one-way relaying strategies and AF/DF based two-way relaying protocols with single-rate coding. The results illustrate the effect of the bandwidth ratio and cooperation strategies on the optimal distortion exponents. Jing Wang 0029, Jie Liang 0001 |
ISIT | 2 |
| 2010 | An improved depth map estimation algorithm for view synthesis and multiview video codingabstractIn this paper, an improved algorithm to generate a smooth and accurate depth map for view synthesis and multiview video coding is developed. For each block in the target view, the algorithm first uses epipolar geometry to find its matched block in the reference view, from which an initial depth is obtained using the triangulation method and depth projection. 3D warping is then applied to refine the depth. In addition, a structural similarity and maximum likelihood-based approach is developed to fuse the depth estimations from multiple references. Finally, the depth map is smoothed via segmentation and plane fitting. Compared to existing 3D warping-based depth estimation, the proposed algorithm can achieve up to 4 dB improvement in view synthesis, while requires much fewer bits to encode the depth map. Experimental results in multiview video coding show that the proposed method can outperform the H.264 JMVC software by more than 1 dB. Xiaoyu Xiu, Jie Liang 0001 |
VCIP | 2 |
| 2010 | Outage Probability of Selection Cooperation with Channel Estimation ErrorsabstractIn this paper, we investigate the performance of selection cooperation in the presence of channel estimation errors. In particular, we consider a cooperative scenario with multiple relays and amplify-and-forward protocol over frequency flat fading channels. In the selection scheme, only the "best" relay which maximizes the effective signal-to-noise ratio (SNR) at the receiver end is selected. We present a lower bound on the effective SNR and derive closed-form expressions for the outage probability of the transmitted signal in the presence of channel estimation error. A simulation study is presented to corroborate the analytical results and to demonstrate the performance of relay selection with imperfect channel estimation. Mehdi Seyfi, Sami Muhaidat, Jie Liang 0001 |
VTC Spring | 3 |
| 2010 | Exploiting Reception Diversity in Adaptive Packet Scheduling over Multimedia Broadcast/Multicast NetworksabstractWe propose a cross-layer optimization framework for multi-session broadcast/multicast (BC/MC) over multimedia network. The proposed framework seeks to perform simultaneous adaptations via intelligent co-operations between satellite gateway, terrestrial gap-fillers and respective BC/MC receivers to cope with highly vibrating satellite link, which is severely constrained by propagation delay, transponder power and channel bandwidth. By jointly optimizing multiple performance criteria across protocol stacks, the scheme mitigates the adverse impacts induced by queuing dynamics, channel variations and user diversities, thereby encompassing channel-dependant and network-friendly features. Meanwhile, it well accommodates return link diversity and the imperfect feedbacks, whilst ensuring fairness, scalability and robustness. We evaluate its performance over diverse network and media configurations in comparison with the state-of-the-art approaches. Numerical results show that simultaneous performance gains can be obtained on multiple essential performance metrics. Jie Liang 0001, Jiangchuan Liu |
WCNC | 2 |
| 2010 | On the Performance of Pilot Symbol Assisted Modulation for Cooperative Systems with Imperfect Channel EstimationabstractIn this paper, we analyze the impact of imperfect channel estimation on the performance of pilot symbol assisted modulation (PSAM) used in a relay communication system with distributed space time block code (STBC) and amplify-and-forward protocol. We derive correlation coefficients between the channel coefficients and their estimates when the relay-to-destination link is either non-fading or fading, in terms of Doppler frequency, number of pilot symbols and SNR. This enables us to choose the optimum number of pilot symbols to compensate for the estimation error when the fading and Doppler effects are severe. Our performance analysis demonstrates that the presence of fading in the relay-to-destination link manifests itself by introducing additional Doppler frequency terms. Furthermore, we derive a tight lower bound for the bit error rate (BER) of BPSK modulation with channel estimation errors in terms of cross correlation coeffcients and the number of pilot symbols. Simulation results are also presented to validate our analytical results. M. Jafar Taghiyar, Sami Muhaidat, Jie Liang 0001 |
WCNC | 3 |
| 2010 | Distortion Exponents for Multi-Relay Cooperative Networks with Limited FeedbackabstractIn this paper, we study the transmission of a Gaussian signal in a multi-relay cooperative system, where each relay is half-duplex and employs the amplify-and-forward (AF) relaying protocol. We focus on the analysis of the distortion exponent that characterizes the high signal-to-noise ratio (SNR) behavior of the end-to-end distortion of the received signal. Specifically, we investigate the feedback scheme where the limited channel state feedback is combined with separate source and channel coding to help the transmission. The feedback scheme is followed by three AF-based multi-relay cooperation protocols, respectively, namely the orthogonal AF protocol, the nonorthogonal AF protocol, and the slotted AF protocol. We derive the optimal distortion exponents of all three cases, and illustrate the effect of the feedback resolution, bandwidth ratio, and number of relays on the optimal distortion exponent. It is shown that the feedback scheme outperforms the best known non-feedback strategies for multi-relay cooperative systems with only a few bits of feedback information. Jing Wang 0029, Jie Liang 0001, Sami Muhaidat |
WCNC | 2 |
| 2010 | M-Channel Multiple Description Coding With Two-Rate Coding and Staggered QuantizationabstractA low complexityM-channel multiple description coding scheme is developed in this paper, in which each description carries one subset of the input with a higher bit rate and the rest with a lower bit rate. The lower-rate codings in different descriptions are designed to be mutually refinable using staggered scalar quantizers. For correlated sources, a two-rate predictive coding is used in each description. Closed-form expressions of the distortions are derived when different numbers of descriptions are received. The application of the proposed scheme in lapped transform based image coding is also investigated, and the optimal transform is obtained. Experimental results using both 1-D memoryless sources and 2-D images demonstrate the superior performance of the proposed scheme. Upul Samarawickrama, Jie Liang 0001, Chao Tian 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Directional Lapped Transforms for Image CodingabstractIn this paper, we present the design of directional lapped transforms for image coding. A lapped transform, which can be implemented by a prefilter followed by a discrete cosine transform (DCT), can be factorized into elementary operators. The corresponding directional lapped transform is generated by applying each elementary operator along a given direction. The proposed directional lapped transforms are not only nonredundant and perfectly reconstructed, but they can also provide a basis along an arbitrary direction. These properties, along with the advantages of lapped transforms, make the proposed transforms appealing for image coding. A block-based directional transform scheme is also presented and integrated into HD Phtoto, one of the state-of-the-art image coding systems, to verify the effectiveness of the proposed transforms. Jizheng Xu, Feng Wu 0001, Jie Liang 0001, Wenjun Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2010 | Cross-layer quality-driven adaptation for scheduling heterogeneous multimedia over 3G satellite networks
Xiaozheng Huang, Jie Liang 0001, Jiangchuan Liu, Barry G. Evans, Imrich Chlamtac |
Wirel. Networks | 3 |
| 2009 | Multiview video coding using projective rectification-based view extrapolation and synthesis bias correctionabstractCurrent view synthesis prediction (VSP) techniques for multiview video coding (MVC) rely on disparity-based view interpolation or depth-based 3D warping. The former cannot be applied to every camera view, whereas the latter may require coding of the depth information of a scene. To avoid these constraints, we propose an improved VSP-based MVC scheme based on the following three techniques: 1) view extrapolation, which allows VSP to be applicable to almost all camera views, 2) projective rectification, which improves the synthesis quality when neighboring camera planes are not parallel, and 3) synthesis bias correction, which uses the past synthesis biases to improve the synthesis quality of the current frame. Experimental results demonstrate that our scheme offers PSNR gains of up to 1.6 dB compared to the current MVC standard. Derek Pang, Xiaoyu Xiu, Jie Liang 0001 |
ICME | 3 |
| 2009 | Rate-distortion analysis of rectification-based view interpolation for multiview video codingabstractView interpolation has been applied in multiview view coding. However, existing schemes assume all cameras are aligned. These methods may not perform well when neighboring cameras point to different directions. In this paper, we apply the rectification based view interpolation to MVC. We first derive the theoretical performance gain of the rectification based view interpolation over existing interpolation method. We then analyze the rate-distortion performance of the proposed MVC method, and compare with the disparity compensation based MVC and view interpolation based MVC without rectification. The analyses show that view rectification can offer additional coding gain over existing view interpolation based MVC, especially when the disparity estimation is accurate and when the neighboring cameras are close to each others. Preliminary experimental results are provided to verify the theoretical analysis. Xiaoyu Xiu, Jie Liang 0001 |
ICME | 2 |
| 2009 | Distortion exponents for decode-and-forward multi-relay cooperative networksabstractIn this paper, we consider the transmission of a Gaussian signal in a multi-relay cooperative system, where each relay is half-duplex and employs the decode-and-forward relaying protocol. We focus on the analysis of the distortion exponent, which characterizes the high signal-to-noise ratio (SNR) behavior of the end-to-end distortion. Specifically, we investigate the layered source coding with progressive or broadcast transmission. Each transmission scheme is further combined with the repetition-based or relay-selection-based multi-relay cooperation protocol. We derive the distortion exponents of all four cases and illustrate the effect of the bandwidth expansion ratio, number of relays and cooperation protocols on the optimal distortion exponent. We also establish the successive refinability of the diversity-multiplexing tradeoff of the repetition-based and relay-selection-based cooperation protocols in multi-relay cooperative systems. Jing Wang 0029, Jie Liang 0001, Sami Muhaidat |
ISIT | 2 |
| 2009 | GPU-aided directional image/video interpolation for real time resolution upconversionabstractImage/video spatial resolution upconversion aims to obtain a high resolution output from the original low resolution input. Fast resolution upconversion is desired in many applications. In this paper, we develop a GPU-friendly two-pass directional image/video resolution upconversion algorithm and present a GPU implementation of the method, using the NVIDIA CUDA (Compute Unified Device Architecture) technology. Design considerations to speed up the execution are discussed, by taking full advantage of the properties of the CUDA framework and the upconversion scheme. Experimental results show that using a mid-range GPU card, the GPU-optimized resolution upconversion implementation can be more than five times as fast as the original method. Ming-Chao Che, Xiaolin Wu 0001, Jie Liang 0001 |
MMSP | 4 |
| 2009 | Projective rectification-based view interpolation for multiview video coding and free viewpoint generationabstractA projective rectification-based view interpolation algorithm is developed for multiview video coding and free viewpoint video. It first calculates the fundamental matrix between two views without using any camera parameter. The two views are then resampled to have horizontal and matched epipolar lines. One-dimensional disparity is estimated next, which is used to interpolate the image for an intermediate viewpoint. After unrectification, the interpolated view can be displayed directly for free viewpoint video purpose. It can also be used as a reference to encode data of an intermediate camera. Experimental results show that the interpolated views can be 3 dB better than existing method. Video coding results illustrate that the method can provide up to 1.3 dB improvement over JMVC. Xiaoyu Xiu, Jie Liang 0001 |
PCS | 2 |
| 2009 | Multiple Description Coding With Prediction CompensationabstractA new multiple description coding paradigm is proposed by combining the time-domain lapped transform, block level source splitting, linear prediction, and prediction residual encoding. The method provides effective redundancy control and fully utilizes the source correlation. The joint optimization of all system components and the asymptotic performance analysis are presented. Image coding results demonstrate the superior performance of the proposed method, especially at low redundancies. Guoqian Sun, Upul Samarawickrama, Jie Liang 0001, Chao Tian 0002, Chengjie Tu, Trac D. Tran |
IEEE Trans. Image Process. | 3 |
| 2008 | M-Channel Multiple Description Coding with Two-Rate Predictive Coding and Staggered QuantizationabstractA low complexity multiple description (MD) coding method is proposed to generate M descriptions. Consider the MD coding of a stationary correlated source. We first fictitiously partition the source into sample blocks of size M, i.e., M polyphases. Each description encodes all input samples, but with a variable bit rate that depends on the indices of the sample and the description. A special DPCM encoder is used in each description, where each sample is predicted from the reconstructed samples in the same description. The prediction error is uniformly scalar-quantized and entropy coded. Upul Samarawickrama, Jie Liang 0001 |
DCC | 2 |
| 2008 | Filter Banks for Prediction-Compensated Multiple Description CodingabstractThis paper investigates the design and application of the optimal filter banks for a prediction-compensated multiple description coding (PC-MDC) scheme, where the coefficients in each subband are split into two descriptions. Each description also includes the prediction residuals of the data in the other description. The optimal designs of orthogonal and biorthogonal filter banks with multiple-level decompositions are formulated in a unified framework. The optimal results in all cases are found to be very close to the optimal filter banks in traditional single description coding. This allows us to apply the proposed method to existing systems with single-description-optimized filter banks and still enjoy near-optimal performance. Image coding results in the JPEG 2000 framework show that the proposed method achieves similar or better performance than other methods. It also has lower complexity and is more compatible to the JPEG 2000 standard. Jing Wang 0029, Jie Liang 0001 |
DCC | 2 |
| 2008 | Directional Lapped Transforms for Image CodingabstractThis paper presents a scheme to design directional lapped transforms. Lapped transforms can be factorized into lifting steps. By introducing directional operator into each lifting step, the directional lapped transform is constructed. The directional lapped transform proposed not only preserves the advantages of lapped transforms, it also can represent directional signals more efficiently. An image coding scheme using the directional lapped transform is also described. Compared to the state-of-the-art image coding using lapped transform, HD photo, the proposed scheme shows more than 20 dB's gain for artificial images with strong directional correlations. And for natural images, up to 1.5 dB's gain can also be observed. Jizheng Xu, Feng Wu 0001, Jie Liang 0001, Wenjun Zhang 0001 |
DCC | 3 |
| 2008 | Lowcomplexity M-channel multiple description coding with two-rate predictive coding and staggered quantizationabstractThis paper presents a new low complexity multiple description coding (MDC) method that can generate any number of descriptions. For correlated sources, a special DPCM encoder is used in each description, such that it carries higher rate information of a subset of the samples and lower rate information of the rest. The lower rate codings in different descriptions are designed to be mutually refinable using staggered scalar quantizers. The closed-form expression of the expected distortion is derived when an arbitrary subset of the descriptions are received. Experimental results on natural images using lapped transform show that the proposed method is competent with the state-of-the-art multiple description image coders. Upul Samarawickrama, Jie Liang 0001 |
ICIP | 2 |
| 2007 | Multiple Description Image Codingwith Prediction CompensationabstractA new multiple description image coding paradigm is presented in this paper by combining the lapped transform, block level source splitting, inter-description prediction, and coding of the prediction residual. Jointly optimal designs of all system components are discussed. Compared with the best multiple description image coding algorithm in the literature, the new method can achieve significant improvement when one description is lost, given the same bit rate and the same central distortion. Guoqian Sun, Upul Samarawickrama, Jie Liang 0001, Chengjie Tu, Trac D. Tran |
ICIP (6) | 3 |
| 2007 | Undersampled Boundary Pre-/Postfilters for Low Bit-Rate DCT-Based Block CodersabstractIt has been well established that critically sampled boundary pre-/postfiltering operators can improve the coding efficiency and mitigate blocking artifacts in traditional discrete cosine transform-based block coders at low bit rates. In these systems, both the prefilter and the postfilter are square matrices. This paper proposes to use undersampled boundary pre- and postfiltering modules, where the pre-/postfilters are rectangular matrices. Specifically, the prefilter is a "fat" matrix, while the postfilter is a "tall" one. In this way, the size of the prefiltered image is smaller than that of the original input image, which leads to improved compression performance and reduced computational complexities at low bit rates. The design and VLSI-friendly implementation of the undersampled pre-/postfilters are derived. Their relations to lapped transforms and filter banks are also presented. Two design examples are also included to demonstrate the validity of the theory. Furthermore, image coding results indicate that the proposed undersampled pre-/postfiltering systems yield excellent and stable performance in low bit-rate image coding. Lu Gan 0002, Chengjie Tu, Jie Liang 0001, Trac D. Tran, Kai-Kuang Ma |
IEEE Trans. Image Process. | 3 |
| 2007 | Wiener Filter-Based Error Resilient Time-Domain Lapped TransformabstractIn this paper, the design of the error resilient time-domain lapped transform is formulated as a linear minimal mean-squared error problem. The optimal Wiener solution and several simplifications with different tradeoffs between complexity and performance are developed. We also prove the persymmetric structure of these Wiener filters. The existing mean reconstruction method is proven to be a special case of the proposed framework. Our method also includes as a special case the linear interpolation method used in DCT-based systems when there is no pre/postfiltering and when the quantization noise is ignored. The design criteria in our previous results are scrutinized and improved solutions are obtained. Various design examples and multiple description image coding experiments are reported to demonstrate the performance of the proposed method. Jie Liang 0001, Chengjie Tu, Lu Gan 0002, Trac D. Tran, Kai-Kuang Ma |
IEEE Trans. Image Process. | 1 |
| 2006 | Two-Dimensional Wiener Filters for Error Resilient Time Domain Lapped TransformabstractThis paper presents the design of two-dimensional Wiener filters for error resilient time domain lapped transform. Two solutions are discussed, and a multi-pass approach is also proposed to make the algorithm adaptive to input statistics. Design examples and image coding experiments show that the adaptive 2-D Wiener filters provide significant improvement over the existing 1-D Wiener filtering method. Jie Liang 0001, Xin Li 0005, Guoqian Sun, Trac D. Tran |
ICASSP (3) | 1 |
| 2006 | Error resilient pre/post-filtering for DCT-based block coding systemsabstractBlock coding based on the discrete cosine transform (DCT) is very popular in image and video compression. Pre/post-filtering can be attached to a DCT-based block coding system to improve coding efficiency as well as to mitigate blocking artifacts. Previously designed pre/post-filters are optimized to maximize coding efficiency solely. For image and video communication over unreliable channels, those pre/post-filters are sensitive to transmission errors. This paper addresses the problem of designing pre/post-filters which are more error resilient. Reconstruction performance is measured by how low the average reconstruction error is, and how uniformly the reconstruction error is distributed. A family of pre/post-filters is designed to provide desired tradeoffs between coding efficiency and robustness to transmission errors. Experiments show that these filtering operators can achieve superior reconstruction performance without sacrificing much coding performance. Chengjie Tu, Trac D. Tran, Jie Liang 0001 |
IEEE Trans. Image Process. | 3 |
| 2005 | Wiener Filtering for Generalized Error Resilient Time Domain Lapped TransformabstractIn this paper, we revisit the design of the time-domain lapped transform for error resilient image transmission. A general structure is first proposed whose solution is given by a Wiener filter. Two simplified schemes with different tradeoffs between complexity and performance are then developed, for which Wiener filter solutions also exist. We show that the existing method is a special case of the general scheme. Design examples and image coding experiments verify that the performance of our new approach is significantly better than existing techniques. Jie Liang 0001, Chengjie Tu, Trac D. Tran, Lu Gan 0002 |
ICASSP (2) | 1 |
| 2005 | Optimal block boundary pre/postfiltering for wavelet-based image and video compressionabstractThis paper presents a pre/postfiltering framework to reduce the reconstruction errors near block boundaries in wavelet-based image and video compression. Two algorithms are developed to obtain the optimal filter, based on boundary filter bank and polyphase structure, respectively. A low-complexity structure is employed to approximate the optimal solution. Performances of the proposed method in the removal of JPEG 2000 tiling artifact and the jittering artifact of three-dimensional wavelet video coding are reported. Comparisons with other methods demonstrate the advantages of our pre/postfiltering framework. Jie Liang 0001, Chengjie Tu, Trac D. Tran |
IEEE Trans. Image Process. | 1 |
| 2004 | Optimal block boundary pre/post-filtering for wavelet-based image and video compressionabstractThis paper presents a pre/post-filtering method to reduce the reconstruction errors near block boundaries in wavelet-based image and video compression. It can be used effectively to mitigate the tiling artifact in JPEG2000 and the jittering artifact in 3D wavelet-based video compression. In this method, a short prefilter is applied across the boundaries of image tiles or video frame groups before wavelet compression, and a post-filter is applied at the same place after wavelet reconstruction. The optimal pre/postfilter is obtained by formulating and solving the corresponding rate-distortion optimization problem. A low-complexity structure is then proposed to approximate the optimal solution. The performance of the proposed method is demonstrated by both image and video coding examples. Jie Liang 0001, Chengjie Tu, Trac D. Tran |
ICIP | 1 |
| 2004 | Over-sampled and under-sampled Pre/post-filters for block DCT codersabstractPre-post-filtering operators have been shown to improve the coding efficiency as well as to mitigate blocking artifacts in traditional DCT-based block coders. Pre-post-filters which preserve the system sampling rate have been extensively investigated. This paper explores the two noncritically-sampled signal decomposition cases - under-sampling and over-sampling - via the pre-post-processing perspective. We discuss various design issues, efficient structures, and present two application examples: under-sampled pre-post-filters for very low bit-rate image coding and over-sampled pre-post-filters for error resilient image transmission. Preliminary experimental results illustrate that noncritically-sampled pre-post-filtering does provide much improved coding performances than its critically-sampled counterpart for the aforementioned specific applications. Chengjie Tu, Trac D. Tran, Jie Liang 0001 |
ICIP | 3 |
| 2003 | On efficient implementation of oversampled linear phase perfect reconstruction filter banksabstractIn this paper, we first present an alternative way of generating oversampled linear phase perfect reconstruction filter banks (OSLPPRFB). We show that this method provides the minimal factorization of a subset of existing OSLPPRFB. The combination of the new structure and the conventional one leads to efficient implementations of a general class of OSLPPRFB. Possible application of the new scheme is discussed. Jie Liang 0001, Lu Gan 0002, Chengjie Tu, Trac D. Tran, Kai-Kuang Ma |
ICASSP (6) | 1 |
| 2003 | Error resilient pre-/post-filtering for DCT-based block coding systemsabstractPre-/post-filtering can be attached to a DCT-based block coding system to improve coding efficiency as well as to mitigate blocking artifacts. Previously designed pre-/post-filters are optimized to maximize coding efficiency solely. For image and video communication over unreliable channels, those pre-/post-filters are sensitive to transmission errors. This paper addresses the problem of designing pre-/post-filters which are more error resilient. A family of pre-/post-filters are designed to provide desired trade-offs between coding efficiency and robustness to transmission errors. These filters achieve superior reconstruction performance without sacrificing much coding performance. Chengjie Tu, Trac D. Tran, Jie Liang 0001 |
ICASSP (3) | 3 |
| 2003 | On efficient implementation of oversampled linear phase perfect reconstruction filter banksabstractIn this paper, we first present an alternative way of generating over-sampled linear phase perfect reconstruction filter banks (OSLP-PRFB). We show that this method provides the minimal factorization of a subset of existing OSLPPRFB. The combination of the new structure and the conventional one leads to efficient implementations of a general class of OSLPPRFB. Possible application of the new scheme is discussed. Jie Liang 0001, Lu Gan 0002, Chengjie Tu, Trac D. Tran, Kai-Kuang Ma |
ICME | 1 |
| 2003 | Error resilient pre-/post-filtering for DCT-based block coding systemsabstractPre-/post filtering can be attached to a DCT-based block coding system to improve the encoding efficiency as well as to mitigate blocking artifacts. Previously designed pre-/post-filters are optimized to maximize coding efficiency solely. For image and video communication over unreliable channels, those pre-/post-filters are sensitive to transmission errors. This paper addresses the problem of designing pre-/post-filters, which are more error resilient. A family of pre-/post-filters is designed to provide desired trade-offs between coding efficiency and robustness to transmission errors. These filters achieve superior reconstruction performance without sacrificing much coding performance. Chengjie Tu, Trac D. Tran, Jie Liang 0001 |
ICME | 3 |
| 2003 | Adaptive runlength codingabstractRunlength coding is the standard coding technique for block transform-based image/video compression. A block of quantized transform coefficients is first represented as a sequence of RUN/LEVEL pairs that are then entropy coded-RUN being the number of consecutive zeros and LEVEL being the value of the following nonzero coefficient. We point out the inefficiency of conventional runlength coding and introduce a novel adaptive runlength (ARL) coding scheme that encodes RUN and LEVEL separately using adaptive binary arithmetic coding and simple context modeling. We aim to maximize compression efficiency by adaptively exploiting the characteristics of block transform coefficients and the dependency between RUN and LEVEL. Coding results show that with the same level of complexity, the proposed ARL coding algorithm outperforms the conventional runlength coding scheme by a large margin in the rate-distortion sense. Chengjie Tu, Jie Liang 0001, Trac D. Tran |
IEEE Signal Process. Lett. | 2 |
| 2002 | DCT-based general structure for linear phase paraunitary filter banksabstractThe factorization of linear-phase paraunitary filter banks (LPPUFB) has been well studied. In this paper, we show that it can be further simplified by fixing the last stage without losing its completeness. The structure can be viewed as the dual of the GenLOT when the last stage is chosen to be the DCT. However, the new structure is more flexible since it can generate basis functions of arbitrary length. The implementation for finite-length signals is discussed. A DCT-oriented initialization method for filter bank optimization is developed to improve its convergence. The proposed method leads to an effective way of handling the sign parameters when modeling orthogonal matrices via Givens rotations. As a result, better optimization results can be obtained. Jie Liang 0001, Trac D. Tran |
ICASSP | 1 |
| 2002 | Further results on DCT-based linear phase paraunitary filter banksabstractA DCT-based simplified general structure for a linear phase paraunitary filter bank (LPPUFB) was developed by the pre- and post-processing of the DCT in the time domain. The new structure can be viewed as the dual of the generalized LOT (GenLOT). We generalize the result to odd-channel LPPUFB and LPPUFB with pair-wise mirror image property. Design examples and their application in image compression are presented. Jie Liang 0001, Trac D. Tran |
ICIP (2) | 1 |