Zixiang Xiong

dblp:70/2057 · DBLP profile ↗
← Back
239ranked-venue papers
21as first author
24since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 133 · 19 first-author · 10 since 2021Computer networks · 44 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 1 first-author · 4 since 2021Theory of computation · 19 · 1 first-authorArtificial intelligence and machine learning · 14 · 8 since 2021Databases, data management, data science and information retrieval · 11Systems, architecture and hardware · 5Security and privacy · 3
YearPublicationVenuePosition
2025 Luminance decomposition and reconstruction for high dynamic range Video Quality Assessment
Jifan Yang, Zhongyuan Wang 0001, Baojin Huang, Jiaxin Ai, Yuhong Yang 0001, Jing Xiao 0004, Zixiang Xiong
Pattern Recognit.7
2025 Finite Blocklength Analysis of Energy-Delay Tradeoff in Uncoordinated MAC
abstract
Polyanskiy [2] proposed a framework for the MAC problem where users employ a common codebook in the finite blocklength regime. In this work, we extend [2] to the case when each user generates a packet independently and asynchronously according to identical renewal processes. Each packet must be decoded with a delay no greater than d, which is a strict delay constraint. We derive a general converse bound for our system model before conducting a comparative analysis of several achievable transmission schemes. First, we consider a transmission scheme where each user transmits a packet as soon as it is generated and suffers interference from other users. We treat interference as noise (TIN). Then, we investigate block transmission where users transmit jointly in fixed slots. Finally, we explore packet splitting, allowing users to divide each packet into two parts with different blocklengths. Using optimized information bit allocation, we discuss two analytical approaches: dependence testing and Gaussian ball. Our numerical results indicate that, when dealing with large delays, TIN performs well compared to block transmission. Conversely, block transmission has the advantage as the system spectral efficiency increases. Packet splitting outperforms other schemes consistently.
Yuming Han, Zixiang Xiong, Anders Høst-Madsen
IEEE Trans. Commun.2
2025 Bursty Versus Continuous Transmission for Wireless Streaming
abstract
This paper analyzes energy consumption in wireless streaming with delay constraints. One fundamental question is whether the transmitter should transmit a steady stream of bits, continuous transmission, or transmit data in bursts and switch off while not transmitting. In the latter case, a question is also for what fraction of time the transmitter should transmit, the duty cycle. In this paper we use traditional and finite blocklength information theory to see whether bursty transmission would be better than transmitting data continuously. When doing this analysis we consider latency and energy efficiency, two fundamental parameters that characterize communication systems. We take into account realistic models of hardware, including overhead power, power amplifier inefficiency and the receiver noise factor. With these models we find that the energy consumption and latency tradeoff behaves quite differently from the ideal case.
Nicholas Whitcomb, Anders Høst-Madsen, Zixiang Xiong, Jeffrey A. Weldon
IEEE Trans. Commun.3
2025 Dynamic Control Authority Allocation in Indirect Shared Control for Steering Assistance
abstract
The concept of shared control has garnered significant attention within the realm of human-machine hybrid intelligence research. This study introduces a novel approach, specifically a dynamic control authority allocation method, for implementing shared control in autonomous vehicles. Unlike conventional mixed-initiative control techniques that blend human and vehicle inputs with weights determined by predefined index, the proposed method utilizes optimization-based techniques to obtain an optimal dynamic allocation for human and vehicle inputs that satisfies safety constraints. Specifically, a convex quadratic programm (QP) is constructed incorporating control barrier functions (CBF) for safety and control Lyapunov functions (CLF) for satisfying automated control objectives. The cost function of the QP is designed such that human weight increases with the magnitude of human input. A smooth control authority transition is obtained by optimizing over the change rate of the weight instead of the weight itself. The proposed method is verified in lane-changing scenarios with human-in-the-loop (HmIL) and hardware-in-the-loop (HdIL) experiments. Results show that the proposed method outperforms index-based control authority allocation method in terms of agility, safety and comfort.
Haocong Chen, Jie Huang 0007, Zixiang Xiong, Yuyi Wang 0001, Xiwen Yuan
IEEE Trans. Intell. Transp. Syst.6
2024 Latency and Energy Minimization in NOMA-Assisted MEC Network: A Federated Deep Reinforcement Learning Approach
abstract
Multi-access edge computing (MEC) is seen as a vital component of forthcoming 6G wireless networks, aiming to support emerging applications that demand high service reliability and low latency. However, ensuring the ultra-reliable and low-latency performance of MEC networks poses a significant challenge due to uncertainties associated with wireless links, constraints imposed by communication and computing resources, and the dynamic nature of network traffic. Enabling ultra-reliable and low-latency MEC mandates efficient load balancing jointly with resource allocation. In this paper, we investigate the joint optimization problem of offloading decisions, computation and communication resource allocation to minimize the expected weighted sum of delivery latency and energy consumption in a non-orthogonal multiple access (NOMA)-assisted MEC network. Given the formulated problem is a mixed-integer non-linear programming (MINLP), a new multi-agent federated deep reinforcement learning (FDRL) solution based on double deep Q-network (DDQN) is developed to efficiently optimize the offloading strategies across the MEC network while accelerating the learning process of the Internet-of-Thing (IoT) devices. Simulation results show that the proposed FDRL scheme can effectively reduce the weighted sum of delivery latency and energy consumption of IoT devices in the MEC network and outperform the baseline approaches.
Arian Ahmadi, Anders Høst-Madsen, Zixiang Xiong
ISCC3
2024 Domain generalized person reidentification based on skewness regularity of higher-order statistics
Mingfu Xiong, Ruimin Hu, Zhongyuan Wang 0001, Javier Del Ser, Khan Muhammad 0001, Zixiang Xiong
Knowl. Based Syst.7
2024 Boosting integral-based human pose estimation through implicit heatmap learning
Congju Du, Zengqiang Yan, Zixiang Xiong, Li Yu 0003
Neural Networks3
2024 WSPTGAN for Global Ocean Surface Wind Speed Generation With High Temporal Resolution and Spatial Coverage
abstract
Obtaining global ocean surface wind speed data with high temporal resolution and spatial coverage is a challenging task. Due to the lack of widely applicable direct measurement methods and algorithms, current research and data products can only achieve good performance in a small spatial range or at low temporal resolution. In this article, a generative adversarial network (GAN) with a transformer structure called Wind Speed Prediction transformer-GAN (WSPTGAN) is proposed to generate wind speed data with good spatial coverage and high temporal resolution for areas. The WSPTGAN is trained with the proposed image-like wind speed data combined partial missing dataset (CPMD), which is combined with the fifth generation of the European Center for Medium-Range Weather Forecast (ECMWF) reanalysis data and Advanced Scatterometer (ASCAT) data from Meteorological Operational satellites. Thanks to the defective data learning mechanism (DDLM), sequential-wise multihead self-attention mechanism (SMSM), and sequence feature adaptive verification mechanism (SFAVM) in the proposed algorithm, the obtained model has good wind speed prediction accuracy with root mean square error (RMSE) of 0.8984 m/s and can achieve multistep 10-min wind speed data generation within the global ocean. After comparison with five state-of-the-art prediction models, it is confirmed that the algorithm in this article is able to make better use of the defective data for learning and prediction of wind field trends in global ocean regions.
Yonghong Hou, Xiaowei Song 0001, Chunping Hou, Zixiang Xiong, Dan Ma 0003
IEEE Trans. Geosci. Remote. Sens.6
2024 Self-Attention-Guided Multiindicator Retrieval for Ocean Surface Wind Field With Multimodal Data Augmentation and Fusion
abstract
The deployment of global navigation satellite system reflectometry (GNSS-R) emerges as a compelling approach for the extraction of ocean surface wind field, primarily due to its exceptional cost-effectiveness, all-weather robustness, and excellent spatiotemporal coverage. Despite these advantages, the insufficient use of various data and the lack of ability to perform multiindicator retrieval limit the performance of existing methods in practical ocean wind field retrieval. To overcome these limitations, this article introduces a novel self-attention-guided ocean surface wind field multiindicator retrieval algorithm based on multimodal observation data augmentation and fusion. Initially, data generation modules are employed to complement high-quality observation data that are not fully provided by GNSS-R system. Subsequently, the multiscale data fusion encoder (MDFE) is implemented to extract and fuse multiple data features of different scales to enhance the utilization ability of data and improve the accuracy of wind field retrieval. Finally, the self-attention multiindicator predictor (SAMIP) is put to use for optimizing the feature attention strategy, achieving accurate retrieval of ocean surface wind speed and direction simultaneously. The proposed method provides a novel solution for the comprehensive utilization of various data products in the GNSS-R system, simultaneously achieves synchronous retrieval of multiple wind field indicators, which are wind speed and direction. In the context of recent advancements in algorithmic development, the proposed algorithm exhibits a notable enhancement in the precision of wind speed and direction retrieval. Benchmarking against ERA5 wind field data, the proposed algorithm achieves a root-mean-square error (RMSE) of 1.23 m/s in wind speed retrieval, demonstrating at least a 9% accuracy improvement compared to five state-of-the-art algorithms from recent years. Furthermore, the RMSE in wind direction retrieval stands at 20.7°, surpassing comparison algorithms by achieving a reduction of over 8% in error. Collectively, these metrics robustly validate the efficacy of the proposed algorithm.
Yonghong Hou, Xiaowei Song 0001, Chunping Hou, Zixiang Xiong, Dan Ma 0003
IEEE Trans. Geosci. Remote. Sens.5
2024 Hyper-Laplacian Prior for Remote Sensing Image Super-Resolution
abstract
Image explicit prior has made breakthrough progress in the super-resolution (SR) due to the additional supervisory information provided. However, existing explicit prior-guided SR methods directly use the Gaussian gradient or Laplacian gradient prior, which cannot fit the gradient distribution of remote sensing images. Through the statistics of gradient probability density distribution of the remote sensing image dataset, we found that the hyper-Laplacian prior can fit the heavy-tailed distribution better, which aroused us to use the hyper-Laplacian before facilitating the SR reconstruction. We propose a novel hyper-Laplacian prior SR method for remote sensing images in this manuscript. Specifically, our model consists of three components: rough reconstruction subnetwork (RRS), hyper-Laplacian prior subnetwork (HPS), and image refinement enhancement subnetwork (RES). In the RRS, we reconstruct low-resolution (LR) images into rough SR images by a set of resblocks. In the HPS, we first introduce the hyper-Laplacian prior for LR images to provide an additional texture. Hereafter, we set up a prior loss which imposes a second-order supervision on the SR image. Like the previous image space loss function, it helps the model to gather the geometric structure of the image. Finally, the outputs of the RRS and HPS are fused and then fed to the RES for high-quality image reconstruction. Numerous studies of SR reconstruction and segmentation on UCMerced, PatternNet, and OpenBayes datasets confirm that our method is superior compared to state-of-the-art methods.
Kanghui Zhao, Tao Lu 0001, Jiaming Wang 0001, Yanduo Zhang, Junjun Jiang, Zixiang Xiong
IEEE Trans. Geosci. Remote. Sens.6
2024 Learning to Hallucinate Face in the Dark
abstract
Face hallucination in low-light environments is an extremely challenging task due to the significant loss of facial structure and facial texture information. Although cascading image relighting and face hallucination tasks is a feasible strategy, simply cascading these two tasks does not achieve satisfactory results because they do not fit into each other naturally. In this article, we propose a novel duplex fusing-embedding learning approach to tackle this challenge in low-light environments. The core of the proposed approach is the duplexity of feature fusion and embedding between relighting and hallucination tasks. In the feature fusion phase, the shallow features from two tasks are bidirectionally fused and activated into a consistent feature space. In the feature embedding phase, the fused features from the previous iteration are fed back and bidirectionally embedded into the deep features of two tasks in the current iteration so that they can learn feature representations that consistently represent both tasks, thereby boosting the performance of relighting and hallucination to generate photorealistic HR face images. Experimental results show that the proposed approach allows current face hallucination methods to learn to hallucinate face in the dark.
Yuanzhi Wang, Tao Lu 0001, Yanduo Zhang, Zixiang Xiong
IEEE Trans. Multim.5
2024 Rethinking Prior-Guided Face Super-Resolution: A New Paradigm With Facial Component Prior
abstract
Recently, facial priors (e.g., facial parsing maps and facial landmarks) have been widely employed in prior-guided face super-resolution (FSR) because it provides the location of facial components and facial structure information, and helps predict the missing high-frequency (HF) information. However, most existing approaches suffer from two shortcomings: 1) the extracted facial priors are inaccurate since they are extracted from low-resolution (LR) or low-quality super-resolved (SR) face images and 2) they only consider embedding facial priors into the reconstruction process from LR to SR face images, thus failing to explore facial priors to generate LR face image. In this article, we propose a novel pre-prior guided approach that extracts facial prior information from original high-resolution (HR) face images and embeds them into LR ones to obtain HF information-rich LR face images, thereby improving the performance of face reconstruction. Specifically, a novel component hybrid method is proposed, which fuses HR facial components and LR facial background to generate new LR face images (namely, LRmix) via facial parsing maps extracted from HR face images. Furthermore, we design a component hybrid network (CHNet) that learns the LR to LRmix mapping function to ensure that the LRmix can be obtained from LR face images in testing and real-world datasets. Experimental results show that our proposed scheme significantly improves the reconstruction performance for FSR.
Tao Lu 0001, Yuanzhi Wang, Yanduo Zhang, Junjun Jiang, Zhongyuan Wang 0001, Zixiang Xiong
IEEE Trans. Neural Networks Learn. Syst.6
2024 Joint Distortion Restoration and Quality Feature Learning for No-reference Image Quality Assessment
abstract
No-reference image quality assessment (NR-IQA) methods, inspired by the free energy principle, improve the accuracy of image quality prediction by simulating the human brain’s repair process for distorted images. However, existing methods use separate optimization schemes for distortion restoration and quality prediction, which undermines the accurate mapping of feature representations to quality scores. To address this issue, we propose a joint restoration and quality feature learning NR-IQA (RQFL-IQA) method to jointly tackle distortion image restoration and quality prediction within a unified framework. To accurately establish the quality reconstruction relationship between distorted and restored images, a hybrid loss function based on pixel-wise and structure-wise representations is used to improve the restoration capability of the image restoration network. The proposed RQFL-IQA exploits rich labels, including restored images and quality scores, to enable the model to learn more discriminative features and establish a more accurate mapping from feature representation to quality scores. In addition, to avoid the impact of poor restoration on quality prediction, we propose a module with a cleaning function to reweight the fusion of restored and primitive features to achieve more perceptual consistency in feature fusion. Experimental results on public IQA datasets show that the proposed RQFL-IQA is superior over existing methods.
Jifan Yang, Zhongyuan Wang 0001, Baojin Huang, Jiaxin Ai, Yuhong Yang 0001, Zixiang Xiong
ACM Trans. Multim. Comput. Commun. Appl.6
2023 PIDNet: A Real-time Semantic Segmentation Network Inspired by PID Controllers
abstract
Two-branch network architecture has shown its efficiency and effectiveness in real-time semantic segmentation tasks. However, direct fusion of high-resolution details and low-frequency context has the drawback of detailed features being easily overwhelmed by surrounding contextual information. This overshoot phenomenon limits the improvement of the segmentation accuracy of existing two-branch models. In this paper, we make a connection between Convolutional Neural Networks (CNN) and Proportional-Integral-Derivative (PID) controllers and reveal that a two-branch network is equivalent to a Proportional-Integral (PI) controller, which inherently suffers from similar overshoot issues. To alleviate this problem, we propose a novel three-branch network architecture: PIDNet, which contains three branches to parse detailed, context and boundary information, respectively, and employs boundary attention to guide the fusion of detailed and context branches. Our family of PIDNets achieve the best trade-off between inference speed and accuracy and their accuracy surpasses all the existing models with similar inference speed on the Cityscapes and CamVid datasets. Specifically, PIDNet-S achieves 78.6% mIOU with inference speed of 93.2 FPS on Cityscapes and 80.1% mIOU with speed of 153.7 FPS on CamVid.
Jiacong Xu, Zixiang Xiong, Shankar P. Bhattacharyya
CVPR2
2023 Self-attention learning network for face super-resolution
Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Jiaming Wang 0001, Zixiang Xiong
Neural Networks6
2023 Hierarchical Associative Encoding and Decoding for Bottom-Up Human Pose Estimation
abstract
Bottom-up human pose estimation decouples computational complexity from the number of people but requires additional operations to match the detected keypoints to each human instance. Existing approaches treat all keypoints equally while ignoring the relationships among keypoints, which in turn limit the performance ceilings. In this work, we propose a hierarchical associative encoding and decoding framework for bottom-up human pose estimation by introducing additional prior knowledge. Specifically, in addition to keypoint-level and instance-level associations, we further divide keypoints into groups and explore group-level associations. This way, prior knowledge is incorporated to determine the keypoint groups for better associative encoding. To deal with complex poses, we introduce a focal pulling loss to focus more on the hard-to-associate keypoints. Moreover, instead of using a pre-defined order for keypoint grouping, we propose a progressive associative decoding method to dynamically determine the order of keypoints for grouping, which helps reduce isolated keypoints. Experimental results on the MS-COCO, CrowdPose and MPII datasets show superior performance of our proposed associative encoding and decoding algorithms. More importantly, we prove, through validation, that hierarchical associative encoding and decoding can be used as a plug-n-play module for performance improvement regardless of backbone architecture. Our source code and pretrained models are available athttps://github.com/ducongju/HAE.
Congju Du, Zengqiang Yan, Li Yu 0003, Zixiang Xiong
IEEE Trans. Circuits Syst. Video Technol.5
2023 FaceFormer: Aggregating Global and Local Representation for Face Hallucination
abstract
Recently, face hallucination methods either feed whole face image into convolutional neural networks (CNNs) or utilize extra facial priors (e.g., facial parsing maps and landmarks) to focus on global facial structure and constrain facial texture generation. However, the limited receptive fields of CNNs and inaccurate facial priors will reduce the naturalness and fidelity of restored face. In this paper, we propose a FaceFormer that aggregates global representation of Transformers and local representation of CNNs to maintain the consistency of facial structure while restoring local facial details. The reason for this design is that the Transformer can capture global facial information by exploiting the long-distance visual relation modeling, while the local modeling capability of CNNs can recover fine-grained facial details. Therefore, aggregating these two independent representations can help to maximize their merits and reconstruct high-quality and high-fidelity face images. Experimental results of face reconstruction and recognition verify that the proposed FaceFormer significantly outperforms current state-of-the-arts.
Yuanzhi Wang, Tao Lu 0001, Yanduo Zhang, Zhongyuan Wang 0001, Junjun Jiang, Zixiang Xiong
IEEE Trans. Circuits Syst. Video Technol.6
2023 Multi-Scale Hybrid Fusion Network for Single Image Deraining
abstract
Deep learning models have been able to generate rain-free images effectively, but the extension of these methods to complex rain conditions where rain streaks show various blurring degrees, shapes, and densities has remained an open problem. Among the major challenges are the capacity to encode the rain streaks and the sheer difficulty of learning multi-scale context features that preserve both global color coherence and exactness of detail. To address the first problem, we design a non-local fusion module (NFM) and an attention fusion module (AFM), and construct the multi-level pyramids' architecture to explore the local and global correlations of rain information from the rain image pyramid. More specifically, we apply the non-local operation to fully exploit the self-similarity of rain streaks and perform the fusion of multi-scale features along the image pyramid. To address the latter challenge, we additionally design a residual learning branch that is capable of adaptively bridging the gaps (e.g., texture and color information) between the predicted rain-free image and the clean background via a hybrid embedding representation. Extensive results have demonstrated that our proposed method is able to generate much better rain-free images on several benchmark datasets than the state-of-the-art algorithms. Moreover, we conduct the joint evaluation experiments with respect to deraining performance and the detection/segmentation accuracy to further verify the effectiveness of our deraining method for downstream vision tasks/applications. The source code is available at https://github.com/kuihua/MSHFN.
Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Guangcheng Wang, Zhen Han 0002, Junjun Jiang, Zixiang Xiong
IEEE Trans. Neural Networks Learn. Syst.8
2022 Spatiotemporal two-stream LSTM network for unsupervised video summarization
Ruimin Hu, Zhongyuan Wang 0001, Zixiang Xiong
Multim. Tools Appl.4
2022 Body size measurement based on deep learning for image segmentation by binocular stereovision system
Xiaowei Song 0001, Xianli Song, Lei Yang 0050, Chunping Hou, Zixiang Xiong
Multim. Tools Appl.6
2022 Rethinking Lightweight: Multiple Angle Strategy for Efficient Video Action Recognition
abstract
Video action recognition task involves modeling spatiotemporal information, and efficiency is critical to capture spatiotemporal dependencies in the video. Most existing models rely on optical flow information to capture the dynamic visual tempos between consecutive video frames. Although impressive performance can be achieved by combining optical flow with RGB, the time-consuming nature of optical flow computation cannot be ignored. Moreover, 3D CNN has successfully modeled spatiotemporal information, yet the enormous computational volume is unsuitable for real-time action recognition. In this letter, we propose a novel lightweight video feature extraction strategy that achieves better recognition performance with lower FLOPs. In particular, we perform convolution on the video cube from three orthogonal angles to learn its appearance and motion features. Compared with the computational volume of 3D CNN, our proposed method is more economical and thus meets the lightweight requirements. Extensive experimental results on public Something Something-V1$\&$V2 and Diving48 datasets show our approach achieves the state-of-the-art performance.
Jianyu Chen 0008, Zhongyuan Wang 0001, Kangli Zeng, Zheng He 0001, Zixiang Xiong
IEEE Signal Process. Lett.5
2022 Dual-Path Deep Fusion Network for Face Image Hallucination
abstract
Along with the performance improvement of deep-learning-based face hallucination methods, various face priors (facial shape, facial landmark heatmaps, or parsing maps) have been used to describe holistic and partial facial features, making the cost of generating super-resolved face images expensive and laborious. To deal with this problem, we present a simple yet effective dual-path deep fusion network (DPDFN) for face image super-resolution (SR) without requiring additional face prior, which learns the global facial shape and local facial components through two individual branches. The proposed DPDFN is composed of three components: a global memory subnetwork (GMN), a local reinforcement subnetwork (LRN), and a fusion and reconstruction module (FRM). In particular, GMN characterize the holistic facial shape by employing recurrent dense residual learning to excavate wide-range context across spatial series. Meanwhile, LRN is committed to learning local facial components, which focuses on the patch-wise mapping relations between low-resolution (LR) and high-resolution (HR) space on local regions rather than the entire image. Furthermore, by aggregating the global and local facial information from the preceding dual-path subnetworks, FRM can generate the corresponding high-quality face image. Experimental results of face hallucination on public face data sets and face recognition on real-world data sets (VGGface and SCFace) show the superiority both on visual effect and objective indicators over the previous state-of-the-art methods.
Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Tao Lu 0001, Junjun Jiang, Zixiang Xiong
IEEE Trans. Neural Networks Learn. Syst.6
2021 Learning Skip Map For Efficient Ultra-High Resolution Image Segmentation
abstract
The pattern distribution of ultra-high resolution images is usually unbalanced. While part of an image contains complex and fine-grained patterns such as boundaries, most areas are composed of simple and repeated patterns. In this work, we propose to learn a skip map, which can guide a segmentation network to skip simple patterns and hence reduce computational complexity. Specifically, the skip map highlights simple-pattern areas that can be down-sampled for processing at a lower resolution, while the remaining complex part is still segmented at the original resolution. Applied on the state-of-the-art ultra-high resolution image segmentation network GLNet, our proposed skip map saves more than 30% computation while maintaining comparable segmentation performance.
Pengcheng Pi, Ziyu Jiang, Zixiang Xiong
ICIP3
2021 Trajectory is not Enough: Hidden Following Detection
abstract
In outdoor crimes such as robbery and kidnapping, suspects generally secretly follow their victims in public places and then look for opportunities to commit crimes. Video anomaly detection (VAD) has achieved fruitful results through deep neural networks (DNN). However, as an abnormal behavior without obvious abnormal physical features, hidden following is highly similar to ordinary walking and accompanying behaviors, so it is difficult to effectively detect hidden dangerous followers using video anomaly detection methods or traditional trajectory analysis methods. We propose "hidden follower'' detection (HFD) task and a HFD model based on gaze pattern extraction. It extracts gaze pattern features of pedestrians from gaze-interval-series and introduces a time series classification model to classify pedestrians with or without hidden following purposes. Based on this model, we propose a hidden follower detection framework (HFDF) to detect hidden followers from normal pedestrians, which utilizes the trajectories and gaze patterns extracted from videos. To cope with the lack of test data, we construct a dataset of 1200 pedestrians from the crowd simulation model to simulate scenes including hidden followers, and we also collected a surveillance video dataset including the hidden following behaviors. The experiments conducted on these two datasets show that HFDF can consistently outperform the state-of-the-art method by a notable margin in the HFD task on the commonly-used F1 benchmark.
Danni Xu, Ruimin Hu, Zixiang Xiong, Zheng Wang 0007, Linbo Luo 0001, Dengshi Li
ACM Multimedia3
2020 Peeking into Occluded Joints: A Novel Framework for Crowd Pose Estimation
Lingteng Qiu, Xuanye Zhang, Yanran Li, Guanbin Li, Zixiang Xiong, Xiaoguang Han 0001, Shuguang Cui
ECCV (19)6
2020 Attention Mechanism Enhanced Kernel Prediction Networks for Denoising of Burst Images
abstract
Deep learning based image denoising methods have been extensively investigated. In this paper, attention mechanism enhanced kernel prediction networks (AME-KPNs) are proposed for burst image denoising, in which, nearly cost-free attention modules are adopted to first refine the feature maps and to further make a full use of the inter-frame and intra-frame redundancies within the whole image burst. The proposed AME-KPNs output per-pixel spatially-adaptive kernels, residual maps and corresponding weight maps, in which, the predicted kernels roughly restore clean pixels at their corresponding locations via an adaptive convolution operation, and subsequently, residuals are weighted and summed to compensate the limited receptive field of predicted kernels. Simulations and real-world experiments are conducted to illustrate the robustness of the proposed AME-KPNs in burst image denoising.
Shenyao Jin, Yili Xia, Yongming Huang 0001, Zixiang Xiong
ICASSP5
2020 Latency-Energy Tradeoff with Realistic Hardware Models
Anders Høst-Madsen, Nicholas Whitcomb, Jeffrey A. Weldon, Zixiang Xiong
ISITA4
2020 JAFPro: Joint Appearance Fusion and Propagation for Human Video Motion Transfer from Multiple Reference Images
abstract
We present a novel framework for human video motion transfer. Deviating from recent studies that use only single source image, we propose to allow users to supply multiple source images by simply imitating some poses in the desired target video. To aggregate the appearance from multiple input images, we propose a JAFPro framework that incorporates two modules: an appearance fusion module that adaptively fuses the information in the supplied images and an appearance propagation module that propagates textures through flow-based warping to further improve the result. An attractive feature of JAFPro is that the quality of its results progressively improves as more imitating images are supplied. Furthermore, we build a new dataset containing a large variety of dancing videos in the wild. Extensive experiments conducted on this dataset demonstrate JAFPro outperforms state-of-the-art methods both qualitatively and quantitatively. We will release our code and dataset upon publication of this work.
Xianggang Yu, Haolin Liu 0004, Xiaoguang Han 0001, Zhen Li 0026, Zixiang Xiong, Shuguang Cui
ACM Multimedia5
2020 Improperness Based SINR Analysis of GFDM Systems Under Joint Tx and Rx I/Q Imbalance
abstract
Adverse impacts of in-phase and quadrature-phase (I/Q) imbalance in both the transmitter (Tx) and receiver (Rx) are quantified for the generalized frequency division multiplexing (GFDM) based transmission over frequency selective fading channels. To this end, we first equip the standard signal-to-interference-plus-noise (SINR) performance evaluation with the ability to consider second-order noncircular (improper) signals, and thus precisely evaluate performance deterioration caused by I/Q distortions over the in-phase (I) and quadrature-phase (Q) channels of a transmission system. Next, we propose a novel means to evaluate the individual SINR contributions from both the channels of GFDM, and hence, provide more meaningful insights into the underlying wireless transmission in the presence of complex non-circularity. This is accompanied by an account of complete augmented second-order statistics of I/Q imbalanced GFDM waveforms which caters for various sources of complex improperness. Simulations in the GFDM system setting support our analysis.
Hao Cheng 0006, Yili Xia, Yongming Huang 0001, Luxi Yang, Zixiang Xiong, Danilo P. Mandic
WCNC5
2020 On the Energy-Delay Tradeoff in Streaming Data: Finite Blocklength Analysis
abstract
This paper investigates basic trade-offs between energy and delay in wireless communication systems using finite blocklength theory. We first assume that data arrive in constant stream of bits, which are put into packets and transmitted over a communications link. Our results show that depending on exactly how energy is measured, in general energy depends √d-1or √Vd-1log d, where d is the delay. This means on that the energy decreases quite slowly with increasing delay. Furthermore, to approach the absolute minimum of -1.59 dB on energy, bandwidth has to increase very rapidly, much more than what is predicted by infinite blocklength theory. We then consider the scenario when data arrive stochastically in packets and can be queued. We devise a scheduling algorithm based on finite blocklength theory and develop bounds for the energy-delay performance. Our results again show that the energy decreases quite slowly with increasing delay.
Mirza Uzair Baig, Lei Yu 0003, Zixiang Xiong, Anders Høst-Madsen, Houqiang Li, Weiping Li 0003
IEEE Trans. Inf. Theory3
2019 Deep Reinforcement Learning of Volume-Guided Progressive View Inpainting for 3D Point Scene Completion From a Single Depth Image
abstract
We present a deep reinforcement learning method of progressive view inpainting for 3D point scene completion under volume guidance, achieving high-quality scene reconstruction from only a single depth image with severe occlusion. Our approach is end-to-end, consisting of three modules: 3D scene volume reconstruction, 2D depth map inpainting, and multi-view selection for completion. Given a single depth image, our method first goes through the 3D volume branch to obtain a volumetric scene reconstruction as a guide to the next view inpainting step, which attempts to make up the missing information; the third step involves projecting the volume under the same view of the input, concatenating them to complete the current view depth, and integrating all depth into the point cloud. Since the occluded areas are unavailable, we resort to a deep Q-Network to glance around and pick the next best view for large hole completion progressively until a scene is adequately reconstructed while guaranteeing validity. All steps are learned jointly to achieve robust and consistent results. We perform qualitative and quantitative evaluations with extensive experiments on the SUNCG data, obtaining better results than the state of the art.
Xiaoguang Han 0001, Zhaoxuan Zhang, Dong Du 0002, Mingdai Yang, Jingming Yu, Xin Yang 0011, Ligang Liu 0001, Zixiang Xiong, Shuguang Cui
CVPR9
2019 Separability and Compactness Network for Image Recognition and Superresolution
abstract
Convolutional neural networks (CNNs) have wide applications in pattern recognition and image processing. Despite recent advances, much remains to be done for CNNs to learn a better representation of image samples. Therefore, constant optimizations should be provided on CNNs. To achieve a good performance on classification, intuitively, samples' interclass separability, or intraclass compactness should be simultaneously maximized. Accordingly, in this paper, we propose a new network, named separability and compactness network (SCNet) to rectify this problem. SCNet minimizes the softmax loss and the distance between features of samples from the same class under a jointly supervised framework, resulting in simultaneous maximization of interclass separability and intraclass compactness of samples. Furthermore, considering the convenience and the efficiency of the cosine similarity in face recognition tasks, we incorporate it into SCNet's distance metric to enable sample features from the same class to line up in the same direction and those from different classes to have a large angle of separation. We apply SCNet to three different tasks: visual classification, face recognition, and image superresolution. Experiments on both public data sets and real-world satellite images validate the effectiveness of our SCNet.
Liguo Zhou, Zhongyuan Wang 0001, Yimin Luo, Zixiang Xiong
IEEE Trans. Neural Networks Learn. Syst.4
2018 iLDPC coded RCM scheme with optimised interleaver
abstract
In this study, an irregular low‐density parity‐check (iLDPC) coded rate compatible modulation (iLDPC‐RCM) is proposed. A particularly designed interleaver is inserted between LDPC and RCM such that a fast convergence and a low bit‐error rate (BER) can be achieved. The combination of strong error correction capability of irregular LDPC code and the seamless adaptation of RCM enables the proposed scheme to achieve a robust and spectrally efficient transmission over time‐varying channels. In addition, a joint belief propagation algorithm is proposed to lower the BER of iLDPC‐RCM and speed up the convergence of iterative decoding. Simulation results show that the proposed iLDPC‐RCM improves the BER and throughput performance while maintaining an acceptable computational complexity at the same time, validating its advantages over the regular LDPC coded RCM with the same coding rate.
Yongqiang Cui, Shaoping Chen, Zixiang Xiong
IET Commun.3
2018 Robust and efficient face recognition via low-rank supported extreme learning machine
Tao Lu 0001, Yingjie Guan, Yanduo Zhang, Shenming Qu, Zixiang Xiong
Multim. Tools Appl.5
2018 Jointly Optimizing User Association and BS Muting for Cache-Enabled Networks With Network-Coded Multicast and Reconstructed Interference Cancelation
abstract
In this paper, we strive to improve the throughput of heterogeneous cellular networks by exploiting the pre-cached files at user end to manage interference. We consider a transmission scheme, where network-coded multicast is employed to help cancel multi-user interference, and reconstructed interference cancellation (RIC) is used to help eliminate inter-cell interference. Because RIC is opportunistic, base station (BS) muting is used to coordinate the residual strong interference. Since user association affects residual interference and is coupled with BS muting while both are operated in a very different timescale from content caching, we jointly optimize user association and BS muting for a given caching policy to maximize the number of users simultaneously served by the transmission scheme. By transforming the formulated problem into a maximal independent set problem with constructed conflict graph, the global optimal solution is found with graph theory methods. By exploiting the topology feature of heterogeneous networks, we proceed to propose two low-complexity algorithms, respectively, implemented in a centralized and distributed manner, which are viable for large-scale networks. Simulation results show that the optimized transmission scheme achieves a remarkable performance gain over the existing schemes.
Kaiyang Guo, Chenyang Yang 0001, Tingting Liu 0001, Zixiang Xiong
IEEE Trans. Commun.4
2018 A Noncoverage Field Model for Improving the Rendering Quality of Virtual Views
abstract
Rendering quality optimization theory is one of the most basic and fascinating components of image-based rendering (IBR). The rendering quality of virtual views is related to the information in the source images. Capturing the image information depends on the geometric configuration of the camera (GCC) and, in particular, the positions and shooting directions of the cameras. Therefore, the rendering quality of virtual views can be improved by optimizing the GCC. This paper investigates the relationship between the GCC and the geometric configurations of the virtual view (GCVV). The influence of the GCC on the rendering quality is also analyzed. Based on the relationships among the GCC, GCVV, and rendering quality, a mathematical model of the noncoverage field (NCF) is proposed to quantify the rendering quality. The performance of the NCF is also analyzed using a set of GCVVs and GCCs. Furthermore, an optimization algorithm based on the NCF that simultaneously optimizes the position and direction of the GCC is presented. The proposed technique can be applied to obtain the optimal rendering quality of IBR for the linear case of camera positions and virtual views. Finally, experimental results are presented and compared with the theoretical results.
Changjian Zhu, Li Yu 0003, Zixiang Xiong
IEEE Trans. Multim.3
2017 Face hallucination using region-based deep convolutional networks
abstract
Most deep learning based face hallucinations exploit random patch prior from training samples, then to learn the mapping functions between low-resolution (LR) and high-resolution (HR) images, and achieve satisfactory reconstruction performance. However, most of them do not take into account the prior information on facial structure, which is pivotal for face hallucination. Different from random patch prior based deep learning approaches, in this paper, we utilize facial structural prior and develop a simple yet powerful face hallucination, named region-based deep convolutional networks (RDCN). Firstly, we divide facial image into several regions of interest, then to train multiple parallel subnetworks of these regions for exacting better structure priors, finally HR output is reconstructed by stitching facial parts. Experiments on the FEI database demonstrate that the proposed region-based convolution networks outperform other state-of-the-art, including recently proposed deep learning based approaches, both in subjective and objective reconstruction qualities.
Tao Lu 0001, Hao Wang 0237, Zixiang Xiong, Junjun Jiang, Yanduo Zhang, Huabing Zhou, Zhongyuan Wang 0001
ICIP3
2017 HEVC-based compression of high bit-depth 3D seismic data
abstract
In a previous work [1] we applied the idea of HEVC intra coding to compression of 32 b/p seismic images. Results are significantly better than those from a licensed commercial wavelet-based codec that is currently used at Shell for seismic image compression, which performs on par with JPEGXR. Building upon [1], this paper exploits HEVC-based inter predictive coding for 3D 32 b/p seismic data. We propose a new model for the Lagrange multiplier in R-D optimization to accommodate 32 b/p bit-depth and extended quantization parameter range. We also focus on reducing the complexity of motion estimation to meet application needs. Experiments with our new codec show a 95% complexity reduction of encoding time at 10:1 compression ratio with only 1 dB loss on average PSNR. The compression performance and subjective quality of our codec received high evaluation marks from Shell geologists.
Milos Radosavljevic, Zixiang Xiong, Ligang Lu, Detlef Hohl, Dejan Vukobratovic
ICIP2
2017 DLML: Deep linear mappings learning for face super-resolution with nonlocal-patch
abstract
Learning-based face super-resolution approaches rely on representative dictionary as self-similarity prior from training samples to estimate the relationship between the low-resolution (LR) and high-resolution (HR) image patches. The most popular approaches, learn mapping function directly from LR patches to HR ones but neglects the multi-layered nature of image degradation process (resolution down-sampling) which means observed LR images are gradually formed from HR version to lower resolution ones. In this paper, we present a novel deep linear mappings learning framework for face super-resolution to learn the complex relationship between LR features and HR ones by alternately updating multi-layered embedding dictionaries and linear mapping matrices instead of directly mapping. Furthermore, in contrast to existing position based studies that only use local patch for self-similarity prior, we develop a feature-induced nonlocal dictionary pair embedding method to support hierarchical multiple linear mappings learning. With coarse-to-fine nature of deep learning architecture, cascaded incremental linear mappings matrices can be used to exploit the complex relationship between LR and HR images. Experimental results demonstrate that such framework outperforms state-of-the-art (including both general super-resolution approaches and face super-resolution approaches) on FEI face database.
Tao Lu 0001, Lanlan Pan, Junjun Jiang, Yanduo Zhang, Zixiang Xiong
ICME5
2017 Face hallucination using deep collaborative representation for local and non-local patches
abstract
Patch-based face hallucination algorithms utilize either local patches (e.g., position-patch approaches) or nonlocal patches (e.g., dictionary-learning approaches) to exploit self-similarity prior from training samples. Although they yield decent results, solo source patches limit their performance due to not fully taking self-similarity prior from both local and nonlocal ones. In order to overcome this shortcoming, we propose a novel and efficient deep collaborative representation (DCR) based approach, to exploit both local and nonlocal self-similarity patches, for boosting face hallucination performance. First we learn a feature-inducing dictionary pair to represent local and nonlocal self-similarity prior, then deep (multiple-layer) representation weights and corresponding support dictionaries are iteratively updated to exploit accurate prior from coarse to fine. Finally, the high resolution (HR) output are optimized layer by layer. Experimental results outperform some state-of-the-art (e.g. Convolutional Neural Network based deep learning approach) which verify the validity of the proposed approach.
Tao Lu 0001, Lanlan Pan, Hao Wang 0237, Yanduo Zhang, Zixiang Xiong
ISCAS6
2017 Low-rank constrained collaborative representation for robust face recognition
abstract
Recently, sparse representation based classifiers (SRC) and collaborative representation based classifiers (CRC) have been shown to give very good performance under controlled scenarios. However, in practical applications, face recognition often encounters variations in illumination, expression, noise and occlusion, which cause severe performance degradation (due to the outliers in testing). In this paper, we present a novel robust face recognition algorithm based on class-wise low-rank constrained collaborative representations. We impose a low-rank constraint on the representation coefficient matrix to discriminate against outliers. The resulting low-rank constrained collaborative representation based classifier (LCRC) jointly minimizes the class-wise reconstruction error and rank of coefficient matrix. Experiments show that LCRC outperforms popular classifiers such as SRC, CRC, SVM, PROCRC on the AR, CMU PIE and LFW databases.
Tao Lu 0001, Yingjie Guan, Deng Chen, Zixiang Xiong
MMSP4
2017 Fault-tolerance based block-level bit allocation and adaptive RDO for depth video coding
abstract
Depth videos affect the visual quality of virtual view greatly, while conventional encoders are not adapted to encode depth videos. To solve this problem, we propose, in the paper, a fault-tolerance based joint block-level bit allocation scheme for depth video coding. The scheme classifies depth blocks into two classes by the fault-tolerance, and constructs virtual view perceptual quality based rate-distortion (R-D) model for each class. Based on the model, a joint block-level bit allocation scheme is proposed to obtain the optimized quantization parameter (QP) to encode each class of blocks. Further, an adaptive RDO algorithm is designed to determine λMODE, in order to achieve better visual quality of virtual views. Experimental results demonstrate that, compared with 3D-HEVC, the proposed method improves the visual quality of virtual views, especially at preserving details and boundaries of objects.
Li Yu 0003, Sen Xiang, Zixiang Xiong
MMSP4
2017 Accurate Depth Extraction Method for Multiple Light-Coding-Based Depth Cameras
abstract
Using multiple depth cameras can extend the field of view, which is beneficial to three-dimensional (3D) reconstruction. However, multiple active depth cameras would interfere with each other, which degrades the quality of depth images significantly. In this paper, we first present a depth extraction method based on sampling of the space at camera viewpoints to address the interference problem. Compared with the previous plane-sweeping method based on structured-light stereo and multiview stereo, depth images with less depth errors and better shape of objects are generated. Then, we analyze factors affecting depth accuracy for light-coding-based depth cameras and demonstrate the proposed depth extraction method also improves the accuracy of depth images by utilizing multiple projected patterns. Based on the analysis, subpixel-based matching (SPM) is finally applied to further improve the depth accuracy obtained by the depth extraction method. In the image matching process of our method, nearest neighbor interpolation is utilized to each projector's image plane to increase their resolution, which leads to the improved depth accuracy at subpixel level and avoids the drawbacks of changing parameters of light-coding-based depth cameras. Experimental results under simulated and real-world examples validate our analysis and show that the proposed method not only mitigates the impacts of interference effectively but also further improves the accuracy of depth images by using SPM.
Yu Pan 0002, Rongke Liu, Boshen Guan, Qiuchen Du, Zixiang Xiong
IEEE Trans. Multim.5
2016 Face hallucination via locality-constrained low-rank representation
abstract
Face hallucination (FH) based on sparse representation (SR) and locality-constrained representation (LCR) gives reasonably good performance. However, neither SR-nor LCR-based methods make full use of the structure information in the training data. On the other hand, low-rank representation (LRR) has been utilized to cluster samples into their respective classes by exploiting low-rank structures of the data. In this paper, we propose a locality-constrained low-rank representation (LCLRR) method to take advantage of both LCR and LRR for FH. LCLRR first enforces a low-rank constraint on choosing the dictionary atoms that belong to a subspace that correspond to the same cluster, it then imposes a locality constraint on selecting atoms that are in the vicinity of test samples. Experiments show that LCLRR outperforms both SR- and LCR-based methods on subjectively and objectively, proving that exploiting the structure information in the training data is feasible in face hallucination.
Tao Lu 0001, Zixiang Xiong, Yongjing Wan
ICASSP2
2016 Full-duplex machine-to-machine communication for wireless-powered Internet-of-Things
abstract
This paper considers machine-to-machine (M2M) communication for wireless-powered Internet-of-Things (IoT) based networking systems. Motivated by the observation that transmitting signals generally requires more energy than receiving signals for most IoT-based systems, we study a special wireless-powered M2M communication system in which the receiver can send its surplus energy to the transmitter. We propose a framework of wireless powered full-duplex M2M communication (WP-FD-M2M) in which the energy transfer from the receiver to the transmitter and the data transmission from the transmitter to the receiver take place at the same time over the same frequency. We establish a stochastic game-based model, referred to as the M2M game, to characterize the interaction between autonomous M2M transmitter and receiver. We prove that, if the transmitter and receiver can sequentially optimize their data transmission and energy transfer based on the Markov strategy, it is possible to achieve the maximum long-term performance for M2M communication without a centralized controller or coordination between the transmitter and receiver. Numerical results show that our proposed approach can significantly improve the performance for M2M communication under various situations.
Yong Xiao 0001, Zixiang Xiong, Dusit Niyato, Zhu Han 0001, Luiz A. DaSilva
ICC2
2016 Very Low-Resolution Face Recognition via Semi-Coupled Locality-Constrained Representation
abstract
Recognition tasks in very low-resolution (VLR) images are more challenging than those in high-resolution (HR) due to lack of adequate discriminative information. Previous VLR and HR coupled learning scheme limits both the representation and discriminative ability of features. In this work, we propose a semi-coupled locality-constrained representation (SLR) approach to learn the discriminative representations and the mapping relationship between VLR and HR features simultaneously. Both VLR and HR local manifold geometries are coded during representation, while the learned mapping function improves the manifold consistency by transforming VLR features to HR ones. Finally, the resolutionrobust features are fed into a sparse representation based classifier (SRC) to predict the face labels. The proposed algorithm gives better performance than many state-of-the-art VLR recognition algorithms.
Tao Lu 0001, Yanduo Zhang, Zixiang Xiong
ICPADS5
2016 Adaptive boosting for image denoising: Beyond low-rank representation and sparse coding
abstract
In the past decade, much progress has been made in image denoising due to the use of low-rank representation and sparse coding. In the meanwhile, state-of-the-art algorithms also rely on an iteration step to boost the denoising performance. However, the boosting step is fixed or non-adaptive. In this work, we perform rank-1 based fixed-point analysis, then, guided by our analysis, we develop the first adaptive boosting (AB) algorithm, whose convergence is guaranteed. Preliminary results on the same image dataset show that AB uniformly outperforms existing denoising algorithms on every image and at each noise level, with more gains at higher noise levels.
Tao Lu 0001, Zixiang Xiong
ICPR3
2016 Efficient low-rank supported extreme learning machine for robust face recognition
abstract
Recently, deep learning based face recognition algorithms have achieved great success in recognition performance. However, designing and training complex learning models suffer from time and labor efficiency. In this paper, we propose a novel three-layer low-rank supported extreme learning machine (LSELM) algorithm to take advantage of both robust feature representation and fast classification for efficient recognition. Every given probe sample is first clustered into a sub-class spanned by linear representation. With this sub-class, low-rank and robust features that are insensitive to disguise, noise, variant expression or illumination are recovered. These discriminative features are then coded to support a forward neural network for efficient prediction. Experimental results show that LSELM is on par with other deep learning based face recognition algorithms in recognition performance but has less time complexity on both AR and extend Yale-B datasets.
Yingjie Guan, Tao Lu 0001, Yanduo Zhang, Zixiang Xiong
VCIP6
2016 High bit-depth image compression with application to seismic data
abstract
Inspired by the high performance of High Efficiency Video Coding (HEVC), this paper reports our work on applying the ideas of HEVC intra coding to compression of high-depth images such as 32 bits per pixel (b/p) seismic data. Compared to a licensed commercial wavelet-based codec that is currently used for seismic image compression, which performs on par with JPEG-XR, our new image codec significantly improves the PSNR vs. compression ratio performance. The codec's subject performance is rated by geologist as highly satisfactory.
Milos Radosavljevic, Zixiang Xiong, Ligang Lu, Dejan Vukobratovic
VCIP2
2016 Performance Advantage of Joint Source-Channel Decoder over Iterative Receiver under M-ary Differential Chaotic Shift Keying Systems
abstract
In order to improve the error performance of M-ary differential chaotic shift keying (DCSK) systems over multi-path fading channel, two types of iterative structures have been employed in receiver design. One is joint source-channel decoder (JSCD), the other is iterative receiver (IR). Although several works have separately addressed the advantages of IR and JSCD, it is not clear which design provides more iterative gain for M-ary DCSK systems over multi-path fading channels. In this work, we employ the extrinsic information transfer (EXIT) chart technique to analyze the iteration behavior of JSCD and IR. Simulation results suggest that with enough source redundancy, JSCD outperforms IR under M-ary DCSK systems over multi-paths fading channels.
Yibo Lyu, Lin Wang 0003, Zixiang Xiong
VTC Spring3
2016 Energy-Saving Predictive Resource Planning and Allocation
abstract
Predictive resource allocation is an emerging approach to improve the performance of mobile systems as human behavior is reported predictable by leveraging big data analytics. Yet what information can be predicted by big data, what information need to be predicted for wireless access optimization, how to translate the information, and how to exploit the synthetic knowledge for allocating radio resources are not well understood and largely explored. In this paper, we are concerned with the latter two issues. In particular, we devise an energy-saving resource planning and allocation policy for multiple base stations (BSs) to serve mobile users with non-real-time (NRT) traffic by exploiting the user, network, and application levels of context information, where RT traffic may occupy partial resources of each BS. Inspired by the solution from an energy minimization problem with future instantaneous information, a low complexity multi-timescale predictive policy is proposed. Upon the arrival of each NRT user request, the resource planning is made with the user and network level context information, defined as the average channel gains of the NRT users and the statistics of residual bandwidth after serving RT traffic, with which the scheduling, power allocation, and BS sleeping can be accomplished after instantaneous channel information and residual network resource are available at each BS in each time slot. Simulation results show that the proposed policy can dramatically reduce the energy consumed by the BSs for serving the NRT traffic.
Chuting Yao, Chenyang Yang 0001, Zixiang Xiong
IEEE Trans. Commun.3
2016 Knowledge-Based Coding of Objects for Multisource Surveillance Video Data
abstract
Global object redundancy (GOR), as opposed to local spatial/temporal redundancies in a single video clip, is a new form of redundancy common in multisource surveillance video data (MSVD). GOR is induced by the repetition of foreground objects across multiple cameras, and becomes influential as the number of objects increases. Eliminating GOR considerably improves MSVD coding efficiency. In an effort to accomplish this, this study first proposes a knowledge-based representation of objects based on careful analysis of GOR composition. The representation contains a constant part and a variational part: the former is used to represent the common knowledge shared by an object across multiple cameras, while the latter is used to represent local variations on the object's surfaces. Based on the proposed representation, a knowledge-based coding (KBC) method is then proposed in which each foreground object is encoded with a hybrid prediction scheme, where the constant part of the object is generated via global prediction from a model library and the variational part is predicted via local reference frames with pose-based, short-term prediction. Experimental results showed that the KBC method saves more than 39% bits on average for encoding foreground objects in high-resolution video clips (compared to 16% for the entire videos). Applying the proposed coding method to surveillance videos in large spatial and temporal scale allows storage savings at the PB level.
Jing Xiao 0004, Ruimin Hu, Yu Chen 0021, Zhongyuan Wang 0001, Zixiang Xiong
IEEE Trans. Multim.6
2015 Large-area depth recovery for RGB-D camera
abstract
In this paper, a large-area depth recovery method for RGB-D camera is proposed. Considering that pixels along edges between different regions usually share similar depth values, we first select reliable pixels along edges of large-area depth missing regions and project them into the world coordinate system. Then, by examining the distribution of these pixels, we apply a weighted least squares method to approximate the surface function. With the help of the surface function, missing depth values can be recovered correctly. To the best of our knowledge, this is the first recovery method focusing on large areas of missing depth information. Qualitative evaluation demonstrates the effectiveness of the proposed method.
Zengqiang Yan, Li Yu 0003, Zixiang Xiong
ICIP3
2015 Texture-free large-area depth recovery for planar surfaces
abstract
This paper presents a texture-free depth enhancement method for large-area depth recovery. The proposed algorithm identifies a large-area depth missing region, and iteratively segments its contour by setting different initial pixels in each iteration. Coordinate transformation is used to analyze the distribution of each contour segment. By examining distributions of all contour segments, statistical histogram analysis is applied in our approach to select contour pixels. Then, selected pixels are projected into the world coordinate system, and multiple linear regression is utilized for surface function approximation. Missing depth values of a large-area depth missing region can be recovered with guidance of the approximated surface function. Quantitative and qualitative evaluations over state-of-the-art depth enhancement methods demonstrate the effectiveness and superiority of our method. Being texture-free, the proposed method has the flexibility of being merged into traditional depth enhancement methods.
Zengqiang Yan, Li Yu 0003, Zixiang Xiong
MMSP3
2015 Energy-saving resource allocation by exploiting the context information
abstract
Improving energy efficiency of wireless systems by exploiting the context information has received attention recently as the smart phone market keeps expanding. In this paper, we devise energy-saving resource allocation policy for multiple base stations serving non-real-time traffic by exploiting three levels of context information, where the background traffic is assumed to occupy partial resources. Based on the solution from a total energy minimization problem with perfect future information, a context-aware BS sleeping, scheduling and power allocation policy is proposed by estimating the required future information with three levels of context information. Simulation results show that our policy provides significant gains over those without exploiting any context information. Moreover, it is seen that different levels of context information play different roles in saving energy and reducing outage in transmission.
Chuting Yao, Chenyang Yang 0001, Zixiang Xiong
PIMRC3
2014 Achieving the degrees of freedom of 2×2×2 interference network with arbitrary antenna configurations
abstract
This paper studies the degrees of freedom (DoF) region of the 2 × 2 × 2 interference network, which is comprised of two sources, two relays and two destinations, each with arbitrary number of antennas. We prove that with linear transceivers, the cut-set outer bound can be achieved without any symbol extensions, except for one specific system setup, which has one-DoF gap to the cut-set bound. We show that to achieve the outer-bound, the transceivers include interference avoidance, cancelation, neutralization and alignment, depending on the antenna configuration.
Chenyang Yang 0001, Zixiang Xiong
ICASSP3
2014 Nonlocal image denoising via collaborative spatial-domain LMMSE estimation
abstract
In recent years, the performance of image denoising has been boosted drastically by nonlocal algorithms and sparse coding techniques. In this paper, we also take a nonlocal approach to image denoising and formulate the problem as one of collaborative LMMSE estimation from grouped image patches. We show that our optimal LMMSE solution amounts to shrinking the singular values of the matrix representation of the grouped image patches. This interpretation of our solution allows us to relate our estimation-theoretic approach to other nonlocal algorithms and sparse coding techniques in the literature. In addition, we develop an iterative algorithm to find the best LMMSE estimate. Experimental results show that our proposed denoising algorithm achieves better PSNR and subjective performance than the state of the art.
Zixiang Xiong, Dong-Qing Zhang, Hong Heather Yu
ICIP2
2014 Low-Complexity Encoding of Quasi-Cyclic Codes Based on Galois Fourier Transform
abstract
This paper presents two novel low-complexity encoding algorithms for quasi-cyclic (QC) codes based on Galois Fourier transform. The key idea behind them is making use of the block diagonal structure of the transformed generator matrix. The first one, named encoding by Galois Fourier transform, is equivalent to the fast implementations of the traditional encoding by Galois Fourier transform. The second one, named encoding in the transform domain (ETD), requires much less computational complexity for encoding binary QC codes. It skips the first step of the first algorithm and applies post-processing to save a large number of Galois field multiplications. Its application to QC-LDPC codes is also studied in this paper. Particularly, the hardware cost of the ETD for RS-based LDPC codes can be greatly reduced by short linear-feedback shift registers.
Qin Huang 0002, Shanbao He, Zixiang Xiong, Zulin Wang
IEEE Trans. Commun.4
2014 On the Minimum Energy of Sending Correlated Sources Over the Gaussian MAC
abstract
In this paper, we investigate the minimum energy of transmitting correlated sources over the Gaussian multiple-access channel. Compared to other works on joint source-channel coding, we consider the fundamental problem of the minimum transmission energy, where the source and channel bandwidths are not naturally matched. Different models of correlated sources are studied. We first treat lossy transmission of Gaussian sources, including multiterminal sources and CEO sources. We then consider lossless transmission of correlated binary sources. In all cases, we lower bound the minimum energy using a cut-set argument that couples transmission energy and the distortions for the Gaussian cases (or source entropy for the discrete case). For the achievable schemes, separate source and channel coding and uncoded transmission are studied as benchmarks. In addition, we show that hybrid digital/analog transmission achieves the best known energy efficiency.
Nan Jiang 0019, Yang Yang 0003, Anders Høst-Madsen, Zixiang Xiong
IEEE Trans. Inf. Theory4
2014 Distributed Compression of Linear Functions: Partial Sum-Rate Tightness and Gap to Optimal Sum-Rate
abstract
We consider the problem of distributed compression of the difference Z = Y1- cY2of two jointly Gaussian sources Y1and Y2under an MSE distortion constraint D on Z. The rate region for this problem is unknown if the correlation coefficient ρ and the weighting factor c satisfy cρ > 0. Inspired by Ahlswede and Han's scheme for the problem of distributed compression of the modulo-2 sum of two binary sources, we first propose a hybrid random-structured coding scheme that is capable of saving the sum-rate over both the random quantize-and-bin (QB) coding scheme and Krithivasan and Pradhan's structured lattice coding scheme. The main idea is to use a random coding component in the first layer to adjust the source correlation so that the structured coding component in the second layer can be more efficient with the outputs from the first layer as decoder side information. We then provide a new sum-rate lower bound for the problem in hand by connecting it to the Gaussian two-terminal source coding problem with covariance matrix distortion constraint. Our lower bound not only improves existing bounds in many cases, but also allows us to prove sum-rate tightness of the QB scheme when c is either relatively small or large and D is larger than some threshold. Furthermore, our lower bound enables us to show that our new hybrid scheme performs within two b/s from the optimal sum-rate for all values of ρ, c, and D.
Yang Yang 0003, Zixiang Xiong
IEEE Trans. Inf. Theory2
2013 Face hallucination via weighted sparse representation
abstract
By incorporating the priors of image positions, position-patch based face hallucination methods can produce high-quality results and save computation time. These methods represent the test image patch as a linear combination of the same position patches in a training dictionary, and the key issue is how to obtain the optimal coefficients. Due to stability and accuracy issues, methods based on least square estimation or sparse representation (SR) proposed so far are not satisfactory. In this paper, we improve existing SR methods by exploiting similarity between the test and training patches. In particular, we impose a similarity constraint (in terms of the distance between the test patch and bases in the dictionary) on the ℓ1minimization regularization term and obtain the coefficients by solving a weighted SR problem. We also provide a new prospective on weighted SR and investigate its robustness to illumination variations. Experiments on commonly used database demonstrate that our method outperforms state of the art.
Zhongyuan Wang 0001, Junjun Jiang, Zixiang Xiong, Ruimin Hu
ICASSP3
2013 Feasibility of interference neutralization in relay-aided MIMO interference broadcast channel with partial connectivity
abstract
In this paper, we study the feasibility of interference neutralization (IN) in partially connected relay-aided multi-input-multi-output (MIMO) interference broadcast channel (IBC), where each relay only communicates with users in the same cell. Due to partial connectivity, there is no need to exchange channel information among the relays in different cells. We first present the necessary and sufficient condition for IN feasibility using linear transceiver. We then provide the minimum relay configuration required to support the maximal number of data streams without interference. Finally, the achievable degrees-of-freedomregion is obtained.
Chenyang Yang 0001, Zixiang Xiong
ICASSP3
2013 Joint Slepian-Wolf/Dirty-paper coding
abstract
We consider a joint Slepian-Wolf/Dirty-paper coding (SW-DPC) setup where a binary source needs to be transmitted over an additive white-Gaussian channel in the presence of interference known only at the encoder, as well as correlated side-information available only at the decoder. We propose a practical joint SW-DPC framework that uses a single low-density parity-check code coupled with trellis coded quantization to simultaneously provide error protection (in the face of noise and interference) and compression (in the face of correlated side-information available at the decoder). Simulation results indicate that the proposed joint SW-DPC scheme with finite-length codes outperforms a scheme with separate SW and DPC by 0.3 dB. In addition, we show that our joint SW-DPC code designed for one set of channel conditions is capable of achieving negligible bit-error rates for many other channel conditions as well. More specifically, it achieves successful decoding for a wide variety of channel conditions for which a separation-based scheme fails.
Momin Uppal, Khalid A. Qaraqe, Zixiang Xiong
ICC3
2013 Support-driven sparse coding for face hallucination
abstract
By incorporating the prior of positions, position patch based face hallucination methods can produce high-quality results and save computation time. Given a low-resolution face image, the key issue of these methods is how to encode the input low-resolution patch. However, due to stability and accuracy issues, the coding approaches proposed so far are not satisfactory. In this paper, we present a novel sparse coding method via exploiting the support information on the coding coefficients. In particular, the support information is characterized by the locality of the image patch manifold, which has been shown to be critical in data representation and analysis. According to the distances between the input patch and bases in the dictionary, we first assign different weights to the coding coefficients and then obtain the coding coefficients by solving a weighted sparse problem. Our proposed method exploits the non-linear manifold structure of patch samples and the sparse property of the redundant data, leading to stable and accurate representation. Experiments on commonly used databases demonstrate that our method outperforms state of the art.
Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zixiang Xiong, Zhen Han 0002
ISCAS4
2013 A gradient-based approach for interference cancelation in systems with multiple Kinect cameras
abstract
Microsoft Kinect cameras provide a fast and convenient way to acquire depth information. However, the interference problem of multiple Kinect cameras dramatically degrades the depth quality. In this paper, we study the interference problem and propose an interference cancelation approach based on the statistical properties of depth maps. The gradient values are investigated and propagated from the interference-free region to the interfered region. The gradient values are obtained based on the statistic gradient features of the depth maps and the optimal solution of depth values is derived with a least error criterion. Experiment results demonstrate that our proposed method can eliminate interference efficiently, and lead to better qualities of depth maps and rendered virtual views.
Sen Xiang, Li Yu 0003, Qiong Liu 0001, Zixiang Xiong
ISCAS4
2013 On the minimum energy of sending Gaussian multiterminal sources over the Gaussian MAC
abstract
We study the minimum energy of sending Gaussian multiterminal sources over the Gaussian multiple access channel (MAC). Distributed transmitters observe Gaussian multiterminal sources and describe their observations to a central decoder, which desires to reconstruct the sources under MSE constraints. We first lower bound the minimum energy by a cut-set argument which couples the transmitted signals and reconstruction errors. For achievability, separate source-channel coding is first studied as a benchmark. We then find out the minimum energy that can be achieved uncoded transmission. A hybrid digital/analog scheme is proposed to achieve the best known energy performance.
Nan Jiang 0019, Yang Yang 0003, Anders Høst-Madsen, Zixiang Xiong
ISIT4
2013 Low-complexity encoding of binary quasi-cyclic codes based on Galois Fourier transform
abstract
This paper presents a novel low-complexity encoding algorithm for binary quasi-cyclic (QC) codes based on matrix transformation. First, a message vector is encoded into a transformed codeword in the transform domain. Then, the transmitted codeword is obtained from the transformed codeword by the inverse Galois Fourier transform. Moreover, a simple and fast mapping is devised to post-process the transformed codeword such that the transmitted codeword is binary as well. The complexity of our proposed encoding algorithm is less than ek(n-k)log2e+ne(log22e+log2e)+ n/2 elog32e bit operations for binary codes. This complexity is much lower than its traditional complexity 2e2(n - k)k. In the examples of encoding the binary (4095, 2016) and (15500, 10850) QC codes, the complexities are 12.09% and 9.49% of those of traditional encoding, respectively.
Qin Huang 0002, Zulin Wang, Zixiang Xiong
ISIT4
2013 Diversity analysis for space-time-frequency (STF) coded MIMO system with a general correlation model
abstract
Previous works on space-time-frequency (STF) codes have focused on frequency-selective but independent fading channels for MIMO-OFDM systems. However, insufficient antenna space, mobile scenarios, and multipath in practice cause correlation in the spatial, temporal, and frequency domains. This paper studies the effect of STF coding on the performance of MIMO-OFDM systems with general spatial, temporal, and frequency/path correlated channels. Specifically, we first derive an upper bound on the maximum achievable diversity and then prove achievability by giving an STF code design example. Finally, we show that our general diversity result recovers those in the existing literature for special correlation structures.
Mang Liao, Youguang Zhang, Zixiang Xiong
WCNC3
2013 A Robust Multi-Level Design for Dirty-Paper Coding
abstract
We propose a robust close-to-capacity dirty-paper coding (DPC) design framework in which multi-level low density parity check (LDPC) codes and trellis coded quantization (TCQ) are employed as the channel and source coding components, respectively. The proposed design framework is robust in the sense that it yields close to capacity solutions in the high-, medium-, and low-rate regimes. This is in contrast to existing practical DPC schemes that perform well only in one or two of these regimes, but not all three. We design codes for transmission rates of 0.5, 1.0, 1.5, and 2.0 bits/sample (b/s) using one, two, three, and four LDPC levels; at a block length of 2×105, the codes perform 0.95, 0.58, 0.55, and 0.54 dB from the corresponding information theoretic limits, respectively. We also propose a low-complexity decoding scheme that does not involve iterative message passing between the source and channel decoders; the low-complexity scheme performs only 1.08, 0.85, and 0.79 dB away from the theoretical limits at transmission rates of 1.0, 1.5, and 2.0 b/s, respectively.
Momin Uppal, Guosen Yue, Yan Xin 0001, Xiaodong Wang 0001, Zixiang Xiong
IEEE Trans. Commun.5
2013 A New Sufficient Condition for Sum-Rate Tightness in Quadratic Gaussian Multiterminal Source Coding
abstract
This paper considers the quadratic Gaussian multiterminal (MT) source coding problem and provides a new sufficient condition for the Berger–Tung (BT) sum-rate bound to be tight. The converse proof utilizes a set of virtual remote sources given which the observed sources are block independent with a maximum block size of 2. The given MT source coding problem is then related to a set of two-terminal problems with matrix-distortion constraints, for which a new lower bound on the sum-rate is given. By formulating a convex optimization problem over all distortion matrices, a sufficient condition is derived for the optimal BT scheme to satisfy the subgradient-based Karush–Kuhn–Tucker condition. The subset of the quadratic Gaussian MT problem satisfying our new sufficient condition subsumes all previously known tight cases, and our proof technique opens a new direction for more general partial solutions.
Yang Yang 0003, Zixiang Xiong
IEEE Trans. Inf. Theory3
2012 3D scene reconstruction by multiple structured-light based commodity depth cameras
abstract
Commodity depth cameras have attracted a lot of research interest recently, in particular the structured-light based Kinect cameras available on the mass market. One important application of such cameras is 3D scene reconstruction and view synthesis. However, a single depth camera often has limited field of view and there is missing depth information when synthesizing a virtual view from a new viewpoint. In this paper, we study the problem of 3D scene reconstruction from multiple structured-light based depth cameras. Since multiple cameras may cause severe interference in the regions where the projected light overlaps, we present a novel planesweeping based algorithm to handle such interference. The proposed algorithm takes into account the correlation between multiple projectors and the infrared images as well as the correlation between the infrared images, thereby recovering the depth information for both overlapped and non-overlapped regions. Simulation results demonstrate that the proposed solution is very effective on various scenes.
Cha Zhang, Wenwu Zhu 0001, Zhengyou Zhang, Zixiang Xiong, Philip A. Chou
ICASSP5
2012 Achieving the degrees of freedom of relay-aided interference broadcast channels
abstract
This paper studies the degrees of freedom (DoF) of relay-aided interference broadcast channels and the relay resource to achieve them. We show that, with the aid of half-duplex relays, a G-cell system can achieve GMBS/2 DoF, where MBSis the number of transmit antennas in each cell. By studying the interference-free constraints, we obtain a lower bound of the relay resource that ensures interference-free transmission. Instead of directly solving the complicated multivariate problem with cubic interference-free constraints, we propose to relax the problem to linear equations by randomly initiating certain variables, and derive an achievable bound on relay resource. Numerical results show that the relay-aided interference broadcast channels require less relay resource than relay-aided interference channels and transmission protocol has large impact on the required resources.
Chenyang Yang 0001, Tingting Liu 0001, Zixiang Xiong
ICASSP4
2012 A density evolution based framework for dirty paper code design using TCQ and multilevel LDPC codes
abstract
We propose a density evolution based dirty-paper code design framework that combines trellis coded quantization with multi-level low-density parity-check (LDPC) codes. Unlike existing design techniques based on Gaussian approximation and EXIT charts, the proposed framework tracks the empirically collected log-likelihood ratio (LLR) distributions at each iteration, and employs density evolution and differential evolution algorithms to design each LDPC component code. In order for the approximated LLR distributions obtained using density evolution to better match the true LLR distributions, a novel decorrelator is added to the decoder to make channel LLRs and check node LLRs almost independent of each other. Simulation results show that at 1 bit per sample transmission rate, the dirty-paper codes designed using the proposed method operate within 0.53 dB of the theoretical limit, reducing the best known result (with a 0.58 dB gap to the limit) by 0.05 dB at the same complexity and block length, while this improvement can be as large as 0.21 dB (corresponding to a 0.37 dB gap to the limit) when we use higher complexity encoder/decoder with longer block length.
Yang Yang 0003, Zixiang Xiong, Yu-chun Wu, Philipp Zhang
ICC2
2012 Edge-preserving interpolation for down/up sampling-based depth compression
abstract
Preserving the edges in depth compression is important for improving the synthesized view quality, this paper presents a novel edge-preserving depth up-sampling method for down/up sampling-based depth coding using both the texture and depth information. We take into account the edge similarity between depth maps and their corresponding texture images as well as the structural similarity among depth maps to build a weight model. Based on the weight model, the optimal MMSE up-sampling coefficients are estimated from the local covariance coefficients of the down-sampled depth map. Experimental results show that our proposed interpolation method for down/up sampling-based depth coding improves both the coding efficiency and synthesized view quality.
Huiping Deng, Li Yu 0003, Zixiang Xiong
ICIP3
2012 Block-based variable density compressed image sampling
abstract
Compressed sampling (CS) is a technique that enables signal reconstruction at sub-Nyquist sampling rate. A key problem in CS is how to design the sampling scheme. In this paper, we propose a novel sampling method for compressed image sampling, which exploits a priori information and uses a block-based strategy to improve image reconstruction. Our block-based sampling scheme assigns more samples to blocks with more high-frequency contents while making sure that important coefficients of each block are sampled. Simulation results show that our proposed method outperforms existing methods on both reconstruction quality and running time.
Bin Liu 0016, Zixiang Xiong, Gonzalo R. Arce, Javier Garcia-Frías, Wenwu Zhu 0001, Zhisheng Yan
ICIP3
2012 Reliable versus unreliable transmission for energy efficient transmission in relay networks
abstract
A network code is said to be reliable when all transmissions in the network are (deterministic) functions of the source messages; well-known examples include decode-forward for relay networks. It is said to be unreliable when transmissions depend on the noise realization at nodes; examples include compress-forward and amplify-forward. The deterministic capacity of a network is defined as the supremum of the rates achievable by reliable codes. In this paper we derive the deterministic capacity of some relay networks in the low power regime. The resulting energy per bit is then compared with the one achievable by arbitrary transmission.
Anders Høst-Madsen, Nan Jiang 0019, Yang Yang 0003, Zixiang Xiong
ISIT4
2012 Rateless coded hybrid amplify/decode-forward cooperation for wireless multicast
abstract
We consider a cooperative wireless multicast channel with no channel state information available at the transmitters, and where all wireless links undergo independent identically distributed quasi-static Rayleigh fading. For cooperation, we propose a rateless coded approach and identify its benefits over a conventional fixed-rate two-hop strategy. In addition to providing robust communications under fading conditions, rateless coding is able to exploit the cooperation benefits to a greater degree. A hybrid rateless coded cooperation scheme is proposed which employs amplify-and-forward (AF) in conjunction with decode-and-forward (DF) strategy. Results based on Monte-Carlo simulations indicate that the rateless coded hybrid scheme is able to significantly outperform two-hop cooperation, as well as the individual AF and DF strategies.
Momin Uppal, Anders Høst-Madsen, Zixiang Xiong
ISIT4
2012 Sending two Gaussians over the Gaussian MAC with bandwidth expansion
Nan Jiang 0019, Yang Yang 0003, Anders Høst-Madsen, Zixiang Xiong
ISITA4
2012 Block-based compressed sampling with non-linear coding for image transmission
abstract
We propose a novel block-based image transmission system, which exploits the a prior information existing in the DCT domain of images and combines both linear and non-linear coding schemes accommodated to a block-based DCT domain compressed sampling method. An image is firstly divided into blocks and each block is separately sampled in DCT domain. Different coding schemes are used to transmit the samples based on their properties. With block-based strategy, each image block can be processed and transmitted separately, which reduces a lot of latency. Besides, an efficient system optimization algorithm is proposed by jointly optimizing the power allocation scheme and the transmission parameters to search for the maximum peak signal-to-noise ratio (PSNR) of the reconstructed image. Simulation results show that the proposed system provides a good performance with less latency.
Bin Liu 0016, Zixiang Xiong, Gonzalo R. Arce, Javier Garcia-Frías
MMSP3
2012 Recursive bilateral filter for encoder-integrated video denoising
abstract
Video denoising based on temporal or spatiotemporal filtering is highly effective but computationally expensive due to the requirement of motion estimation. Encoder-integrated denoising is an efficient framework that embeds the filtering process into the encoding pipeline so that motion estimation for denoising can be avoided. State-of-the-arts encoder-integrated methods use Least Minimum Mean Square Error (LMMSE) optimal filters to reduce noise based on additive noise models. However, in practice, the LMMSE filters could result in blurred edges and visual artifacts due to inaccurate image signal estimation caused by outlier input samples and non-adaptivity within macroblocks. This paper presents new recursive bilateral filters that can be easily integrated into video encoders to tackle the limitations of the LMMSE filters. The robustness of the bilateral filters results in substantially better objective and subjective quality with marginal computation cost increase.
Dong-Qing Zhang, Zixiang Xiong, Hong Heather Yu
VCIP3
2012 On Outage Capacity in the Low Power Regime
abstract
This paper derives a formula for the wideband slope for outage capacity and consider its application to some specific wireless channels. For the broadcast channel, the formula indicates that superposition is always superior to time-division multiple access (TDMA). On the other hand, for the interference channel, we show that TDMA is better than superposition for most realistic situations.
Anders Høst-Madsen, Momin Uppal, Zixiang Xiong
IEEE Trans. Inf. Theory3
2012 The Sum-Rate Bound for a New Class of Quadratic Gaussian Multiterminal Source Coding Problems
abstract
In this paper, we show tightness of the Berger-Tung (BT) sum-rate bound for a new class of quadratic Gaussian multiterminal (MT) source coding problems dubbed bi-eigen equal-variance with equal distortion (BEEV-ED), where theL×Lsource covariance matrix has equal diagonal elements with two distinct eigenvalues, and theLtarget distortions are equal. LetK(KL) be the number of repetitions of the larger eigenvalue, the BEEV covariance structure allows us to connectKi.i.d. virtual Gaussian sources with theLgiven MT sources via anL×Ksemiorthogonal transform whose rows have equal Euclidean norm plus additive i.i.d. Gaussian noises, resulting in the two sets of sources being mutually conditional i.i.d. By relating the given MT source coding problem to a generalized Gaussian CEO problem with theKvirtual sources as remote sources and theLMT sources as observations, we obtain a lower bound on the MT sum-rate, and show its achievability by BT schemes under the equal distortion constraints. Our BEEV-ED class of quadratic Gaussian MT source coding problems subsumes both the positive-symmetric case considered by Wagner et al. and the negative-symmetric case. Other examples, including a subclass of sources with BE circulant symmetric covariance matrices and equal distortion constraints, are also provided to highlight tightness of the sum-rate bound.
Yang Yang 0003, Zixiang Xiong
IEEE Trans. Inf. Theory2
2012 On the Generalized Gaussian CEO Problem
abstract
This paper considers a distributed source coding (DSC) problem where L encoders observe noisy linear combinations of K correlated remote Gaussian sources, and separately transmit the compressed observations to the decoder to reconstruct the remote sources subject to a sum-distortion constraint. This DSC problem is referred to as the generalized Gaussian CEO problem since it can be viewed as a generalization of the quadratic Gaussian CEO problem where the number of remote source K=1. First, we provide a new outer region obtained using the entropy power inequality and an equivalent argument (in the sense of having the same rate-distortion region and Berger-Tung inner region) among a certain class of generalized Gaussian CEO problems. We then give two sufficient conditions for our new outer region to match the inner region achieved by Berger-Tung schemes, where the second matching condition implies that in the low-distortion regime, the Berger-Tung inner rate region is always tight, while in the high-distortion regime, the same region is tight if a certain condition holds. The sum-rate part of the outer region is also studied and shown to meet the Berger-Tung sum-rate upper bound under a certain condition, which is obtained using the Karush-Kuhn-Tucker conditions of the underlying convex semidefinite optimization problem, and is in general weaker than the aforesaid two for rate region tightness.
Yang Yang 0003, Zixiang Xiong
IEEE Trans. Inf. Theory2
2011 A Multi-Level Design for Dirty-Paper Coding with Applications to the Cognitive Radio Channel
abstract
We propose a close-to-capacity dirty-paper coding framework which employs multi-level low density parity-check (LDPC) and trellis coded quantization. The proposed coding framework is robust in the sense that it performs close to capacity in the high as well as the low rate regimes. This is in contrast to existing practical DPC schemes which perform well at one of these regimes, but never both. In order to evaluate the performance of our scheme, we consider its application to a cognitive radio channel. At a block length of 2 × 105, the designed dirty-paper coding scheme operates within 0.95, 0.58 and 0.6 dB of the theoretical limit at transmission rates of 0.5, 1.0 and 1.5 bits/sample, respectively. As far as the authors are aware, this is the best performance reported in the literature so far.
Momin Uppal, Guosen Yue, Yan Xin 0001, Xiaodong Wang 0001, Zixiang Xiong
GLOBECOM5
2011 Distributed compression of linear functions: Partial sum-rate tightness and gap to optimal sum-rate
abstract
We consider the problem of distributed compression of the difference Z = Y1-cY2of two jointly Gaussian sources Y1and Y2(with positive correlation coefficient ρ and positive c) under an MSE distortion constraint D on Z. The rate region for this problem is unknown. We provide a new lower bound on the minimum sum-rate by utilizing the connection of the above problem with the two-terminal source coding problem with matrix-distortion constraint. Our lower bound not only improves existing bounds in many cases, but also allows us to prove sum-rate tightness of the Berger-Tung scheme when c is either relatively small or large and D is larger than some threshold. Furthermore, our lower bound enables us to show that the improved lattice-based scheme recently introduced in [1] (with the smallest achievable sum-rate) performs within 1.18 b/s from the optimal sum-rate for all values of ρ, c, and D.
Yang Yang 0003, Zixiang Xiong
ISIT2
2011 Depth camera assisted multiterminal video coding
abstract
This paper addresses multiterminal video coding with the help of a low-resolution depth camera. In this setup, the depth sequence, together with the high-resolution texture sequences, are collected, compressed and transmitted separately to the joint decoder in order to obtain more accurate depth information and consequently better rate-distortion performance. At the decoder end, side information (for Wyner-Ziv coding) is generated based on successive refinement of the decompressed low-resolution depth map and texture frame warped from other terminals. Experimental results show sum-rate savings with the depth camera than without for the same PSNR performance. Comparisons to simulcast and JMVM coding are also provided. Although the sum-rate gain of our multiterminal video coding scheme (with or without the depth camera) over simulcast is relatively small, this work is the first that incorporates depth camera in a multiterminal setting.
Yang Yang 0003, Zixiang Xiong
MMSP3
2011 Witsenhausen-Wyner Video Coding
abstract
Inspired by Witsenhausen and Wyner's 1980 (now expired) patent on “interframe coder for video signals,” this paper presents a Witsenhausen-Wyner video codec, where the motion-compensated previously decoded video frame is used at the decoder as side information for joint decoding. Specifically, we replace predictive Inter coding in H.264/AVC by the syndrome-based coding scheme of Witsenhausen and Wyner, while keeping the Intra and Skip modes of H.264/AVC unchanged. We employ forward motion estimation at the encoder and send the motion vectors to help generate side information at the decoder, since our focus is not on low-complexity encoding. We also examine the tradeoff between the motion vector resolution and coding efficiency. Within the Witsenhausen-Wyner coding mode, we optimize the decision between syndrome coding and entropy coding among different discrete cosine transform (DCT) bands and among different bit-planes within each DCT coefficient. Extensive simulations of video transmission over wireless networks show that Witsenhausen-Wyner video coding is more robust against channel errors than H.264/AVC. The price paid for enhanced error-resilience with Witsenhausen-Wyner coding is a small loss in compression efficiency.
Mei Guo, Zixiang Xiong, Feng Wu 0001, Debin Zhao, Xiangyang Ji, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2011 On the Sum-Rate Loss of Quadratic Gaussian Multiterminal Source Coding
abstract
This work studies the sum-rate loss of quadratic Gaussian multiterminal source coding, i.e., the difference between the minimum sum-rates of distributed encoding and joint encoding (both with joint decoding) of correlated Gaussian sources subject to MSE distortion constraints on individual sources. It is shown that under the nondegraded assumption, i.e., all target distortions are simultaneously achievable by a Gaussian Berger-Tung scheme, the supremum of the sum-rate loss of distributed encoding over joint encoding ofLjointly Gaussian sources increases almost linearly in the number of sourcesL, with an asymptotic slope of 0.1083 bit per sample per source asLgoes to infinity. This result is obtained even though we currently do not have the full knowledge of the minimum sum-rate for the distributed encoding case. The main idea is to upper-bound the minimum sum-rate of multiterminal source coding by that achieved by parallel Gaussian test channels while lower-bounding the minimum sum-rate of joint encoding by a reverse water-filling solution to a relaxed joint encoding problem of the same set of Gaussian sources with a sum-distortion constraint (that equals the sum of the individual target distortions). We show that under the nondegraded assumption, the supremum difference between the upper bound for distributed encoding and the lower bound for joint encoding is achieved in the bi-eigen equal-variance with equal distortion case, in which both bounds are known to be tight.
Yang Yang 0003, Zixiang Xiong
IEEE Trans. Inf. Theory3
2010 A Dirty-Paper Coding Scheme for the Cognitive Radio Channel
abstract
We implement a dirty-paper coded framework for the cognitive radio channel. We assume that the cognitive user has non-causal knowledge about the primary user's transmissions. Thus the secondary receiver can employ dirty-paper coding to counter the effect of any interference from the primary user. In addition, we consider a situation where the introduction of the cognitive user should not affect the performance of the primary system -- nor should the primary system have to change its encoding/decoding process. For the primary user we use a low-density parity-check code and a 4-ary pulse amplitude modulation format. For the cognitive user, we propose a dirty-paper coding scheme which employs trellis-coded quantization as the source code and an irregular repeat-accumulate code as the channel code. At a transmission rate of 1.0 bits/sample, the designed dirty-paper coding scheme operates within 1.23 dB of the theoretical limit.
Momin Uppal, Guosen Yue, Yan Xin 0001, Xiaodong Wang 0001, Zixiang Xiong
ICC5
2010 Outage capacity of the broadcast channel in the low power regime
abstract
We consider outage capacity in the broadcast channel when the base-station has none or little channel knowledge. We find the minimum energy per bit needed to achieve a certain outage probability and the corresponding wideband slope.
Momin Uppal, Anders Høst-Madsen, Zixiang Xiong
ISIT3
2010 A rateless coded protocol for half-duplex wireless relay channels
abstract
We propose a rateless coded protocol for a half-duplex wireless relay channel where all links experience independent quasi-static Rayleigh fading. The protocol utilizes a combination of rateless coded decode-forward and compress-forward relaying schemes. Assuming very limited feedback from the destination, we derive the theoretical performance limits specifically with BPSK modulation. We then implement the rateless coded relaying protocol using carefully designed Raptor codes.
Momin Uppal, Guosen Yue, Xiaodong Wang 0001, Zixiang Xiong
ISIT4
2010 The generalized quadratic Gaussian CEO problem: New cases with tight rate region and applications
abstract
We study the generalized quadratic Gaussian CEO problem where K jointly Gaussian remote sources are transformed and independently-Gaussian-corrupted to form L observations, which are separately compressed by L encoders while the decoder attempts to reconstruct the K remote sources subject to a sum-distortion constraint. We first give a sufficient condition for the existing inner and outer regions to coincide, and then show that the new condition contains more matching cases than Oohama's. The main novelty lies in the fact that rotating the remote sources via orthogonal transformation while keeping a constant observation covariance matrix and the same observation noises does not change the rate region. Further, by exploiting the relationship between the sum-rate bound of the generalized quadratic Gaussian CEO problem and that of the quadratic Gaussian multiterminal source coding problem, we obtain new cases of the latter with tight sum-rate bound.
Yang Yang 0003, Zixiang Xiong
ISIT3
2010 On the sum-rate loss of quadratic Gaussian multiterminal source coding
abstract
This work studies the sum-rate loss of quadratic Gaussian multiterminal source coding, i.e., the difference between the minimum sum-rates of distributed encoding and joint encoding (both with joint decoding) of correlated Gaussian sources subject to MSE distortion constraints on individual sources. It is shown that under the non-degraded assumption, i.e., all target distortions are simultaneously achievable by a Berger-Tung scheme, the supremum of the sum-rate loss of distributed encoding over joint encoding of L jointly Gaussian sources increases almost linearly in the number of sources L, with an asymptotic slope of 0.1083 b/s per source as L goes to infinity. This result is obtained even though we currently do not have the full knowledge of the minimum sum-rate for the distributed encoding case. The main idea is to upper-bound the minimum sum-rate of multiterminal source coding by that achieved by parallel Gaussian test channels while lower-bounding the minimum sum-rate of joint encoding by a reverse water-filling solution to a relaxed joint encoding problem of the same set of Gaussian sources with a sum-distortion constraint (that equals the sum of the individual target distortions). We show that under the non-degraded assumption, the supremum difference between the upper bound for distributed encoding and the lower bound for joint encoding is achieved in the bi-eigen equal-variance with equal distortion case, in which both bounds are known to be tight.
Yang Yang 0003, Zixiang Xiong
ISIT3
2010 Automated Digital Dental Articulation
James J. Xia, Yu-Bing Chang, Jaime Gateno, Zixiang Xiong, Xiaobo Zhou 0001
MICCAI (3)4
2010 An Automatic and Robust Algorithm of Reestablishment of Digital Dental Occlusion
abstract
In the field of craniomaxillofacial (CMF) surgery, surgical planning can be performed on composite 3-D models that are generated by merging a computerized tomography scan with digital dental models. Digital dental models can be generated by scanning the surfaces of plaster dental models or dental impressions with a high-resolution laser scanner. During the planning process, one of the essential steps is to reestablish the dental occlusion. Unfortunately, this task is time-consuming and often inaccurate. This paper presents a new approach to automatically and efficiently reestablish dental occlusion. It includes two steps. The first step is to initially position the models based on dental curves and a point matching technique. The second step is to reposition the models to the final desired occlusion based on iterative surface-based minimum distance mapping with collision constraints. With linearization of rotation matrix, the alignment is modeled by solving quadratic programming. The simulation was completed on 12 sets of digital dental models. Two sets of dental models were partially edentulous, and another two sets have first premolar extractions for orthodontic treatment. Two validation methods were applied to the articulated models. The results show that using our method, the dental models can be successfully articulated with a small degree of deviations from the occlusion achieved with the gold-standard method.
Yu-Bing Chang, James J. Xia, Jaime Gateno, Zixiang Xiong, Xiaobo Zhou 0001, Stephen T. C. Wong
IEEE Trans. Medical Imaging4
2009 Cooperation in the MAC channel using frequency division multiplexing
abstract
This paper considers cooperation in the low power/SNR regime for the multiple access channel. We assume that transmitters have no channel state information, and consider the outage capacity. We further assume that the nodes operate in half duplex. The nodes are multiplexed by assigning to each node a unique frequency band via frequency division multiplexing (FDM). A node transmits only on this band while listening on the others. We perform outage wideband analysis on FDM based cooperation, and show perhaps surprisingly that FDM loses very little in performance compared to the idealized situation where nodes can operate in full duplex. We then develop practical rateless coding methods for FDM using Raptor codes.
Anders Høst-Madsen, Momin Uppal, Zixiang Xiong
ISIT3
2009 Distributed source coding without Slepian-Wolf compression
abstract
Slepian-Wolf (SW) coding, which is concerned with separate near-lossless compression of correlated sources (with joint decoding), forms the basis of distributed source coding (DSC) and can be used to exploit the correlation among quantized sources in lossy DSC problems such as Wyner-Ziv (WZ) coding and multiterminal (MT) source coding. However, SW coding is in general lossy, especially at short block length, and practical implementation is not nearly as well understood as entropy coding. This paper studies distributed source coding without SW coding. We employ entropy coding (after quantization if necessary) at each encoder while relying on joint estimation at the decoder to exploit the source correlation. We start from the simple lossless case before giving single-letter characterizations of the rate-distortion function for WZ coding without SW compression, and achievable rate region for MT source coding without SW compression. Examples on the binary symmetric and quadratic Gaussian cases are given.
Yang Yang 0003, Zixiang Xiong
ISIT2
2009 Code design for quadratic Gaussian multiterminal source coding: The symmetric case
abstract
Whereas the theory and practice of two-terminal quadratic Gaussian multiterminal (MT) source coding is complete, the theory with more than two terminals is only partial, with the sum-rate limit only known in the symmetric case where all sources are positively symmetric and all target distortions equal. This paper proposes the first code design for quadratic Gaussian MT source coding in this symmetric setup. The aim is to approach corner points of the rate region via TCQ for quantization and LDPC codes for Slepian-Wolf compression. We provide high-rate analysis of our code design. Simulations with three and four terminals show a very small sum-rate loss.
Zixiang Xiong, Yang Yang 0003
ISIT2
2009 Three-terminal video coding
abstract
Following recent works on the sum-rate of quadratic Gaussian multi-terminal coding of symmetric sources, this paper presents a practical three-terminal video coding scheme that saves the sum-rate over independent coding of three correlated video sequences. The first video sequence is coded by H.264/AVC, while the other two sequences are compressed sequentially using Wyner-Ziv coding with side information generated from the already decoded sequences. To improve the performance of Wyner-Ziv coding, a depth-map-based view interpolation approach and a novel soft-decision side information generation method are proposed. Experimental results show that our scheme can achieve a lower sum-rate than separate H.264/AVC coding. Comparison with JMVM joint encoding is also provided.
Yang Yang 0003, Zixiang Xiong
MMSP3
2009 Compress-spread-forward with multiterminal source coding and complete complementary sequences
abstract
We propose a new technique, compress-spread forward (CSF), for high-performance wireless streaming from two base stations in parallel. CSF uses multiterminal source coding for efficient source compression and complete complementary sequences for error-free multiple access and synchronization. Our practical design shows significant performance gains due to spatial diversity and distributed source coding.
Chadi Khirallah, Vladimir Stankovic 0001, Lina Stankovic, Yang Yang 0003, Zixiang Xiong
IEEE Trans. Commun.5
2009 Code design for MIMO broadcast channels
abstract
Recent information-theoretic results show the optimality of dirty-paper coding (DPC) in achieving the full capacity region of the Gaussian multiple-input multiple-output (MIMO) broadcast channel (BC). This paper presents a DPC based code design for BCs. We consider the case in which there is an individual rate/signal-to-interference-plus-noise ratio (SINR) constraint for each user. For a fixed transmitter power, we choose the linear transmit precoding matrix such that the SINRs at users are uniformly maximized, thus ensuring the best bit-error rate performance. We start with Cover's simplest two-user Gaussian BC and present a coding scheme that operates 1.44 dB from the boundary of the capacity region at the rate of one bit per real sample (b/s) for each user. We then extend the coding strategy to a two-user MIMO Gaussian BC with two transmit antennas at the base-station and develop the first limit-approaching code design using nested turbo codes for DPC. At the rate of 1 b/s for each user, our design operates 1.48 dB from the capacity region boundary. We also consider the performance of our scheme over a slow fading BC. For two transmit antennas, simulation results indicate a performance loss of only 1.4 dB, 1.64 dB and 1.99 dB from the theoretical limit in terms of the total transmission power for the two, three and four user case, respectively.
Momin Uppal, Vladimir Stankovic 0001, Zixiang Xiong
IEEE Trans. Commun.3
2009 Wyner-Ziv coding based on TCQ and LDPC codes
abstract
This paper considers trellis coded quantization (TCQ) and low-density parity-check (LDPC) codes for the quadratic Gaussian Wyner-Ziv coding problem. After TCQ of the source X, LDPC codes are used to implement Slepian-Wolf coding of the quantized source Q(X) with side information Y at the decoder. Assuming 256-state TCQ and ideal Slepian-Wolf coding in the sense of achieving the theoretical limit H(Q(X)|Y ), we experimentally show that Slepian-Wolf coded TCQ performs 0.2 dB away from the Wyner-Ziv distortion-rate function DWZ(R) at high rate. This result mirrors that of entropy-constrained TCQ in classic source coding of Gaussian sources. Furthermore, using 8,192-state TCQ and assuming ideal Slepian-Wolf coding, our simulations show that Slepian-Wolf coded TCQ performs only 0.1 dB away from DWZ(R) at high rate. These results establish the practical performance limit of Slepian-Wolf coded TCQ for quadratic Gaussian Wyner-Ziv coding. Practical designs give performance very close to the theoretical limit. For example, with 8,192-state TCQ, irregular LDPC codes for Slepian-Wolf coding and optimal non-linear estimation at the decoder, our performance gap to DWZ(R) is 0.20 dB, 0.22 dB, 0.30 dB, and 0.93 dB at 3.83 bit per sample (b/s), 1.83 b/s, 1.53 b/s, and 1.05 b/s, respectively. When 256-state 4-D trellis-coded vector quantization instead of TCQ is employed, the performance gap to DWZ(R) is 0.51 dB, 0.51 dB, 0.54 dB, and 0.80 dB at 2.04 b/s, 1.38 b/s, 1.0 b/s, and 0.5 b/s, respectively.
Yang Yang 0003, Samuel Cheng 0001, Zixiang Xiong, Wei Zhao 0001
IEEE Trans. Commun.3
2009 Two-Terminal Video Coding
abstract
Following recent works on the rate region of the quadratic Gaussian two-terminal source coding problem and limit-approaching code designs, this paper examines multiterminal source coding of two correlated, i.e., stereo, video sequences to save the sum rate over independent coding of both sequences. Two multiterminal video coding schemes are proposed. In the first scheme, the left sequence of the stereo pair is coded by H.264/AVC and used at the joint decoder to facilitate Wyner-Ziv coding of the right video sequence. The first I-frame of the right sequence is successively coded by H.264/AVC Intracoding and Wyner-Ziv coding. An efficient stereo matching algorithm based on loopy belief propagation is then adopted at the decoder to produce pixel-level disparity maps between the corresponding frames of the two decoded video sequences on the fly. Based on the disparity maps, side information for both motion vectors and motion-compensated residual frames of the right sequence are generated at the decoder before Wyner-Ziv encoding. In the second scheme, source splitting is employed on top of classic and Wyner-Ziv coding for compression of both I-frames to allow flexible rate allocation between the two sequences. Experiments with both schemes on stereo video sequences using H.264/AVC, LDPC codes for Slepian-Wolf coding of the motion vectors, and scalar quantization in conjunction with LDPC codes for Wyner-Ziv coding of the residual coefficients give a slightly lower sum rate than separate H.264/AVC coding of both sequences at the same video quality.
Yang Yang 0003, Vladimir Stankovic 0001, Zixiang Xiong, Wei Zhao 0001
IEEE Trans. Image Process.3
2009 Near-capacity dirty-paper code design: a source-channel coding approach
abstract
This paper examines near-capacity dirty-paper code designs based on source–channel coding. We first point out that the performance loss in signal-to-noise ratio (SNR) in our code designs can be broken into the sum of the packing loss from channel coding and a modulo loss, which is a function of the granular loss from source coding and the target dirty-paper coding rate (or SNR). We then examine practical designs by combining trellis-coded quantization (TCQ) with both systematic and nonsystematic irregular repeat–accumulate (IRA) codes. Like previous approaches, we exploit the extrinsic information transfer (EXIT) chart technique for capacity-approaching IRA code design; but unlike previous approaches, we emphasize the role of strong source coding to achieve as much granular gain as possible using TCQ. Instead of systematic doping, we employ two relatively shifted TCQ codebooks, where the shift is optimized (via tuning the EXIT charts) to facilitate the IRA code design. Our designs synergistically combine TCQ with IRA codes so that they work together as well as they do individually. By bringing together TCQ (the best quantizer from the source coding community) and EXIT chart-based IRA code designs (the best from the channel coding community), we are able to approach the theoretical limit of dirty-paper coding. For example, at 0.25 bit per symbol (b/s), our best code design (with 2048-state TCQ) performs only 0.630 dB away from the Shannon capacity.
Yang Yang 0003, Angelos D. Liveris, Vladimir Stankovic 0001, Zixiang Xiong
IEEE Trans. Inf. Theory5
2009 On Practical Design for Joint Distributed Source and Network Coding
abstract
This paper considers the problem of communicating correlated information from multiple source nodes over a network of noiseless channels to multiple destination nodes, where each destination node wants to recover all sources. The problem involves a joint consideration of distributed compression and network information relaying. Although the optimal rate region has been theoretically characterized, it was not clear how to design practical communication schemes with low complexity. This work provides a partial solution to this problem by proposing a low-complexity scheme for the special case with two sources whose correlation is characterized by a binary symmetric channel. Our scheme is based on a careful combination of linear syndrome-based Slepian-Wolf coding and random linear mixing (network coding). It is in general suboptimal; however, its low complexity and robustness to network dynamics make it suitable for practical implementation.
Yunnan Wu, Vladimir Stankovic 0001, Zixiang Xiong, Sun-Yuan Kung
IEEE Trans. Inf. Theory3
2009 Scalable Video Multicast Using Expanding Window Fountain Codes
abstract
Fountain codes were introduced as an efficient and universal forward error correction (FEC) solution for data multicast over lossy packet networks. They have recently been proposed for large scale multimedia content delivery in practical multimedia distribution systems. However, standard fountain codes, such as LT or Raptor codes, are not designed to meet unequal error protection (UEP) requirements typical in real-time scalable video multicast applications. In this paper, we propose recently introduced UEP expanding window fountain (EWF) codes as a flexible and efficient solution for real-time scalable video multicast. We demonstrate that the design flexibility and UEP performance make EWF codes ideally suited for this scenario, i.e., EWF codes offer a number of design parameters to be “tuned” at the server side to meet the different reception criteria of heterogeneous receivers. The performance analysis using both analytical results and simulation experiments of H.264 scalable video coding (SVC) multicast to heterogeneous receiver classes confirms the flexibility and efficiency of the proposed EWF-based FEC solution.
Dejan Vukobratovic, Vladimir Stankovic 0001, Dino Sejdinovic, Lina Stankovic, Zixiang Xiong
IEEE Trans. Multim.5
2009 Bandwidth efficient multi-station wireless streaming based on complete complementary sequences
abstract
Data streaming from multiple base stations to a client is recognized as a robust technique for multimedia streaming. However the resulting transmission in parallel over wireless channels poses serious challenges, especially multiple access interference, multipath fading, noise effects and synchronization. Spread spectrum techniques seem the obvious choice to mitigate these effects, but at the cost of increased bandwidth requirements. This paper proposes a solution that exploits complete complementary spectrum spreading and data compression techniques jointly to resolve the communication challenges whilst ensuring efficient use of spectrum and acceptable bit error rate. Our proposed spreading scheme reduces the required transmission bandwidth by exploiting correlation among information present at multiple base stations. Results obtained show 1.75 Mchip/sec (or 25%) reduction in transmission rate, with only up to 6 dB loss in frequency-selective channel compared to a straightforward solution based solely on complete complementary spectrum spreading.
Chadi Khirallah, Vladimir Stankovic 0001, Lina Stankovic, Yang Yang 0003, Zixiang Xiong
IEEE Trans. Wirel. Commun.5
2008 Expanding Window Fountain codes for scalable video multicast
abstract
Digital Fountain (DF) codes have recently been suggested as an efficient forward error correction (FEC) solution for video multicast to heterogeneous receiver classes over lossy packet networks. However, to adapt DF codes to low-delay constraints and varying importance of scalable multimedia content, unequal error protection (UEP) DF schemes are needed. Thus, in this paper, Expanding Window Fountain (EWF) codes are proposed as a FEC solution for scalable video multicast. We demonstrate that the design flexibility and UEP performancemake EWF codes ideally suited for this scenario, i.e., EWF codes offer a number of design parameters to be “tuned” at the server side to meet the different reception conditions of heterogeneous receivers. Performance analysis of H.264 Scalable Video Coding (SVC) multicast to heterogeneous receiver classes confirms the flexibility and efficiency of the proposed EWF-based FEC solution.
Dejan Vukobratovic, Vladimir Stankovic 0001, Dino Sejdinovic, Lina Stankovic, Zixiang Xiong
ICME5
2008 Compress-forward coding with BPSK modulation for the half-duplex Gaussian relay channel
abstract
Cover and El Gamal derived the tightest bounds on the capacity of the relay channel using random coding and suggested two coding strategies, namely, decode-forward (DF) and compress-forward (CF), to provide the best known lower bound of the achievable rate. Practical code designs proposed recently mainly exploit DF to approach the lower bound. Following the latest development in practical distributed source-channel coding, this paper studies CF coding with BPSK modulation for the relay channel. In CF scheme, Wyner-Ziv coding is applied at the relay to exploit the joint statistics between signals at the relay and the destination. We employ Slepian-Wolf coded nested scalar quantization (SWCNSQ) in practical Wyner-Ziv coding at the relay and compute the achievable rates of this scheme with BPSK modulation for the half-duplex Gaussian relay channel. We present a code design based on LDPC codes for error protection at the source and NSQ and IRA codes for CF at the relay. Simulation results show that our design comes within 1.48-1.92 dB of the SWCNSQ limit.
Zhixin Liu 0007, Momin Uppal, Vladimir Stankovic 0001, Zixiang Xiong
ISIT4
2008 Nested turbo codes for the Costa problem
abstract
Driven by applications in data-hiding, MIMO broadcast channel coding, precoding for interference cancellation, and transmitter cooperation in wireless networks, Costa coding has lately become a very active research area. In this paper, we first offer code design guidelines in terms of source- channel coding for algebraic binning. We then address practical code design based on nested lattice codes and propose nested turbo codes using turbo-like trellis-coded quantization (TCQ) for source coding and turbo trellis-coded modulation (TTCM) for channel coding. Compared to TCQ, turbo-like TCQ offers structural similarity between the source and channel coding components, leading to more efficient nesting with TTCM and better source coding performance. Due to the difference in effective dimensionality between turbo-like TCQ and TTCM, there is a performance tradeoff between these two components when they are nested together, meaning that the performance of turbo-like TCQ worsens as the TTCM code becomes stronger and vice versa. Optimization of this performance tradeoff leads to our code design that outperforms existing TCQ/TCM and TCQ/TTCM constructions and exhibits a gap of 0.94, 1.42 and 2.65 dB to the Costa capacity at 2.0, 1.0, and 0.5 bits/sample, respectively.
Momin Uppal, Angelos D. Liveris, Samuel Cheng 0001, Vladimir Stankovic 0001, Zixiang Xiong
IEEE Trans. Commun.6
2008 On Multiterminal Source Code Design
abstract
Multiterminal (MT) source coding refers to separate lossy encoding and joint decoding of multiple correlated sources. Recently, the rate region of bothdirectandindirectMT source coding in the quadratic Gaussian setup with two encoders was determined. We are thus motivated to design practical MT source codes that can potentially achieve the entire rate region. In this paper, we present two practical MT coding schemes under the framework of Slepian–Wolf coded quantization (SWCQ) for both direct and indirect MT problems. The first,asymmetricSWCQ scheme relies on quantization and Wyner–Ziv coding, and it is implemented via source splitting to achieve any point on the sum–rate bound. In the second, conceptually simpler scheme,symmetricSWCQ, the two quantized sources are compressed using symmetric Slepian–Wolf coding via a channel code partitioning technique that is capable of achieving any point on the Slepian–Wolf sum–rate bound. Our practical designs employ trellis-coded quantization and turbo/low-density parity-check (LDPC) codes for both asymmetric and symmetric Slepian–Wolf coding. Simulation results show a gap of only 0.139–0.194 bit per sample away from the sum–rate bound for both direct and indirect MT coding problems.
Yang Yang 0003, Vladimir Stankovic 0001, Zixiang Xiong, Wei Zhao 0001
IEEE Trans. Inf. Theory3
2007 Facial Feature Extraction from Range Images using a 3D Morphable Model
abstract
In this paper, a novel scheme is introduced for human facial feature extraction. Unlike previous methods that fit a 3D morphable model to 2D intensity images, our scheme utilizes 3D range images to extract features without requiring manually-defined initial landmark points. A linear transformation is used to achieve the mapping between the 3D model and a 3D range image, which makes the computation simple and fast. Moreover, our scheme is robust to the illumination and pose variations. In addition to features from range images, extra features can be obtained by examining optional 2D texture images. Using our scheme, we can also perform automatic eye/mouth corner localization. Experimental results show the high accuracy and robustness of our scheme.
Le Zou, Samuel Cheng 0001, Zixiang Xiong, Mi Lu, Kenneth R. Castleman
ICASSP (2)3
2007 An FPGA Implementation of Dirty Paper Precoder
abstract
Dirty paper code (DPC) can be used in a number of communication network applications; broadcast channels, multiuser interference channels and ISI channels to name a few. We study various implementation bottlenecks and issues with implementing a DPC pre-coder based on nested trellis technique. The aim is to achieve a practical hardware realization of the precoder for wireless LAN/DSL applications. We describe the architectural development process and realization of the precoder on a Xilinx Virtex 2V8000 FPGA. To the best of our knowledge this is the first reported DPC pre-coder hardware implementation.
Pankaj Bhagawat, Weihuang Wang, Momin Uppal, Gwan S. Choi, Zixiang Xiong, Mark B. Yeary, Alan Harris
ICC5
2007 Efficient Multimedia Multicast Using Distributed Source Coding
abstract
We propose a system for real-time multimedia multicast over heterogeneous wireless-wireline networks. The encoded source is transmitted from the base station over the wireless radio link to numerous Internet servers. Each server performs distributed source coding (as quantization followed by Slepian-Wolf coding) by exploiting mutual correlation among packets received at different servers. The resulting packets are forwarded to the clients for joint decoding. We provide an algorithm for optimal nonuniform scalar quantizers design at the server side that minimizes the required rate under the decoder bit error rate constraint. For scalable multimedia codes, we develop joint source-channel coding scheme which combines error-protection at the base station and distributed source coding at the servers. Our experimental results show significant performance improvements over conventional solutions due to spatial diversity and distributed source coding gains.
Vladimir Stankovic 0001, Yang Yang 0003, Zixiang Xiong
ICC3
2007 Multiterminal Video Coding
abstract
Following recent works on the rate region of the quadratic Gaussian two-terminal source coding problem and limit-approaching code designs, this paper examines multiterminal source coding of two correlated video sequences to save the sum rate over independent coding. Specifically, the first video sequence is coded by H.264 and used at the joint decoder to facilitate Wyner-Ziv coding of the second video sequence. The first I-frame of the right sequence is successively coded by H.264 and Slepian-Wolf coding. An efficient stereo matching algorithm based on loopy belief propagation is then adopted at the decoder to produce pixel-level disparity maps between the corresponding frames of the two decoded video sequences on the fly. Based on the disparity maps, side information for both motion vectors and motion-compensated residual frames of the second sequence are generated at the decoder before Wyner-Ziv encoding. Experimental results on stereo video sequences using H.264, LDPC codes for Slepian-Wolf coding of the motion vectors and scalar quantization in conjunction with LDPC codes for Wyner-Ziv coding of the residual coefficients show savings in terms of the sum-rate when compared to separate H.264 coding at the same video quality.
Yang Yang 0003, Vladimir Stankovic 0001, Wei Zhao 0001, Zixiang Xiong
ICIP (3)4
2007 Practical rateless cooperation in multiple access channels using multiplexed Raptor codes
abstract
In this paper we develop practical rateless coded cooperation strategies for a two-user multiple access channel. At the heart of our practical strategies lie Raptor codes concatenated with a space-time code, and an iterative decoding procedure to jointly recover the two users' messages at the base station. Since cooperation in multiple access channels is known to be particularly beneficial when considering the outage probability at low SNRs, or equivalently at low transmission rates, we simulate our strategies at a transmission rate of 0.25 bits/sample. Experiments indicate that our schemes perform very close to the theoretical limit, with the performance gap being less than 0.55 dB.
Momin Uppal, Anders Høst-Madsen, Zixiang Xiong
ISIT3
2007 Distributed Joint Source-Channel Coding of Video Using Raptor Codes
abstract
Extending recent works on distributed source coding, this paper considers distributed source-channel coding and targets at the important application of scalable video transmission over wireless networks. The idea is to use a single channel code for both video compression (via Slepian-Wolf coding) and packet loss protection. First, we provide a theoretical code design framework for distributed joint source-channel coding over erasure channels and then apply it to the targeted video application. The resulting video coder is based on a cross-layer design where video compression and protection are performed jointly. We choose Raptor codes - the best approximation to a digital fountain - and address in detail both encoder and decoder designs. Using the received packets together with a correlated video available at the decoder as side information, we devise a new iterative soft-decision decoder for joint Raptor decoding. Simulation results show that, compared to one separate design using Slepian-Wolf compression plus erasure protection and another based on FGS coding plus erasure protection, the proposed joint design provides better video quality at the same number of transmitted packets. Our work represents the first in capitalizing the latest in distributed source coding and near-capacity channel coding for robust video transmission over erasure channels.
Qian Xu 0001, Vladimir Stankovic 0001, Zixiang Xiong
IEEE J. Sel. Areas Commun.3
2007 Nested Turbo Codes for the Costa Problem
abstract
Driven by the applications in data hiding, multiple-input multiple-output broadcast channel coding, precoding for interference cancellation, and transmitter cooperation in wireless networks, Costa coding has lately become a very active research area. In this paper, we first offer code design guidelines in terms of source channel coding for algebraic binning. We then address a practical code design based on nested lattice codes, and propose nested turbo codes by using turbo-like trellis-coded quantization (TCQ) for source coding and turbo trellis-coded modulation (TTCM) for channel coding. Compared to the TCQ, the turbo-like TCQ offers a structural similarity between the source and channel coding components, leading to a more efficient nesting with TTCM and a better source coding performance. Due to the difference in effective dimensionality between turbo-like TCQ and TTCM, there is a performance tradeoff between these two components when they are nested together, meaning that the performance of the turbo-like TCQ worsens as the TTCM code becomes stronger and vice versa. The optimization of this performance tradeoff leads to our code design that outperforms existing TCQ/TCM and TCQ/TTCM constructions, and exhibits a gap of 0.94, 1.42, and 2.65 dB to the Costa capacity at 2.0, 1.0, and 0.5 b/s, respectively.
Momin Uppal, Angelos D. Liveris, Samuel Cheng 0001, Vladimir Stankovic 0001, Zixiang Xiong
IEEE Trans. Commun.6
2007 Wyner-Ziv Video Compression and Fountain Codes for Receiver-Driven Layered Multicast
abstract
The increasing popularity of video streaming applications that distribute data to a large number of clients motivates the design of reliable multimedia delivery systems capable of adapting to diverse transmission conditions. Receiver-driven layered multicast (RLM) efficiently addresses the issue of heterogeneity in clients' available bandwidths and packet loss rates by shifting rate control to the receiver side. We propose a system for RLM over the Internet and 3G wireless networks based on layered Wyner-Ziv video coding and digital fountain codes. Layered Wyner-Ziv video coding improves robustness to packet loss compared to current scalable video coders, such as MPEG-4 FGS coder, while generating a scalable output bit stream. Digital fountain codes are near-capacity erasure protection codes that are ideally suited for multicast applications due to their rateless property. By combining an error-resilient Wyner-Ziv video coder and rateless fountain codes, our system allows reliable video multicast to an arbitrary number of heterogeneous receivers without the requirement of feedback channels. Simulation results show performance improvements over a previous scheme that exploits multiple description and layered coding.
Qian Xu 0001, Vladimir Stankovic 0001, Zixiang Xiong
IEEE Trans. Circuits Syst. Video Technol.3
2007 3-D Face Recognition Based on Warped Example Faces
abstract
In this paper, we describe a novel 3-D face recognition scheme for 3-D face recognition that can automatically identify faces from range images, and is insensitive to holes, facial expression, and hair. In our scheme, a number of carefully selected range images constitute a set of example faces, and another range image is chosen as a ldquogeneric face.rdquo The generic face is then warped to match each of the example faces in the least mean square sense. Each such warp is specified by a vector of displacement values. In feature extraction operation, when a target face image comes in, the generic face is warped to match it. The geometric transformation used in the warping is a linear combination of the example face warping vectors. The coefficients in the linear combination are adjusted to minimize the root mean square error. After the matching process is complete, the coefficients of the composite warp are used as features and passed to a Mahalanobis-distance-based classifier for face recognition. Our technique is tested on a data set containing more than 600 range images. Experimental results in the access-control scenario show the effectiveness of the extracted features.
Le Zou, Samuel Cheng 0001, Zixiang Xiong, Mi Lu, Kenneth R. Castleman
IEEE Trans. Inf. Forensics Secur.3
2007 Multiuser resource allocation for video transmission over a chip-interleaved multicarrier system
abstract
Abstract We propose a 4G system for transmission of video from a server at the base station to numerous wireless clients. We employ the latest technology in scalable video compression (3‐D wavelet video coding) and in channel coding (punctured turbo codes); for the physical layer, we resort to the multicarrier chip‐interleaved system with two‐layer interleaving, which achieves high spectral efficiency and is very suitable for downlink applications. We develop fast algorithms for a cross‐layer resource allocation that minimize the expected distortion of the reconstructed video averaged over all clients. The algorithms find a near‐optimal power, bandwidth, and subcarrier allocation at the physical layer and a source‐channel symbol allocation at the application layer. Our experimental results demonstrate that such a cross‐layer optimization framework leads to higher quality performance of the overall system. Copyright © 2007 John Wiley & Sons, Ltd.
Kai Yang 0001, Vladimir Stankovic 0001, Zixiang Xiong, Xiaodong Wang 0001
Wirel. Commun. Mob. Comput.3
2006 Video Multicast over Heterogeneous Networks Based on Distributed Source Coding Principles
abstract
Real-time multimedia multicast over wireless networks is an exciting application that has generated a lot of interest recently. Its main challenge lies in the stringent bandwidth and time-delay requirements of real-time multimedia and severe impairments of the wireless channels. We develop a system for multimedia multicast over wireless-wireline networks, that leverages the knowledge on network information theory, multimedia processing, error control, and networking. In particular, the encoded multimedia data are broadcast to multiple Internet servers over a wireless channel. Each server merely compresses the signal it has received using distributed source coding. The receiver collects bitstreams from the servers before performing joint decoding. Due to spatial diversity gain and distributed source coding, our system significantly outperforms conventional solutions.
Vladimir Stankovic 0001, Yang Yang 0003, Zixiang Xiong
ICIP3
2006 Code Designs for MIMO Broadcast Channels
abstract
Recent information-theoretical results show the optimality of dirty-paper coding (DPC) in achieving the capacity of the Gaussian multiple-input multiple-output (MIMO) broadcast channel (BC). This paper presents the first practical limit approaching DPC-based design for the MIMO BC. We start with Cover's simplest two-user Gaussian BC and present a code design that operates 1.44 dB away from the capacity region boundary at a transmission rate of 1.0 bit per sample (b/s). Then we consider the non-degraded two-user MIMO fading BC with two transmit antennas. For this setup, the performance loss of our code design is 3.7 dB and 2.45 dB from the the sum-rate capacity when the transmission rate for each user is 1.0 b/s and 2.0 b/s, respectively.
Momin Uppal, Vladimir Stankovic 0001, Zixiang Xiong
ISIT3
2006 Noise-injected neural networks show promise for use on small-sample expression data
abstract
BACKGROUND: Overfitting the data is a salient issue for classifier design in small-sample settings. This is why selecting a classifier from a constrained family of classifiers, ones that do not possess the potential to too finely partition the feature space, is typically preferable. But overfitting is not merely a consequence of the classifier family; it is highly dependent on the classification rule used to design a classifier from the sample data. Thus, it is possible to consider families that are rather complex but for which there are classification rules that perform well for small samples. Such classification rules can be advantageous because they facilitate satisfactory classification when the class-conditional distributions are not easily separated and the sample is not large. Here we consider neural networks, from the perspectives of classical design based solely on the sample data and from noise-injection-based design. RESULTS: This paper provides an extensive simulation-based comparative study of noise-injected neural-network design. It considers a number of different feature-label models across various small sample sizes using varying amounts of noise injection. Besides comparing noise-injected neural-network design to classical neural-network design, the paper compares it to a number of other classification rules. Our particular interest is with the use of microarray data for expression-based classification for diagnosis and prognosis. To that end, we consider noise-injected neural-network design as it relates to a study of survivability of breast cancer patients. CONCLUSION: The conclusion is that in many instances noise-injected neural network design is superior to the other tested methods, and in almost all cases it does not perform substantially worse than the best of the other methods. Since the amount of noise injected is consequential, the effect of differing amounts of injected noise must be considered.
Jianping Hua, James Lowey, Zixiang Xiong, Edward R. Dougherty
BMC Bioinform.3
2006 Distributed source coding
Zixiang Xiong, Javier Garcia-Frías, Bernd Girod
Signal Process.1
2006 Layered Wyner-Ziv video coding for transmission over unreliable channels
Qian Xu 0001, Vladimir Stankovic 0001, Zixiang Xiong
Signal Process.3
2006 Source-Optimized Irregular Repeat Accumulate Codes With Inherent Unequal Error Protection Capabilities and Their Application to Scalable Image Transmission
abstract
The common practice for achieving unequal error protection (UEP) in scalable multimedia communication systems is to design rate-compatible punctured channel codes before computing the UEP rate assignments. This paper proposes a new approach to designing powerful irregular repeat accumulate (IRA) codes that are optimized for the multimedia source and to exploiting the inherent irregularity in IRA codes for UEP. Using the end-to-end distortion due to the first error bit in channel decoding as the cost function, which is readily given by the operational distortion-rate function of embedded source codes, we incorporate this cost function into the channel code design process via density evolution and obtain IRA codes that minimize the average cost function instead of the usual probability of error. Because the resulting IRA codes have inherent UEP capabilities due to irregularity, the new IRA code design effectively integrates channel code optimization and UEP rate assignments, resulting in source-optimized channel coding or joint source-channel coding. We simulate our source-optimized IRA codes for transporting SPIHT-coded images over a binary symmetric channel with crossover probability p. When p = 0.03 and the channel code length is long (e.g., with one codeword for the whole 512 x 512 image), we are able to operate at only 9.38% away from the channel capacity with code length 132380 bits, achieving the best published results in terms of average peak signal-to-noise ratio (PSNR). Compared to conventional IRA code design (that minimizes the probability of error) with the same code rate, the performance gain in average PSNR from using our proposed source-optimized IRA code design is 0.8759 dB when p = 0.1 and the code length is 12800 bits. As predicted by Shannon's separation principle, we observe that this performance gain diminishes as the code length increases.
Chingfu Lan, Zixiang Xiong, Krishna Narayanan 0001
IEEE Trans. Image Process.2
2006 Layered Wyner-Ziv Video Coding
abstract
Following recent theoretical works on successive Wyner-Ziv coding (WZC), we propose a practical layered Wyner-Ziv video coder using the DCT, nested scalar quantization, and irregular LDPC code based Slepian-Wolf coding (or lossless source coding with side information at the decoder). Our main novelty is to use the base layer of a standard scalable video coder (e.g., MPEG-4/H.26L FGS or H.263+) as the decoder side information and perform layered WZC for quality enhancement. Similar to FGS coding, there is no performance difference between layered and monolithic WZC when the enhancement bitstream is generated in our proposed coder. Using an H.26L coded version as the base layer, experiments indicate that WZC gives slightly worse performance than FGS coding when the channel (for both the base and enhancement layers) is noiseless. However, when the channel is noisy, extensive simulations of video transmission over wireless networks conforming to the CDMA2000 1X standard show that H.26L base layer coding plus Wyner-Ziv enhancement layer coding are more robust against channel errors than H.26L FGS coding. These results demonstrate that layered Wyner-Ziv video coding is a promising new technique for video streaming over wireless networks.
Qian Xu 0001, Zixiang Xiong
IEEE Trans. Image Process.2
2006 Slepian-Wolf Coded Nested Lattice Quantization for Wyner-Ziv Coding: High-Rate Performance Analysis and Code Design
abstract
Nested lattice quantization provides a practical scheme for Wyner-Ziv coding. This paper examines the high-rate performance of nested lattice quantizers and gives the theoretical performance for general continuous sources. In the quadratic Gaussian case, as the rate increases, we observe an increasing gap between the performance of finite-dimensional nested lattice quantizers and the Wyner-Ziv distortion-rate function. We argue that this is because the boundary gain decreases as the rate of the nested lattice quantizers increases. To increase the boundary gain and ultimately boost the overall performance, a new practical Wyner-Ziv coding scheme called Slepian-Wolf coded nested lattice quantization (SWC-NQ) is proposed, where Slepian-Wolf coding is applied to the quantization indices of the source for the purpose of compression with side information at the decoder. Theoretical analysis shows that for the quadratic Gaussian case and at high rate, SWC-NQ performs the same as conventional entropy-coded lattice quantization with the side information available at both the encoder and the decoder. Furthermore, a nonlinear minimum mean-square error (MSE) estimator is introduced at the decoder, which is theoretically proven to degenerate to the linear minimum MSE estimator at high rate and experimentally shown to outperform the linear estimator at low rate. Practical designs of one- and two-dimensional nested lattice quantizers together with multilevel low-density parity-check (LDPC) codes for Slepian-Wolf coding give performance close to the theoretical limits of SWC-NQ
Zhixin Liu 0007, Samuel Cheng 0001, Angelos D. Liveris, Zixiang Xiong
IEEE Trans. Inf. Theory4
2006 On dualities in multiterminal coding problems
abstract
It has been shown recently that under certain conditions there exist dualities between different multiterminal (MT) source and channel coding problems. Following these results, we study lossless MT source coding and deterministic MT channel coding problems and point out different dualities between them. In particular, we show that there exists a functional duality between a Slepian-Wolf (SW) coding problem and a deterministic broadcast channel (DBC) coding problem and between a lossless multiple-description (MD) coding problem and a deterministic multiple-access channel (DMAC) coding problem. In analogy to the duality established between DBC and DMAC coding problems, we further propose a similar duality between SW and lossless MD coding problems; in this way, we form a closed "duality loop" of four MT coding problems, which imposes the existence of a single common rate point in the achievable rate regions of all four dual problems. We also consider duality in zero-error MT coding and shed light on practical code design with an example. Finally, extension to the case with only one lossless/deterministic component in the source/channel coding problem is provided.
Vladimir Stankovic 0001, Samuel Cheng 0001, Zixiang Xiong
IEEE Trans. Inf. Theory3
2006 On code design for the Slepian-Wolf problem and lossless multiterminal networks
abstract
A Slepian-Wolf coding scheme for compressing two uniform memoryless binary sources using a single channel code that can achieve arbitrary rate allocation among encoders was outlined in the work of Pradhan and Ramchandran. Inspired by this work, we address the problem of practical code design for general multiterminal lossless networks where multiple memoryless correlated binary sources are separately compressed and sent; each decoder receives a set of compressed sources and attempts to jointly reconstruct them. First, we propose a near-lossless practical code design for the Slepian-Wolf system with multiple sources. For two uniform sources, if the code approaches the capacity of the channel that models the correlation between the sources, then the system will approach the theoretical limit. Thus, the great advantage of this design method is its possibility to approach the theoretical limits with a single channel code for any rate allocation among the encoders. Based on Slepian-Wolf code constructions, we continue with providing practical designs for the general lossless multiterminal network which consists of an arbitrary number of encoders and decoders. Using irregular repeat-accumulate and turbo codes in our designs, we obtain the best results reported so far and almost reach the theoretical bounds.
Vladimir Stankovic 0001, Angelos D. Liveris, Zixiang Xiong, Costas N. Georghiades
IEEE Trans. Inf. Theory3
2006 Progressive Image Transmission over Space-Time Coded OFDM-Based MIMO Systems with Adaptive Modulation
abstract
This paper considers a progressive image transmission system over wireless channels by combining joint source-channel coding (JSCC), space-time coding, and orthogonal frequency division multiplexing (OFDM). The BER performance of the space-time coded OFDM-based MIMO system based on a newly built broadband MIMO fading model is first evaluated by assuming perfect channel state information at the receiver for coherent detection. Then, for a given average SNR (hence, BER), a fast local search algorithm is applied to optimize the unequal error protection design in JSCC, subjected to fixed total transmitted energy for various constellation sizes. This design allows the measurement of the expected reconstructed image quality. With this end-to-end system performance evaluation, an adaptive modulation scheme is proposed to pick the constellation size that offers the best reconstructed image quality for each average SNR. Simulation results of practical image transmissions confirm the effectiveness of our proposed adaptive modulation scheme.
Zixiang Xiong
IEEE Trans. Mob. Comput.2
2006 Progressive Video Delivery over Wideband Wireless Channels Using Space-Time Differentially Coded OFDM Systems
abstract
Progressive video delivery over wireless networks is very challenging due to the time-varying nature of wireless channels and limited power in the mobile devices. This paper proposes an end-to-end architecture for multilayer progressive video delivery over space-time differentially coded orthogonal frequency division multiplexing (STDC-OFDM) systems. An input video sequence is compressed by 3D-ESCOT into a layered bitstream. We input multiple layers of the bitstream in series to a STDC-OFDM channel. Different video source layers are protected by different error protection schemes in order to achieve unequal error protection. In progressive transmission, the reconstruction quality is important not only at the target transmission rate but also at the intermediate rates. So, the error protection strategy needs to optimize the average performance over the set of intermediate rates. We propose to use progressive joint source-channel coding to generate operational transmission distortion-rate (TD-R) functions and operational transmission distortion-power (TD-P) functions for multiple layers before forming the operational transmission distortion-power-rate (TD-PR) surfaces. Lagrange multipliers are then employed on the fly to obtain the optimal power allocation and optimal rate allocation among multiple layers, subject to constraints on the total transmission rate and the total power level. Progressive joint source-channel coding offers the scalability feature to handle bandwidth variations and changes in channel conditions. By extending the rate-distortion function in source coding to the TD-PR surface in joint source-channel coding, our work can use the "equal slope" argument to effectively solve the transmission rate allocation problem as well as the transmission power allocation problem for multilayer video transmission. Experiments show that our scheme achieves significant improvement over a nonoptimal system with the same total power level and total transmission rate.
Shengjie Zhao 0001, Zixiang Xiong, Xiaodong Wang 0001, Jianping Hua
IEEE Trans. Mob. Comput.2
2006 Progressive image transmission over differentially space-time coded OFDM systems
abstract
In this paper, we consider progressive image transmission over differentially space-time coded orthogonal frequency-division multiplexing (OFDM) systems and treat the problem as one of optimal joint source-channel coding (JSCC) in the form of unequal error protection (UEP), as necessitated by embedded source coding (e.g., SPIHT and JPEG 2000). We adopt a product channel code structure that is proven to provide powerful error protection and employ low-complexity decision-feedback decoding for differentially space-time coded OFDM without assuming channel state information. For a given SNR, the BER performance of the differentially space-time coded OFDM system is treated as the channel condition in the JSCC/UEP design via a fast product code optimization algorithm so that the end-to-end quality of reconstructed images is optimized in the average minimum MSE sense. Extensive image transmission experiments show that SNR/BER improvements can be translated into quality gains in reconstructed images. Moreover, compared to another non-coherent detection algorithm, i.e., the iterative receiver based on expectation-maximization algorithm for the space-time coded OFDM systems, differentially space-time coded OFDM systems suffer some quality loss in reconstructed images. With the efficiency and simplicity of decision-feedback differential decoding, differentially space-time coded OFDM is thus a feasible modulation scheme for applications such as wireless image over mobile devices (e.g., cell phones). Copyright © 2006 John Wiley & Sons, Ltd.
Zixiang Xiong, Xiaodong Wang 0001
Wirel. Commun. Mob. Comput.2
2005 Distributed Joint Source-Channel Coding of Video Using Raptor Codes
abstract
Summary form only given. In this paper, we consider the case of a noisy channel in Wyner-Ziv coding (WZC) and address distributed joint source-channel coding (JSCC), while targeting at video transmission over packet erasure channels. Our idea is to use a single Raptor code, for both SWC and erasure protection. Raptor codes are the latest addition to a family of low-complexity rateless fountain codes which consist of a high-rate precode and an LT code. We use IRA codes as the precode of our Raptor code, as IRA codes are well suited for distributed JSCC. For the decoder design, due to the presence of side information, we develop a new iterative soft-decision Raptor decoder for joint decoding that combines the received packets and the side information.
Qian Xu 0001, Vladimir Stankovic 0001, Zixiang Xiong
DCC3
2005 On Multiterminal Source Code Design
abstract
Multiterminal (MT) source coding refers to separate lossy encoding and joint decoding of multiple correlated sources. This paper presents two practical MT coding schemes under the same general framework of Slepian-Wolf coded quantization (SWCQ) for both direct and indirect quadratic Gaussian MT source coding problems with two encoders. The first asymmetric SWCQ scheme relies on quantization and Wyner-Ziv coding, and is implemented via source-splitting to achieve any point on the inner sum-rate bound for both direct and indirect MT coding problems. In the second symmetric SWCQ scheme, the two quantization outputs are compressed using multilevel symmetric Slepian-Wolf coding. This scheme is conceptually simpler and can potentially achieve most of the points on the inner sum-rate bound. Our practical designs employ trellis coded quantization, LDPC code based asymmetric Slepian-Wolf code, and arithmetic code and turbo code based symmetric Slepian-Wolf code. Simulation results show a gap of only 0.24-0.29 bit per sample away from the inner sum-rate bound for both direct and indirect MT coding problems.
Yang Yang 0003, Vladimir Stankovic 0001, Zixiang Xiong, Wei Zhao 0001
DCC3
2005 Distributed source and joint source-channel coding: from theory to practice
abstract
This paper reviews the theory behind distributed source coding and distributed source-channel coding and surveys the progress made in code design during the last years.
Javier Garcia-Frías, Zixiang Xiong
ICASSP (5)2
2005 Wyner-Ziv coding for the half-duplex relay channel
abstract
Cover and El Gamal derived the tightest bounds on the capacity of the relay channel using random coding and proposed two coding strategies, namely decode-and-forward (DF) and compress-and-forward (CF), to provide the best known lower bounds of the achievable rate region. Depending on transmission parameters, either DF or CF could be superior. Several practical code designs based on DF have appeared recently. We present the first practical CF design for the half-duplex Gaussian relay channel based on Wyner-Ziv coding of the received source signal at the relay. Assuming ideal source and channel coding, our design achieves the lower bound of CF. It thus realizes the performance gain of CF over DF promised by the theory when the relay is close to the destination. Our practical implementation based on LDPC codes for error protection at the source and nested scalar quantization and IRA (irregular repeat-accumulate) codes for Wyner-Ziv coding at the relay comes as close as 0.76 dB to the theoretical limit of CF.
Zhixin Liu 0007, Vladimir Stankovic 0001, Zixiang Xiong
ICASSP (5)3
2005 Distributed joint source-channel coding of video
abstract
Based on recent works on source-channel coding for Wyner-Ziv coding, we consider the case with noisy channel in Wyner-Ziv coding and address distributed joint source-channel coding, while targeting at the important application of video transmission over packet erasure channels. The idea is to use a single channel code for both Slepian-Wolf coding (or source coding with side information at the decoder) and erasure protection. We choose Raptor codes - the best approximation to a digital fountain - for the targeted application and study both encoder and decoder designs under the new setting of distributed joint source-channel coding. Our work represents the first in capitalizing the latest in distributed source coding (e.g., Wyner-Ziv video coding) and near-capacity channel coding (e.g., fountain codes) for robust video transmission over erasure channels.
Qian Xu 0001, Vladimir Stankovic 0001, Angelos D. Liveris, Zixiang Xiong
ICIP (2)4
2005 Near-capacity dirty-paper code designs based on TCQ and IRA codes
abstract
This paper addresses near-capacity dirty-paper code designs based on TCQ and IRA codes, where the former is employed as the most efficient means of vector quantization and the latter for their capacity-approaching performance. By bringing together TCQ - the best quantizer from the source coding community and EXIT chart based IRA code designs - the best from the channel coding community, we are able to approach the theoretical limit of dirty-paper coding. For example, at 0.25 b/s, one of our code designs (with 1024-state TCQ) performs 0.83 dB away from the capacity
Angelos D. Liveris, Vladimir Stankovic 0001, Zixiang Xiong
ISIT4
2005 Optimal number of features as a function of sample size for various classification rules
abstract
MOTIVATION: Given the joint feature-label distribution, increasing the number of features always results in decreased classification error; however, this is not the case when a classifier is designed via a classification rule from sample data. Typically (but not always), for fixed sample size, the error of a designed classifier decreases and then increases as the number of features grows. The potential downside of using too many features is most critical for small samples, which are commonplace for gene-expression-based classifiers for phenotype discrimination. For fixed sample size and feature-label distribution, the issue is to find an optimal number of features. RESULTS: Since only in rare cases is there a known distribution of the error as a function of the number of features and sample size, this study employs simulation for various feature-label distributions and classification rules, and across a wide range of sample and feature-set sizes. To achieve the desired end, finding the optimal number of features as a function of sample size, it employs massively parallel computation. Seven classifiers are treated: 3-nearest-neighbor, Gaussian kernel, linear support vector machine, polynomial support vector machine, perceptron, regular histogram and linear discriminant analysis. Three Gaussian-based models are considered: linear, nonlinear and bimodal. In addition, real patient data from a large breast-cancer study is considered. To mitigate the combinatorial search for finding optimal feature sets, and to model the situation in which subsets of genes are co-regulated and correlation is internal to these subsets, we assume that the covariance matrix of the features is blocked, with each block corresponding to a group of correlated features. Altogether there are a large number of error surfaces for the many cases. These are provided in full on a companion website, which is meant to serve as resource for those working with small-sample classification. AVAILABILITY: For the companion website, please visit http://public.tgen.org/tamu/ofs/ CONTACT: [email protected].
Jianping Hua, Zixiang Xiong, James Lowey, Edward Suh, Edward R. Dougherty
Bioinform.2
2005 Optimal robust classifiers
Edward R. Dougherty, Jianping Hua, Zixiang Xiong, Yidong Chen 0002
Pattern Recognit.3
2005 Determination of the optimal number of features for quadratic discriminant analysis via the normal approximation to the discriminant distribution
Jianping Hua, Zixiang Xiong, Edward R. Dougherty
Pattern Recognit.2
2005 Robust layered multiple description coding of scalable media data for multicast
abstract
Layered multiple description codes allow robust transmission of scalable media data over packet erasure networks, while providing simple rate adaptation and bandwidth savings for shared bottleneck links. We show how to efficiently design layered multiple description codes for multicast and broadcast applications in memoryless packet erasure networks. Our approach offers a significantly better quality tradeoff among clients than the best previous solution.
Vladimir Stankovic 0001, Raouf Hamzaoui, Zixiang Xiong
IEEE Signal Process. Lett.3
2005 EM-based iterative receiver design with carrier-frequency offset estimation for MIMO OFDM systems
abstract
In this letter, we study the design of expectation-maximization (EM)-based iterative receivers for multiple-input multiple-output (MIMO) orthogonal frequency-division multiplexing systems with the presence of carrier-frequency offset (CFO). Motivated by the spirit of maximum-likelihood estimation in the EM algorithm, we first present a pilot-aided CFO estimation scheme that allows fast Fourier transform-based fast implementation. Then this CFO estimation is incorporated into the initialization step of the iterative receiver. Experimental results show the effectiveness of our receiver design in combating CFO.
Zixiang Xiong, Xiaodong Wang 0001
IEEE Trans. Commun.2
2005 Fast Algorithm for Distortion-Based Error Protection of Embedded Image Codes
abstract
We consider a joint source-channel coding system that protects an embedded bitstream using a finite family of channel codes with error detection and error correction capability. The performance of this system may be measured by the expected distortion or by the expected number of correctly decoded source bits. Whereas a rate-based optimal solution can be found in linear time, the computation of a distortion-based optimal solution is prohibitive. Under the assumption of the convexity of the operational distortion-rate function of the source coder, we give a lower bound on the expected distortion of a distortion-based optimal solution that depends only on a rate-based optimal solution. Then, we propose a local search (LS) algorithm that starts from a rate-based optimal solution and converges in linear time to a local minimum of the expected distortion. Experimental results for a binary symmetric channel show that our LS algorithm is near optimal, whereas its complexity is much lower than that of the previous best solution.
Raouf Hamzaoui, Vladimir Stankovic 0001, Zixiang Xiong
IEEE Trans. Image Process.3
2005 Subspace-Based Prototyping and Classification of Chromosome Images
abstract
Chromosomes are essential genomic information carriers. Chromosome classification constitutes an important part of routine clinical and cancer cytogenetics analysis. Cytogeneticists perform visual interpretation of banded chromosome images according to the diagrammatic models of various chromosome types known as the ideograms, which mimic artists' depiction of the chromosomes. In this paper, we present a subspace-based approach for automated prototyping and classification of chromosome images. We show that 1) prototype chromosome images can be quantitatively synthesized from a subspace to objectively represent the chromosome images of a given type or population, and 2) the transformation coefficients (or projected coordinate values of sample chromosomes) in the subspace can be utilized as the extracted feature measurements for classification purposes. We examine in particular the formation of three well-known subspaces, namely the ones derived from principal component analysis (PCA), Fisher's linear discriminant analysis, and the discrete cosine transform (DCT). These subspaces are implemented and evaluated for prototyping two-dimensional (2-D) images and for classification of both 2-D images and one-dimensional profiles of chromosomes. Experimental results show that previously unseen prototype chromosome images of high visual quality can be synthesized using the proposed subspace-based method, and that PCA and the DCT significantly outperform the well-known benchmark technique of weighted density distribution functions in classifying 2-D chromosome images.
Qiang Wu 0007, Zhongmin Liu, Zixiang Xiong, Kenneth R. Castleman
IEEE Trans. Image Process.4
2005 Computing the channel capacity and rate-distortion function with two-sided state information
abstract
In this correspondence, we present iterative algorithms that numerically compute the capacity-power and rate-distortion functions for coding with two-sided state information. Numerical examples are provided to demonstrate efficiency of our algorithms.
Samuel Cheng 0001, Vladimir Stankovic 0001, Zixiang Xiong
IEEE Trans. Inf. Theory3
2005 Optimal Resource Allocation for Wireless Video over CDMA Networks
abstract
We present a multiple-channel video transmission scheme in wireless CDMA networks over multipath fading channels. We map an embedded video bitstream, which is encoded into multiple independently decodable layers by 3D-ESCOT video coding technique, to multiple CDMA channels. One video source layer is transmitted over one CDMA channel. Each video source layer is protected by a product channel code structure. A product channel code is obtained by the combination of a row code based on rate compatible punctured convolutional code (RCPC) with cyclic redundancy check (CRC) error detection and a source-channel column code, i.e., systematic rate-compatible Reed-Solomon (RS) style erasure code. For a given budget on the available bandwidth and total transmit power, the transmitter determines the optimal power allocations and the optimal transmission rates among multiple CDMA channels, as well as the optimal product channel code rate allocation, i.e., the optimal unequal Reed-Solomon code source/parity rate allocations and the optimal RCPC rate protection for each channel. In formulating such an optimization problem, we make use of results on the large-system CDMA performance for various multiuser receivers in multipath fading channels. The channel is modeled as the concatenation of wireless BER channel and a wireline packet erasure channel with a fixed packet loss probability. By solving the optimization problem, we obtain the optimal power level allocation and the optimal transmission rate allocation over multiple CDMA channels. For each CDMA channel, we also employ a fast joint source-channel coding algorithm to obtain the optimal product channel code structure. Simulation results show that the proposed framework allows the video quality to degrade gracefully as the fading worsens or the bandwidth decreases, and it offers improved video quality at the receiver.
Shengjie Zhao 0001, Zixiang Xiong, Xiaodong Wang 0001
IEEE Trans. Mob. Comput.2
2005 Audio coding and image denoising based on the nonuniform modulated complex lapped transform
abstract
Xiong and Malvar recently introduced a nonuniform modulated complex lapped transform (NMCLT) with good time-localization and controllable frequency resolution by using an oversampled nonuniform filter bank to generate its real and the imaginary components. In this paper, we first show that oversampling in the NMCLT is not necessary in theory but a by-product of fast implementation in practice. We also point out that the amount of oversampling, which can be flexibly controlled, depends on the application. We then describe in detail the implementation of the inverse transform, which was not addressed clearly by Xiong and Malvar. We present the first applications of the NMCLT to audio coding and image denoising. A scalable audio coder has been implemented by controlling the amount of oversampling and exploiting redundancy among the NMCLT coefficients via predictive coding. Experimental results show that the audio coder reduces pre-echoes and improves the sound quality of audio clips with transient sounds. A simple denoising algorithm based on the NMCLT has also been devised to provide images with better visual quality than those obtained with wavelet-based soft thresholding.
Samuel Cheng 0001, Zixiang Xiong
IEEE Trans. Multim.2
2004 Successive Refinement for the Wyner-Ziv Problem and Layered Code Design
abstract
This paper focuses on successive refinement for the Wyner-Ziv problem and layered code refinement coding scheme for the Wyner-Ziv problem consists of multistage encoders and decoders where each decoder uses all the information generated from previous encoding stages and the side information, which could be different from stage to stage. On extending successive refinability from jointly Gaussian source to the more general type of sources, the difference between the source and the side information is Gaussian and independent of the side information. A layered (successive) coding scheme using nested scalar quantization and Slepian-Wolf coding of bit planes based on LDPC codes is also described.
Samuel Cheng 0001, Zixiang Xiong
Data Compression Conference2
2004 Slepian-Wolf Coding of Multiple M-ary Sources Using LDPC Codes
abstract
This paper presents a Slepian-Wolf coding of n correlated m-ary sources for LDPC codes. On applying the syndrome concept, multilevel codes with low-density parity-check (LDPC) codes can be used to approach the Slepian-Wolf limit at each level. The advantage of LDPC codes is that they can be designed for different correlation models between source outputs and the side information and approach the Slepian-Wolf limits. Specifically, for Slepian-Wolf coding of three sources, a design rule of rates for coding each source, which facilitates code design and allows multistage decoding was proposed.
Chingfu Lan, Angelos D. Liveris, Krishna Narayanan 0001, Zixiang Xiong, Costas N. Georghiades
Data Compression Conference4
2004 Slepian-Wolf Coded Nested Quantization (SWC-NQ) for Wyner-Ziv Coding: Performance Analysis and Code Design
abstract
This paper examines the high-rate performance of low-dimensional nested lattice quantizers for the quadratic Gaussian Wyner-Ziv problem, using a pair of nested lattices with the same dimensionality. As the rate increases, the gap increases between the performances of low dimensional nested lattice quantizers and the Wyner-Ziv rate-distortion function. This gap is due to the relatively weak channel coding component (or coarse lattice) in the nested lattice pair. To enhance the lattice channel code and boost the overall performance, Slepian-Wolf coding is applied to the quantization indices to achieve further compression. Thereby a Wyner-Ziv coding paradigm is introduced using Slepian-Wolf coded nested lattice quantization (SWC-NQ). Theoretical analysis and simulation results show that, for the quadratic Gaussian source and at high rate, SWC-NQ performs the same as traditional entropy-constrained lattice quantization with side information available at both the encoder and decoder.
Zhixin Liu 0007, Samuel Cheng 0001, Angelos D. Liveris, Zixiang Xiong
Data Compression Conference4
2004 Design of Slepian-Wolf Codes by Channel Code Partitioning
abstract
A Slepian-Wolf coding scheme that can achieve arbitrary rate allocation among two encoders was outlined in the work of Pradhan and Ramchandran. Inspired by this work, we start with a detailed solution for general (asymmetric or symmetric) Slepian-Wolf coding based on partitioning a single systematic channel code, and continue with practical code designs using advanced channel codes. By using systematic IRA and turbo codes, we devise a powerful scheme that is capable of approaching any point on the Slepian-Wolf bound. We further study an extension of the technique to multiple sources, and show that for a particular correlation model among the sources, a single practical channel code can be designed for coding all the sources in symmetric and asymmetric scenarios. If the code approaches the capacity of the channel that models the correlation between the sources, then the system will approach the Slepian-Wolf limit. Using systematic IRA and punctured turbo codes for coding two binary sources, each being independent identically distributed, with correlation modeled by a binary symmetric channel, we obtain results which are 0.04 bits away from the theoretical limit in both symmetric and asymmetric Slepian-Wolf settings.
Vladimir Stankovic 0001, Angelos D. Liveris, Zixiang Xiong, Costas N. Georghiades
Data Compression Conference3
2004 Packet Erasure Protection for Multicasting
abstract
Priority encoding transmission is an efficient forward error correction system for the robust transmission of scalable image and video data over packet erasure networks. For a memoryless packet erasure channel we first study the sensitivity of an optimal protection solution to a change in the packet erasure rate and in the number of packets. We then propose a practical error protection algorithm for multicasting and broadcasting applications. Instead of computing an optimal solution for each client independently, we show that comparable results can be obtained much faster by refining protections already computed for other clients. We also consider the situation where clients share a bottleneck link and develop layered multiple description codes that provide a better quality trade-off among all clients than previous solutions.
Vladimir Stankovic 0001, Zixiang Xiong
Data Compression Conference2
2004 Asymmetric Code Design for Remote Multiterminal Source Coding
abstract
Asymmetric code design for remote multiterminal source coding in the quadratic Gaussian case is presented in this paper. For remote multiterminal source coding of X, to achieve the minimum sum rate of the two independent encoders subject to a fidelity criterion d, theoretical bounds were derived independently. The main idea is to quantize the first observation Y/sub 1/ and apply Wyner-Ziv coding on Y/sub 2/ by using the quantized version of Y/sub 1/ as side information in an efficient asymmetric coding scheme. The practical code design gives results that are very close to the sum-rate bound.
Yang Yang 0003, Vladimir Stankovic 0001, Zixiang Xiong, Wei Zhao 0001
Data Compression Conference3
2004 PAGER: A Distributed Algorithm for the Dead-end Problem of Location-based Routing in Sensor Networks
abstract
The dead-end problem is an importance issue of location-based routing in sensor networks, which occurs when a message falls into a local minimum using greedy forwarding. Current methods for this problem are insufficient either in eliminating traffic/path memorization or finding satisfied short paths. We propose a novel algorithm, named partial-partition avoiding geographic routing (PAGER), to solve the problem. The basic idea of PAGER is to divide a sensor network graph into functional sub-graphs, and provide each sensor node with message forwarding directions based on these sub-graphs. That results in loop-free short paths without memorization of traffics/paths in sensor nodes. We implement our algorithm in a protocol and evaluate it in sensor networks with different parameters. Results show that PAGER generates considerably shorter paths, higher delivery ratio and lower energy consumption than the greedy perimeter stateless routing protocol. At the same time, PAGER achieves better performance in handling large-scale networks than the ad-hoc on-demand distance vector protocol.
Le Zou, Mi Lu, Zixiang Xiong
ICCCN3
2004 Code design for lossless multiterminal networks
abstract
This paper considers a general multiterminal (MT) system, which consists of L encoders and P decoders. Let X/sub 1/,..., X/sub L/ be memoryless, uniform, correlated random binary vectors of length n, and let x/sub 1/,..., x/sub L/ denote their realizations. Let further /spl Sigma/ = {1,...,L}. The i-th encoder compresses X/sub i/ independently from other encoders. The j-th decoder receives the bitstreams from a set of encoders /spl Sigma//sub j//spl sube/ /spl Sigma/ and jointly decodes them. It should reconstruct the received source messages with arbitrarily small probability of error. To construct a practical coding scheme for this network, we exploit the fact that such a network can be split into P subnetworks, each being regarded as a Slepian-Wolf (SW) coding system with multiple sources. This SW subnetwork consists of a decoder which receives encodings of all X/sub k/'s such that k/spl isin//spl Sigma//sub sw//spl sube//spl Sigma/ and attempts to reconstruct them perfectly. Based on (V. Stankovic et al. 2004), we first provide a code design for this setting, and then extend it to the general case.
Vladimir Stankovic 0001, Angelos D. Liveris, Zixiang Xiong, Costas N. Georghiades
ISIT3
2004 Channel symmetry in Slepian-Wolf code design based on LDPC codes with application to the quadratic Gaussian Wyner-Ziv problem
abstract
We consider Slepian-Wolf code design (or source coding with side information) based on LDPC codes. We show that density evolution defined in conventional channel coding can be used in analyzing the Slepian-Wolf coding performance provided that a certain symmetry condition, dubbed dual symmetry, is satisfied by the hypothetical channel between the source and the side information. Exploiting such an analysis. we design an efficient LDPC code based Slepian-Wolf coding scheme and apply it to the quadratic Gaussian Wyner-Ziv problems.
Samuel Cheng 0001, Zixiang Xiong
ITW2
2004 Source-channel coding for algebraic multiterminal binning
abstract
This paper addresses practical code design problems for multiterminal communication networks. The basic element of a multiterminal network code is the binning scheme. We aim to develop a unified practical code design paradigm for related problems in multiterminal networks based on source-channel coding for algebraic binning. First, a framework based on Slepian-Wolf coded quantization is highlighted for Wyner-Ziv coding and multiterminal source coding (e.g., the CEO problem). Then, a nested turbo scheme is proposed for Costa coding by exploiting the duality between Costa coding and Wyner-Ziv coding.
Zixiang Xiong, Vladimir Stankovic 0001, Samuel Cheng 0001, Angelos D. Liveris
ITW1
2004 Exploiting temporal correlation with adaptive block-size motion alignment for 3D wavelet coding
abstract
This paper proposes an adaptive block-size motion alignment technique in 3D wavelet coding to further exploit temporal correlations across pictures. Similar to B picture in traditional video coding, each macroblock can motion align from forward and/or backward for temporal wavelet de-composition. In each direction, a macroblock may select its partition from one of seven modes - 16x16, 8x16, 16x8, 8x8, 8x4, 4x8 and 4x4 - to allow accurate motion alignment. Furthermore, the rate-distortion optimization criterions are proposed to select motion mode, motion vectors and partition mode. Although the proposed technique greatly improves the accuracy of motion alignment, it does not directly bring the coding efficiency gain because of smaller block size and more block boundaries. Therefore, an overlapped block motion alignment is further proposed to cope with block boundaries and to suppress spatial high-frequency components. The experimental results show the proposed adaptive block-size motion alignment with the overlapped block motion alignment can achieve up to 1.0 dB gain in 3D wavelet video coding. Our 3D wavelet coder outperforms the MC-EZBC for most sequences by 1~2dB and we are doing up to 1.5 dB better than H.264.
Ruiqin Xiong, Feng Wu 0001, Shipeng Li 0001, Zixiang Xiong, Ya-Qin Zhang
VCIP4
2004 Layered Wyner-Ziv video coding
abstract
Wyner-Ziv coding refers to lossy source coding with side information at the decoder. Recently some practical applications of Wyner-Ziv coding to video compression have been studied due to its advantage of error robustness over standard video coding standards. Based on recent theoretical result on successive Wyner-Ziv coding, we propose in this paper a practical layered Wyner-Ziv video codec using the DCT, nested scalar quantizer (NSQ), and irregular LDPC code based Slepian-Wolf coding (or lossless source coding with side information). The DCT is applied as an approximation to the conditional KLT, which makes the components of the transformed block conditionally independent given the side information. NSQ is a binning scheme that facilitates layered bit-plane coding of the bin indices while reducing the bit rate. LDPC code based Slepian-Wolf coding exploits the correlation between the quantized version of the source and the side information to achieve further compression. Different from previous works, an attractive feature of our proposed system is that video encoding is done only once but decoding allowed at many lower bit rates without quality loss.
Qian Xu 0001, Zixiang Xiong
VCIP2
2004 Advanced motion threading for 3D wavelet video coding
Lin Luo 0004, Feng Wu 0001, Shipeng Li 0001, Zixiang Xiong, Zhenquan Zhuang
Signal Process. Image Commun.4
2004 Iterative decoding of differentially space-time coded multiple descriptions of images
abstract
We consider transporting images over wireless fading channels rather than the on-off channels considered in most works related to multiple description (MD) coding. The new idea is to use iterative (turbo) decoding techniques developed for serially concatenated coding systems to improve the performance of the receiver in successive decoding iterations. We treat MD coding, realized by embedded image coding plus unequal error protection (e.g., product code structure), as the outer constituent code and use a differential space-time code as the inner constituent code. Experimental results show that our iterative scheme can effectively improve the system performance. Furthermore, most of the gain in PSNR is achieved with only one iteration in the low SNR range, e.g., 0-20 dB.
Zixiang Xiong, Xiaodong Wang 0001
IEEE Signal Process. Lett.2
2004 Scalable image and video transmission using irregular repeat-accumulate codes with fast algorithm for optimal unequal error protection
abstract
This paper considers designing and applying punctured irregular repeat-accumulate (IRA) codes for scalable image and video transmission over binary symmetric channels. IRA codes of different rates are obtained by puncturing the parity bits of a mother IRA code, which uses a systematic encoder. One of the main ideas presented here is the design of the mother code such that the entire set of higher rate codes obtained by puncturing are good. To find a good unequal error protection for embedded bit streams, we employ the fast joint source-channel coding algorithm in Hamzaoui et al. to minimize the expected end-to-end distortion. We test with two scalable image coders (SPIHT and JPEG-2000) and two scalable video coders (3-D SPIHT and H.26L-based PFGS). Simulations show better results with IRA codes than those reported in Banister et al. with JPEG-2000 and turbo codes. The IRA codes proposed here also have lower decoding complexity than the turbo codes used by Banister et al.
Chingfu Lan, Tianli Chu, Krishna Narayanan 0001, Zixiang Xiong
IEEE Trans. Commun.4
2004 Real-time error protection of embedded codes for packet erasure and fading channels
abstract
Reliable real-time transmission of packetized embedded multimedia data over noisy channels requires the design of fast error control algorithms. For packet erasure channels, efficient forward error correction is obtained by using systematic Reed-Solomon (RS) codes across packets. For fading channels, state-of-the-art performance is given by a product channel code where each column code is an RS code and each row code is a concatenation of an outer cyclic redundancy check code and an inner rate-compatible punctured convolutional code. For each of these two systems, we propose a low-memory linear-time iterative improvement algorithm to compute an error protection solution. Experimental results for the two-dimensional and three-dimensional set partitioning in hierarchical trees coders showed that our algorithms provide close to optimal average peak signal-to-noise ratio performance, and that their running time is significantly lower than that of all previously proposed solutions.
Vladimir Stankovic 0001, Raouf Hamzaoui, Zixiang Xiong
IEEE Trans. Circuits Syst. Video Technol.3
2004 Efficient channel code rate selection algorithms for forward error correction of packetized multimedia bitstreams in varying channels
abstract
We study joint source-channel coding systems for the transmission of images over varying channels without feedback. We consider the situation where the channel statistics are unknown to the transmitter and focus on systems that enable good performance over a wide range of channel conditions. We first propose a linear-time channel code rate selection algorithm for a hybrid transmission system that combines packetization of an embedded wavelet bitstream into independently decodable packets and forward error correction with a concatenated cyclic redundancy check/rate-compatible punctured convolutional (RCPC) channel coder. We then consider an extension of this hybrid system with additional Reed-Solomon (RS) coding across the packets and give a linear-time algorithm for the efficient selection of both the RS and RCPC code rates. Experimental results for a wireline/wireless link modeled as the combination of a packet erasure channel and a Rayleigh flat-fading channel showed that our schemes significantly outperformed the best previous forward error correction systems in many situations where the actual channel parameter values deviated from the ones used in the optimization of the source-channel rate allocation.
Vladimir Stankovic 0001, Raouf Hamzaoui, Zixiang Xiong
IEEE Trans. Multim.3
2004 Wavelet-based VBR video traffic smoothing
abstract
In a typical video application, such as video-on-demand, videos are continuously streamed from a video server to a distributed set of receivers. The constant-quality video compression technique commonly used, variable bit rate (VBR) encoding, produces flows with multiple time-scale rate variability, so smoothing the VBR video traffic within an entire distribution tree presents a challenging task. This paper proposes a novel wavelet-based traffic smoothing (WTS) algorithm. Unlike existing algorithms, the WTS algorithm considers traffic smoothing at multiple resolutions. It results in a pruned version of a full tree, which corresponds to the original VBR traffic. Theoretical analysis and numerical evaluation demonstrate that: 1) WTS performs well across several metrics in smoothing bursty traffic and 2) for a video bit stream with N frames, the computational complexity of WTS is O(NlogN).
Dejian Ye, J. C. Barker, Zixiang Xiong, Wenwu Zhu 0001
IEEE Trans. Multim.3
2004 Error robust scalable audio streaming over wireless IP networks
abstract
Streaming high-fidelity audio over wireless Internet protocol (IP) networks is a challenging task because the networks present not only packet losses, but also residual bit errors. These losses and errors have severe adverse effect on the compressed audio bitstream. To solve this problem, this paper introduces error resilience in conjunction with error protection for scalable audio streaming over wireless networks. Specifically, error resilience is achieved by performing bitstream data partitioning and reversible variable length coding in the audio coder. Error protection is provided by layered product channel code to simultaneously handle packet losses and residual bit errors. Both the row and column codes of the product code provide unequal error protection for different layers of the audio bitstream by considering the characteristics of the scalable audio. Rate-distortion optimization is performed to determine the best source-channel coding tradeoff that minimizes the average expected end-to-end distortion. Simulation results demonstrate the effectiveness of our proposed approach.
Qian Zhang 0001, Guijin Wang, Zixiang Xiong, Jianping Zhou 0001, Wenwu Zhu 0001
IEEE Trans. Multim.3
2003 Distributed Compression of Binary Sources Using Conventional Parallel and Serial Concatenated Convolutional Codes
abstract
It is shown how conventional parallel (turbo) and serial concatenated convolutional codes can be used to compress close to the Slepian-Wolf limit for the correlated binary sources. Conventional refers to codes already used in channel coding. Focusing on the asymmetric case of compression of an equipolarable memoryless binary source with side information at the decoder, the approach is based on modeling the correlation as a channel and using syndromes. The encoding and decoding procedures are explained in detail. The performance achieved is seen to be better than the recently published results using nonconventional turbo codes and close to the Slepian-Wolf limit.
Angelos D. Liveris, Zixiang Xiong, Costas N. Georghiades
DCC2
2003 Optimal error protection for progressive image transmission over finite-state Markov channels
abstract
This paper presents a unified framework for addressing progressive image transmission over both memoryless and fading channels based on finite-state Markov channels (FSMC) models. The main advantage of using FSMC models is that they allow analytical derivation of optimal joint source-channel coding solutions in the form of unequal error protection without the burden of simulating wireless channels. Our analyses and experiments confirm that 1) FSMC models work for fading channels and 2) PSNR results obtained with FSMC models compare favorably with those published in the literature.
Zhongmin Liu, Minyi Zhao, Zixiang Xiong
ICC3
2003 Fast forward error protection of packetized multimedia bitstreams for transmission over varying channels
abstract
We propose a real-time optimization algorithm that selects an appropriate channel code for hybrid systems that combine packetization of an embedded wavelet bitstream into independently decodable packets and forward error correction using a family of channel codes with error detection and error correction capability. Such systems are very powerful for the transmission of audio, images, and video over fading and erasure channels with varying statistics. We also give an implementation that uses an optimal packetization technique and a concatenated cyclic redundancy check/rate-compatible punctured convolutional coder. Experimental results show that the peak signal-to-noise ratio of the average mean square error of our system is up to 1.74 dB higher than that of the previous best hybrid system for a Rayleigh fading channel and a transmission rate of 0.25 bits per pixel. Finally, we compare the hybrid approach to a state-of-the-art approach that uses a product code to protect the information bitstream.
Vladimir Stankovic 0001, Raouf Hamzaoui, Zixiang Xiong
ICC3
2003 Microarray BASICA: background adjustment, segmentation, image compression and analysis of microarray images
abstract
This paper presents Microarray BASICA: an integrated image processing tool for background adjustment, segmentation, image compression and analysis of microarray images. BASICA uses the fast Mann-Whitney test-based algorithm introduced in J. Hua et al., (2002) to segment microarray images, and post-processing to eliminate the segmentation irregularities. The segmentation results, along with the foreground and background intensities obtained with background adjustment, are then used for the independent compression of foreground and background. We introduce a new distortion measure for microarray image compression and devise a coding scheme by modifying the object-based embedded block coding with optimized truncation (object-based EBCOT) algorithm J. Hua et al., (2002) to achieve optimal rate-distortion performance in lossy coding while still maintaining outstanding lossless compression performance. Experimental results show that BASICA can extract sufficiently accurate genetic information at bitrates as low as 43bpp.
Jianping Hua, Zhongmin Liu, Zixiang Xiong, Qiang Wu 0007, Kenneth R. Castleman
ICIP (1)3
2003 Nested convolutional/turbo codes for the binary Wyner-Ziv problem
abstract
We show how concatenated (convolutional) codes can be used to compress close to the Wyner-Ziv limit for binary sources. Focusing on the case of lossy compression of an equiprobable memoryless binary source with side information at the decoder, the approach is based on nested binary linear codes, which is the extension of Wyner's lossless compression scheme to the lossy case proposed by Shamai, Verdu and Zamir. Based on our previous work on lossless compression with concatenated codes, we are able to combine the only two previously suggested nested schemes into a novel turbo scheme with improved performance. Our scheme can come within 0.09 bits from the theoretical limit, which to our knowledge is the first result ever reported for the binary Wyner-Ziv problem.
Angelos D. Liveris, Zixiang Xiong, Costas N. Georghiades
ICIP (1)2
2003 Product code error protection of packetized multimedia bitstreams
abstract
Sherwood and Zeger (1997) proposed a source-channel coding system where the source code is an embedded bitstream and the channel code is a product code such that each row code is a concatenation of a cyclic redundancy check (CRC) and rate-compatible punctured convolutional codes (RCPC) and the column codes are Reed-Solomon (RS) codes. We improve this system for wireless applications by efficiently reorganizing the source code into a set of independently decodable packets, which makes it more robust in varying channels. We also give a linear-time algorithm for finding an optimal equal error protection for the resulting system. Experimental results show that the performance of our system significantly outperforms that of the current state-of-the-art in fading channels with varying statistics.
Vladimir Stankovic 0001, Raouf Hamzaoui, Zixiang Xiong
ICIP (1)3
2003 Generation of sports highlights using motion activity in combination with a common audio feature extraction framework
abstract
In our past work we have used temporal patterns of motion activity to extract sports highlights. We have also used audio classification based approaches to develop a common audio-based platform for feature extraction that works across three different sports. In this paper, we combine the two aforementioned complementary approaches so as to get higher accuracy. We propose a framework for mining the semantic audio-visual labels in order to detect "interesting" events. Our results show that the proposed techniques work well across our three sports of interest, soccer, golf and baseball.
Zixiang Xiong, Regunathan Radhakrishnan, Ajay Divakaran
ICIP (1)1
2003 Optimal rate allocation in progressive joint source-channel coding for image transmission over CDMA networks
abstract
This paper presents an optimal rate allocation scheme in joint source-channel coding (JSCC) for image transmission over CDMA networks. Our scheme first uses progressive joint source-channel codes [V. Stankovic et al., Sep. 2002] to generate operational transmission rate-distortion (TR-D) functions of different multiple-access channels with different BERs. The Lagrange multiplier is then employed to set the optimal rate allocation among different channels, subject to a total transmission rate constraint. Progressive JSCC offers the scalability feature that is very attractive for handling bandwidth variations in wireless communications. By extending the rate-distortion function in source coding to the TR-D function in JSCC, our work borrows the "equal slope" argument [Y. Shoham and A. Gersho, Sep. 1988] to effectively solve the rate allocation problem in JSCC for multi-channel image transmission. Experiments show that our scheme runs in real time and provides competitive performance at all intermediate transmission rates.
Jianping Hua, Zixiang Xiong
ICME2
2003 Iterative decoding of differentially space-time coded multiple descriptions of images
abstract
We consider transporting images over wireless fading channels rather than the on-off channels considered in most works related to multiple description (MD) coding. The idea is to treat MD coding as a special case of distributed source coding with co-located sources, and use iterative (turbo) decoding techniques developed for serially concatenated coding systems to improve the performance of the receiver with successive source decoding iterations. MDs of an input image are generated by embedded coding plus unequal error protection. The proposed receiver consists of a maximum a posterior (MAP) differential space-time decoder, an MD source decoder, an interleaver and a deinterleaver. The two decoders exchange extrinsic values or a priori probabilities of transmitted bits between themselves in successive iterations. Experimental results show that our iterative scheme can effectively improve the system performance. Furthermore, most of the gain in PSNR is achieved with only two iterations.
Zixiang Xiong, Xiaodong Wang 0001
ICME2
2003 Optimal resource allocation for wireless video over CDMA networks
abstract
We present a multiple-channel video transmission scheme in wireless CDMA networks over multipath fading channels. We map an embedded video bitstream, which is encoded into multiple independently decodable layers by 3D-ESCOT video coding techniques, to multiple CDMA channels. Each video source layer is protected by a product channel code structure. For a given budget on the available bandwidth and total transmit power, the transmitter determines the optimal power allocations and the optimal transmission rates among multiple CDMA channels, as well as the optimal product channel code rate allocation. We make use of results on the large-system CDMA performance for various multiuser receivers in multipath fading channels. Simulation results show that the proposed framework allows the video quality to degrade gracefully as the fading worsens or the bandwidth decreases, and it offers improved video quality at the receiver.
Shengjie Zhao 0001, Zixiang Xiong, Xiaodong Wang 0001
ICME2
2003 Real-time unequal error protection for distortion-optimal progressive image transmission
abstract
For optimal progressive transmission of an embedded image code over a noisy channel, we consider an unequal error protection strategy that minimizes the average of the expected distortion over a set of intermediate rates. In contrast to previous work, we find a near-optimal solution in real-time. For a binary symmetric channel, two state-of-the-art source coders (SPIHT and JPEG200), and a rate-compatible punctured turbo coder as a channel coder, we compare our solution to the strategy that optimizes the end-to-end performance.
Vladimir Stankovic 0001, Youssef Charfi, Raouf Hamzaoui, Zixiang Xiong
WCNC4
2003 Real-time unequal error protection algorithms for progressive image transmission
abstract
We consider unequal error protection strategies for the efficient progressive transmission of embedded image codes over noisy channels. In progressive transmission, the reconstruction quality is important not only at the target transmission rate but also at the intermediate rates. An adequate error protection strategy may, thus, consist of optimizing the average performance over the set of intermediate rates. The performance can be the expected number of correctly decoded source bits or the expected distortion. For the rate-based performance, we prove some interesting properties of an optimal solution and give an optimal linear-time algorithm to compute it. For the distortion-based performance, we propose an efficient linear-time local search algorithm. For a binary symmetric channel, two state-of-the-art source coders (SPIHT and JPEG2000), we compare the progressive ability of our proposed solutions to that of the strategies that optimize the end-to-end performance of the system. Experimental results showed that the proposed solutions had a slightly worse performance at the target transmission rate and a better performance at most of the intermediate rates, especially at the lowest ones.
Vladimir Stankovic 0001, Raouf Hamzaoui, Youssef Charfi, Zixiang Xiong
IEEE J. Sel. Areas Commun.4
2003 Chromosome Image Enhancement Using Multiscale Differential Operators
abstract
Chromosome banding patterns are very important features for karyotyping, based on which cytogenetic diagnosis procedures are conducted. Due to cell culture, staining, and imaging conditions, image enhancement is a desirable preprocessing step before performing chromosome classification. In this paper, we apply a family of differential wavelet transforms (Wang and Lee, 1998), (Wang, 1999) for this purpose. The proposed differential filters facilitate the extraction of multiscale geometric features of chromosome images. Moreover, desirable fast computation can be realized. We study the behavior of both banding edge pattern and noise in the wavelet transform domain. Based on the fact that image geometrical features like edges are correlated across different scales in the wavelet representation, a multiscale point-wise product (MPP) is used to characterize the correlation of the image features in the scale-space. A novel algorithm is proposed for the enhancement of banding patterns in a chromosome image. In order to compare objectively the performance of the proposed algorithm against several existing image-enhancement techniques, a quantitative criteria, the contrast improvement ratio (CIR), has been adopted to evaluate the enhancement results. The experimental results indicate that the proposed method consistently outperforms existing techniques in terms of the CIR measure, as well as in visual effect. The effect of enhancement on cytogenetic diagnosis is further investigated by classification tests conducted prior to and following the chromosome image enhancement. In comparison with conventional techniques, the proposed method leads to better classification results, thereby benefiting the subsequent cytogenetic diagnosis.
Qiang Wu 0007, Kenneth R. Castleman, Zixiang Xiong
IEEE Trans. Medical Imaging4
2003 Lossy-to-Lossless Compression of Medical Volumetric Data Using Three-dimensional Integer Wavelet Transforms
abstract
We study lossy-to-lossless compression of medical volumetric data using three-dimensional (3-D) integer wavelet transforms. To achieve good lossy coding performance, it is important to have transforms that are unitary. In addition to the lifting approach, we first introduce a general 3-D integer wavelet packet transform structure that allows implicit bit shifting of wavelet coefficients to approximate a 3-D unitary transformation. We then focus on context modeling for efficient arithmetic coding of wavelet coefficients. Two state-of-the-art 3-D wavelet video coding techniques, namely, 3-D set partitioning in hierarchical trees (Kim et al., 2000) and 3-D embedded subband coding with optimal truncation (Xu et al., 2001), are modified and applied to compression of medical volumetric data, achieving the best performance published so far in the literature-both in terms of lossy and lossless compression.
Zixiang Xiong, Xiaolin Wu 0001, Samuel Cheng 0001, Jianping Hua
IEEE Trans. Medical Imaging1
2002 Compression of binary sources with side information using low-density parity-check codes
abstract
It is shown how low-density parity-check (LDPC) codes can be used as an application of the Slepian-Wolf (1973) theorem for correlated binary sources. We focus on the asymmetric case of compression with side information. The approach is based on viewing the correlation as a channel and applying the syndrome concept. The encoding and decoding procedures, i.e. the compression and decompression, are explained in detail. The simulated performance results are better than most of the existing turbo code results available in the literature and very close to the Slepian-Wolf limit.
Angelos D. Liveris, Zixiang Xiong, Costas N. Georghiades
GLOBECOM2
2002 Scalable image transmission over differentially space-time coded OFDM systems
abstract
This paper combines joint source-channel coding, differential space-time coding and OFDM for scalable image transmission over differentially space-time coded OFDM (DSTC-OFDM) systems. The SNR versus BER performance of these systems are measured by simulation. We consider transmitting embedded source bitstreams over such systems using product channel codes that are designed to combat both bit errors and fading. For a given average SNR (hence BER), a fast joint source-channel coding algorithm is used to optimize the product code rates. Experiments show that the diversity gain in our DSTC-OFDM system can be translated into quality gains (in PSNR) in transmitted images.
Zixiang Xiong, Xiaodong Wang 0001
GLOBECOM2
2002 Enhanced spread spectrum watermarking of MPEG-2 AAC audio
abstract
A novel AAC watermarking scheme with low structural and computational complexity is proposed to achieve real-time AAC audio watermark embedding capability. In this scheme, an enhanced spread spectrum watermarking method that increases the robustness and information hiding capacity of the watermarking channel is developed and employed. Furthermore, this scheme improves the speed of the encoding process in that we directly use the quantized indices rather than the dequantized coefficients, thus no explicit dequantization and requantization are performed. Our embedded watermark is tested to be robust against AAC/MP3 transcoding at a rate of 10 bps. The scheme is suitable for speedy encoding and decoding of watermark with low cost hardware.
Samuel Cheng 0001, Hong Heather Yu, Zixiang Xiong
ICASSP3
2002 A distributed source coding technique for highly correlated images using turbo-codes
abstract
According to the Slepian-Wolf theorem [1], the output of two correlated sources can be compressed to the same extent without loss, no matter if they communicate with each other or not, provided that the decompression takes place at a common decoder having both compressed outputs available. In this paper, as an application of the Slepian-Wolf theorem, an advanced distributed source coding scheme for correlated images is presented. Assuming that the correlated image is a noisy version of the original, the scheme involves modulo encoding of the pixel values and encoding (compression) of the resulting symbols with binary and nonbinary turbo-codes, so that rate savings are achieved practically without loss.
Angelos D. Liveris, Zixiang Xiong, Costas N. Georghiades
ICASSP2
2002 Joint error control and power allocation for video transmission over CDMA networks with multiuser detection
abstract
Error control and power allocation for transmitting wireless video over CDMA networks are considered in conjunction with multiuser detection. We map a layered video bitstream to several CDMA fading channels and inject multiple source/parity layers into each of these channels at the transmitter. At the receiver, we employ a linear minimum mean-square error (MMSE) multiuser detector in the uplink and two types of blind linear MMSE detectors in the downlink, for demodulating the received data. For given constraints on the available bandwidth and transmit power, the transmitter determines the optimal power allocation among different CDMA fading channels and the optimal number of source and parity packets to send that offer the best video quality. Simulation results show a performance gain of up to 1.5 dB with joint optimization over that with rate optimization only.
Shengjie Zhao 0001, Zixiang Xiong, Xiaodong Wang 0001
ICC2
2002 Scalable image transmission using rate-compatible irregular repeat accumulate (IRA) codes
abstract
This paper considers designing and applying rate-compatible irregular repeat accumulate (IRA) codes for scalable image transmission over binary symmetric channels. IRA codes of different rates are obtained by puncturing the parity bits of a mother IRA code which uses a systematic encoder. One of the main ideas presented here is the design of the mother code such that the entire set of higher rate codes obtained by puncturing are good. Using the Viterbi algorithm for finding the optimal unequal error protection of JPEG2000 bit streams, we present better results with IRA codes than those reported with turbo codes. The proposed IRA codes also have lower decoding complexity than the turbo codes.
Chingfu Lan, Krishna Narayanan 0001, Zixiang Xiong
ICIP (3)3
2002 On optimal subspaces for appearance-based object recognition
abstract
On the subject of optimal subspaces for appearance-based object recognition, it is generally believed that algorithms based on LDA (linear discriminant analysis) are superior to those based on PCA (principal components analysis), provided that relatively large training data sets are available. In this paper, we show that while this is generally true for classification with the nearest-neighbor classifier, it is not always the case with a maximum-likelihood classifier. We support our claim by presenting both intuitively plausible arguments and actual results on a large data set of human chromosomes. Our conjecture is that perhaps only when the underlying object classes are linearly separable would LDA be truly superior to other known subspaces of equal dimensionality.
Zhongmin Liu, Zixiang Xiong, Jie Chen 0008, Qiang Wu 0007, Kenneth R. Castleman
ICIP (3)2
2002 Packet loss protection of embedded data with fast local search
abstract
Unequal loss protection with systematic Reed-Solomon codes allows reliable transmission of embedded multimedia over packet erasure channels. The design of a fast algorithm with low memory requirements for the computation of an unequal loss protection solution is essential in real-time systems. Because the determination of an optimal solution is time-consuming, fast suboptimal solutions have been used. In this paper, we present a fast iterative improvement algorithm with negligible memory requirements. Experimental results for the JPEG2000, 2D, and 3D set partitioning in hierarchical trees (SPIHT) coders showed that our algorithm provided close to optimal peak signal-to-noise ratio (PSNR) performance, while its time complexity was significantly lower than that of all previously proposed algorithms.
Vladimir Stankovic 0001, Raouf Hamzaoui, Zixiang Xiong
ICIP (2)3
2002 Joint product code optimization for scalable multimedia transmission over wireless channels
abstract
State-of-the-art systems for the transmission of images over wireless channels generate an embedded bitstream and protect it with a product code where the row code is a concatenation of an outer cyclic redundancy check (CRC) code and an inner rate-compatible punctured convolutional (RCPC) code, and the column code is a Reed-Solomon (RS) code. In previous works, the product code was optimized by searching for the best RS protection for each RCPC code rate. We present a local search algorithm that jointly optimizes the RS and the RCPC codes. Experimental results show that our algorithm provides an approximately optimal solution, while its time complexity is much lower than that of the previous works.
Vladimir Stankovic 0001, Raouf Hamzaoui, Zixiang Xiong
ICME (1)3
2002 Optimal rate allocation for progressive fine granularity scalable video coding
abstract
We examine the enhancement-layer rate allocation problem in progressive fine granularity scalable (PFGS) video coding. The problem arises from the fact that different frames in the enhancement layer have different rates in PFGS coding. A rate-distortion (R-D) function for a multiframe group is first established for enhancement-layer PFGS coding, followed by experiments to verify its validity using real test sequences. Optimal rate allocation among frames in the group is then given based on the R-D function, together with a simple implementation that is suitable for applications such as streaming video. Experiments show that, compared with uniform bit allocation, optimal bit allocation not only makes the quality variation in decoded video much smoother, but also improves the average PSNR of PFGS coding by 0.3-0.5 dB.
Zixiang Xiong, Feng Wu 0001, Shipeng Li 0001
IEEE Signal Process. Lett.2
2002 Progressive trellis-coded space-frequency quantization for wavelet image coding
abstract
This paper addresses progressive wavelet image coding within the trellis-coded space-frequency quantization (TCSFQ) framework (Xiong et al., 1999). A method similar to that in Bilgin et al. (1999), is used to approximately invert TCSFQ when decoding at rates lower than the encoding rate. Our experiments show that the loss incurred for progressive transmission is within 1 dB in peak signal-to-noise ratio and that the progressive coding performance of TCSFQ is competitive with that of the celebrated SPIHT coder (Said et al., 1996) at all rates.
Pierre Seigneurbieux, Zixiang Xiong
IEEE Trans. Circuits Syst. Video Technol.2
2002 Memory-constrained 3D wavelet transform for video coding without boundary effects
abstract
Three-dimensional (3D) wavelet-based scalable video coding provides a viable alternative to standard MC-DCT coding. However, many current 3D wavelet coders experience severe boundary effects across group of pictures (GOP) boundaries. This paper proposes a memory-efficient transform technique via lifting that effectively computes wavelet transforms of a video sequence continuously on the fly, thus eliminating the boundary effects due to limited length of individual GOPs. Coding results show that the proposed scheme completely eliminates the boundary effects and gives superb video playback quality.
Jizheng Xu, Zixiang Xiong, Shipeng Li 0001, Ya-Qin Zhang
IEEE Trans. Circuits Syst. Video Technol.2
2002 Joint error control and power allocation for video transmission over CDMA networks with multiuser detection
abstract
Error control and power allocation for transmitting wireless video over CDMA networks are considered in conjunction with multiuser detection. We map a layered video bitstream to several CDMA fading channels and inject multiple source/parity layers into each of these channels at the transmitter. At the receiver, we employ a linear minimum mean-square error (MMSE) multiuser detector in the uplink and two types of blind linear MMSE detectors, i.e., the direct-matrix-inversion blind detector and the subspace blind detector, in the downlink, for demodulating the received data. For given constraints on the available bandwidth and transmit power, the transmitter determines the optimal power allocation among different CDMA fading channels and the optimal number of source and parity packets to send that offer the best video quality. We formulate a combined optimization problem and give the optimal joint rate and power allocation for each of these three receivers. Simulation results show a performance gain of up to 3.5 dB with joint optimization over with rate optimization only.
Shengjie Zhao 0001, Zixiang Xiong, Xiaodong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2001 Joint UEP and layered source coding with application to transmission of JPEG-2000 coded images
abstract
This paper presents a joint source-channel coding framework based on layered source coding and Reed-Solomon channel coding for unequal error protection. An iterative procedure is described to search for the best source coding rate and the optimal UEP of layered bitstreams. We apply our JSCC technique to transmission of JPEG-2000 coded images over binary symmetric channels. Compared to results reported in the literature, our UEP based approach gives better results while having lower complexity.
Tianli Chu, Zhongmin Liu, Zixiang Xiong, Xiaolin Wu 0001
GLOBECOM3
2001 Combined wavelet video coding and error control for Internet streaming and multicast
abstract
This paper proposes an integrated approach to Internet video streaming and multicast (e.g., receiver-driven layered multicast (RLM)) based on combined wavelet video coding and error control. We design a packetized wavelet video (PWV) coder to facilitate its integration with error control. The PWV coder produces packetized layered bitstreams that are independent among layers while being embedded within each layer. Thus a lost packet only renders the following packets in the same layer useless. Based on the PWV coder, we search for a multi-layered error control strategy that optimally trades off source and channel coding for each layer under a given transmission rate to mitigate the effects of packet loss. Our integrated approach extends the single-layered approach for RLM. Theoretical analysis shows a gain of up to one dB on a channel with 20% packet loss. This is also substantiated by our simulations with a gain of up to 2.2 dB.
Tianli Chu, Zixiang Xiong
GLOBECOM2
2001 Wavelet-based smoothing and multiplexing of VBR video traffic
abstract
Although VBR coding is more efficient than CBR coding, the burstiness of VBR video traffic brings great difficulty to network resource management. To address this problem, we propose a novel wavelet-based traffic smoothing (WTS) algorithm. Unlike existing algorithms, which only have one resolution, the WTS algorithm has a multiresolution property that is preferable for smoothing VBR video traffic that exhibits a self-similar behavior. WTS allows traffic smoothing at multiple resolutions and the best transmission schedule is searched as a pruned subtree of a full binary tree, which corresponds to the original VBR video traffic. WTS optimizes several metrics simultaneously for both the single and multiple flow cases while traditional algorithms only optimize one or two metrics for a single flow. The computational complexity of the WTS algorithm is also lower.
Dejian Ye, Zixiang Xiong, Huai-Rong Shao, Qiufeng Wu, Wenwu Zhu 0001
GLOBECOM2
2001 Scalable audio coding using the nonuniform modulated complex lapped transform
abstract
This paper introduces a scalable audio coder using the nonuniform modulated complex lapped transform (NMCLT), which is a new nonuniform oversampled filter bank with a better combination of time- and frequency-domain localization than previous designs. Masking functions for different critical Bark bands are first calculated directly from the NMCLT coefficients as perceptual weights and arithmetic coding is then used to compress bit planes of the weighted NMCLT coefficients to generate a perceptually scalable audio bitstream. The loss in coding performance due to oversampling is offset by limiting the amount of redundancy in the transform and exploiting the correlations among the NMCLT basis functions. Experiments show that our new coder outperforms a coder with the modulated lapped transform (MLT) both objectively and subjectively.
Anne-Sophie Scheuble, Zixiang Xiong
ICASSP2
2001 Image enhancement using multiscale differential operators
abstract
Differential operators have been widely used for multiscale geometric descriptions of images. Efficient computation of these differential operators can be obtained by taking advantage of the spline techniques. We make use of a special class of these operators for image enhancement, with a particular application to chromosome image enhancement. These operators constitute a translation invariant wavelet transform well suited for the structural description of chromosome geometry. Based on the fact that the geometrical features like edges are correlated between different scales in the representation, a novel algorithm is designed to enhance the salient features of the image. Comparisons of this algorithm with other approaches are presented.
Qiang Wu 0007, Kenneth R. Castleman, Zixiang Xiong
ICASSP4
2001 Progressive trellis-coded space-frequency quantization for wavelet image coding
abstract
This paper addresses progressive wavelet image coding within the trellis-coded space-frequency quantization (TCSFQ) framework (Xiong and Wu 1999). A method similar to that in Bilgin et al. (1999) is used to approximately invert TCSFQ when decoding at rates lower than the encoding rate. Our experiments show that the loss incurred for progressive coding is within one dB in PSNR and that the progressive coding performance of TCSFQ is competitive with that of the celebrated SPIHT coder (Said and Pearlman 1996) at all rates.
Pierre Seigneurbieux, Zixiang Xiong
ICIP (1)2
2001 Image enhancement using multiscale oriented wavelets
abstract
We describe a novel method of enhancing image geometric features using the oriented wavelets introduced in our earlier work (see Yu-Ping Wang, IEEE Trans. Image Proc., vol.8, no.12, p.1757-71, 1999). The poor directional selectivity of the conventional 2D wavelet transform has been circumvented by using this class of oriented wavelets. By taking advantage of directional wavelet decomposition, both the directional and scale correlation information are utilized for enhancing salient structures in images. To remove phase dependence, a pair of quadrature filters are used. The proposed algorithm has been applied to chromosome image enhancement. Comparisons of this algorithm with other approaches are presented.
Qiang Wu 0007, Kenneth R. Castleman, Zixiang Xiong
ICIP (1)4
2001 Compression of M-FISH images using 3-D ESCOT
abstract
This paper introduces a lossy to lossless coding technique for compression of multitarget fluorescence in situ hybridization (M-FISH) images using 3-D embedded subband coding with optimal truncation (3-D ESCOT) (Xu et al.). With a lifting-based integer wavelet decomposition, 3-D ESCOT achieves about twice as much compression as Lempel-Ziv (WinZip) coding-the current method for archiving M-FISH images. The lossy coding performance of 3-D ESCOT is significantly better than that of 2-D based JPEG-2000.
Jizheng Xu, Zixiang Xiong, Qiang Wu 0007, Shipeng Li 0001
ICIP (2)2
2001 Packetized wavelet video coding and error control for receiver-driven layered multicast: an integrated approach
abstract
This paper proposes an integrated approach to receiver-driven layered multicast (RLM) based on packetized wavelet video (PWV) coding and error control. The PWV coder produces packetized layered bitstreams that are independent among layers while being embedded within each layer. Error control determines the optimal joint source-channel coding trade-off given the available transmission rate to mitigate the effects of packet loss. The PWV coder is designed to facilitate its integration with error control for RLM such that packets from different source layers are generated in a rate-distortion optimized way and a lost packet only renders the following packets in the same layer useless. Our work extends the single-layered approach in Chou et al., (2001). Theoretical analysis shows a gain of up to 1 dB on a channel with 20% packet loss over results presented in Chou. This gain is also substantiated by our simulations with a gain of up to 2.2 dB.
Tianli Chu, Jianping Hua, Zixiang Xiong
MMSP3
2001 High-performance 3-D embedded wavelet video (EWV) coding
abstract
This paper presents a rate-distortion (R-D) optimized 3-D embedded wavelet video (EWV) coder by extending the concept of EBCOT from 2-D to 3-D. After a lifting based 3-D wavelet transform, different subbands are coded independently using bit plane coding with different context models to provide flexible scalability in both spatial and temporal domain. A global R-D optimization procedure is used to generate an embedded bitstream for a target bit rate. Experiments show that, even without motion estimation, the EWV coder outperforms both MPEG-4 and 3-D ESCOT for most low motion video sequences.
Jianping Hua, Zixiang Xiong, Xiaolin Wu 0001
MMSP2
2001 Error resilient scalable audio coding (ERSAC) for mobile applications
abstract
Delivering high-fidelity audio over wireless channels is a challenging task because the wireless channel presents not only erasure errors, but also random bit errors. These errors have severe adverse effect on decompressing the received audio bitstream and may crash the decoder completely if not handled properly. To solve this problem, this paper addresses error resilient scalable audio coding (ERSAC) for wireless audio. Specifically, we perform data partition and reversible variable length coding, in the scalable audio bitstream. Simulation results show that ERSAC has very effective error-resilience in addition to bitstream scalability.
Jianping Zhou 0001, Qian Zhang 0001, Zixiang Xiong, Wenwu Zhu 0001
MMSP3
2001 A nonuniform modulated complex lapped transform
abstract
A nonuniform modulated complex lapped transform (NMCLT) is introduced in this paper as a two-stage extension of the modulated complex lapped transform (MCLT). The NMCLT is a new nonuniform oversampled filter bank with a better combination of time- and frequency-domain localization than previous designs, Adaptive nonuniform subband decompositions can be easily generated by varying the number of coefficients brought to the second stage on a frame-by-frame basis. The NMCLT is ideally suited for processing wideband signals such as audio, with reduced time-spreading artifacts.
Zixiang Xiong, Henrique S. Malvar
IEEE Signal Process. Lett.1
2001 Progressive source coding for a power constrained Gaussian channel
abstract
We consider the progressive transmission of a lossy source across a power constrained Gaussian channel using binary phase-shift keying modulation. Under the theoretical assumptions of infinite bandwidth, arbitrarily complex channel coding, and lossless transmission, we derive the optimal channel code rate and the optimal energy allocation per transmitted bit. Under the practical assumptions of a low complexity class of algebraic channel codes and progressive image coding, we numerically optimize the choice of channel code rate and the energy per bit allocation. This model provides an additional degree of freedom with respect to previously proposed schemes, and can achieve a higher performance for sources such as images. It also allows one to control bandwidth expansion or reduction.
Marc P. C. Fossorier, Zixiang Xiong, Kenneth Zeger
IEEE Trans. Commun.2
2001 3-D wavelet coding of video with arbitrary regions of support
abstract
We examine 3-D wavelet coding of video with arbitrary regions of support (AROS). A critically sampled wavelet transform is applied to the AROS and a modified 3-D set partitioning in hierarchical trees (SPIHT) algorithm is used to quantize and code the wavelet coefficients in the AROS only. Experiments show that, for typical MPEG-4 pre-segmented sequences, our proposed method can achieve a gain of up to 5.6 dB in average PSNR at the same rate over 3-D SPIHT coding of regular volumes that embed the AROS of the given video sequences.
Gavin Minami, Zixiang Xiong, Albert Wang 0005, Sanjeev Mehrotra
IEEE Trans. Circuits Syst. Video Technol.2
2001 On packetization of embedded multimedia bitstreams
abstract
We study the problem of packetizing embedded multimedia bitstreams to improve the error resilience of source (compression) codes. This problem is important because of the increasing popularity of embedded compression methodology and its suitability for scalable streaming media over IP or/and mobile IP. We study various packetization schemes against packet erasure at both low and high bit rates. Maximizing packetization efficiency for embedded bitstreams is formulated as a discrete optimization problem and globally optimal packetization (OP) algorithms are proposed under different settings. Suboptimal packetization algorithms are also devised to reduce the complexity of the OP algorithms. In order to assess their effectiveness, the proposed packetization algorithms are used to packetize embedded image and video bitstreams with simulated packet loss. Experimental results show that our OP algorithms slightly outperforms suboptimal ones. In addition to confirming the superiority of the OP algorithms, these results also provide justification of heuristic packetization methods published in the literature.
Xiaolin Wu 0001, Samuel Cheng 0001, Zixiang Xiong
IEEE Trans. Multim.3
2000 Optimal Packetization of Embedded Bitstreams
abstract
Summary form only given. To achieve error resilience, digital communication systems typically partition a data file into blocks of samples in the time domain which represent small cohesive segments of the input source, be it image, video or audio. We consider the problem of packing a set of embedded bitstreams B/sub k/ into M packets of payload L. Optimal packetization is to select ML bits to fill in M packets while satisfying certain alignment constraints that are imposed by error resilience designs, and at the same time minimizing the distortion. Just as source and channel coding have conflicting objectives of removing and adding redundancy, error resilience via packetization will somewhat reduce the rate distortion performance of the compression code when the transmission is error free. Our goal is to minimize such losses of coding efficiency.
Xiaolin Wu 0001, Zixiang Xiong
Data Compression Conference2
2000 Object-based multiresolution watermarking of images and video
abstract
This paper proposes a new approach to digital watermarking of image and video objects with arbitrary regions of support (AROS) based on the 2D and 3D shape adaptive wavelet transforms. The hierarchical nature of the wavelet representation of objects allows detection of the digital watermark at various resolutions. We show that, when subjected to image/video compression, the corresponding watermark of objects can still be correctly identified at each resolution (excluding the lowest one) in the wavelet domain. Such a multiresolution watermarking scheme for objects has computational advantage, especially for the video case. Potential applications of our proposed scheme include watermark-based object searching and indexing.
Xiaoyun Wu, Wenwu Zhu 0001, Zixiang Xiong, Ya-Qin Zhang
ISCAS3
2000 Low bit-rate scalable video coding with 3-D set partitioning in hierarchical trees (3-D SPIHT)
abstract
We propose a low bit-rate embedded video coding scheme that utilizes a 3-D extension of the set partitioning in hierarchical trees (SPIHT) algorithm which has proved so successful in still image coding. Three-dimensional spatio-temporal orientation trees coupled with powerful SPIHT sorting and refinement renders 3-D SPIHT video coder so efficient that it provides comparable performance to H.263 objectively and subjectively when operated at the bit rates of 30 to 60 kbits/s with minimal system complexity. Extension to color-embedded video coding is accomplished without explicit bit allocation, and can be used for any color plane representation. In addition to being rate scalable, the proposed video coder allows multiresolutional scalability in encoding and decoding in both time and space from one bit stream. This added functionality along with many desirable attributes, such as full embeddedness for progressive transmission, precise rate control for constant bit-rate traffic, and low complexity for possible software-only video applications, makes the proposed video coder an attractive candidate for multimedia applications.
Beong-Jo Kim, Zixiang Xiong, William A. Pearlman
IEEE Trans. Circuits Syst. Video Technol.2
1999 Three-Dimensional Wavelet Coding of Video with Global Motion Compensation
abstract
Three-dimensional (2D+T) wavelet coding of video using SPIHT has been shown to outperform standard predictive video coders on complex high-motion sequences, and is competitive with standard predictive video coders on simple low-motion sequences. However, on a number of typical moderate-motion sequences characterized by largely rigid motions, 3D SPIHT performs several dB worse than motion-compensated predictive coders, because it is does not take advantage of the real physical motion underlying the scene. We introduce global motion compensation for 3D subband video coders, and find 0.5 to 2 dB gain on sequences with dominant background motion. Our approach is a hybrid of video coding based on sprites, or mosaics, and subband coding.
Albert Wang 0005, Zixiang Xiong, Philip A. Chou, Sanjeev Mehrotra
Data Compression Conference2
1999 Trellis Coded Color Quantization of Images
abstract
We examine color quantization of images using trellis coded quantization (TCQ). Together with a simple dithering scheme, an 8-bit trellis coded color quantizer reproduces images that are visually indistinguishable from the 24-bit originals. The proposed algorithm can be viewed as a predictive trellis coded color quantization scheme. It is universal in the sense that no training or lookup table is needed. The complexity of TCQ is linear with respect to image size, making trellis coded color quantization suitable for interactive graphics and a window-based display environment.
Samuel Cheng 0001, Zixiang Xiong, Jian Qiao Huang
ICIP (4)2
1999 Progressive Video Coding for Noisy Channels
Beong-Jo Kim, Zixiang Xiong, William A. Pearlman, Youngseop Kim
J. Vis. Commun. Image Represent.2
1999 Wavelet image coding using trellis coded space-frequency quantization
abstract
The progress in wavelet image coding have brought the field into its maturity. Major developments in the process are rate-distortion (R-D) based wavelet packet transformation, zerotree quantization, subband classification and trellis-coded quantization, and sophisticated context modeling in entropy coding. Drawing from past experience and recent in sights, we propose a new wavelet image coding technique with trellis coded space-frequency quantization (TCSFQ). TCSFQ aims to explore space-frequency characterizations of wavelet image representations via R-D optimized zerotree pruning, trellis-coded quantization, and context modeling in entropy coding. Experiments indicate that the TCSFQ coder achieves twice as much compression as the baseline JPEG coder does at the same peak signal to noise ratio (PSNR), making it better than all other coders described in the literature.
Zixiang Xiong, Xiaolin Wu 0001
IEEE Signal Process. Lett.1
1999 A comparative study of DCT- and wavelet-based image coding
abstract
We undertake a study of the performance difference of the discrete cosine transform (DCT) and the wavelet transform for both image and video coding, while comparing other aspects of the coding system on an equal footing based on the state-of-the-art coding techniques. The studies reveal that, for still images, the wavelet transform outperforms the DCT typically by the order of about 1 dB in peak signal-to-noise ratio. For video coding, the advantage of wavelet schemes is less obvious. We believe that the image and video compression algorithm should be addressed from the overall system viewpoint: quantization, entropy coding, and the complex interplay among elements of the coding system are more important than spending all the efforts on optimizing the transform.
Zixiang Xiong, Kannan Ramchandran, Michael T. Orchard, Ya-Qin Zhang
IEEE Trans. Circuits Syst. Video Technol.1
1999 Multiresolution watermarking for images and video
abstract
This paper proposes a unified approach to digital watermarking of images and video based on the two- and three-dimensional discrete wavelet transforms. The hierarchical nature of the wavelet representation allows multiresolutional detection of the digital watermark, which is a Gaussian distributed random vector added to all the high-pass bands in the wavelet domain. We show that when subjected to distortion from compression or image halftoning, the corresponding watermark can still be correctly identified at each resolution (excluding the lowest one) in the wavelet domain. Computational savings from such a multiresolution watermarking framework is obvious, especially for the video case.
Wenwu Zhu 0001, Zixiang Xiong, Ya-Qin Zhang
IEEE Trans. Circuits Syst. Video Technol.2
1999 Inverse halftoning using wavelets
abstract
This work introduces a new approach to inverse halftoning using nonorthogonal wavelets. The distinct features of this wavelet-based approach are: 1) edge information in the highpass wavelet images of a halftone image is extracted and used to assist inverse halftoning, 2) cross-scale correlations in the multiscale wavelet decomposition are used for removing background halftoning noise while preserving important edges in the wavelet lowpass image, and 3) experiments show that our simple wavelet-based approach outperforms the best results obtained from inverse halftoning methods published in the literature, which are iterative in nature.
Zixiang Xiong, Michael T. Orchard, Kannan Ramchandran
IEEE Trans. Image Process.1
1998 Multiresolutional encoding and decoding in embedded image and video coders
abstract
We address multiresolutional encoding and decoding within the embedded zerotree wavelet (EZW) framework for both images and video. By varying a resolution parameter, one can obtain decoded images at different resolutions from one single encoded bitstream, which is already rate scalable for EZW coders. Similarly one can decode video sequences at different rates and different spatial and temporal resolutions from one bitstream. Furthermore, a layered bitstream can be generated with multiresolutional encoding, from which the higher resolution layers can be used to increase the spatial/temporal resolution of the images/video obtained from the low resolution layer. In other words, we have achieved full scalability in rate and partial scalability in space and time. This added spatial/temporal scalability is significant for emerging multimedia applications such as fast decoding, image/video database browsing, telemedicine, multipoint video conferencing, and distance learning.
Zixiang Xiong, Beong-Jo Kim, William A. Pearlman
ICASSP1
1998 Joint Source-Channel Image Coding for a Power Constrained Noisy Channel
abstract
We study joint source-channel coding for a power constrained Gaussian channel and its application to progressive image compression. For a given power constrained, we consider the optimum allocation of energy per bit for a BPSK transmitter and the best choice of channel code rate, when the performance is measured by end-to-end average quantizer distortion. Choosing the average energy per transmitted bit in conjunction with both the source rate and the channel code rate provides an additional degree of freedom with respect to previously proposed schemes, and therefore can achieve higher overall PSNRs for images.
Marc P. C. Fossorier, Zixiang Xiong, Kenneth Zeger
ICIP (2)2
1998 Progressive Video Coding for Noisy Channels
abstract
We extend the work of Sherwood and Zeger (IEEE Signal Processing Letters, vol.4, p.189-91, 1997) to progressive video coding for noisy channels. By utilizing a three dimensional (3D) extension of the set partitioning in hierarchical trees (SPIHT) algorithm, we cascade the resulting 3D SPIHT video coder with the rate-compatible punctured convolutional (RCPC) channel coder for transmission of video over a binary symmetric channel (BSC). Progressive coding is achieved by increasing the target rate of the 3D embedded SPIHT video coder as the channel condition improves. The performance of our proposed coding system is acceptable at a low transmission rate and bad channel conditions. Its low complexity makes it suitable for emerging applications such as video over wireless channels.
Zixiang Xiong, Beong-Jo Kim, William A. Pearlman
ICIP (1)1
1998 Multiresolution Watermarking for Images and Video: A Unified Approach
abstract
This paper proposes a unified approach to digital watermarking of images and video based on the 2D and 3D discrete wavelet transforms. The hierarchical nature of the wavelet representation allows multiresolutional detection of the digital watermark, which is a Gaussian distributed random vector added to all the high pass bands in the wavelet domain. We show that, when subjected to distortion from compression or image halftoning, the corresponding watermark can still be correctly identified at each resolution (excluding the lowest one) in the wavelet domain. Computational saving from such a multiresolution watermarking framework is obvious, especially for the video case.
Wenwu Zhu 0001, Zixiang Xiong, Ya-Qin Zhang
ICIP (1)2
1998 Wavelet image coding using trellis coded space-frequency quantization
abstract
We propose a new wavelet image coding technique with trellis coded space-frequency quantization (TCSFQ). Experiments indicate that the TCSFQ coder is better than all other coders described in the literature.
Zixiang Xiong, Xiaolin Wu 0001
MMSP1
1998 Progressive coding of medical volumetric data using three-dimensional integer wavelet packet transform
abstract
We examine progressive lossy to lossless compression of medical volumetric data using three-dimensional (3D) integer wavelet packet transforms and set partitioning in hierarchical trees (SPIHT). To achieve good lossy coding performance, we describe a 3D integer wavelet packet transform that allows implicit bit shifting of wavelet coefficients to approximate a 3D unitary transformation. We also address context modeling for efficient entropy coding within the SPIHT framework. Both lossy and lossless coding performance are better than those previously reported.
Zixiang Xiong, Xiaolin Wu 0001, David Y. Y. Yun, William A. Pearlman
MMSP1
1998 Wavelet packet image coding using space-frequency quantization
abstract
We extend our previous work on space-frequency quantization (SFQ) for image coding from wavelet transforms to the more general wavelet packet transforms. The resulting wavelet packet coder offers a universal transform coding framework within the constraints of filterbank structures by allowing joint transform and quantizer design without assuming a priori statistics of the input image. In other words, the new coder adaptively chooses the representation to suit the image and the quantization to suit the representation. Experimental results show that, for some image classes, our new coder gives excellent coding performance.
Zixiang Xiong, Kannan Ramchandran, Michael T. Orchard
IEEE Trans. Image Process.1
1997 Global motion compensation for low bitrate video coding
abstract
A global motion compensation (GMC) scheme is described and implemented for low bitrate coding. Experiments show that GMC gives excellent coding results. When coupled with the H.263 low bitrate video coding standard, GMC is capable of achieving up to 27% bitrate saving (or 1.8 dB gain in PSNR) for a high motion sequence such as Stefan. For a moderate motion sequence such as Cost-guard, we obtained an 8% bitrate saving over H.263 at the same quality using GMC.
Zixiang Xiong, Tihao Chiang, Ya-Qin Zhang
MMSP1
1997 A deblocking algorithm for JPEG compressed images using overcomplete wavelet representations
abstract
This paper introduces a new approach to deblocking of JPEG compressed images using overcomplete wavelet representations. By exploiting cross-scale correlations among wavelet coefficients, edge information in the JPEG compressed images is extracted and protected, while blocky noise in the smooth background regions is smoothed out in the wavelet domain. Compared with the iterative methods reported in the literature, our simple wavelet-based method has much lower computational complexity, yet it is capable of achieving the same peak signal-to-noise ratio (PSNR) improvement as the best iterative method and giving visually very pleasing images as well.
Zixiang Xiong, Michael T. Orchard, Ya-Qin Zhang
IEEE Trans. Circuits Syst. Video Technol.1
1997 Joint space-frequency segmentation using balanced wavelet packet trees for least-cost image representation
abstract
We examine the question of how to choose a space varying filterbank tree representation that minimizes some additive cost function for an image. The idea is that for a particular cost function, e.g., energy compaction or quantization distortion, some tree structures perform better than others. While the wavelet tree represents a good choice for many signals, it is generally outperformed by the best tree from the library of wavelet packet frequency-selective trees. The double-tree library of bases performs better still, by allowing different wavelet packet trees over all binary spatial segments of the image. We build on this foundation and present efficient new pruning algorithms for both one- and two-dimensional (1-D and 2-D) trees that will find the best basis from a library that is many times larger than the library of the single-tree or double-tree algorithms. The augmentation of the library of bases overcomes the constrained nature of the spatial variation in the double-tree bases, and is a significant enhancement in practice. Use of these algorithms to select the least-cost expansion for images with a rate-distortion cost function gives a very effective signal adaptive compression scheme. This scheme is universal in the sense that, without assuming a model for the signal or making use of training data, it performs very well over a large class of signal types. In experiments it achieves compression rates that are competitive with the best training-based schemes.
Cormac Herley, Zixiang Xiong, Kannan Ramchandran, Michael T. Orchard
IEEE Trans. Image Process.2
1997 Space-frequency quantization for wavelet image coding
abstract
A new class of image coding algorithms coupling standard scalar quantization of frequency coefficients with tree-structured quantization (related to spatial structures) has attracted wide attention because its good performance appears to confirm the promised efficiencies of hierarchical representation. This paper addresses the problem of how spatial quantization modes and standard scalar quantization can be applied in a jointly optimal fashion in an image coder. We consider zerotree quantization (zeroing out tree-structured sets of wavelet coefficients) and the simplest form of scalar quantization (a single common uniform scalar quantizer applied to all nonzeroed coefficients), and we formalize the problem of optimizing their joint application. We develop an image coding algorithm for solving the resulting optimization problem. Despite the basic form of the two quantizers considered, the resulting algorithm demonstrates coding performance that is competitive, often outperforming the very best coding algorithms in the literature.
Zixiang Xiong, Kannan Ramchandran, Michael T. Orchard
IEEE Trans. Image Process.1
1996 Inverse halftoning using wavelets
abstract
This paper introduces a new approach to inverse halftoning using nonorthogonal wavelets. The distinct features of this wavelet-based approach are: a) edge information in the highpass wavelet images of a halftone is extracted and used to assist inverse halftoning, b) cross-scale correlations in the multiscale wavelet decomposition are used for removing background halftoning noise while preserving important edges in the wavelet lowpass image, c) experiments show that our simple wavelet-based approach outperforms the best results obtained from inverse halftoning methods published in the literature, which are iterative in nature.
Zixiang Xiong, Michael T. Orchard, Kannan Ramchandran
ICIP (1)1
1996 A DCT-based embedded image coder
abstract
Since Shapiro (see ibid., vol.41, no.12, p. 445, 1993) published his work on embedded zerotree wavelet (EZW) image coding, there have been increased research activities in image coding centered around wavelets. We first point out that the wavelet transform is just one member in a family of linear transformations, and the discrete cosine transform (DCT) can also be coupled with an embedded zerotree quantizer. We then present such an image coder that outperforms any other DCT-based coder published in the literature, including that of the Joint Photographers Expert Group (JPEG). Moreover, our DCT-based embedded image coder gives higher peak signal-to-noise ratios (PSNR) than the quoted results of Shapiro's EZW coder.
Zixiang Xiong, Onur G. Guleryuz, Michael T. Orchard
IEEE Signal Process. Lett.1
1996 Adaptive transforms for image coding using spatially varying wavelet packets
abstract
We introduce a novel, adaptive image representation using spatially varying wavelet packets (WPs), Our adaptive representation uses the fast double-tree algorithm introduced previously (Herley et al., 1993) to optimize an operational rate-distortion (R-D) cost function, as is appropriate for the lossy image compression framework. This involves jointly determining which filter bank tree (WP frequency decomposition) to use, and when to change the filter bank tree (spatial segmentation). For optimality, the spatial and frequency segmentations must be done jointly, not sequentially. Due to computational complexity constraints, we consider quadtree spatial segmentations and binary WP frequency decompositions (corresponding to two-channel filter banks) for application to image coding. We present results verifying the usefulness and versatility of this adaptive representation for image coding using both a first-order entropy rate-measure-based coder as well as a powerful space-frequency quantization-based (SPQ-based) wavelet coder introduced by Xiong et al. (1993).
Kannan Ramchandran, Zixiang Xiong, Kohtaro Asai, Martin Vetterli
IEEE Trans. Image Process.2
1995 An efficient algorithm to find a jointly optimal time-frequency segmentation using time-varying filter banks
abstract
We examine the question of how to choose a time-varying filter bank representation for a signal which is optimal with respect to an additive cost function. We present in detail an efficient algorithm for the Haar filter set which finds the optimal basis, given the constraint that the time and frequency segmentations are binary. Extension to multiple dimensions is simple, and the use of arbitrary filter sets is also possible. We verify that the algorithm indeed produces a lower cost representation than any of the wavelet packet representations for compression of images using a simple rate-distortion cost.
Cormac Herley, Zixiang Xiong, Kannan Ramchandran, Michael T. Orchard
ICASSP2
1995 Space-frequency quantization for a space-varying wavelet packet image coder
abstract
We introduce a new image coding algorithm which exploits the idea of space-varying wavelet packets, where the best filter bank representation is chosen from a large library. The filter bank tree representations in the library are free to vary in structure over different segments of the image, and a fast search algorithm is given. In addition we employ the idea of space-frequency quantization, which is a rate-distortion optimized extension of the zero-tree wavelet coder of Shapiro to wavelet packets. The coder thus adaptively chooses the representation to suit the image and adaptively chooses the quantization to suit the representation. We present coding results that confirm the excellent performance of the scheme.
Zixiang Xiong, Cormac Herley, Kannan Ramchandran, Michael T. Orchard
ICIP1
1994 Wavelet Packets-Based Image Coding Using Joint Space-frequency Quantization
abstract
A novel quantization scheme targeted at jointly optimizing the spatial and frequency characterization of the wavelet representation of images was introduced in Xiong et al. (1993) for image compression applications. The present authors extend the concept of joint space-frequency quantization (SFQ) to the more flexible class of wavelet packet representations (Coifman and Wickerhauser, 1992), which are a generalization of the multiresolution decomposition using the wavelet transform. They propose a fast algorithm to jointly search for the best wavelet packet basis and space-frequency quantizer, presenting empirical evidence of its high performance (e.g., for the "Barbara" image coded at 0.5 b/p, they get a 0.7 dB gain in PSNR over the fixed-wavelet based SFQ of Xiong et al. and 1.5 dB over Shapiro's embedded wavelet coder (Shapiro, 1993)).>
Zixiang Xiong, Kannan Ramchandran, Michael T. Orchard, Kohtaro Asai
ICIP (3)1
1993 Marginal analysis prioritization for image compression based on a hierarchical wavelet decomposition
Zixiang Xiong, Nikolas P. Galatsanos, Michael T. Orchard
ICASSP (5)1